Serialized image segmentation method based on space-time Swin Transform fusion
By adopting the spatial-temporal SwinTransformer fusion method in image segmentation task, the problem that convolutional neural networks cannot learn the context dependence between serialized image slices is solved, and richer spatial and temporal feature interaction is achieved, segmentation performance is improved, and features are efficiently fused in linear time.
Patent Information
- Application Number
- CN202510191562.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-23
AI Technical Summary
卷积神经网络由于其操作的局部性,无法很好地学习序列化图像切片间的上下文依赖关系,且原始Vision Transformer在捕捉补丁之间长距离依赖关系时忽视了单张切片中局部特征的提取,导致时间成本上升。
Using the serialized image segmentation method based on spatiotemporal SwinTransformer fusion, spatial features are extracted from a single image through the Spatial-SwinTransformer model and arranged them in chronological order, and the information between image slices is captured in combination with the Temporal-ViT model. Finally, feature fusion is performed between the spatial feature extraction module and the temporal feature extraction module using the cross-attention mechanism.
This method not only compensates for the possible spatial feature loss problem in Vision Transformer in the time dimension, but also improves the overall segmentation performance of the model, can more comprehensively capture complex changes in key parts of the image, and efficiently integrate spatiotemporal features in linear time.
Smart Images

Figure CN120032131A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a serialized image segmentation method based on spatiotemporal Swin Transformer fusion, and belongs to the field of image segmentation. Background Art
[0002] In image recognition tasks, including object detection, image classification and segmentation, activity recognition and other problems, excellent performance can be easily achieved by using deep neural networks (full name in English, DNN). Among them, image segmentation is the task of classifying the pixels of an image into different labels, which is divided into three types: semantic segmentation, instance segmentation and panoramic segmentation. This paper focuses on exploring the semantic segmentation task. In semantic segmentation, pixels are assigned labels according to the objects or structures they belong to, without distinguishing instances of the same category. Among them, the segmentation method of serialized images refers to the technology that uses the temporal or spatial correlation between adjacent slices in the image sequence to improve the segmentation accuracy. These methods attempt to capture and utilize the information between multiple consecutive images to enhance the understanding and segmentation of the target structure.
[0003] Recurrent Neural Network (RNN) is a design for processing sequence data, and its typical representative is Long Short Term Memory (LSTM). Tseng et al. [1] proposed a deep encoder-decoder structure with a cross-modal convolutional layer, using convolutional LSTM to model two-dimensional slice sequences, integrating different modalities and modeling 2D slice sequences in an end-to-end manner, and solving the label imbalance problem through a reweighting scheme and two-stage training. Cai et al. [2] proposed a new recurrent neural network architecture, which combines a deep convolutional subnetwork with a long short-term memory recurrent layer to perform context learning and segmentation consistency optimization in an end-to-end manner. The results show that this method performs better than the existing state-of-the-art methods in segmentation. These studies show that convolutional long short-term memory networks have advantages and potential in image sequence processing.
[0004] Although convolutional neural networks perform well, they are unable to learn global and long-range semantic information interactions well due to the local nature of their operations. Vaswani et al. [3] proposed a deep learning architecture for natural language processing tasks, the Transformer. Unlike traditional recurrent neural networks, the Transformer uses a self-attention mechanism to capture the dependencies between elements in a sequence. The Transformer architecture consists of multiple identical layers, each of which contains two sublayers: a multi-head self-attention mechanism and a feedforward network. The multi-head self-attention mechanism allows the network to measure the importance of different elements in a sequence. This is achieved by calculating an attention score for each element in the sequence relative to all other elements. The feedforward network processes the information from the attention mechanism, which includes two linear transformations followed by a nonlinear activation function. The Transformer's self-attention mechanism and feedforward network work together to process the input sequence and produce an output sequence. The output of the final layer is used to make predictions, such as language modeling or machine translation. The Transformer's use of the self-attention mechanism and its ability to process input sequence data in parallel have been proven to be very effective, leading to its widespread application in various natural language processing (NLP) tasks and becoming an important cornerstone in the field of natural language processing.
[0005] Inspired by the Transformer architecture, Dosovitskiy et al. [4] applied the architecture to computer vision tasks, resulting in the creation of the Vision Transformer (ViT). Unlike traditional convolutional neural networks that process image data using convolutional layers, ViT processes image data as a sequence of feature vectors extracted from image patches. This enables ViT to leverage the Transformer's ability to process sequential data and handle long-range dependencies between patches within an image. ViT's network architecture is similar to the Transformer and consists of multiple identical layers. Each layer contains two sublayers: a multi-head self-attention mechanism and a feedforward network. The ViT architecture has been successfully applied to a variety of computer vision tasks, including image classification, object detection, and segmentation. It is able to process image data as a sequence of feature vectors, thereby effectively capturing long-range dependencies between patches within an image, and performs better than traditional CNNs.
[0006] The original Vision Transformer is good at capturing long-range dependencies between patches, but neglects local feature extraction because the two-dimensional patches are projected into vectors through simple linear layers. Therefore, some researchers began to focus on improving the modeling ability of local information. Han et al. [5] proposed the TNT model, further dividing the patch into multiple sub-patches and introducing a novel inside-outside Transformer architecture, using the inner Transformer block to model the relationship between sub-patches and the outer Transformer block to exchange information at the patch level, thereby enhancing the feature representation ability. Liu et al. [6] improved the ViT model to obtain the SwinTransformer model, which efficiently handles local and cross-window connections in the image through a layered architecture and a shifting window mechanism, realizes the modeling of visual entities of different scales, and significantly surpasses the previous state-of-the-art level in a variety of computer vision tasks. These works demonstrate the benefits of local information exchange and global information exchange in the Vision Transformer. In addition, as a key component of the Transformer, the self-attention layer provides the ability to globally interact between image patches. Improving the calculation of the self-attention layer has attracted many researchers. Chen et al. [7] proposed a dual-branch Transformer architecture that processes image patches of different sizes separately and fuses these patches through an attention mechanism at multiple stages to generate stronger multi-scale feature representations. For efficient fusion, a token fusion module based on cross-attention is introduced, which significantly reduces computational and memory complexity.
[0007] At present, the Vision Transformer provides a new solution to the problem that convolutional neural networks cannot learn the contextual dependencies between serialized image slices well due to the locality of their operations. For example, the Vision Transformer can extract key features from information obtained from multiple cameras and integrate this information to help the car make correct driving decisions. This mechanism allows the model to consider the information of the entire input sequence when processing the data at each position, so that it can learn richer contextual dependencies. Unlike CNN, the self-attention mechanism in ViT can pay attention to all pixels or patches at once, which enables it to better capture long-range dependencies in the image. At the same time, since the computational complexity of ViT depends mainly on the sequence length rather than the image size, they also have good performance when processing high-resolution images. However, the original Vision Transformer often ignores the extraction of local features in a single slice when capturing long-range dependencies between patches. In addition, the computational complexity of the self-attention mechanism is quadratic with the sequence length, which means that the amount of computation increases sharply when the input sequence becomes longer. This will lead to a significant increase in time cost for processing long time series or high-resolution images. Therefore, for serialized images, the present invention first trains a SwinTransformer model for the spatial dimension of a single image on the preprocessed data set. Based on the spatial feature maps of multiple single images obtained, they are arranged in chronological order, and the Swin Transformer model for serialized images is trained. Secondly, this method designs a serialized image segmentation model that integrates spatiotemporal feature maps, and uses the self-attention mechanism in SwinTransformer to capture the spatial features of a single image and the temporal features of a serialized image. Finally, the cross-attention mechanism fully integrates the advantages of the spatial feature extraction module and the temporal feature extraction module, and efficiently integrates the two in linear time, which not only makes up for the loss of spatial features of serial images in the VisionTransformer of the temporal dimension, but also greatly reduces the time cost.
[0008] Therefore, the present invention proposes a serialized image segmentation method based on spatiotemporal SwinTransformer fusion. For the original image sequence, the Spatial-SwinTransformer model is first used to extract the spatial feature map of each image from a single image. The spatial features obtained in this process are linearly projected and added to the temporal feature extraction model, which enhances the representation of the serialized image so that it contains richer spatial information features. Subsequently, the Temporal-ViT model is used to process the serialized image to model the information between image slices and capture the dependencies of discriminative regions in images at different time points. Finally, the cross-attention mechanism is used to perform feature fusion between the spatial feature extraction module and the temporal feature extraction module to exchange information, and the attention map of the fusion module is efficiently generated. The interaction process helps to fuse mask area features of different dimensions in linear time. Summary of the invention
[0009] Aiming at the problem that Vision Transformer ignores the extraction of local features in a single slice when capturing long-distance dependencies between slices and the long input sequence leads to increased time cost, the present invention proposes a serialized image segmentation method based on spatiotemporal SwinTransformer fusion.
[0010] The present invention includes three steps, as follows: Step 1 uses the Spatial-SwinTransformer model to extract spatial features from a single image. The model captures local and global structural information in a single image through a multi-head self-attention mechanism and a multi-layer perceptron to generate a spatial feature map for each image; Step 2 is to obtain a temporal feature map in the time dimension. The spatial features obtained in the previous step are arranged in time order, linearly projected and added to the temporal feature extraction model, which enhances the representation of the serialized image so that it contains richer spatial information features. Subsequently, the Temporal-ViT model is used to process the serialized image to model the information between image slices and capture the dependencies of discriminative regions in images at different time points; Step 3 performs feature fusion between the spatial feature extraction module and the temporal feature extraction module through a cross-attention mechanism to exchange information, and efficiently generates an attention map of the fusion module, and its interactive process helps to fuse mask area features of different dimensions in linear time.
[0011] The specific scheme of the present invention is as shown in the attached Figure 1 shown.
[0012] Step 1: Spatial feature extraction module
[0013] The spatiotemporal Swin Transformer fusion model proposed in the present invention adopts the Swin-UNet architecture, which consists of an encoder, a bottleneck layer, a decoder and skip connections.
[0014] Given an image sequence, which contains n single images. In the model, a single image is regarded as a visual sentence. In order to convert the input into block embedding, the single image is divided into 3×3=9 non-overlapping blocks, and the feature dimension of each block becomes 3×3×3=27. In addition, the feature dimension is projected to an arbitrary dimension (denoted as C) through a linear embedding layer. Among them, a visual sentence consists of a series of visual words.
[0015] The spatiotemporal fusion model has two data streams: one operates between visual sentences, and the other processes visual words within each sentence. For word embedding, this method utilizes a SwinTransformer block to explore the relationship between visual words.
[0016] For the encoder, in order to convert the input into block embeddings, the image is segmented into non-overlapping blocks. In addition, the feature dimension is projected to an arbitrary dimension (denoted as C) through a linear embedding layer. The converted blocks are passed through multiple SwinTransformer modules and patch merging layers to generate hierarchical feature representations. Specifically, the block merging layers are responsible for downsampling and increasing the dimension, while the SwinTransformer modules are responsible for feature representation learning. The converted blocks are passed through multiple SwinTransformer modules and patch merging layers to generate hierarchical feature representations. Specifically, the block merging layers are responsible for downsampling and increasing the dimension, while the SwinTransformer modules are responsible for feature representation learning.
[0017] Step 2: Temporal feature extraction module
[0018] Based on the hierarchical spatial feature representation of multiple time points obtained by the spatial feature extraction module, the spatial feature maps of all time points are arranged in the order of the image data acquisition, and the position code is added in this order. The time series feature matrix with position code is put into the Swin Transformer module of the time dimension as input, and the rest is the same as the Swin Transformer in the spatial feature extraction module.
[0019] After the SwinTransformer block in the time dimension, a new set of temporal feature maps is obtained. This feature map not only retains the original spatial information, but also combines the changing process of the image in the time dimension, fully considering the temporal relationship. This chapter adopts a symmetrical Transformer-based decoder. The decoder consists of a SwinTransformer module and a block expansion layer. The extracted image features are fused with multi-scale features from the encoder through jump connections to supplement the spatial information loss caused by downsampling. In contrast to the block merging layer, this paper uses a block expansion layer in the decoder to perform upsampling. The block expansion layer reshapes the feature maps of adjacent dimensions into large feature maps with a resolution magnified by 2 times. This method uses the last patch expansion layer to perform 4x upsampling to restore the resolution of the feature map to the input resolution (W×H), and then applies a linear projection layer on these upsampled features to output pixel-level segmentation predictions. Such fusion ensures that the model can consider both spatial and temporal information and improve the recognition accuracy of discriminative areas.
[0020] Step 3: Spatiotemporal feature fusion module
[0021] Based on the above two modules, the present invention proposes a simple and effective cross attention module to fuse the information between the spatial feature extraction module and the temporal feature extraction module. Each Transformer module exchanges information with other modules through attention, generating an attention map in the fusion process in linear time.
[0022] In order to fuse spatiotemporal features more efficiently and effectively, this method first uses the patchtoken of each module to exchange information with the patchtokens of other modules, and then reprojects it to its own module. Since the patchtoken has learned the abstract information of all the patchtokens in its own module, the interaction with the patchtokens of another module helps to contain information of different dimensions. After being fused with the tokens of other modules, the patchtoken interacts with its own patchtokens again in the encoder of the next Transformer, which is able to pass the information learned from other branches to its own patchtokens, enriching the representation of each patchtoken.
[0023] The spatial feature extraction module and the temporal feature extraction module can be regarded as branches Spa and Tem, respectively. For branch Spa, the same process can be performed by simply exchanging the indexes Tem and Spa.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] First, ordinary CNN-based segmentation methods can usually only process information in the spatial or temporal dimensions separately, and cannot fully capture the complex change patterns of the key parts of the image, and it is difficult to obtain long-distance contextual dependencies. Therefore, the present invention proposes a serialized segmentation model fused with spatiotemporal SwinTransformer. First, the spatial feature map of a single image is obtained in the spatial dimension, and the global and local features of the single image are comprehensively considered. Then, the obtained multiple spatial feature maps are arranged in chronological order to obtain the temporal feature map of the sequence image. Finally, the feature map in the spatial dimension and the temporal feature map with contextual dependencies are fused through cross attention. The advantages of this method are as follows: (1) Comprehensively capture the changes in key areas: By combining information in the spatial and temporal dimensions, the model can more comprehensively capture the complex change patterns of the key parts of the image. This includes not only the local and global features within a single image, but also the changing trends of these features over time. (2) Efficient spatiotemporal feature fusion: It not only makes up for the spatial feature loss problem that may occur in the Vision Transformer in the temporal dimension, but also improves the overall segmentation performance of the model.
[0026] Second, the present invention proposes a method for fusing spatiotemporal features through cross-attention, which can achieve enhanced spatial and temporal feature interaction on the one hand: spatial or temporal features alone may not be able to fully describe key areas, and cross-attention allows the model to establish a direct connection between spatial and temporal dimensions, thereby more accurately capturing key image areas and their changing patterns over time. This two-way information exchange ensures that both local and global features at each point in time are fully considered. It enables the model to not only focus on the details in the current image, but also take into account the impact of contextual dependencies. On the other hand, it can be efficiently fused in linear time: the cross-attention mechanism can achieve efficient feature fusion while maintaining computational efficiency. This means that even in the face of large-scale data sets or high-resolution images, the model can respond quickly and generate accurate results. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is the overall model diagram of the present invention.
[0028] Figure 2 For cross-attention.
[0029] Figure 3 Cross attention module for Tem branch.
[0030] Figure 4 Figure 2 is an image segmentation method based on spatiotemporal Swin Transformer fusion. DETAILED DESCRIPTION
[0031] The following is a detailed description of the implementation examples of the present invention in conjunction with the accompanying drawings:
[0032] The present invention is a serialized image segmentation method based on spatiotemporal Swin Transformer fusion. Step 1 uses a spatial-dimensional Spatial-Swin Transformer model to extract spatial features from a single image. The model captures local and global structural information in a single image through a multi-head self-attention mechanism and a multi-layer perceptron to generate a spatial feature map of each image; Step 2 is to obtain a temporal feature map in the temporal dimension. The spatial features obtained in the previous step are arranged in chronological order, linearly projected and added to the temporal feature extraction model, which enhances the representation of the serialized image so that it contains richer spatial information features. Subsequently, the Temporal-ViT model in the temporal dimension is used to process the serialized image to model the information between image slices and capture the dependency of discriminative regions in images at different time points; Step 3 performs feature fusion between the spatial feature extraction module and the temporal feature extraction module through a cross-attention mechanism to exchange information, and efficiently generates an attention map of the fusion module, and its interactive process helps to fuse mask area features of different dimensions in linear time.
[0033] Specifically, the method comprises the following steps:
[0034] Step 1: Spatial feature extraction module
[0035] The spatiotemporal Swin Transformer fusion model proposed in this paper adopts the Swin-UNet architecture, which consists of an encoder, a bottleneck layer, a decoder, and skip connections.
[0036] Given an image sequence consisting of n single images:
[0037] E=[L 1 ,L 2 ,...,L n ]∈R n×p×p×3 (1)
[0038] Among them, (p, p) is the resolution of each image, n is the number of single images in the sequence, R is a real number array, and E is an image sequence.
[0039] In the model, a single image is regarded as a visual sentence. In order to convert the input into block embedding, the single image is divided into 3×3=9 non-overlapping blocks, and the feature dimension of each block becomes 3×3×3=27. In addition, the feature dimension is projected to an arbitrary dimension (denoted as C) through a linear embedding layer. Among them, a visual sentence consists of a series of visual words:
[0040] L i =[l i,1 ,l i,2,...,l i,9 ] (2)
[0041] Among them, l i,j ∈R s×s×3 is the jth visual word of the ith visual sentence, (s, s) is the spatial size of the visual words (non-overlapping blocks generated by segmentation), j = 1, 2, ..., 9.
[0042] The converted blocks are passed through multiple Swin Transformer modules and patch merging layers to generate hierarchical feature representations. Specifically, the patch merging layer is responsible for downsampling and increasing the dimension, while the Swin Transformer module is responsible for feature representation learning.
[0043] The spatiotemporal fusion model has two data streams: one operates between visual sentences, and the other processes visual words within each sentence. For word embedding, the present invention uses a Swin Transformer block to explore the relationship between visual words.
[0044] Convert visual words into word embedding sequences via linear projection:
[0045] Q i =[q i,1 ,q i,2 ,...,q i,9 ],q i,j =FC(Vec(l i,j )) (3)
[0046] Among them, q i,j ∈R C is the j-th word embedding, C is the dimension of word embedding, Vec(·) is the vectorization operation, and FC represents the output of the fully connected layer.
[0047] Position encoding: Spatial information is an important factor in image recognition. For the visual words in a sentence, this paper adds the corresponding word position encoding to each word embedding to retain the spatial information, as shown in formula (4), using the standard one-dimensional position encoding.
[0048]
[0049] Among them, E word ∈R 9×C It is the word position code shared in each sentence. The word position code is used to save the local relative position. It contains the codes of 9 positions (corresponding to the 9 visual words in the sentence), and each code is a C-dimensional vector. i Represents the original query vector of the i-th visual word, before adding the position encoding, usually generated by the word embedding feature extraction method. Denote the query vector of the $i$-th visual word after adding the position encoding, which is the result of adding the original query vector and the position encoding, integrating the content and position information of the word.
[0050] Next, the present invention will introduce each module in detail:
[0051] Swin Transformer block: Different from the traditional multi-head self-attention (MSA) module, the Swin Transformer block is constructed based on shifted windows. As Figure 1 shown, each Swin Transformer block consists of a LayerNorm (LN) layer, a multi-head self-attention module, a residual connection, and a two-layer MLP with GELU non-linearity. Two consecutive Swin Transformer blocks apply a window-based multi-head self-attention (W-MSA) module and a shifted-window-based multi-head self-attention (SW-MSA) module respectively. Among them, MSA is a standard global self-attention mechanism, calculating the global relationship between all input patches. While W-MSA divides the input feature map into non-overlapping windows and calculates self-attention within each window. Further, SW-MSA introduces a sliding window mechanism on the basis of W-MSA, enabling information interaction between different windows by sliding the window positions between different layers. According to this window partitioning mechanism, consecutive Swin Transformer blocks can be regarded as Spatial-Swin Transformer blocks, formulated as:
[0052]
[0053] where and $Q$ m represent the outputs of the $m$-th (S)W-MSA module and the $m$-th MLP module respectively. W-MSA is a window-based multi-head self-attention operation, SW-MSA is a shifted-window-based multi-head self-attention operation, LN is the layer normalization process, and MLP represents a multi-layer perceptron.
[0054] The self-attention formula is as follows:
[0055]
[0056] where, as shown in formula (9), in the self-attention module, the input is linearly transformed into three parts, namely the query the key and the value $M$ 2 represents the number of patches in the window, $d, d$ k , $d$ vare the dimensions of input, query (key), and value, respectively. W-MSA and SW-MSA introduce the relative position information between windows through an additional matrix X in formula (9), and obtain formula (10). The value of X is taken from the bias matrix This enables the model to capture the local features within the window and the spatial relationship between different windows. M is the size of the window. For a window of size M×M, the relative position range between patches is [-M+1,M-1] in both horizontal and vertical directions. R is a set of real numbers.
[0057] As shown in formula (10), the spatial feature extraction model applies an attention mechanism to the input features and adjusts the importance weights by calculating the dot product of the query matrix B, the key matrix K, and the value matrix V. In order to enhance the information interaction between windows, the Swin Transformer introduces a relative position bias matrix X. Specifically, by calculating An attention score is obtained, which represents the correlation between each local area in the image and other areas, and then it is converted into a probability distribution using a softmax function, and finally multiplied with the value matrix V to obtain a weighted sum output. Since the purpose of the present invention is image segmentation, the spatial feature extraction model gradually extracts and refines features at different levels by stacking 24 Swin Transformer blocks. The Swin Transformer architecture contains 4 stages, stage 1: 2 Swin Transformer blocks. Stage 2: 2 Swin Transformer blocks. Stage 3: 18 Swin Transformer blocks. Stage 4: 2 Swin Transformer blocks, a total of 24 Swin Transformer blocks. In the shallow blocks (Swin Transformer blocks 1 to 4), the window is small, and the attention mechanism mainly operates in the local area, which is mainly responsible for capturing local features. In the deep blocks (Swin Transformer blocks 5 to 24), as the number of layers increases, the receptive field of the feature map gradually expands, and the model can capture a wider range of contextual information. The MLP of each layer further transforms the features, enhances the expressive power of the model, and enables the model to learn more complex patterns. The above feature extraction process calculates the interaction between any two visual words, that is, calculates their semantic relevance through the self-attention mechanism (focusing on the key areas), so that the model can accurately capture the local features of the key areas in the spatial dimension and the global features of the entire pneumonia image.
[0058] Finally, a linear layer is used to generate the output. Multi-head self-attention splits the query, key, and value and performs the attention function in parallel, then concatenates the output values of each head and forms the final output through linear projection.
[0059] Multilayer Perceptron (MLP): MLP is used for feature transformation and nonlinear operations between self-attention layers.
[0060] MLP(X)=FC(σ(FC(X))),FC(X)=XT+b (11)
[0061] Where T and b are the weight and bias of the fully connected layer, respectively; σ(·) is the activation function, such as GELU; FC represents the output of the fully connected layer.
[0062] Layer Normalization (LN): Layer Normalization is a key part of Transformer for stable training and faster convergence. Layer Normalization is applied to each sample x∈R d ,as follows:
[0063]
[0064] Where μ∈R, δ∈R are the mean and standard deviation of the feature respectively, represents element-wise multiplication, γ∈R d ,β∈R d are the affine transformation parameters, and R represents the set of real numbers.
[0065] Encoder: In the encoder, the C-dimensional tokenized input with a resolution of H / 4×W / 4 is fed into two consecutive SwinTransformer blocks for feature learning, and the feature dimension and resolution remain unchanged. At the same time, the block merging layer reduces the number of patches (2x downsampling) and increases the feature dimension to 2x. This process is repeated three times in the encoder.
[0066] Bottleneck layer: Since Transformer is too deep to converge, the bottleneck layer only uses two consecutive SwinTransformer blocks to learn deep feature representations, and the feature dimension and resolution remain unchanged.
[0067] Decoder: Corresponding to the encoder, the decoder builds a symmetrical structure based on SwinTransformer blocks. To this end, block expansion layers are used in the decoder instead of block merging layers in the encoder for upsampling. The block expansion layer reshapes the feature maps of adjacent dimensions into higher resolution feature maps (2x upsampling) and reduces the feature dimension by half. For example, the first block expansion layer applies a linear layer to increase the feature dimension of the input feature (W / 32×H / 32×8C) to twice the original (W / 32×H / 32×16C) before upsampling, and then uses a rearrange operation to expand the resolution of the input feature by 2 times and reduce the feature dimension to one-fourth of the input dimension (W / 32×H / 32×16C→W / 16×H / 16×4C).
[0068] Skip connections: Skip connections are used to fuse the multi-scale features in the encoder with the upsampled features, concatenating shallow features with deep features to reduce the loss of spatial information caused by downsampling. A linear layer is then applied to make the dimension of the concatenated features the same as the dimension of the upsampled features.
[0069] Step 2: Temporal feature extraction module
[0070] Based on the hierarchical spatial feature representation of multiple time points obtained by the spatial feature extraction module, the spatial feature maps of all time points are arranged in the order of the image data acquisition, and the position code is added in this order. The time series feature matrix with position code is put into the Swin Transformer module of the time dimension as input, and the rest is the same as the Swin Transformer in the spatial feature extraction module.
[0071] For visual sentences, this paper creates sentence embedding memory to store sentence-level representation sequences:
[0072]
[0073] Among them, the category label P class Initialized to zero, used to aggregate global sentence information; n is the number of words in the sentence, d is the embedding dimension, and the extra dimension comes from the category label P class , R represents the set of real numbers.
[0074] At each layer, the word embedding sequence is transformed to the sentence embedding domain via linear projection and added to the sentence embedding:
[0075] P m-1 =P′ m-1 +FC(Vec(Q m+1 ))(14)
[0076] Among them, P m-1 , P′ m-1 ∈R d , R represents the set of real numbers, d is the embedding dimension, and the fully connected layer FC ensures that the dimensions of the addition operation match, Q m+1 is the feature representation of the spatial dimension obtained in the spatial feature extraction module, Vec(·) represents the vectorization operation, and P′ m-1 represents the original m-1th sentence embedding vector, P m-1 Represents the m-1th sentence embedding vector after the word embedding sequence conversion. Through the above addition operation, the sentence embedding representation is enhanced by word-level features.
[0077] In addition, this method transforms sentence embeddings using the standard SwinTransformer block to model the relationship between sentence embeddings:
[0078]
[0079] Among them, and P m respectively represent the outputs of the m-th (S)WMSA module and the m-th MLP module. W-MSA is window-based multi-head self-attention operation, and SW-MSA is shifted window-based multi-head self-attention operation, as shown in formula (10). LN is the layer normalization process, and MLP represents multi-layer perceptron.
[0080] In summary, the input and output of the model in the present invention include visual word embedding and sentence embedding. The Spatial-SwinTransformer block is used to model the relationships between visual words to extract local features, while the Temporal-SwinTransformer block captures the inherent information in the sentence sequence.
[0081] Position encoding: For sentence embedding, in this paper, corresponding sentence position encoding is added to each visual sentence to preserve spatial information, as shown in formula (19), using standard one-dimensional position encoding.
[0082] P sentence ←P + E sentence (19)
[0083] Among them, E sentence ∈R (n+1)×d is the sentence position encoding, n represents that the image sequence contains n single images, plus a class token P class , d is the embedding dimension, P represents the original sentence embedding matrix, that is, the state before adding the position encoding, which contains the initial word embedding and the embedding of the class token, but does not contain any position encoding information yet. P sentence represents the updated sentence embedding matrix. After the addition operation of the position encoding, the new sentence embedding contains the embedding information of the original words and class tokens as well as additional position encoding information. In this way, the sentence position encoding is used to save the global time series information.
[0084] After the SwinTransformer block in the time dimension, a new set of temporal feature maps is obtained. This feature map not only retains the original spatial information, but also integrates the change process of the image in the time dimension, fully considering the temporal relationship. This chapter adopts a symmetrical Transformer-based decoder. The decoder consists of a SwinTransformer module and a block expansion layer. The extracted image features are fused with multi-scale features from the encoder through jump connections to supplement the spatial information loss caused by downsampling. In contrast to the block merging layer, this paper uses a block expansion layer in the decoder to perform upsampling. The block expansion layer reshapes the feature maps of adjacent dimensions into large feature maps with a resolution magnified by 2 times. This chapter uses the last patch expansion layer to perform 4x upsampling to restore the resolution of the feature map to the input resolution (W×H), and then applies a linear projection layer on these upsampled features to output pixel-level segmentation predictions. Such fusion ensures that the model can consider both spatial and temporal information and improve the recognition accuracy of discriminative areas.
[0085] Step 3: Spatiotemporal feature fusion module
[0086] Based on the above two modules, the present invention proposes a simple and effective cross attention module to fuse the information between the spatial feature extraction module and the temporal feature extraction module. Each Transformer module exchanges information with other modules through attention, generating an attention map in the fusion process in linear time. Figure 2 The basic idea of cross attention adopted in this chapter is demonstrated.
[0087] Among them, the fusion involves the patchtoken of one module and the patchtokens of another module. Specifically, in order to fuse spatiotemporal features more efficiently and effectively, this chapter first uses the patchtoken of each module to exchange information with the patchtokens of other modules, and then reprojects it to its own module. Since the patchtoken has learned the abstract information of all the patchtokens in its own module, the interaction with the patchtokens of another module helps to contain information of different dimensions. After being fused with the tokens of other modules, the patchtoken interacts with its own patchtokens again in the encoder of the next Transformer, which is able to pass the information learned from other branches to its own patchtokens, enriching the representation of each patchtoken.
[0088] The spatial feature extraction module and the temporal feature extraction module can be regarded as branches Spa and Tem respectively.
[0089]
[0090] Among them, each branch Spa represents a spatial feature extraction module, that is, a single image is divided into 9 non-overlapping blocks, and each branch Tem represents a temporal feature extraction module, that is, the image sequence contains n single images.
[0091] like Figure 3 As shown in the figure, this method describes the schematic diagram of the cross attention module of the branch Tem. For the branch Spa, Figure 3 The cross attention process of branch Spa can be performed by exchanging the index Tem and Spa in the branch. For branch Tem, it first collects patchtokens from the Spa branch and connects its own patchtokens with it to get y′ Tem , as shown in formula (22):
[0092]
[0093] Among them, y′ Tem Indicates the fusion vector obtained by splicing the information exchanged between the Spa branch and the Tem branch. Tem (·) is the projection function used for dimension alignment, represents the feature vector in the temporal feature extraction branch, Represents the feature vector in the spatial feature extraction branch; || is the concatenation operator, which means concatenating two vectors together.
[0094] Then, the module and y′ Tem Cross attention (CA) is performed between the two branches, where patchtoken is the only query item and the patchtokens information of the Spa branch is merged into its own patchtoken. Mathematically, CA can be expressed as:
[0095]
[0096] Among them, W q ,W k ,W v ∈R C×(Ch) It is a parameter optimized by the model during training through back-propagation and gradient descent, which is used to map input features to the space of queries, keys, and values. Specifically, W q Used to generate query vector q, W k Used to generate key vector k, W v Used to generate the value vector v. C is the dimension of the input feature, h is the number of attention heads, and R represents a set of real numbers; Represents the projection function f after dimension alignment operation Tem (·) Generated eigenvector; y′Tem It indicates that the Spa branch and the Tem branch exchange information and obtain the fusion vector through splicing.
[0097] It should be noted that this method only uses blocks as queries, and the computational and memory complexity of generating the attention map is in the linear range, rather than quadratic growth like full attention, making the whole process more efficient. In addition, like self-attention, this chapter also uses multiple heads in CA and denotes it as MCA. Specifically, given The output z of the crisscross attention module after layer normalization and residual shortcut operation Tem , defined as follows:
[0098]
[0099]
[0100] Among them, f Tem (·) and g Tem (·) are the projection and back-projection functions for dimension alignment, represents the feature vector in the temporal feature extraction branch, represents the feature vector in the spatial feature extraction branch; || is the concatenation operator, which means concatenating two vectors together; LN is the layer normalization process; MCA means using multiple heads in the cross attention to form a multi-head cross attention, as shown in formula (9). represents the output vector after layer normalization, multi-head crisscross attention (MCA) and residual shortcut operations, z Tem Represents the final output of the criss-cross attention module after back-projection and concatenation.
[0101] References:
[0102] [1]Tseng KL, Lin YL, Hsu W, et al.Joint Sequence Learning and Cross-Modality Convolution for 3D Biomedical Segmentation[J]. IEEE Computer Society, 2017. DOI: 10.1109 / CVPR.2017.398.
[0103] [2]Cai J,Lu L,Xie Y,et al.Improving Deep Pancreas Segmentation in CTand MRI Images via Recurrent Neural Contextual Learning and Direct LossFunction[J].2017.DOI:10.48550 / arXiv.1707.04912.
[0104] [3]Vaswani A,Shazeer N,Parmar N,et al.Attention Is All You Need[J].arXiv,2017.DOI:10.48550 / arXiv.1706.03762.
[0105] [4]Dosovitskiy A,Beyer L,Kolesnikov A,et al.An Image is Worth16x16Words:Transformers for Image Recognition at Scale[C] / / InternationalConference on Learning Representations.2021.
[0106] [5]Han,Kai,et al."Transformer in transformer."Advances in neuralinformationprocessingsystems 34(2021):15908-15919.
[0107] [6]Liu Z,Lin Y,Cao Y,et al.Swin Transformer:Hierarchical VisionTransformer using Shifted Windows[J].2021.DOI:10.48550 / arXiv.2103.14030.
[0108] [7]Chen C F,Fan Q,Panda R.CrossViT:Cross-Attention Multi-Scale VisionTransformer for Image Classification[J].2021.DOI:10.48550 / arXiv.2103.14899。
Claims
1. A serialized image segmentation method based on spatiotemporal Swin Transformer fusion, characterized by: Step 1: Spatial feature extraction module The spatiotemporal Swin Transformer fusion model adopts the Swin-UNet architecture, which consists of an encoder, a bottleneck layer, a decoder, and skip connections; Given an image sequence, which contains n single images; in the model, a single image is regarded as a visual sentence. In order to convert the input into a block embedding, the single image is divided into 3×3=9 non-overlapping blocks, and the feature dimension of each block becomes 3×3×3=27; in addition, the feature dimension is projected to an arbitrary dimension represented as C through a linear embedding layer; where a visual sentence consists of a series of visual words; The spatiotemporal fusion model has two data streams: one operates between visual sentences, and the other processes visual words within each sentence. For word embedding, a SwinTransformer block is used to explore the relationship between visual words. For the encoder, in order to convert the input into a block embedding, the image is split into non-overlapping blocks; in addition, the feature dimension is projected to an arbitrary dimension through a linear embedding layer; the converted blocks pass through multiple SwinTransformer modules and patch merging layers to generate hierarchical feature representations; the block merging layer is responsible for downsampling and increasing the dimension, while the SwinTransformer module is responsible for feature representation learning; the converted blocks pass through multiple SwinTransformer modules and patch merging layers to generate hierarchical feature representations; the block merging layer is responsible for downsampling and increasing the dimension, while the SwinTransformer module is responsible for feature representation learning; Step 2: Temporal feature extraction module Based on the hierarchical spatial feature representation of multiple time points obtained by the spatial feature extraction module, the spatial feature maps of all time points are arranged in the order of the time when the image data is obtained, and the position code is added in this order; the time series feature matrix with position code is put into the Swin Transformer module of the time dimension as input, and the rest of the parts are the same as the Swin Transformer in the spatial feature extraction module; After the SwinTransformer block in the time dimension, a new set of temporal feature maps is obtained; this feature map not only retains the original spatial information, but also combines the changing process of the image in the time dimension, fully considering the temporal relationship; a symmetrical Transformer-based decoder is adopted; the decoder consists of a SwinTransformer module and a block expansion layer; the extracted image features are fused with the multi-scale features from the encoder through jump connections to supplement the spatial information loss caused by downsampling; the block expansion layer in the decoder is used to perform upsampling; the block expansion layer reshapes the feature maps of adjacent dimensions into large feature maps with a resolution magnified by 2 times; a 4x upsampling is performed using the last patch expansion layer to restore the resolution of the feature map to the input resolution (W×H), and then a linear projection layer is applied on these upsampled features to output pixel-level segmentation predictions; Step 3: Spatiotemporal feature fusion module Based on the above two modules, a cross attention module is proposed to fuse the information between the spatial feature extraction module and the temporal feature extraction module; each Transformer module exchanges information with other modules through attention, generating an attention map in the fusion process in linear time; The patchtoken of each module is used to exchange information with the patchtokens of other modules, and then reprojected to its own module; after being fused with the tokens of other modules, the patchtoken interacts with its own patchtokens again in the encoder of the next Transformer, and is able to pass the information learned from other branches to its own patchtokens.
2. The method according to claim 1, characterized in that: The following steps are involved: Step 1: Spatial feature extraction module The spatiotemporal Swin Transformer fusion model adopts the Swin-UNet architecture, which consists of an encoder, a bottleneck layer, a decoder, and skip connections; Given an image sequence consisting of n single images: And=[L 1 ,THE 2 ,...,THE n ]∈R n×p×p×3 (1) Where (p, p) is the resolution of each image, n is the number of single images in the sequence, R is a real number array, and E is an image sequence; In the model, a single image is regarded as a visual sentence. In order to convert the input into block embedding, the single image is divided into 3×3=9 non-overlapping blocks, and the feature dimension of each block becomes 3×3×3=27; in addition, the feature dimension is projected to an arbitrary dimension (denoted as C) through a linear embedding layer; where a visual sentence consists of a series of visual words: L i =[l i,1 ,l i,2 ,...,l i,9 ] (2) Among them, l i,j ∈R s×s×3 is the jth visual word of the i-th visual sentence, (s, s) is the size of the visual word space, j = 1, 2, ..., 9; The transformed blocks are passed through multiple Swin Transformer modules and patch merging layers to generate hierarchical feature representations; The spatiotemporal fusion model has two data streams: one operates between visual sentences, and the other processes visual words within each sentence. For word embedding, a Swin Transformer block is used to explore the relationship between visual words. Convert visual words into word embedding sequences via linear projection: Q i =[q i,1 ,q i,2 ,...,q i,9 ],q i,j =FC(Vec(l i,j )) (3) Among them, q i,j ∈R C is the jth word embedding, C is the dimension of the word embedding, Vec(·) is the vectorization operation, and FC represents the output of the fully connected layer; Position encoding: Spatial information is an important factor in image recognition. For the visual words in a sentence, this paper adds the corresponding word position encoding to each word embedding to retain the spatial information, as shown in formula (4), using the standard one-dimensional position encoding. Among them, E word ∈R 9×C It is the word position code shared in each sentence. The word position code is used to save the local relative position. It contains 9 positions corresponding to the codes of 9 visual words in the sentence. Each code is a C-dimensional vector. i The original query vector representing the i-th visual word, before adding the position encoding, is usually generated by the word embedding feature extraction method; It represents the query vector of the i-th visual word after adding the word position encoding. It is the result of adding the original query vector and the position encoding, integrating the content and position information of the word. Next, each module is introduced: SwinTransformer blocks: Each SwinTransformer block consists of a LayerNorm (LN) layer, a multi-head self-attention module, a residual connection, and a two-layer MLP with GELU nonlinearity; two consecutive SwinTransformer blocks apply a window-based multi-head self-attention (W-MSA) module and a shifted window-based multi-head self-attention (SW-MSA) module respectively; among them, MSA is a standard global self-attention mechanism that calculates the global relationship between all input patches; while WMSA divides the input feature map into non-overlapping windows and calculates self-attention within each window; SWMSA introduces a sliding window mechanism based on WMSA, which enables information interaction between different windows by sliding the position of the window between different layers; according to this window division mechanism, consecutive SwinTransformer blocks can be regarded as Spatial-Swin Transformer blocks, which are formulated as: in, and Q m Represent the outputs of the m-th (S)WMSA module and the m-th MLP module respectively; W-MSA is a window-based multi-head self-attention operation, SW-MSA is a shifted window-based multi-head self-attention operation, LN is a layer normalization process, and MLP stands for multi-layer perceptron; The self-attention formula is as follows: As shown in formula (9), in the self-attention module, the input is linearly transformed into three parts, namely the query key Sum M 2 represents the number of blocks in the window, d,d k ,d v are the dimensions of input, query (key), and value respectively; while W-MSA and SW-MSA introduce the relative position information between windows through an additional matrix B in formula (9), and obtain formula (10), where the value of B is taken from the bias matrix This enables the model to capture the spatial relationship between different windows. M is the size of the window. For a window of size M×M, the relative position range between patches is [-M+1,M-1] in both horizontal and vertical directions. R is a set of real numbers. Use a linear layer to generate the output; multi-head self-attention splits the query, key, and value and performs the attention function in parallel, then concatenates the output values of each head and forms the final output through linear projection; Multilayer Perceptron (MLP): MLP is used for feature transformation and nonlinear operations between self-attention layers; MLP(X)=FC(σ(FC(X))),FC(X)=XT+b(11) Where T and b are the weight and bias term of the fully connected layer respectively; σ(·) is the activation function, and FC represents the output of the fully connected layer; Layer Normalization (LN): Layer normalization is applied to each sample x∈R d ,as follows: Where μ∈R, δ∈R are the mean and standard deviation of the feature respectively, represents element-wise multiplication, γ∈R d ,β∈R d is the affine transformation parameter, R represents the set of real numbers; Encoder: In the encoder, the C-dimensional tokenized input with a resolution of H / 4×W / 4 is fed into two consecutive SwinTransformer blocks for feature learning, and the feature dimension and resolution remain unchanged; at the same time, the block merging layer reduces the number of patches (2x downsampling) and increases the feature dimension to 2x; this process is repeated three times in the encoder; Bottleneck layer: Two consecutive SwinTransformer blocks are used to learn deep feature representations, with feature dimension and resolution kept constant; Decoder: Corresponding to the encoder, the decoder builds a symmetrical structure based on the SwinTransformer block; to this end, the decoder uses a block expansion layer instead of the block merging layer in the encoder for upsampling; the block expansion layer reshapes the feature map of the adjacent dimension into a higher resolution feature map, that is, 2x upsampling, and halves the feature dimension; Skip connections: Skip connections are used to fuse the multi-scale features in the encoder with the upsampled features; a linear layer is applied to make the dimension of the concatenated features the same as the dimension of the upsampled features; Step 2: Temporal feature extraction module Based on the hierarchical spatial feature representation of multiple time points obtained by the spatial feature extraction module, the spatial feature maps of all time points are arranged in the order of the time when the image data is obtained, and the position code is added in this order; the time series feature matrix with position code is put into the Swin Transformer module of the time dimension as input, and the rest of the parts are the same as the Swin Transformer in the spatial feature extraction module; For visual sentences, a sentence embedding memory is created to store the sentence-level representation sequence: Among them, the category label P class Initialized to zero, used to aggregate global sentence information; n is the number of words in the sentence, d is the embedding dimension, and the extra dimension comes from the category label P class , R represents the set of real numbers; At each layer, the word embedding sequence is transformed to the sentence embedding domain via linear projection and added to the sentence embedding: P m-1 =P′ m-1 +FC(Things(Q m+1 )) (14) Among them, P m-1 , P′ m-1 ∈R d , R represents the set of real numbers, d is the embedding dimension, and the fully connected layer FC ensures that the dimension of the addition operation matches, Q m+1 is the feature representation of the spatial dimension obtained in the spatial feature extraction module, Vec(·) represents the vectorization operation, and P′ m-1 represents the original m-1th sentence embedding vector, P m-1 Represents the m-1th sentence embedding vector after the word embedding sequence conversion; through the above addition operation, the sentence embedding representation is enhanced by word-level features; In addition, this method transforms sentence embeddings using the standard SwinTransformer block to model the relationship between sentence embeddings: in, and P m They represent the outputs of the m-th (S)WMSA module and the m-th MLP module respectively. W-MSA is a window-based multi-head self-attention operation, and SW-MSA is a shifted window-based multi-head self-attention operation, as shown in formula (10). LN is the layer normalization process, and MLP stands for multi-layer perceptron. The input and output of the model include visual word embeddings and sentence embeddings; the Spatial-SwinTransformer block is used to model the relationship between visual words to extract local features, while the Temporal-SwinTransformer block captures the intrinsic information in the sentence sequence; Position encoding: For sentence embedding, a corresponding sentence position encoding is added to each visual sentence to preserve spatial information, as shown in formula (19), using a standard one-dimensional position encoding; P sentence ←P+E sentence (19) Among them, E sentence ∈R (n+1)×d is the sentence position encoding, n means that the image sequence contains n single images, plus a category label P class , d is the embedding dimension, P represents the original sentence embedding matrix, that is, the state before adding the positional encoding, which contains the initial word embedding and the embedding of the category tag, but does not yet contain any positional encoding information; P sentence represents the updated sentence embedding matrix. After adding the positional encoding, the new sentence embedding contains the original word and category tag embedding information as well as the additional positional encoding information. In this way, the sentence position encoding is used to save the global time series information. After the SwinTransformer block in the time dimension, a new set of time feature maps is obtained; Step 3: Spatiotemporal feature fusion module Based on the above two modules, the cross attention module is used to fuse the information between the spatial feature extraction module and the temporal feature extraction module; each Transformer module exchanges information with other modules through attention, and generates an attention map of the fusion process in linear time; Among them, the fusion involves the patchtoken of one module and the patchtokens of another module; the patchtoken of each module is used to exchange information with the patchtokens of other modules, and then reprojected to its own module; after being fused with the tokens of other modules, the patchtoken interacts with its own patchtokens again in the encoder of the next Transformer, passing the information learned from other branches to its own patchtokens; The spatial feature extraction module and the temporal feature extraction module are regarded as branches Spa and Tem respectively; Among them, each branch Spa represents a spatial feature extraction module, that is, a single image is divided into 9 non-overlapping blocks, and each branch Tem represents a temporal feature extraction module, that is, the image sequence contains n single images; For branch Tem, it first collects patchtokens from the Spa branch and connects its own patchtokens with it to get y′ Tem , as shown in formula (22): Among them, y′ Tem represents the fusion vector obtained by splicing the information exchanged between the Spa branch and the Tem branch; f Tem (·) is the projection function used for dimension alignment, represents the feature vector in the temporal feature extraction branch, represents the feature vector in the spatial feature extraction branch; || is the concatenation operator, which means concatenating two vectors together; Then, the module and y′ Tem Cross attention (CA) is performed between the two branches, where patchtoken is the only query item and the patchtokens information of the Spa branch is merged into its own patchtoken; CA is expressed as: Among them, W q ,W k ,W v ∈R C×(Ch) It is the parameter optimized by the model during training through back propagation and gradient descent, which is used to map the input features to the space of query, key and value; W q Used to generate query vector q, W k Used to generate key vector k, W v Used to generate the value vector v; C is the dimension of the input feature, h is the number of attention heads, and R represents a set of real numbers; Represents the projection function f after dimension alignment operation Tem (·) Generated eigenvector; y′ Tem It indicates that the Spa branch and the Tem branch exchange information and obtain the fusion vector through splicing; Given The output z of the crisscross attention module with layer normalization and residual shortcut operation Tem , defined as follows: Among them, f Tem (·) and g Tem (·) are the projection and back-projection functions for dimension alignment, represents the feature vector in the temporal feature extraction branch, represents the feature vector in the spatial feature extraction branch; || is the concatenation operator, which means concatenating two vectors together; LN is the layer normalization process; MCA means using multiple heads in the cross attention to form a multi-head cross attention, as shown in formula (9); represents the output vector after layer normalization, multi-head crisscross attention (MCA) and residual shortcut operations, z Tem Represents the final output of the criss-cross attention module after back-projection and concatenation.
Citation Information
Cited By
Parking space state monitoring method based on multi-sensor fusion
CN121365355A
Transform classification method based on membrane space-time attention
CN121388729A
Transform classification method based on film space-time attention
CN121388729B