A method for extracting channel and spatial information from feature maps using a convolutional neural network based on a context-aware attention mechanism
By introducing mixed context-aware spatial and channel attention modules into convolutional neural networks, the problem that the existing technology is difficult to take into account both local and global information is solved, and efficient feature extraction and robustness improvement are achieved.
Patent Information
- Application Number
- CN202411745329.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-02
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-12-02
AI Technical Summary
In the field of computer vision, it is difficult for the prior art to take into account both local pixel information and global pixel information at the same time, resulting in the problem of information loss or excessive calculation cost when extracting channel and spatial information in the feature map.
A convolutional neural network based on perceived context attention mechanism is designed to extract channel and spatial information in feature maps by introducing mixed context-aware spatial attention modules (HCA-S) and channel attention modules (HCA-C) into the convolutional neural network. These modules dynamically adjust attention weights through local enhancement perception units and dimensionality reduction operations, ensuring that local fine-grained and global coarse-grained information is captured simultaneously.
It realizes that the global context information of the feature map is retained without increasing the computational burden, and at the same time, the fine-grained information of the local area is adaptively perceived, thereby improving the effect of feature extraction and the robustness of the model.
Smart Images

Figure CN119227748B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method for extracting channel and spatial information in a feature graph by using a convolutional neural network based on a perceptual context attention mechanism. Background Art
[0002] Convolutional neural networks are an important type of neural network in the field of computer vision. The attention mechanism is a mechanism that can focus on important information and ignore unimportant information. In recent years, with the emergence of Vision Transformer (ViT) and its variants, ViT can be combined with convolutional neural networks to extract important pixel information, demonstrating the performance of attention-based methods. However, these carefully designed attention modes still have some limitations. The window attention method either involves complex shift operations or sacrifices the ability to model global pixel information; the sparse attention method captures the global receptive field at a low computational cost, but usually causes important pixel information in some areas to be lost in the process of capturing local features; the local perception method only focuses on the local area based on the specific weight of each pixel, which leads to the neglect of other key information. Therefore, it is necessary to design a more effective method to take into account both local pixel information and global pixel information. Summary of the invention
[0003] The purpose of the present invention is to overcome the shortcomings of the prior art and propose a method for designing a convolutional neural network based on a perceptual context attention mechanism to extract channel and spatial information in a feature map.
[0004] The technical solution adopted by the present invention to solve the technical problem is:
[0005] A method for extracting channel and spatial information in a feature map using a convolutional neural network based on a context-aware attention mechanism, comprising:
[0006] The convolutional neural network uses the ResNet-50 network as the backbone network, and divides the residual block after res_conv4_1 into three branches: spatial mixing branch-1, spatial mixing branch-2, and channel mixing branch; the spatial mixing branch-1 and the spatial mixing branch-2 are both embedded with a hybrid context-aware spatial attention HCA-S (Hybrid Context-Aware Attention - Spatial) module to emphasize the important spatial position information of the feature map; the channel mixing branch is embedded with a hybrid context-aware channel attention HCA-C (Hybrid Context-Aware Attention - Channel) module to obtain important information of the feature map from the channel dimension;
[0007] The HCA-S module uses convolution to obtain three tensors from the input feature map: a query matrix, a key matrix, and a value matrix; the query matrix and the key matrix are evenly divided into a number of non-overlapping windows, and the two-dimensional space of each window feature map is flattened into a one-dimensional sequence, and a number of window query matrices and window key matrices are obtained respectively, and a local enhancement key matrix and a local enhancement value matrix in the spatial dimension are established based on the spatial attention matrix generated by the window query matrix and the window key matrix; a global perception weight is constructed based on the local enhancement key matrix and the local enhancement value matrix; the global perception weight is used to reconstruct the value matrix, thereby obtaining a global context perception matrix with local enhancement, and the global context perception matrix is multiplied and added with the input feature map to obtain the final output feature map of the HCA-S module;
[0008] The two-dimensional space of the input feature map of the HCA-C (Hybrid Context-Aware Attention - Channel) module is flattened into a one-dimensional sequence, and then passed through a fully connected layer to obtain a query matrix, a key matrix and a value matrix; the query matrix and the key matrix are evenly divided into a number of groups along the channel dimension, and a local enhanced key matrix and a local enhanced value matrix on the channel dimension are established based on a multi-channel group fusion attention matrix generated based on the channel group query matrix and the channel group key matrix, and a global perception weight is constructed based on the local enhanced key matrix and the local enhanced value matrix; the global perception weight is used to reconstruct the value matrix to obtain a global context perception matrix with local enhancement, and the global context perception matrix is multiplied and added with the input feature map to obtain the final output feature map of the HCA-C module.
[0009] Furthermore, in the HCA-S module, the process of establishing the local enhancement key matrix and the local enhancement value matrix in the spatial dimension based on the spatial attention matrix generated by the window query matrix and the window key matrix is as follows:
[0010] The attention scores of the divided independent windows are calculated in parallel, and the window query matrix is multiplied by the transposed window key matrix to obtain the window attention matrix;
[0011] The window attention matrix is compressed along the column dimension to obtain the window attention vector. By using the compression operation on the window, the most prominent key semantic information in the window is captured, while the redundant semantic information is removed.
[0012] Aggregate the compressed window attention vectors to obtain a multi-window fused spatial attention matrix;
[0013] After the spatial attention matrix is reshaped, it is normalized to obtain the weighted perceptron in the spatial dimension.
[0014] Based on the generated weighted perceptron, the key matrix and the value matrix are readjusted respectively by using scalar product to obtain the local enhanced key matrix and the local enhanced value matrix in the spatial dimension, so that the key matrix and the value matrix can enhance the attention to the pixels in the salient area.
[0015] Furthermore, in the HCA-C module, the process of establishing the local enhancement key matrix and the local enhancement value matrix on the channel dimension based on the multi-channel group fusion attention matrix generated by the channel group query matrix and the channel group key matrix is as follows:
[0016] Reshape the divided groups to obtain a channel group query matrix and a channel group key matrix for each group;
[0017] The matrix obtained by transposing the channel group key matrix and performing matrix multiplication with the channel group query matrix is compressed along the row dimension to obtain the channel group attention vector;
[0018] The channel group attention vectors of several groups are aggregated to obtain a multi-channel group fusion attention matrix;
[0019] After reshaping the multi-channel group fusion attention matrix, normalize it to obtain the weighted perceptron in the channel dimension.
[0020] Based on the generated weighted perceptron, the key matrix and value matrix are resized respectively using scalar product to obtain the local enhanced key matrix and local enhanced value matrix in the channel dimension.
[0021] Furthermore, in the HCA-S module, the local enhancement key matrix and the local enhancement value matrix are subjected to spatial perception dimensionality reduction operations, and the key matrix after dimensionality reduction is multiplied with the query matrix and normalized to obtain the global perception weight Attn_s for:
[0022] ;
[0023] in, X Q S is the query matrix of the HCA-S module, K s SAR It is the matrix after channel-aware dimensionality reduction operation on the local enhanced key matrix in the spatial dimension.
[0024] Furthermore, in the HCA-C module, the local enhancement key matrix and the local enhancement value matrix are respectively subjected to channel-aware dimensionality reduction operations, and the key matrix after dimensionality reduction is multiplied with the query matrix and normalized to obtain the global perception weight Attn_c for:
[0025] ;
[0026] in, X Q C is the query matrix of the HCA-C module, K c CAR It is the matrix after channel-aware dimensionality reduction operation on the local enhanced key matrix in the channel dimension.
[0027] Furthermore, the output feature maps of the spatial mixing branch-1, the spatial mixing branch-2 and the channel mixing branch are compressed into a feature vector of 2048 dimensions through average pooling; the dimensions of the 2048-dimensional feature vectors output by the three branches are reduced to 256 through 1×1 convolution, batch normalization and ReLU activation function, and a 256-dimensional feature embedding is obtained.
[0028] Furthermore, the 256-dimensional feature embedding is used for training with triplet loss, and is used for training with cross entropy loss after conversion through a fully connected layer.
[0029] Furthermore, the final output feature maps of the three branches are divided into two parts horizontally, so that the HCA-S module and the HCA-C module can focus more on mining fine-grained visual clues.
[0030] Furthermore, the final output feature maps of the three branches are concatenated to obtain the final feature representation of the input feature map.
[0031] Furthermore, a position-aware constraint loss is applied between the HCA-S modules of the spatial mixing branch-1 and the spatial mixing branch-2 to force the spatial mixing branch-1 and the spatial mixing branch-2 to focus on different important spatial positions. The position-aware constraint loss is defined as follows:
[0032] B S1’ local , B S2’ local = softmax( B S1 local , B S2 local );
[0033] in, B S1 local and B S2 local They are weighted perceptrons on the spatial dimension in the HCA-S modules of the spatial mixing branch-1 and the spatial mixing branch-2, respectively;B S1’ local and B S2’ local are the corresponding weighted perceptrons generated respectively.
[0034] Technical effects of the present invention:
[0035] Compared with the prior art, the present invention proposes a novel attention mechanism, called Hybrid Context-Aware Adaptive Attention Mechanism (HCAA). The core idea of this attention mechanism is to retain the pixel perception of the global context of the entire feature map without introducing additional computational burden, while adaptively perceiving the fine-grained information of the local area. Specifically, in the space and channel modules designed based on this mechanism, the input feature map is first divided along different dimensions, and all queries and keys within a certain range are interacted to generate local context-aware weights, and then the keys and values encapsulating the local information are compressed and input into the fully connected layer with the original query. In this process, the relationship between each query and the compressed key and value is used to further embed the learned local information into the global information, thereby obtaining a convolutional neural network based on the perceptual context attention mechanism. In other words, the query is kept unchanged, the global information is adaptively perceived by dynamically adjusting the key and value, and the important regional information is efficiently perceived by the local sensor, so as to achieve the effect of taking into account both local pixel information and global pixel information. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a convolutional neural network architecture diagram based on the context-aware attention mechanism of the present invention;
[0037] Figure 2 This is a diagram of the HCA-S module architecture of the present invention;
[0038] Figure 3 This is a diagram of the HCA-C module architecture of the present invention. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solution and advantages of the embodiments of the present invention more clear, the technical solution in the embodiments of the present invention is clearly and completely described below in conjunction with the accompanying drawings.
[0040] Embodiment 1:
[0041] like Figure 1As shown, the present embodiment involves a method for extracting channel and spatial information from a feature graph using a convolutional neural network based on a perceptual context attention mechanism, comprising the following steps:
[0042] The convolutional neural network uses the ResNet-50 network as the backbone network, and divides the residual block after the ResNet-50 network res_conv4_1 into three branches: spatial mixing branch-1, spatial mixing branch-2, and channel mixing branch; at the same time, in order to obtain a higher resolution feature map, the res_conv5_1 residual block does not use downsampling operation;
[0043] In the spatial mixing branch-1 and the spatial mixing branch-2, the same hybrid context-aware spatial attention HCA-S module is embedded, which is used to emphasize the important spatial position information of the feature map. First, the HCA-S is subjected to a global average pooling (GAP) operation to compress the output feature map into a feature vector containing 2048 dimensions. Next, the 2048-dimensional feature vector is subjected to a dimensionality reduction operation to reduce it to 256 dimensions. Finally, the obtained 256-dimensional feature vector is used as the final feature embedding of the spatial mixing branch-1 and the spatial mixing branch-2. The dimensionality reduction operation consists of a 2D convolution with a convolution kernel size of 1, a batch normalization (BN) layer, and a ReLU activation function. In addition, in order to force the two spatial branches to pay attention to the important positions of different regions, this embodiment proposes a position-aware constraint loss, which is applied between the two HCA-S modules. When adjusting the weights of the key matrix and the value matrix in the HCA-S modules in the two branches, their original weight perceptrons are no longer used, but the corresponding weight perceptrons generated by the loss are used.
[0044] In the channel mixing branch, the hybrid context-aware channel attention HCA-C module is embedded. The subsequent operations of this branch are similar to those of the spatial mixing branch-1 and the spatial mixing branch-2. They first undergo a global average pooling operation to obtain a 2048-dimensional feature vector, and then undergo a dimensionality reduction operation to obtain a 256-dimensional feature embedding. The role of the channel mixing branch in FCGHNet is to obtain important information from the feature map from the channel dimension. In addition, the final output feature maps of the three branches are divided into two parts horizontally to make the HCA-S module and the HCA-C module focus more on mining fine-grained visual clues;
[0045] During the testing phase, the final output feature maps of these three branches are embedded and concatenated together as the output feature map of the network.
[0046] 1. HCA-S module architecture
[0047] The hybrid context-aware spatial attention HCA-S module aims to use context-awareness to gradually mix local fine-grained information and global coarse-grained information in the spatial dimension to efficiently learn attention. Figure 2 As shown, the input feature X∈R of the spatial attention module C×H×W , where C, H, and W represent the number of channels, height, and width of X, respectively. In order to use the attention mechanism to capture the context-aware features of adaptive mixing in the spatial dimension, we first use three unshared 1×1 convolutions to obtain three tensors of the same size, which are the query matrix X Q S ∈R C×H×W , key matrix X K S ∈R C×H×W Sum Matrix X V S ∈R C ×H×W :
[0048] ( X Q S , X K S , X V S ) = Conv( X );
[0049] Among them, Conv represents 1×1 convolution. Figure 2 As shown, the query matrix X Q S and key matrix X K S Evenly divided into N non-overlapping windows, this not only refines the context perception of local fine-grained information, but also has a computational cost no higher than that of the Vision Transformer (ViT). Next, the two-dimensional space of each window feature map ( K , K ) is flattened to a size of K 2 One-dimensional sequence, respectively get N window query matrices Q i ∈R (K×K)×C and the window key matrix K i ∈R (K ×K)×C N = (HW) / K2 , i ∈(1,2,...,N), K represents the window size. The process of window division can be described as follows:
[0050] ( Q i , K i ) = Flatten( σ ( X Q S , X K S ));
[0051] in, σ () indicates that the input feature map is divided into multiple windows along the spatial dimension. The attention scores are calculated in parallel for the N independent windows after division. Therefore, i The query vector at the mth position in the window Q i m and i The key vector for the nth position in the window K i n Multiply to get the window pairwise relationship A i m,n :
[0052] A i m,n = Q i m • K i n .
[0053] At the same time, in order to calculate the attention matrix of each window, the window query matrix Q i with the transposed window key matrix K i Multiply to get the window attention matrix A i ∈R (K×K)×(K×K) :
[0054] A i = Q i ( K i ) T ;
[0055] in, represents matrix multiplication. Then, A i Compress along the column dimension to get the window attention vector B i ∈R (K×K)×1 .
[0056] B i = Squeeze( A 1 , A 2 ,..., A i );
[0057] Here, Squeeze represents the compression operation. By using the compression operation on the pixels in the N windows, the most prominent key semantic information in the window is captured, while the redundant semantic information is removed. The N window attention vectors obtained after compression are B i Aggregate to obtain a spatial attention matrix called multi-window fusion B s :
[0058] B s = W inter ([ B 1 , B 2 ,..., B i ]);
[0059] Among them, W inter Represents the vector between windows B i Aggregation operation.
[0060] Keep the query matrix X Q S Unchanged, key matrix X K S Sum Matrix X V S The local enhanced perception unit (LAU) is applied to dynamically adjust the attention level of different regions according to the different attention weights of different regions. Specifically, the spatial attention matrix of multi-window fusion is first B s After the Reshape operation, the normalized Sigmoid operation is performed to obtain the weighted perceptron in the spatial dimension. BS local ∈R H×W , the calculation process is as follows:
[0061] B S local = Sigmoid(Reshape( B s )).
[0062] Then, based on the generated weighted perceptron B S local , and use scalar products to respectively X K S Sum Matrix X V S Re-adjust to obtain the local enhanced bond matrix in the spatial dimension K s ∈R C×H×W and the local enhancement value matrix V s ∈R C ×H×W The purpose of the local enhancement perception unit (LAU) is to adaptively adjust the degree of attention to different areas, retain more fine-grained local pixel information in the salient area, and reduce the interference of irrelevant information and redundant information in the background area, thereby improving its robustness. Next, the matrix obtained by the local enhancement perception unit is K s and V s Perform spatial perception dimensionality reduction operations respectively to obtain a matrix of size C×(H / P)×(W / P) K s SAR and V s SAR This can reduce the high computational cost caused by self-attention calculations. The specific implementation process is as follows:
[0063] ( K s SAR , V s SAR ) = SAR( K s , V s );
[0064] Here, SAR represents the spatially aware dimensionality reduction operation, which is achieved by using a global average pooling layer of fixed size P×P.
[0065] In order to explicitly model the global receptive field with local perception enhancement over the entire feature map, first, K s SAR The shape of becomes C×(HW / P²), and then let X Q S Multiply it with Softmax and normalize it to get a global perceptual weight between 0 and 1 Attn_s ∈R HW×(HW / (p×p)) :
[0066] ;
[0067] Then, V s SAR The shape of becomes (HW / P²)×C, and is multiplied by the global perception weight to obtain the global context perception matrix with local enhancement and deformed into F s ∈R C×H×W .
[0068] F s = Attn_s V s SAR .
[0069] In the context-aware generation process of global self-attention, local fine-grained information and global coarse-grained information in the spatial dimension are mixed. F s Multiply and add the input feature map to get the final output feature map of the spatial module F ∈R C×H×W .
[0070] F = F s ⊙ X + X .
[0071] 2. HCA-C module architecture
[0072] like Figure 3 As shown, the tensor X∈R C×H×WAs the input feature map of the channel attention module, C, H, and W represent the number of channels, height, and width of X, respectively. In order to use the self-attention mechanism to mix local and global information based on context perception in the channel dimension, the two-dimensional space of the input feature map X is flattened into a one-dimensional sequence of size HW, and then the query matrix is obtained through the fully connected layer. X Q C ∈R C×HW , key matrix X K C ∈R C×HW Sum Matrix X V C ∈R C×HW :
[0073] ( X Q C , X K C , X V C ) = FC(Flatten( X ));
[0074] Among them, FC represents the fully connected layer and Flatten represents the flattening operation. First, the query matrix is X Q C and key matrix X K C Divide evenly, X Q C and X K C are divided into G groups. Then, the channel group query matrix is obtained by reshaping operation Q g ∈R (C / G)×HW and channel group key matrix K g ∈R (C / G)×HW :
[0075] ( Q g , K g ) = Reshape(divide ( X Q C , X K C ));
[0076] Where g = 1, 2, ..., G. For any channel group, the channel group query vector at the nth position of the gth channel group Q g n and the channel group key vector at the mth position K g m Multiply them to get the channel pair relationship A g n,m :
[0077] A g n,m = Q g n ( K g m ) T ;
[0078] Next, the channel group key matrix composed of the key vectors of all positions of the g-th channel group is K g After transposition and Q g Perform matrix multiplication to obtain the matrix A g ∈R (C / G)×(C / G) . Then A g Compression operation is performed along the row dimension to obtain the channel group attention vector B g ∈R C / G :
[0079] B g = Squeeze( A g ).
[0080] For G channel groups, the G vectors obtained by compression B g , using a one-dimensional convolution with a kernel size of K×K to capture the information of each channel group and its K neighbors. This local cross-channel group interaction process can be expressed as:
[0081] ( C 1 , C 2 , ..., C g ) = C1D K ( B 1, B 2 , ..., B g );
[0082] Among them, C1D K represents a one-dimensional separable convolution operation with shared parameters, and K represents the range of cross-channel group interaction. The reason why all channel groups use the same parameters to learn channel attention is to make the model invariant to feature map flipping and translation. Next, an aggregation operation is performed to obtain the multi-channel group fusion attention matrix B c ∈R G×(C / G) :
[0083] B c = G inter ([ C 1 , C 2 , ..., C g ]);
[0084] in, G inter Represents the vector between channel groups C g Aggregation operation.
[0085] Keep the query matrix X Q C Unchanged, key matrix X K C Sum Matrix X V C The local enhanced perception units are applied separately to dynamically adjust the attention level of different channels according to the different attention weights of different channel groups. Specifically, the channel attention map is first B c After the reshaping operation, the normalization operation is performed to obtain the weighted perceptron in the channel dimension. B c local ∈R C Then, based on the generated weight perceptron, the key matrix is respectively X K C Sum Matrix X V C Re-adjust to obtain the local enhanced bond matrix on the channel dimension K c ∈RC×H×W and the local enhancement value matrix V c ∈R C×H×W Next, the matrix obtained by the local enhanced perception unit is subjected to channel-aware dimensionality reduction operation to obtain K c CAR and V c CAR , which can reduce the high computational cost caused by self-attention calculation. The specific implementation process is as follows:
[0086] ( K c CAR , V c CAR ) = CAR( K c , V c ) ;
[0087] Among them, CAR represents the channel-aware dimensionality reduction operation, which is achieved by using a 1×1 convolution with a dimensionality reduction factor of r.
[0088] In order to explicitly model the global receptive field with local perception enhancement over the entire feature map, first, K c CAR The shape of is changed to HW×(C / r) and X Q C Multiply and perform a series of transformations on the result to obtain a global perception weight between 0 and 1 Attn_c for:
[0089] ;
[0090] Then, V c CAR The shape of is changed to (C / r)×HW and then multiplied and deformed with the value matrix to obtain a global context-aware matrix with local enhancement F c ∈R C×H×W .
[0091] F c = Attn_c V c CAR .
[0092] In the context-aware generation process of global self-attention, local fine-grained information and global coarse-grained information in the channel dimension are mixed. F c Multiply and add the input feature map to get the final output feature map of the channel module F ∈R C×H×W .
[0093] F = F c ⊙ X + X .
[0094] 3. Position-aware constraint loss
[0095] To further improve the effect of attention learning, the location-aware constraint loss (LAC) forces the spatial mixing branch-1 and the spatial mixing branch-2 to focus on different important spatial locations. The constraint is defined as follows:
[0096] B S1’ local , B S2’ local = softmax( B S1 local , B S2 local );
[0097] in, B S1 local and B S2 local The weight perceptrons on the spatial dimension in the HCA-S modules of the spatial mixing branch-1 and the spatial mixing branch-2 are used. When weighting the key matrix and value matrix in the HCA-S modules in these two branches, the B S1 local and B S2 local , but use the corresponding B S1’ local and B S2’ local , B S1’ local and B S2’local are the corresponding weighted perceptrons generated respectively.
[0098] The present invention introduces a multi-branch network that adaptively focuses on different attention granularities by progressively mixing coarse-grained global information and fine-grained local information. Secondly, in order to progressively mix information of different granularities, two different modules are designed to dynamically perceive surrounding pixels from the spatial dimension and channel dimension, respectively, thereby further enhancing the modeling ability of the network. The present invention can take into account the effects of local information and global information at the same time, and output a feature map with richer information.
Claims
1. A method for extracting channel and spatial information from a feature map using a convolutional neural network based on a context-aware attention mechanism, characterized in that: include: The convolutional neural network uses ResNet-50 as the backbone network, and divides the residual block after the ResNet-50 network res_conv4_1 into three branches: spatial mixing branch-1, spatial mixing branch-2 and channel mixing branch; both spatial mixing branch-1 and spatial mixing branch-2 are embedded with a hybrid context-aware spatial attention HCA-S module to emphasize the important spatial position information of the feature map; the channel mixing branch is embedded with a hybrid context-aware channel attention HCA-C module to obtain important information of the feature map from the channel dimension; The HCA-S module and HCA-C module dynamically perceive surrounding pixels from the spatial dimension and channel dimension respectively; The HCA-S module uses convolution to obtain three tensors of the input feature map: a query matrix, a key matrix and a value matrix; the query matrix and the key matrix are evenly divided into a number of non-overlapping windows, and the two-dimensional space of each window feature map is flattened into a one-dimensional sequence to obtain a number of window query matrices and window key matrices, respectively, and a local enhancement key matrix and a local enhancement value matrix in the spatial dimension are established based on the spatial attention matrix generated by the window query matrix and the window key matrix; Based on the local enhancement key matrix and the local enhancement value matrix, the global perception weights are constructed, and the value matrix is reconstructed using the global perception weights to obtain the global context perception matrix with local enhancement. The global context perception matrix is multiplied and added with the input feature map to obtain the final output feature map of the HCA-S module. The two-dimensional space of the input feature map of the HCA-C module is flattened into a one-dimensional sequence, and then passed through a fully connected layer to obtain a query matrix, a key matrix and a value matrix; the query matrix and the key matrix are evenly divided into a plurality of groups along the channel dimension, a local enhancement key matrix and a local enhancement value matrix on the channel dimension are established based on a multi-channel group fusion attention matrix generated based on the channel group query matrix and the channel group key matrix, a global perception weight is constructed based on the local enhancement key matrix and the local enhancement value matrix, the value matrix is reconstructed using the global perception weight, a global context perception matrix with local enhancement is obtained, and the global context perception matrix is multiplied and added with the input feature map to obtain the final output feature map of the HCA-C module; In the HCA-S module, the process of establishing the local enhancement key matrix and the local enhancement value matrix in the spatial dimension based on the spatial attention matrix generated by the window query matrix and the window key matrix is as follows: The attention scores of the divided independent windows are calculated in parallel, and the window query matrix is multiplied by the transposed window key matrix to obtain the window attention matrix; Compress the window attention matrix along the column dimension to obtain the window attention vector; Aggregate the compressed window attention vectors to obtain a multi-window fused spatial attention matrix; After the spatial attention matrix is reshaped, it is normalized to obtain the weighted perceptron in the spatial dimension. Based on the generated weighted perceptron, the key matrix and the value matrix are resized respectively by using scalar product to obtain a local enhanced key matrix and a local enhanced value matrix in the spatial dimension; In the HCA-C module, the process of establishing the local enhancement key matrix and the local enhancement value matrix on the channel dimension based on the multi-channel group fusion attention matrix generated by the channel group query matrix and the channel group key matrix is: Reshape the divided groups to obtain a channel group query matrix and a channel group key matrix for each group; The matrix obtained by transposing the channel group key matrix and performing matrix multiplication with the channel group query matrix is compressed along the row dimension to obtain the channel group attention vector; The channel group attention vectors of several groups are aggregated to obtain a multi-channel group fusion attention matrix; After reshaping the multi-channel group fusion attention matrix, normalize it to obtain the weighted perceptron in the channel dimension. Based on the generated weighted perceptron, the key matrix and value matrix are resized respectively using scalar product to obtain the local enhanced key matrix and local enhanced value matrix in the channel dimension.
2. The method for extracting channel and spatial information in a feature map using a convolutional neural network based on a perceptual context attention mechanism according to claim 1, characterized in that: In the HCA-S module, the local enhancement key matrix and the local enhancement value matrix are subjected to spatial perception dimensionality reduction operations respectively, and the key matrix after dimensionality reduction is multiplied with the query matrix and normalized to obtain the global perception weight. Attn_s for: ; in, X Q S is the query matrix of the HCA-S module, K s SAR It is the matrix after the channel-aware dimensionality reduction operation of the local enhanced key matrix in the spatial dimension, and C represents the number of channels of the input feature map.
3. The method for extracting channel and spatial information in a feature map using a convolutional neural network based on a perceptual context attention mechanism according to claim 1, characterized in that: In the HCA-C module, the local enhancement key matrix and the local enhancement value matrix are respectively subjected to channel-aware dimensionality reduction operations, and the key matrix after dimensionality reduction is multiplied with the query matrix and normalized to obtain the global perception weight. Attn_c for: ; in, X Q C is the query matrix of the HCA-C module, K c CAR It is the matrix after the channel-aware dimensionality reduction operation of the local enhanced key matrix in the channel dimension, and C represents the number of channels of the input feature map.
4. The method for extracting channel and spatial information in a feature map using a convolutional neural network based on a perceptual context attention mechanism according to claim 1, characterized in that: A position-aware constraint loss is applied between the HCA-S modules of the spatial mixing branch-1 and the spatial mixing branch-2, and the position-aware constraint loss is defined as follows: B S1’ local , B S2’ local = softmax( B S1 local , B S2 local ); in, B S1 local and B S2 local They are weighted perceptrons on the spatial dimension in the HCA-S modules of the spatial mixing branch-1 and the spatial mixing branch-2, respectively; B S1’ local and B S2’ local are the corresponding weighted perceptrons generated respectively.
5. The method for extracting channel and spatial information in a feature map using a convolutional neural network based on a perceptual context attention mechanism according to claim 1, characterized in that: The output feature maps of the spatial mixing branch-1, the spatial mixing branch-2 and the channel mixing branch are compressed into a feature vector of 2048 dimensions through average pooling; the dimensions of the 2048-dimensional feature vectors output by the three branches are reduced to 256 through 1×1 convolution, batch normalization and ReLU activation function, and a 256-dimensional feature embedding is obtained.
6. The method for extracting channel and spatial information in a feature map using a convolutional neural network based on a perceptual context attention mechanism according to claim 5, characterized in that: The 256-dimensional feature embedding is used for training with triplet loss, and is converted through a fully connected layer for training with cross entropy loss.
7. The method for extracting channel and spatial information in a feature map using a convolutional neural network based on a perceptual context attention mechanism according to claim 1, characterized in that: The final output feature maps of the three branches are divided into two parts horizontally.
8. The method for extracting channel and spatial information in a feature map using a convolutional neural network based on a context-aware attention mechanism according to claim 1, characterized in that: The final output feature maps of the three branches are concatenated to obtain the final feature representation of the input feature map.
Citation Information
Patent Citations
Local refinement and global enhancement network for vehicle re-identification
CN116644788A
Convolutional neural network based on neighborhood and grid attention mechanism
CN118839734A