Facial expression recognition method fusing space channel features and fast region convolution
Through the methods of fast regional convolution, spatial channel feature fusion and multi-scale feature fusion, the accuracy and real-time problems of facial expression recognition technology in complex environments are solved, and efficient and fast facial expression recognition is achieved.
Patent Information
- Application Number
- CN202510396660.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
The existing facial expression recognition technology has challenges in processing complexity and diversity, especially in different lighting conditions and shooting angles, and the calculation resources consumed too much in real-time detection, making it difficult to meet the real-time requirements.
The fast regional convolution module is used to extract key regional features of expression changes, combined with spatial channel feature fusion and multi-scale feature fusion module, and the convolution and multi-scale attention mechanism can be separated by depth, and the computing efficiency and feature representation capabilities are optimized.
While maintaining high accuracy, it significantly improves the speed of expression recognition, reduces computing resource consumption, and enhances the robustness and recognition capabilities of the model in complex environments.
Smart Images

Figure CN120340089A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision and artificial intelligence, and particularly relates to a facial expression recognition method that fuses spatial-channel features and fast region convolution. Background Art
[0002] Although deep learning techniques have made remarkable progress in facial expression recognition, the existing methods still have the following problems: (1) Human facial expressions have rich diversity and complexity. For example, expressions such as smiling, anger, and surprise contain various subtle changes, and this diversity makes it challenging to accurately recognize and classify different expressions; (2) Under different lighting conditions and shooting angles, the appearance of facial images may change significantly, which will have a certain impact on the stability and accuracy of the facial expression recognition system; (3) In application scenarios of real-time detection, the existing methods still face challenges in terms of processing speed and resource consumption. Given the limitations of the existing technology, researchers are urgently in need of developing a face recognition algorithm that is both efficient and accurate. Improving efficiency means reducing the consumption of computing resources and accelerating the processing speed, so as to meet the application requirements of real-time or near-real-time. Improving accuracy involves the sensitivity and robustness of the algorithm to expression changes, especially in complex environments. [2] For this reason, this paper proposes a facial expression recognition model that integrates spatial-channel feature fusion, a fast region convolution module, and a multi-scale feature fusion module. The model aims to use spatial-channel feature fusion to enhance the feature representation ability of facial expressions, use fast region convolution to optimize the computing efficiency, and use multi-scale feature fusion to enhance the ability to recognize complex expressions, and finally achieve a significant improvement in the speed of expression recognition while maintaining high accuracy. Summary of the Invention
[0003] The present invention proposes a facial expression recognition method that fuses spatial-channel features and fast region convolution. By combining deep learning techniques and computer vision principles, this method achieves the goal of reducing the model size and improving the processing speed while maintaining high accuracy. (1) A fast region convolution strategy is proposed, which effectively reduces the computational cost and the number of memory accesses by using partial convolution, accelerating the operation speed of the model, so that facial expression recognition can be efficiently performed in an environment close to real-time. (2) A spatial-channel feature fusion strategy is proposed, which enhances the ability to perceive expression details by reshaping and grouping channel information at different scales, thereby improving the accuracy of recognizing complex expressions. (3) A multi-scale feature fusion strategy is proposed, which efficiently integrates cross-level features by fusing low-level fine features and high-level semantic features, improving the feature representation ability while enhancing the robustness in complex scenarios. Brief Description of the Drawings
[0004] Figure 1 is the framework of the proposed method;
[0005] Figure 2 It is a schematic structural diagram of the regional convolution module;
[0006] Figure 3 It is a schematic structural diagram of the spatial channel fusion module;
[0007] Figure 4 It is a schematic structural diagram of the multi-scale attention module. Detailed implementation manners
[0008] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings:
[0009] The purpose of the present invention is to provide a face expression recognition method that fuses spatial channel features and fast regional convolution. The implementation idea is as follows: First, a new fast regional convolution module is designed to obtain local feature information, that is, the full convolution is used to reduce the calculation amount, and the features of the key regions of the expression change are extracted through regional convolution; Then, a spatial channel feature fusion module is designed to reduce the memory access by using depthwise separable convolution, and fuse the channel features and spatial features to extract more discriminative feature representations; Finally, a multi-scale attention module is designed to enhance the feature representation ability of the key regions of the expression change through the multi-scale feature extraction strategy and the spatial information aggregation strategy. A preferred implementation manner of the face expression recognition method based on the fusion of spatial channel features and fast regional convolution of the present invention specifically includes the following steps:
[0010] The model proposed in this paper integrates three modules: fast regional convolution, spatial channel feature fusion, and multi-scale feature fusion to improve the performance of face expression recognition. The fast regional convolution module improves the calculation efficiency and processing speed, meeting the requirements of real-time monitoring and interaction; The spatial channel feature fusion module strengthens the recognition of subtle expression changes; The multi-scale feature fusion module enhances the recognition ability of complex expression changes through the attention mechanism, improving the accuracy and robustness of the model.
[0011] 1.1 Implementation of the fast regional convolution module
[0012] The lightweight Transformer network model reduces the complexity and computational cost of the model while maintaining effective feature extraction capabilities, making it suitable for deployment on resource-constrained devices. The model of ShipGeoNet uses a convolutional neural network to extract the texture information and shape information of synthetic aperture radar images. The deep bidirectional long short-term memory network uses a pre-trained model to extract features and optimize them. This network has advantages in processing time series data and can capture dynamic change patterns. The multi-scale spatial pyramid attention mechanism combines spatial pyramid pooling and the attention mechanism to extract multi-scale and multi-region features. This mechanism improves the model's sensitivity to different scale features and different position features of the image, thus improving the accuracy and robustness in image recognition tasks. Inspired by these modules, we proposed our fast regional convolution module, which is one of the core parts of the present invention and is used to extract the key region features of facial expressions. The computational amount is reduced through a fully convolutional network, and the key regions of expression changes are captured through regional convolution. This module includes a fast regional convolution layer, a 1×1 fully connected layer, and a batch normalization layer. The specific structure is as Figure 2 shown. The fast regional convolution layer reduces the computational and memory requirements through partial convolution technology, improves the inference speed, and maintains the recognition accuracy at the same time.
[0013] The fast regional convolution module includes a fast regional convolution layer, a 1×1 fully connected layer, and a batch normalization layer. For a fast regional convolution layer with K convolution kernels of size k, an input dimension of C in and an output channel of C out , the corresponding weight matrix can be expressed as W∈R Cin×Cout×k×k×K , the bias is expressed as b∈R D , the cumulative mean of the batch normalization layer is μ, the cumulative standard deviation is δ, the scaling coefficient is γ, and the bias is β. Since both convolution and batch normalization during inference are linear operations, the convolution can be folded into a convolution with weights and biases as shown in the following equations:
[0014]
[0015] The weights and biases of the convolution layer during inference are as shown in the following equations:
[0016]
[0017] where i represents the number of branches.
[0018] During the training phase, the fast region convolution module contains one or more fast region convolution layers, which increases the capacity of the model during training and enables it to learn more complex feature representations. Then, the feature maps of these fast region convolution layers are aggregated, and batch normalization and fully convolutional operations are used to merge features in the spatial dimension. Finally, a non-linear activation function is utilized to make the model more flexible and complex. During the inference phase, the model structure is simplified to reduce the computational load and memory access, thereby enabling fast inference.
[0019] The implementation process of fast region convolution is divided into three stages. (1) Fully convolutional operation. A 1×1 fully convolutional operation is used to mix and adjust channel features, enhancing the expressive ability of facial expression features. (2) Region convolution. Different from traditional convolution, region convolution focuses on the feature representation of facial expression changes, selects the key regions that are most important for facial expression recognition, thereby reducing the spatial regions and the number of channels involved in convolution, significantly reducing the computational complexity and the number of memory accesses, and helping to maintain the real-time performance and speed of the model. Figure 3 The schematic diagram of the region convolution sub-module is given, and the self-attention mechanism is used to identify the key regions associated with facial expression changes. (3) Batch normalization. By normalizing the feature maps of small batches of data, adjusting and scaling the non-linear output, accelerating the convergence speed of the model, improving the stability of training, and alleviating overfitting to a certain extent, enabling the network to more efficiently learn the key features of facial expression changes.
[0020] 1.2 Implementation of the spatial channel feature fusion module
[0021] We propose an enhanced framework for brain tumor recognition, which utilizes the powerful feature extraction ability of deep learning networks and feature fusion techniques to enhance the integration ability of different features and improve the classification performance. Based on deep learning-based image enhancement techniques, it uses multi-scale feature fusion to capture feature information at different resolutions, enhancing the details and quality of the image. The self-supervised feature fusion information constraint method proposed in
[15] reduces the dependence on a large amount of labeled data through self-supervised learning and improves the effectiveness of feature fusion through information constraint, thereby obtaining better classification performance. The proposed adaptive multi-scale feature fusion method enables the model to better adapt to objects of different sizes and shapes, improving the accuracy and robustness of the model. Through the proposal of these modules, we propose a multi-scale attention module, which reduces memory access through depthwise separable convolution and fuses channel features and spatial features to extract more discriminative feature representations. The structure of this module is as Figure 3 shown. By using depth convolution, the computational load is reduced and the speed is increased, local features and details are extracted, the feature information of the facial region is fused, and the facial expression detail representation is enhanced.
[0022] As Figure 4As shown, the spatial channel feature fusion module uses channel features and spatial features to determine the important feature representations of expression changes. The specific process of spatial feature fusion is described as follows. First, in the channel branch, global pooling across the spatial dimension is used to aggregate spatial information:
[0023] s c = Concat(Avvg(T1), Max(T1), Avg(T2), Max(T2)) (3)
[0024] where S c represents the aggregated feature, T1 and T2 represent the input features, Avg(·) and Max(·) represent global average pooling and global max pooling respectively. Then, depth convolution is performed on the aggregated feature to determine the channel feature weights W c1 and W c2 :
[0025]
[0026] where Conv1(·) and Conv2(·) represent depth convolution. Secondly, the following formula is used to normalize the channel feature weights W c1 and W c2 :
[0027]
[0028] where and represent the normalized channel feature weights respectively. It should be noted that: using a similar process, the feature weights and of the spatial branch are determined, and the channel feature weights and spatial feature weights are fused to determine the important features corresponding to the expression changes. Finally, the following formula is used to fuse the spatial features and channel features:
[0029] Output = (W c1 + W s1 ) * T1 + (W c2 + W s2 ) * T2 (6)
[0030] Since the sum of the spatial feature weights and the channel feature weights is 1, important bi-temporal features are retained and useless bi-temporal features are discarded, thus achieving effective feature fusion.
[0031] Through this designed temporal feature fusion module, the model's ability to capture expression details is enhanced, the model's adaptability to complex conditions is improved, and the model's generalization ability and robustness are enhanced.
[0032] 1.3 Implementation of the multi-scale attention module
[0033] We improve the segmentation accuracy through the proposed multi-class semantic segmentation network based on edge enhancement and multi-scale attention mechanism. An adaptive mechanism is studied to adjust the weights of multi-scale attention, improving the interpretability and generalization of the model. Using multi-scale Res2Net and coordinate attention mechanism, protein-protein interaction points are predicted. The global attention mechanism and multi-scale feature fusion strategy are adopted to enhance the feature representation of small targets and improve the accuracy of small target detection under complex backgrounds and low signal-to-noise ratios. Therefore, we propose a multi-scale attention module, which enhances the feature representation ability of key regions of expression changes through multi-scale feature extraction strategy and spatial information aggregation strategy. The structure of this module is as Figure 4 shown. By obtaining multi-scale features, parallel attention branches and cross-space information aggregation, the model's perception ability of expression details is improved.
[0034] To solve these problems, this paper proposes a multi-scale attention module. This module first obtains feature representations of different scales to improve the model's perception ability of expression details; then, parallel attention branches are used to extract the attention weights of grouped features to capture the dependencies between spatial features; finally, cross-space information aggregation is adopted to extract significant features at different spatial positions. Most importantly, the multi-scale attention module integrates regional convolution, which only performs convolution on selected regions, significantly reducing the computational complexity and the number of memory accesses, thereby improving the real-time performance of the model. Therefore, the multi-scale attention module not only improves the recognition accuracy but also significantly reduces resource consumption and improves the processing speed, providing a strong technical support for real-time face expression recognition, especially in application scenarios that require rapid processing of complex face expression changes.
[0035] 3.4.1 Feature Grouping
[0036] The given feature X ∈ R C×H×W is divided into G groups of sub-features:
[0037]
[0038] Without loss of generality, assume G << C, and attention weights are used to strengthen the feature representation of regions of interest in each sub-feature.
[0039] 3.4.2 Parallel Attention Branches
[0040] To collect multi-scale spatial information, the multi-scale attention module proposed in this paper uses three branches to extract the attention weights of the feature map, where two branches perform 1×1 full convolution and the third branch performs 3×3 regional convolution.
[0041] To capture the dependencies between spatial features and reduce the computational complexity, this paper models cross - spatial information interaction. Specifically, in the 1×1 global convolution branch, global average pooling encoding of features is performed along two spatial directions respectively, while in the 3×3 partial convolution branch, only one 3×3 partial convolution, one batch normalization layer, and the non - linear activation function Relu are stacked to capture multi - scale feature representations.
[0042] The implementation process of the parallel attention branch is described as follows. On the one hand, this paper first concatenates the two encoded features along the height direction and the width direction respectively and makes them share the same 1×1 global convolution; then, the non - linear Sigmoid function is used to fit the two - dimensional double - norm distribution on the linear convolution; finally, to realize the cross - channel interaction features between the 1×1 global convolution branches, this paper aggregates the attention weights of two channels in each group by element - wise multiplication. On the other hand, the 3×3 regional convolution branch captures local cross - channel feature interactions using regional convolution to expand the feature space. In this way, the proposed multi - scale attention module not only encodes channel features to adjust the weights of different channel features, but also retains the accurate spatial structure information.
[0043] 3.4.3 Cross - spatial Information Aggregation Benefiting from the ability to establish correlations between channels and spaces, cross - spatial learning has been widely studied and widely applied in many computer vision tasks [27 - 28]. To make full use of channel and spatial features, this paper introduces a cross - spatial information aggregation strategy to achieve deeper feature aggregation.
[0044] The method in this paper aims to enhance feature representation by encoding global information. First, an attention map that captures multi - scale spatial information is constructed by taking the dot product of the feature map of the regional convolution branch and the non - linear feature map. Then, global average pooling is performed on the feature map of the regional convolution branch to encode global spatial information, and the feature map of the 1×1 branch is adjusted to the corresponding dimension, thereby constructing an attention map that captures spatial location information. Finally, the two spatial attention weights are aggregated, and the pixel - level pairwise relationship is captured using Sigmoid.
[0045] Through this cross - space information aggregation strategy, precise location information is integrated into the multi - scale attention module. This strategy enables the convolutional neural network to achieve more refined pixel - level attention allocation in high - level feature maps by fusing context information at different scales. In addition, the cross - space information aggregation strategy strengthens the parallelized structure of the convolutional kernels, enabling it to effectively capture the dependencies between cross - space information. Different from the traditional method of gradually forming a limited receptive field, cross - space information aggregation makes full use of context information during the feature mapping process through parallel 3×3 and 1×1 convolutions, improving the representation ability of local features, and thus achieving more accurate feature representation in computer vision tasks. Embodiments of the present invention provide an efficient facial expression recognition method, which can significantly reduce the demand for computing resources and shorten the processing time while maintaining high accuracy. Through experimental verification, the method proposed in the present invention has superior performance compared with other methods, and at the same time, the computational cost and network scale are much smaller than those of the same - type methods, which is particularly suitable for the application scenario of real - time facial expression recognition.
Claims
1. A facial expression recognition method that fuses spatial channel features and fast regional convolution, characterized in that The following steps are involved: Step A: Use the fast regional convolution module to reduce the amount of calculation through full convolution, and use regional convolution to extract the feature representation of the key areas of expression changes: Step B: Use the spatial channel feature fusion module to reduce memory access through depthwise separable convolution, and fuse channel features and spatial features to extract more recognizable feature representations; Step C: Use the multi-scale attention module to enhance the feature representation capability of key areas of expression changes through multi-scale feature extraction strategy and spatial information aggregation strategy.
2. The method according to claim 1, characterized in that, The fast region convolution module includes a fast region convolution layer, a 1×1 fully connected layer, and a batch normalization layer. Although the MobileNet network performs well in feature extraction using a 3×3 convolution layer, this structure has the following problems in application scenarios that require fast response, such as expression recognition. First, the 3×3 convolution is computationally intensive, which will result in slower processing speeds and increased latency on mobile devices, which is not conducive to expression recognition tasks that require real-time feedback. Second, the convolution operation requires frequent access to memory, which not only increases the memory access cost, but may also become a performance bottleneck. To solve the above problems, this paper replaces the 3×3 convolutional layer in the MobileNet network with Fast Region Convolution layers. By using the proposed partial convolution technology, the computational process is optimized and memory access is reduced, thereby improving the inference speed while enhancing the accuracy of facial expression recognition. The core of the fast convolution module lies in its ability to identify the region in the image that is most relevant to facial expression changes and perform convolution on that region, thus avoiding redundant calculations for the entire image region and better meeting the requirements of real-time facial expression recognition on mobile devices. As shown in Figure 2, the fast region convolution module includes a fast region convolution layer, a 1×1 fully connected layer, and a batch normalization layer. For the fast region convolution layer with K convolution kernels of size k, an input dimension of C in , and an output channel of C out , the corresponding weight matrix can be represented as W ∈ R Cin×Cout×k×k×K , the bias is represented as b ∈ R D , the cumulative mean of the batch normalization layer is μ, the cumulative standard deviation is δ, the scaling coefficient is γ, and the bias is β. Since both convolution and batch normalization during inference are linear operations, the convolution can be folded into a convolution with weights and biases as shown in the following equations: The weights and biases of the convolutional layer during inference are shown in the following formulas: Here, i represents the number of branches. During the training phase, the fast region convolution module contains one or more fast region convolution layers, which increases the capacity of the model during training and learns more complex feature representations; Then, the feature maps of these fast regional convolutional layers are aggregated, and batch normalization and full convolution are used to merge features in the spatial dimension; finally, nonlinear activation functions are used to make the model more flexible and complex. In the inference stage, the model structure is simplified to reduce the amount of calculation and memory access, thereby achieving fast inference. In order to make full use of the information from all channels, this paper uses the regional convolution submodule. Compared with conventional convolution, the effective receptive field of the regional convolution submodule on the feature map focuses more on the central position. As shown in Figure 3, first calculate the Frobenius norm at each position p: where i = 1, 2, 3... k 2 , j, and c are the number of channels of the feature map, represents the j-th channel of the convolutional kernel F at position i and then evaluates the importance of each position. If a certain position has a larger Frobenius norm than other positions, then that position is often more important. The implementation process of fast regional convolution is divided into three stages. (1) Full convolution. 1×1 full convolution is used to mix and adjust channel features to enhance the expressiveness of expression features. (2) Regional convolution. Unlike traditional convolution, regional convolution focuses on the feature representation of expression changes and selects the key areas that are most important for expression recognition, thereby reducing the number of spatial areas and channels involved in the convolution, significantly reducing the computational complexity and memory access times, and helping to maintain the real-time and fast performance of the model. Figure 3 shows a schematic diagram of the regional convolution submodule, which uses the self-attention mechanism to identify key areas associated with expression changes. (3) Batch normalization. By normalizing the feature maps of small batches of data, adjusting and scaling nonlinear outputs, accelerating the convergence speed of the model, improving the stability of training, and alleviating overfitting to a certain extent, the network can learn the key features of expression changes more efficiently.
3. The method according to claim 1, characterized in that, The described spatial-channel feature fusion module reduces the computational complexity and improves the speed through depth convolution, extracts local features and details, fuses the facial region feature information, and enhances the expression detail representation. In the field of facial expression recognition, researchers face various challenges, such as subtle expression changes that are difficult to capture, changes in lighting, changes in shooting angles, uneven distribution of dataset samples, insufficient model generalization ability, and computational efficiency when running on resource-constrained terminal devices. These problems directly affect the accuracy and real-time performance of the model. To address these challenges, this paper introduces a spatial-channel feature fusion module for improving the performance of expression recognition. First, the spatial-channel attention module effectively extracts local features and retains more detailed information while reducing the computational complexity and increasing the computational speed by introducing depth convolution, thereby improving the recognition effect for different expression intensities and different expression types. Second, the spatial-channel attention module enhances the feature representation of expression details by fusing the feature information of different regions of the face, especially improving the recognition accuracy when the expression intensity is low or there is partial occlusion of the face. Therefore, the spatial-channel feature fusion module increases the robustness of the model to individual differences, age changes, and expression changes through the fusion of spatial features and channel features, ensuring the accuracy of the model under diverse conditions. The spatial-channel feature fusion module uses channel features and spatial features to determine the important feature representations of expression changes. The specific process of spatial feature fusion is described as follows. First, in the channel branch, global pooling across the spatial dimension is used to transmit and aggregate spatial information: s c = Concat(Avg(T1), Max(T1), Avg(T2), Max(T2)) (4) Among them, S c represents the aggregation feature, T1 and T2 represent the input features, Avg(·) and Max(·) respectively represent global average pooling and global max pooling. Then, perform depth convolution on the aggregation feature to respectively determine the channel feature weights W c1 and W c2 : Where Conv1(·) and Conv2(·) represent depth convolution. Secondly, the channel feature weights W c1 and W c2 are normalized using the following formula: Among them and respectively represent the normalized channel feature weights. It should be noted that: using a similar process, the feature weights of the spatial branch are determined and The channel feature weights and spatial feature weights are fused to determine the important features corresponding to the expression changes. Finally, the spatial features and channel features are fused using the following formula: Output =(W′ c1 +W′ s1 )*T1+(W′ c2 +W′ s2 )*T2 (7) Since the sum of the spatial feature weight and the channel feature weight is 1, important bi-temporal features are retained and useless bi-temporal features are discarded, thus achieving effective feature fusion. Through this designed temporal feature fusion module, the model's ability to capture expression details is enhanced, the adaptability of the model to complex conditions is improved, and the generalization ability and robustness of the model are enhanced.
4. The method according to claim 1, wherein The described multi-scale attention module improves the model's perception ability of expression details by obtaining multi-scale features, parallel attention branches, and cross-spatial information aggregation. The coordinate attention mechanism utilizes the coordinate information of each position in the feature map to enhance the model's attention to specific regions. The coordinate attention mechanism is particularly suitable for visual tasks because it can capture the spatial structure information of images and adaptively focus on the most important image regions using attention weights. Despite the many advantages of the coordinate attention mechanism, it requires high computational resources because it needs to calculate the attention weights of each position in the feature map, resulting in a large model size and a complex training process, and thus insufficient generalization ability. Due to the design limitations of the coordinate attention mechanism, it cannot fully capture the subtle changes in expressions, especially under complex backgrounds or non-ideal lighting conditions, which affects the accurate recognition of expression intensity and expression types by the model. To address these issues, this paper proposes a multi-scale attention module. This module first obtains feature representations at different scales to enhance the model's perception ability of expression details; for fusion, parallel attention branches are used to extract the attention weights of grouped features to capture the dependencies between spatial features; finally, cross-space information aggregation is adopted to extract the significant features at different spatial positions. Most importantly, the multi-scale attention module integrates regional convolution, which only performs convolution on the selected regions, significantly reducing the computational complexity and the number of memory accesses, and thus improving the real-time performance of the model. Therefore, the multi-scale attention module improves the recognition accuracy while significantly reducing resource consumption and increasing the processing speed, providing a strong technical support for real-time face expression recognition, especially in application scenarios that require rapid processing of complex face expression changes. Partition the given feature \(X\in\mathbb{R}\) C×H×W into \(G\) groups of sub - features: Without loss of generality, assume G << C and use the attention weights to enhance the feature representation of the region of interest in each sub-feature. To collect multi-scale spatial information, the multi-scale attention module proposed in this paper uses three branches to extract the attention weights of the feature maps, where two branches perform 1×1 fully convolutional operations and the third branch performs 3×3 regional convolution. To capture the dependencies between spatial features and reduce the computational complexity, this paper models the cross-space information interaction. Specifically, in the 1×1 fully convolutional branch, the features are globally averaged and pooled along two spatial directions for encoding, while in the 3×3 partial convolutional branch, only one 3×3 partial convolution, one batch normalization layer, and the non-linear activation function Relu are stacked to capture multi-scale feature representations. This paper reshapes and permutes the sub-features of the G groups of feature maps into the batch dimension and redefines the tensor in the shape of C / / G×H×W. On the one hand, this paper first concatenates the two encoded features along the height direction and the width direction respectively and makes them share the same 1×1 fully convolutional operation; then, the non-linear Sigmoid function is used to fit the two-dimensional double-norm distribution on the linear convolution; finally, to achieve the cross-channel interactive features between the 1×1 fully convolutional branches, this paper aggregates the attention weights of the two channels within each group using element-wise multiplication. On the other hand, the 3×3 regional convolutional branch uses regional convolution to capture local cross-channel feature interactions and expand the feature space. In this way, the proposed multi-scale attention module not only encodes the channel features to adjust the weights of different channel features, but also retains the accurate spatial structure information. Benefiting from the ability to establish correlations between channels and spaces, cross-space learning has been widely studied and widely applied in many computer vision tasks [27-28]. To make full use of the channel and spatial features, and maintain high resolution in attention learning, this paper introduces a cross-space information aggregation strategy to achieve deeper feature aggregation. The method in this paper aims to enhance feature representation by encoding global information. To perform this process efficiently, a non-linear Softmax function is adopted here to process the feature map of global average pooling; global average pooling helps to fit linear transformation and lays the foundation for subsequent matrix dot product operations; an attention map capturing multi-scale spatial information is constructed by taking the dot product of the feature map of the regional convolution branch and the non-linear feature map. In addition, global average pooling is performed on the feature map of the regional convolution branch to encode global spatial information, and the feature map of the 1×1 branch is adjusted to the corresponding dimension, thereby constructing an attention map capturing spatial location information. Finally, the two spatial attention weights are aggregated, and pixel-level pairwise relationships are captured using Sigmoid. Through this cross-spatial information aggregation strategy, precise location information is integrated into the multi-scale attention module. This strategy enables the convolutional neural network to achieve more refined pixel-level attention allocation in high-level feature maps by fusing context information at different scales. In addition, the cross-spatial information aggregation strategy strengthens the structure of convolutional kernel parallelization, enabling it to effectively capture the dependencies between cross-spatial information. Different from the traditional method of gradually forming a limited receptive field, cross-spatial information aggregation makes full use of context information during the feature map process through parallel 3×3 and 1×1 convolutions, improving the representational ability of local features, and thus achieving more accurate feature representation in computer vision tasks.
Citation Information
Cited By
Intelligent analysis method and system for information extraction of intelligent electric meter
CN120779323A
Lightweight change detection method and device, computer equipment and storage medium
CN121170417A
Improved network model crack identification and segmentation method for complex texture background
CN122156601A
An improved network model crack identification segmentation method for complex texture background
CN122156601B
Automatic corn canopy coverage extraction method based on deep learning
CN122265751A