An image semantic end-to-end transmission method based on a double attention mechanism

CN117746211BActive Publication Date: 2026-09-11BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311766799.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-21
Publication Date
2026-09-11
Estimated Expiration
2043-12-21

AI Technical Summary

Technical Problem

但是该方案仍然存在以下问题:该方案中所使用的注意力机制本质上是一个挤压激励网络,只考虑不同通道中特征的重要性而忽略了各个位置像素的重要性,导致图像压缩和传输过程中丢失了某些有价值的信息,影响图像传输的准确性

Benefits of technology

[0054] Most existing adaptive JSCC schemes for image transmission only consider different channel conditions, neglecting the importance of different features in image processing and transmission. This invention introduces channel attention and spatial attention mechanisms. Considering the interdependence of spatial and channel features, an SNR adaptive CBAM is proposed. By concatenating pooled feature vectors with signal-to-noise ratio state information, and then passing them sequentially through channel attention and spatial attention networks to obtain importance evaluation results, different weights are assigned to different features based on different SNR levels and image content. This enhances the feature representation capability of the image, improves the accuracy of image transmission, achieves better transmission performance, and saves storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117746211B_ABST
    Figure CN117746211B_ABST
Patent Text Reader

Abstract

The application discloses an image semantic end-to-end transmission method based on a double attention mechanism, relates to the field of semantic communication, and comprises the following steps: preprocessing an input original image through a normalization layer; extracting different semantic features of the original image by using a semantic feature extractor; evaluating the importance of the extracted semantic features and reallocating weights by using a semantic importance evaluation module to obtain a semantic importance feature vector; encoding the semantic importance feature vector and a channel feedback signal-to-noise ratio by using a JSCC encoder to obtain a complex-valued channel input symbol vector; decoding by using a JSCC decoder; reconstructing a target image by using a feature reconstruction module; and constructing a loss function. The application introduces a channel attention mechanism and a spatial attention mechanism, proposes an SNR adaptive CBAM, enhances the feature representation capability of an image, realizes better transmission performance, improves the flexibility and robustness of image semantic transmission, and saves storage space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semantic communication technology, and specifically to an end-to-end image semantic transmission method based on a dual attention mechanism. Background Technology

[0002] With the development of artificial intelligence and computing technology, semantic communication focuses on the semantic level of information, aiming to accurately convey the expected meaning of messages. It can effectively reduce communication bandwidth and improve communication robustness, and is widely used in image transmission. With the development of deep learning (DL) technology, joint source-channel coding (JSCC) technology for end-to-end transmission has received widespread attention. Since JSCC technology can jointly design and optimize the source coding and channel coding processes, effectively solve the "cliff effect", reduce communication bandwidth, and improve communication robustness, JSCC technology is widely used in semantic communication. In the paper "Bourtsoulatze E, Kurka DB, Gündüz D. Deep joint source-channel coding for wireless image transmission[J].IEEE Transactions on Cognitive Communications and Networking,2019,5(3):567-579", Bourtsoulatze et al. introduced a deep JSCC scheme, which can directly map pixel values ​​from the image to complex value channel input symbols, surpassing traditional transmission solutions and being particularly effective in low signal-to-noise ratio scenarios. Yang M et al. proposed a deep JSCC scheme with adaptive rate control for wireless image transmission (Yang M, Kim HS. Deepjoint source-channel coding for wireless image transmission with adaptive rate control[C] / / ICASSP 2022-2022IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP).IEEE,2022:5193-5197.), in which the output of the semantic encoder is divided into selective and non-selective features, and a policy network is used to determine the transmission of selective and non-selective features. This scheme can achieve performance comparable to optimization models trained specifically using static target rates.Gunduz D et al. proposed a JSCC scheme that incorporates channel output feedback, which demonstrates good performance in reconstruction quality for end-to-end fixed-length image transmission and reduces the average delay for variable-length image transmission (Kurka DB, Gündüz D. Deep joint source-channel coding of images with feedback[C] / / ICASSP2020-2020IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP).IEEE,2020:5235-5239). In the paper "Bao X, Jiang M, Zhang H. ADJSCC-l: SNR-Adaptive JSCC Networks for Multi-Layer Wireless Image Transmission [C] / / 2021 7th International Conference on Computer and Communications (ICCC). IEEE, 2021: 1812-1816.", Bao X et al. proposed an image transmission architecture that uses an SNR adaptive module and exhibits excellent recovery capability for the mismatch between the SNR of the training and test channels caused by channel variations. This architecture can successfully improve the reconstruction quality of wireless image transmission in scenarios with low signal-to-noise ratio and limited bandwidth.

[0003] Currently, an attention-based deep JSCC (ADJSCC) scheme has been proposed in existing technologies (Xu J, Ai B, Chen W, et al. Wireless image transmission using deep source channel coding with attention modules[J].IEEE Transactions on Circuits and Systems for Video Technology, 2021, 32(4):2315-2328). Inspired by the resource allocation strategy in traditional JSCC, this scheme can successfully train only one network to cover a large number of signal-to-noise ratios at different SNR levels, without having to train a large number of neural networks to cover different SNR scenarios. Furthermore, this scheme will dynamically adjust the compression ratio of source coding and the coding rate of channel coding according to the channel SNR, which can allocate computing resources to more critical tasks. Unlike the resource allocation strategy of traditional JSCC, this scheme uses channel-based soft attention to scale features according to the SNR conditions, which occupies less storage space and has stronger robustness in the presence of channel mismatch. However, the scheme still has the following problems: the attention mechanism used in the scheme is essentially a squeeze-excitation network, which only considers the importance of features in different channels and ignores the importance of pixels at each position, resulting in the loss of some valuable information during image compression and transmission, affecting the accuracy of image transmission. Summary of the Invention

[0004] For image sources in environments with poor signal-to-noise ratios (SNR), extracting semantic features from image content is crucial for efficient information exchange. However, existing attention-based deep JSCC schemes mostly only consider different channel conditions, neglecting the varying importance of different features in image processing and transmission. Uniform compression of different features in an image may lead to the loss of key image details, especially in low SNR scenarios. Therefore, this invention provides an end-to-end image semantic transmission method based on a dual attention mechanism.

[0005] The technical solution adopted by this invention to solve the technical problem is as follows:

[0006] The present invention provides an end-to-end image semantic transmission system based on a dual attention mechanism, which mainly includes:

[0007] The normalization layer, transmitter, channel, receiver, and denormalization layer are connected in sequence.

[0008] The normalization layer is used to preprocess the original input image.

[0009] The transmitter includes a semantic feature extractor, a JSCC encoder, and a first semantic importance evaluation module. The semantic feature extractor is used to extract different semantic features of the original image. The first semantic importance evaluation module is used to evaluate the importance of the extracted semantic features and reassign weights to obtain a semantic importance feature vector. The JSCC encoder is used to encode the semantic importance feature vector and the channel feedback signal-to-noise ratio to obtain a complex-valued channel input symbol vector.

[0010] A channel is used to connect a transmitter and a receiver to transmit images;

[0011] The receiver includes a JSCC decoder, a feature reconstruction module, and a second semantic importance evaluation module. The JSCC decoder and the second semantic importance evaluation module work together to reduce noise interference in the additive white Gaussian noise channel and recover semantic features. The feature reconstruction module is used to deeply mine semantic feature information through an attention mechanism, fuse semantic features, and reconstruct the target image.

[0012] The denormalization layer scales each element before outputting the target image.

[0013] Furthermore, the first semantic importance evaluation module is implemented by an SNR adaptive CBAM module, which is composed of a channel attention mechanism module and a spatial attention mechanism module connected sequentially. The SNR adaptive CBAM module can be integrated into the main layers of the convolutional neural network. The channel attention mechanism module consists of a first average pooling layer, a first maximum pooling layer, a multilayer perceptron with hidden layers, and a first sigmoid activation function layer. The spatial attention mechanism module consists of a second average pooling layer, a second maximum pooling layer, a connection layer, a convolutional layer, and a second sigmoid activation function layer.

[0014] Furthermore, the semantic feature extractor consists of two generated feature blocks, each connected to an SNR adaptive CBAM module, which assigns feature weights based on image content and channel information. Each generated feature block includes a second convolutional layer, a generalized division normalization layer, and a parameter-corrected linear unit layer. The image features generated by the generated feature blocks include edges and textures.

[0015] Furthermore, the JSCC encoder consists of two residual convolutional blocks, a first convolutional layer, and a generalized division normalization layer. Each residual convolutional block is followed by an SNR adaptive CBAM module, which allocates feature weights according to image content and channel information. Each residual convolutional block employs a pre-activated residual neural network, which includes a third convolutional layer, a fourth convolutional layer, two generalized division normalization layers, and two parameter-corrected linear unit layers.

[0016] Furthermore, the structure of the second semantic importance assessment module is the same as that of the first semantic importance assessment module.

[0017] The present invention provides an end-to-end image semantic transmission method based on a dual attention mechanism, implemented using the aforementioned end-to-end image semantic transmission system based on a dual attention mechanism. The method includes the following steps:

[0018] S1: The input original image is composed of vectors express, Let n represent the set of real numbers. The size of the original input image is n = H × W × C, where H is the height of the image, W is the width of the image, and C is the number of channels of the image.

[0019] S2: Preprocess the original input image using a normalization layer;

[0020] S3: Use a semantic feature extractor to extract different semantic features from the original image;

[0021] S4: The importance of the extracted semantic features is evaluated by the first semantic importance evaluation module, and the weights are reassigned to obtain the semantic importance feature vector.

[0022] S5: The semantic importance feature vector s′ and the channel feedback signal-to-noise ratio μ are encoded using a JSCC encoder to obtain the complex-valued channel input symbol vector. k represents the size of the channel input symbols. Representing the set of complex numbers:

[0023] s′=M β (T α (s),μ)

[0024] z = M β (E θ (s′),μ)

[0025] Where T α (·) represents a semantic feature extraction network, whose network parameters are denoted as α; M β (·) denotes the semantic importance estimator, whose network parameters are denoted as β; E θ(·) represents the encoding function of the corresponding JSCC encoder, whose network parameters are represented by θ; power normalization is performed on the vector z. Where z* represents the complex conjugate of z;

[0026] S6: Vector z is transmitted to the receiver through the channel, and the JSCC decoder receives the channel output signal. Represented as:

[0027]

[0028] Where vector From the distribution Composed of independent and identically distributed samples; σ 2 Indicates noise power. Representing a circularly symmetric complex Gaussian distribution, where I represents the identity matrix; the JSCC decoder uses the decoding function R. ξ (·) is used to map the channel output signal The channel feedback signal-to-noise ratio μ,ξ represents the set of decoding function parameters;

[0029] S7: The feature reconstruction module uses the reconstruction function R. η (·) is used to reconstruct the target image, where η represents the reconstruction function R. η The parameter set of (·):

[0030]

[0031] in This indicates the channel output signal received by the JSCC decoder. This represents an estimate of the original image s;

[0032] S8: Construct the loss function;

[0033] Original images and reconstructed images The mean square error distribution between them is as follows:

[0034]

[0035] Where s i and Representing the original image s and the reconstructed image respectively. The color component intensity of each pixel in the dataset is determined; the JSCC encoder and decoder are modeled using a convolutional neural network, with the optimization objective being to find the optimal parameters to minimize the distortion θ. * and ξ * The loss function for a convolutional neural network is:

[0036]

[0037] Where θ *ξ represents the optimal parameters of the JSCC encoder. * This represents the optimal parameters for the JSCC decoder. Represents the original image s and the reconstructed image The joint probability distribution, where p(μ) represents the probability distribution of the signal-to-noise ratio.

[0038] Furthermore, in step S4, for the feature map obtained from the semantic feature extractor... The SNR adaptive CBAM module generates a one-dimensional channel attention map based on the SNR level. and 2D spatial attention map Then, multiply it with the input feature map to generate a new feature map. Used for joint encoding and decoding:

[0039]

[0040]

[0041] in This represents element-wise multiplication; throughout the multiplication process, attention weights are broadcast: channel attention weights are broadcast in the spatial dimension, and spatial attention weights are broadcast in the channel dimension.

[0042] Furthermore, the specific implementation steps of the SNR adaptive CBAM module in generating a one-dimensional channel attention map based on the SNR level are as follows:

[0043] The channel attention mechanism module uses average pooling and max pooling operations to merge spatial information in the feature maps and obtain average pooling features. and max pooling features These features are concatenated with the channel feedback signal-to-noise ratio μ to generate contextual information. and

[0044]

[0045]

[0046] After reshaping the contextual information, the features are input into the same multilayer perceptron with hidden layers to generate a one-dimensional channel attention map. Hide activation size set to r is the reduction ratio; element-wise summation is used to merge output features to generate a channel attention map.

[0047]

[0048] Where δ and σ represent the rectified linear unit and the sigmoid activation function, respectively; the two inputs of the network share MLP weight parameters W0 and W1, where

[0049] Furthermore, the specific steps for obtaining the two-dimensional spatial attention map are as follows:

[0050] The one-dimensional channel attention map is used as the input feature map for the spatial attention mechanism module; average pooling and max pooling operations are applied along the channel axis to aggregate the channel information of the feature map, generating two two-dimensional feature maps. and After concatenating these two 2D feature maps to generate an effective feature descriptor, convolutional layers and sigmoid layers are applied to reduce the dimensionality and generate a 2D spatial attention map.

[0051]

[0052] Where σ represents the Sigmoid activation function, f 7×7 This represents a convolution operation with a kernel size of 7×7.

[0053] The beneficial effects of this invention are:

[0054] Most existing adaptive JSCC schemes for image transmission only consider different channel conditions, neglecting the importance of different features in image processing and transmission. This invention introduces channel attention and spatial attention mechanisms. Considering the interdependence of spatial and channel features, an SNR adaptive CBAM is proposed. By concatenating pooled feature vectors with signal-to-noise ratio state information, and then passing them sequentially through channel attention and spatial attention networks to obtain importance evaluation results, different weights are assigned to different features based on different SNR levels and image content. This enhances the feature representation capability of the image, improves the accuracy of image transmission, achieves better transmission performance, and saves storage space.

[0055] In addition, this invention takes into account the differences in the importance of semantic features and matches semantically important features with channel conditions to improve the flexibility and robustness of image semantic transmission.

[0056] This invention assigns weights based on image content and channel conditions, enabling efficient processing and accurate transmission of image semantic information in complex communication environments. Addressing the future rapid development of information technology and the demand for large-scale, highly adaptable, and low-latency information interaction in complex environments, this invention enhances the efficiency and robust information carrying capacity of communication. Furthermore, SNR-adaptive CBAM can be integrated as a plug-and-play module into many network architectures; even with changes in network deployment, this attention module can still be applied, demonstrating strong portability. Attached Figure Description

[0057] Figure 1 This is a block diagram illustrating the structural composition of an end-to-end image semantic transmission system based on a dual attention mechanism according to the present invention.

[0058] Figure 2 This is a structural diagram of the SNR adaptive CBAM module.

[0059] Figure 3 This is a diagram illustrating the specific architecture of an image semantic encoder.

[0060] Figure 4 A comparison of peak signal-to-noise ratio (PSNR) performance on AWGN channels.

[0061] Figure 5 A comparison of structural similarity (SSIM) performance on AWGN channels. Detailed Implementation

[0062] The present invention will be further described in detail below with reference to the accompanying drawings.

[0063] Firstly, this invention provides an end-to-end image semantic transmission system based on a dual attention mechanism. This system extracts source features using a deep learning-based semantic feature extractor and assigns different weights to different features through joint training of a first semantic importance evaluation module and a JSCC encoder. The JSCC decoder recovers the semantic features and reconstructs the target image at the receiver using a feature reconstruction module. The architecture proposed in this invention uses a semantic importance evaluation module that can simultaneously consider image content and channel SNR conditions for image compression and reconstruction. This invention effectively allocates attention weights to more critical tasks and can adaptively adjust the data rate in different channel environments to achieve efficient semantic transmission.

[0064] The system architecture is as follows Figure 1 As shown, it mainly consists of the following basic modules: normalization layer, transmitter, channel, receiver, and denormalization layer, which are connected in sequence. The composition and function of each module are as follows:

[0065] 1. Normalization layer, used to preprocess the input raw image, mapping the image pixel values ​​to the range of (0-1).

[0066] 2. The transmitter mainly includes an image semantic encoder and a first semantic importance evaluation module. The image semantic encoder mainly includes a semantic feature extractor and a JSCC encoder. The semantic feature extractor is mainly used to extract different semantic features of the original image. The first semantic importance evaluation module is mainly used to evaluate the importance of the extracted semantic features and reassign weights to obtain a semantic importance feature vector. The JSCC encoder is mainly used to encode the semantic importance feature vector and the channel feedback signal-to-noise ratio to obtain a complex-valued channel input symbol vector.

[0067] Specifically, the semantic feature extractor in the image semantic encoder mainly employs a semantic feature extraction network (Convolutional Neural Network (CNN)). Both the semantic feature extraction network and the JSCC encoder consist of multiple non-linear layers. The input to the semantic feature extractor is the source values, and the output of the JSCC encoder is the channel symbols. The semantic feature extractor can use the Convolutional Neural Network (CNN) to extract various features from the image and then encode them in the JSCC encoder before transmitting them to the channel.

[0068] The first semantic importance assessment module is mainly implemented by the SNR adaptive CBAM module, and its specific structure is as follows: Figure 2 As shown, the SNR adaptive CBAM module is mainly composed of a channel attention mechanism module and a spatial attention mechanism module connected sequentially. The channel attention mechanism module mainly consists of a first average pooling layer AvgPool_1, a first maximum pooling layer MaxPool_1, a multi-layer perceptron (MLP) with hidden layers, and a first sigmoid activation function layer. The spatial attention mechanism module mainly consists of a second average pooling layer AvgPool_2, a second maximum pooling layer MaxPool_2, a connection layer, a convolutional layer 2DConv, and a second sigmoid activation function layer. The SNR adaptive CBAM module can be integrated into the main layers of a convolutional neural network, such as convolutional layers and transposed convolutional layers. This invention introduces a dual attention mechanism and proposes a signal-to-noise ratio adaptive deep JSCC mechanism with a Convolutional Block Attention Module (CBAM), which can apply matrix operations to features in both spatial and channel dimensions.

[0069] like Figure 3As shown, the semantic feature extractor mainly consists of two Generate Feature Blocks (GFBs); each GFB is followed by an SNR-adaptive CBAM module, which can allocate feature weights according to image content and channel information. The JSCC encoder mainly consists of two Residual Convolution Blocks (RCBs), a first convolutional layer Conv2D_1, and a Generalized Divisive Normalization (GDN) layer. Each Residual Convolution Block (RCB) is followed by an SNR-adaptive CBAM module, which can allocate feature weights according to image content and channel information.

[0070] Specifically, the Generate Feature Block (GFB) mainly consists of a second convolutional layer (Conv2D_2), a Generalized Divisive Normalization (GDN) layer, and a Parametric Rectified Linear Unit (PReLU) layer. These layers are connected sequentially. The image features generated by the Generate Feature Block (GFB) include edges and textures. The second convolutional layer (Conv2D_2) is composed of parameters m×m×C|↓s t This means the kernel size is m and the number of output channels is C; in this embodiment, the kernel size of the second convolutional layer Conv2D_2 is 7, and the number of output channels is 512; the symbol ↓ represents the downsampling operation, and the parameter s t This represents the stride in the second convolutional layer, Conv2D_2.

[0071] Specifically, the Residual Convolution Block (RCB) mainly consists of a pre-activation residual neural network (Pre-activation ResNet) containing two convolutional layers (Conv2D_3 and Conv2D_4), two generalized divisive normalization (GDN) layers, and two parametric rectified linear unit (PReLU) layers. The first GDN layer, the first PReLU layer, the third convolutional layer (Conv2D_3), the second GDN layer, the second PReLU layer, and the fourth convolutional layer (Conv2D_4) are connected in sequence. The parameter settings for the third and fourth convolutional layers, Conv2D_3 and Conv2D_4, are the same as those for the second convolutional layer, Conv2D_2. Residual convolution blocks (RCBs) can mitigate gradient vanishing and overfitting during network training, enhancing the model's training performance.

[0072] In traditional Joint Source-channel Coding (JSCC), batch normalization (BN) introduces different means and standard deviations for each processed batch, which is equivalent to adding noise. Therefore, this method is unsuitable for generative models such as image reconstruction and compression. This invention introduces a Generalized Divisive Normalization (GDN) layer to normalize features. The GDN layer expands the channels of each module in the first convolutional layer Conv2D_1 to effectively address tasks related to image compression and density modeling. Furthermore, this invention incorporates a Pre-activation Residual Neural Network (Pre-activation ResNet), which enhances the training performance and transmission accuracy of the model by changing the position of the normalization layer within the Pre-activation Residual Neural Network.

[0073] This invention designs an image semantic encoder using a residual structure, which combines a pre-activation residual neural network (Pre-activation ResNet) to change the position of the normalization layer in the residual structure. This is used to alleviate gradient vanishing and overfitting during network training, enhance the training effect of the model, and effectively improve the image reconstruction quality at a specific compression ratio.

[0074] 3. Channel, mainly used to connect the transmitter and receiver to transmit images.

[0075] 4. Receiver, mainly including: an image semantic decoder and a second semantic importance assessment module. The image semantic decoder mainly includes: a JSCC decoder and a feature reconstruction module. The JSCC decoder is mainly used in conjunction with the second semantic importance assessment module to reduce noise interference in the Additive White Gaussian Noise (AWGN) channel and recover semantic features. The second semantic importance assessment module is mainly used in conjunction with the JSCC decoder to reduce noise interference in the Additive White Gaussian Noise (AWGN) channel and recover semantic features. The feature reconstruction module is mainly used to deeply mine semantic feature information through an attention mechanism, fuse semantic features, and reconstruct the target image.

[0076] The second semantic importance assessment module has the same structural composition as the first semantic importance assessment module; that is, the second semantic importance assessment module is also implemented by the SNR adaptive CBAM module, and its specific structural composition is as follows: Figure 2 As shown, please refer to the above description for a detailed introduction.

[0077] The structure of the image semantic decoder is completely symmetrical to that of the image semantic encoder (for example, the convolutional layer in the image semantic encoder becomes the deconvolutional layer, and the GDN layer becomes the anti-GDN layer). Therefore, the image semantic decoder performs the opposite operations of the image semantic encoder in sequence.

[0078] 5. Denormalization layer: Before outputting the target image, each element needs to be scaled back to a value within the range of (0-255) using a denormalization layer.

[0079] Secondly, this invention provides an end-to-end image semantic transmission method based on a dual attention mechanism. This invention jointly considers the importance of image content and different signal-to-noise ratio conditions to establish an adaptive semantic feature extraction and semantic representation method. In the joint coding of the source and channel, image importance is prioritized, and different weights are assigned to different levels of importance. Based on semantic importance, image compression and transmission are performed at the transmitter, and image source reconstruction is performed at the receiver. This invention optimizes the encoding and transmission process of image sources, solving the problem of semantically adaptive image transmission under poor transmission environments.

[0080] The specific operation process of this method is as follows:

[0081] S1: The input original image is composed of vectors It means that, among them Let n represent the set of real numbers. The size of the original input image is n = H × W × C, where H is the height of the image, W is the width of the image, and C is the number of channels of the image.

[0082] S2: Preprocessing of the original image;

[0083] First, the original input image is preprocessed using a normalization layer, which maps the image pixel values ​​to the range of 0 to 1.

[0084] S3: Use a semantic feature extractor to extract different semantic features from the original image;

[0085] Using a semantic feature extractor based on a convolutional neural network (CNN) and channel feedback signal-to-noise ratio Combine different semantic features extracted from the original image.

[0086] S4: The importance of the extracted semantic features is evaluated by the first semantic importance evaluation module, and the weights are reassigned to obtain the semantic importance feature vector.

[0087] The first semantic importance evaluation module achieves this by evaluating the importance weights of features. Specifically, for the feature map obtained from the semantic feature extractor... The SNR adaptive CBAM module generates a one-dimensional channel attention map based on the SNR level. and 2D spatial attention map Then, multiply it with the input feature map to generate a new feature map. Used for joint encoding and decoding. The specific process is as follows:

[0088]

[0089]

[0090] in This represents element-wise multiplication. Throughout the multiplication process, attention weights are broadcast: channel attention weights are broadcast in the spatial dimension, and spatial attention weights are broadcast in the channel dimension.

[0091] In addition, the specific implementation process of the channel attention mechanism module is as follows:

[0092] The channel attention mechanism primarily focuses on the inter-channel relationships of features. The channel attention mechanism module mainly consists of a first average pooling layer (AvgPool_1), a first maximum pooling layer (MaxPool_1), a multi-layer perceptron (MLP) with hidden layers, and a first sigmoid activation function layer.

[0093] To infer more refined channel attention, both average pooling and max pooling are considered during image compression. Average pooling provides feedback to every pixel in the feature map, while max pooling only receives feedback from the location with the highest pixel value during gradient backpropagation. Figure 2 As shown, this invention uses average pooling and max pooling operations to merge spatial information in the feature map, thereby obtaining average pooling features. and max pooling features These features are concatenated with the channel feedback signal-to-noise ratio μ to generate contextual information. and As shown below:

[0094]

[0095]

[0096] After reshaping the contextual information, the features are fed into the same multi-layer perceptron (MLP) with hidden layers to generate one-dimensional channel attention maps. To minimize parameter overhead, the hidden activation size is set to... Where r is the reduction ratio (set to 16 in this invention). The output features are merged using element-wise summation (vector 1 and vector 2) to generate the channel attention map. The specific calculation process of the channel attention mechanism module is shown below:

[0097]

[0098] Where δ and σ represent the Rectified Linear Unit (ReLU) and the Sigmoid activation function, respectively. The two inputs of the network share MLP weight parameters W0 and W1, where...

[0099] In addition, the specific implementation process of the spatial attention mechanism module is as follows:

[0100] The spatial attention mechanism primarily focuses on the spatial relationships of features. The spatial attention mechanism module mainly consists of a second average pooling layer (AvgPool_2), a second maximum pooling layer (MaxPool_2), a connection layer, a convolutional layer (2DConv), and a second sigmoid activation function layer.

[0101] The feature map output by the channel attention mechanism module serves as the input feature map for the spatial attention mechanism module. For example... Figure 2 As shown, firstly, this invention applies average pooling and max pooling operations along the channel axis to aggregate the channel information of the feature map, generating two two-dimensional feature maps: and After concatenating these two 2D feature maps to generate effective feature descriptors [MaxPool, AvgPool], a 2DConv convolutional layer and a second Sigmoid activation function layer are applied to reduce dimensionality and generate a 2D spatial attention map. The specific calculation process of the spatial attention mechanism module is shown below:

[0102]

[0103] Where σ represents the Sigmoid activation function, f 7×7 This represents a convolution operation with a kernel size of 7×7.

[0104] This invention proposes a dual attention module based on SNR adaptive mechanism and CBAM. The channel attention mechanism module and the spatial attention mechanism module can effectively enhance the model's perception ability by prioritizing key features and suppressing irrelevant features.

[0105] S5: Encodes using the JSCC encoder;

[0106] The semantic importance feature vector s′ and the channel feedback signal-to-noise ratio μ are encoded using a JSCC encoder to obtain the complex-valued channel input symbol vector. Where k represents the size of the channel input symbols. The encoding process for representing the set of complex numbers can be represented by the following formula:

[0107] s′=M β (T α (s),μ)

[0108] z = M β (E θ (s′),μ)

[0109] Where T α (·) represents a semantic feature extraction network, whose network parameters are denoted as α; M β (·) denotes the semantic importance estimator, whose network parameters are denoted as β; E θ (·) represents the encoding function of the corresponding JSCC encoder, and its network parameters are represented as θ.

[0110] In general, the transmitter maps the n-dimensional vector *s* of the real-valued image to the k-dimensional vector *z* of the channel input samples. To satisfy the average power constraint of the JSCC encoder, the vector *z* is power normalized. Where z* represents the complex conjugate of z.

[0111] S6: The JSCC decoder performs decoding;

[0112] Vector z is transmitted to the receiver via a physical channel, and the channel output signal received by the JSCC decoder is... It can be expressed using the following formula:

[0113]

[0114] Where vector From the distribution Composed of independent and identically distributed samples; σ 2 Indicates noise power. Let I represent a circularly symmetric complex Gaussian distribution, and let I represent the identity matrix.

[0115] The JSCC decoder uses the decoding function R. ξ(·) is used to map the channel output signal The channel feedback signal-to-noise ratio μ,ξ represents the set of parameters for the decoding function.

[0116] S7: Reconstruct the target image using the feature reconstruction module;

[0117] The feature reconstruction module uses the reconstruction function R. η (·) is used to reconstruct the target image, where η represents the reconstruction function R. η The parameter set of (·); the target image reconstructed by the receiver is shown below:

[0118]

[0119] in This indicates the channel output signal received by the JSCC decoder. This represents the estimate of the original image s.

[0120] Among them, in calculating the estimate of the original image s The decoding function and reconstruction function are combined with the semantic importance evaluation module to perform joint calculation. Specifically, a convolutional neural network, which is the opposite of steps S5 and S3, can be used to evaluate semantic importance in step S4 for joint calculation.

[0121] S8: Construct the loss function;

[0122] Original images and reconstructed images The mean square error (MSE) distribution between them is shown below:

[0123]

[0124] Where s i and Representing the original image s and the reconstructed image respectively. The intensity of the color component of each pixel.

[0125] In this invention, a convolutional neural network (CNN) is used to model the JSCC encoder and JSCC decoder. The optimization objective is to find the optimal parameters to minimize the distortion θ. * and ξ * The loss function for a convolutional neural network is shown below:

[0126]

[0127] Where θ * ξ represents the optimal parameters of the JSCC encoder. * This represents the optimal parameters for the JSCC decoder. Represents the original image s and the reconstructed image The joint probability distribution, p(μ) represents the probability distribution of signal-to-noise ratio (SNR).

[0128] This invention utilizes the CIFAR-10 dataset to train and evaluate a proposed end-to-end image semantic transmission system based on a dual attention mechanism. The system is built and trained using Tensorflow and Keras, a high-level application programming interface designed for building and training deep learning models. The Adaptive Moment Estimation (Adam) optimizer is chosen for optimization, with a learning rate of 0.0001, a batch size of 128, and a fixed training period of 1024 to measure training efficiency. System performance is evaluated using PSNR and SSIM. Peak Signal-to-Noise Ratio (PSNR) is a measure of the ratio between the maximum potential power of a signal and the power of harmful noise that affects its representation accuracy, determined by MSE. The average MSE of N images is defined as follows:

[0129]

[0130] Where s (i) and Let represent the i-th original image and its reconstructed target image, respectively. Accordingly, the formula for calculating PSNR is as follows:

[0131]

[0132] MAX represents the maximum pixel value of the image (MAX is 255 if each pixel in the image is 8 bits).

[0133] SSIM is used to quantify the similarity between two digital images. It uses three criteria to measure the image: brightness, contrast, and structure. Two given images s and s The formula for calculating SSIM between them is as follows:

[0134]

[0135] Where μ s and Let σ represent the mean and standard deviation of the original image s, respectively. s and Representing the reconstructed images The mean and standard deviation, Indicates s and The covariance, C1, and C2 are constants used to maintain brightness, contrast, and structural stability. The expression for the signal-to-noise ratio of the proposed system is shown below:

[0136]

[0137] Among them, P s P represents the power of the signal. n This indicates the power of the noise.

[0138] To verify the effectiveness of the proposed SNR adaptive CBAM architecture, the deep learning-based JSCC architecture used in the latest solutions was selected as a benchmark. This invention focuses on SNR... train Training a deep learning-based JSCC scheme at -10dB, 0dB, 10dB, and 20dB. The SNR-adaptive CBAM architecture uses SNR coverage. train The signal-to-noise ratio (SNR) is trained within the range of [-10, 20] dB. The performance of the proposed SNR adaptive CBAM architecture and the DL-based JSCC scheme is then compared in terms of SNR. test The evaluation was performed at [-10, 20] dB. To mitigate the impact of channel noise randomness on the test results, all images in the test dataset were transmitted 10 times on the AWGN channels.

[0139] like Figure 4 The figure shows the comparison results between the proposed JSCC with SNR-adaptive CBAM architecture and the existing DL-based JSCC scheme when the compression ratio is 1 / 3. The proposed JSCC with SNR-adaptive CBAM architecture outperforms the existing DL-based JSCC scheme trained at a specific SNR. test With the increase of [the required value], the peak signal-to-noise ratio (PSNR) of all schemes gradually increases. For existing DL-based JSCC schemes, when testing the signal-to-noise ratio (SNR)... test Its corresponding training signal-to-noise ratio (SNR) train Optimal performance can be obtained when the training signal-to-noise ratio (SNR) is similar. train =Test signal-to-noise ratio (SNR) test At that time, the performance of the JSCC scheme with SNR adaptive CBAM architecture proposed in this invention is still better than the existing DL-based JSCC scheme. If the signal-to-noise ratio (SNR) is tested... testDeviation from training signal-to-noise ratio (SNR) train The advantages of this invention become even more apparent when the compression ratio is maximized. The compression ratio is defined as CR = k / n, where k represents the number of pixels required for the compressed image, and n represents the number of pixels required for the original image. A smaller compression ratio CR indicates less channel resources occupied by the source, but also higher performance requirements for the model.

[0140] This invention also uses structural similarity (SSIM) to evaluate image quality. The results are as follows: Figure 5 The diagram illustrates a performance comparison of different schemes under the SSIM evaluation criterion. When the compression ratio CR = 1 / 3, the JSCC scheme with an SNR adaptive CBAM architecture proposed in this invention exhibits excellent signal-to-noise ratio (SNR) adaptive performance, outperforming existing DL-based JSCC schemes under various SNR conditions. Even when the training SNR... train =Test signal-to-noise ratio (SNR) test At the same time, the JSCC scheme with SNR adaptive CBAM architecture proposed in this invention can still achieve performance comparable to the existing DL-based JSCC scheme.

[0141] Based on the above analysis, it can be seen that the JSCC with SNR adaptive CBAM proposed in this invention still exhibits good PSNR and SSIM performance in poor channel environments (e.g., SNR = 0dB) or low CR (e.g., CR = 1 / 12). By utilizing a dual attention mechanism, this invention demonstrates excellent SNR adaptive characteristics, capable of adapting to different SNR levels based on a single training model. Furthermore, the SNR adaptive CBAM architecture significantly reduces model complexity and the total number of training parameters, effectively alleviating storage pressure and improving image transmission efficiency.

[0142] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. An end-to-end image semantic transmission system based on a dual attention mechanism, characterized in that, include: The normalization layer, transmitter, channel, receiver, and denormalization layer are connected in sequence. The normalization layer is used to preprocess the original input image. The transmitter includes a semantic feature extractor, a JSCC encoder, and a first semantic importance evaluation module. The semantic feature extractor extracts different semantic features from the original image. The first semantic importance evaluation module evaluates the importance of the extracted semantic features and reassigns weights to obtain a semantic importance feature vector. The JSCC encoder encodes the semantic importance feature vector and the channel feedback signal-to-noise ratio to obtain a complex-valued channel input symbol vector. The first semantic importance evaluation module is implemented by an SNR adaptive CBAM module, which is composed of a channel attention mechanism module and a spatial attention mechanism module connected sequentially. The SNR adaptive CBAM module can be integrated into convolutional layers. In the main layers of the neural network, the channel attention mechanism module consists of a first average pooling layer, a first maximum pooling layer, a multilayer perceptron with hidden layers, and a first sigmoid activation function layer; the spatial attention mechanism module consists of a second average pooling layer, a second maximum pooling layer, a connection layer, a convolutional layer, and a second sigmoid activation function layer; the first semantic importance evaluation module introduces channel attention and spatial attention mechanisms. Considering the interdependence of spatial and channel features, it concatenates pooled feature vectors with signal-to-noise ratio (SNR) state information, and sequentially passes them through the channel attention network and the spatial attention network to obtain importance evaluation results. Different weights are assigned to different features based on different SNR levels and image content. A channel is used to connect a transmitter and a receiver to transmit images; The receiver includes a JSCC decoder, a feature reconstruction module, and a second semantic importance assessment module. The JSCC decoder and the second semantic importance assessment module work together to reduce noise interference in the additive white Gaussian noise channel and recover semantic features. The feature reconstruction module is used to deeply mine semantic feature information through an attention mechanism, fuse semantic features and reconstruct the target image; The denormalization layer scales each element before outputting the target image.

2. The image semantic end-to-end transmission system based on a dual attention mechanism according to claim 1, characterized in that, The semantic feature extractor consists of two generated feature blocks, each connected to an SNR adaptive CBAM module, which assigns feature weights based on image content and channel information. Each generated feature block includes a second convolutional layer, a generalized division normalization layer, and a parameter-corrected linear unit layer. The image features generated by the generated feature blocks include edges and textures.

3. The image semantic end-to-end transmission system based on a dual attention mechanism according to claim 1, characterized in that, The JSCC encoder consists of two residual convolutional blocks, a first convolutional layer, and a generalized division normalization layer. Each residual convolutional block is followed by an SNR adaptive CBAM module, which assigns feature weights based on image content and channel information. Each residual convolutional block employs a pre-activated residual neural network, which includes a third convolutional layer, a fourth convolutional layer, two generalized division normalization layers, and two parameter-corrected linear unit layers.

4. The image semantic end-to-end transmission system based on a dual attention mechanism according to claim 1, characterized in that, The structure of the second semantic importance assessment module is the same as that of the first semantic importance assessment module.

5. An end-to-end image semantic transmission method based on a dual attention mechanism, characterized in that, The image semantic end-to-end transmission system based on a dual attention mechanism, as described in any one of claims 1-4, is implemented using the following steps: S1: The input original image is composed of vectors express, Let represent the set of real numbers, and the size of the original input image be . , The height of the image. The width of the image. The number of channels in the image; S2: Preprocess the original input image using a normalization layer; S3: Use a semantic feature extractor to extract different semantic features from the original image; S4: The importance of the extracted semantic features is evaluated by the first semantic importance evaluation module, and the weights are reassigned to obtain the semantic importance feature vector. ; For feature maps obtained from semantic feature extractors The SNR adaptive CBAM module generates a one-dimensional channel attention map based on the SNR level. and 2D spatial attention map Then, it is multiplied with the input feature map to generate a new feature map. Used for joint encoding and decoding: ; in This represents element-wise multiplication; throughout the multiplication process, attention weights are broadcast: channel attention weights are broadcast in the spatial dimension, and spatial attention weights are broadcast in the channel dimension. S5: Employ the JSCC encoder to process semantically important feature vectors. and channel feedback signal-to-noise ratio Encoding is performed to obtain the complex-valued channel input symbol vector. , Indicates the size of the channel input symbols. Representing the set of complex numbers: , ,in This represents a semantic feature extraction network, whose network parameters are expressed as follows: ; The semantic importance evaluator is represented by the network parameters as follows: ; This represents the encoding function of the corresponding JSCC encoder, and its network parameters are expressed as follows: For vectors Perform power normalization ,in express The complex conjugate; S6: Vector The channel output signal is transmitted to the receiver via the channel and received by the JSCC decoder. Represented as: , where vector From the distribution Composed of independent and identically distributed samples; Indicates noise power. This represents a circularly symmetric complex Gaussian distribution. Represents the identity matrix; the JSCC decoder uses the decoding function. To map the channel output signal and channel feedback signal-to-noise ratio , Indicates the parameter set of the decoding function; S7: The feature reconstruction module uses reconstruction functions. To reconstruct the target image, Represents the reconstruction function Parameter set: ,in This indicates the channel output signal received by the JSCC decoder. Represents the original image The estimate; S8: Construct the loss function; Original image and reconstructed images The mean square error distribution between them is as follows: ,in and These represent the original images. and reconstructed images The color component intensity of each pixel in the image is determined; the JSCC encoder and decoder are modeled using a convolutional neural network, with the optimization objective being to find the optimal parameters to minimize distortion. and The loss function for a convolutional neural network is: ,in This represents the optimal parameters for the JSCC encoder. This represents the optimal parameters for the JSCC decoder. Represents the original image and reconstructed images The joint probability distribution, This represents the probability distribution of the signal-to-noise ratio.

6. The image semantic end-to-end transmission method based on a dual attention mechanism according to claim 5, characterized in that, The specific implementation steps of the SNR adaptive CBAM module in generating a one-dimensional channel attention map based on the SNR level are as follows: The channel attention mechanism module uses average pooling and max pooling operations to merge spatial information in the feature maps and obtain average pooling features. and max pooling features ; These characteristics are related to the channel feedback signal-to-noise ratio. Connect to generate context information and : , After reshaping the contextual information, the features are input into the same multilayer perceptron with hidden layers to generate a one-dimensional channel attention map. The hidden activation size is set to... , It is the reduction ratio; element-wise summation is used to merge output features to generate a channel attention map. : ,in and Representing the rectified linear unit and the sigmoid activation function, respectively; the two inputs of the network share MLP weight parameters, respectively. and ,in , .

7. The image semantic end-to-end transmission method based on a dual attention mechanism according to claim 6, characterized in that, The specific steps for obtaining the two-dimensional spatial attention map are as follows: The one-dimensional channel attention map is used as the input feature map for the spatial attention mechanism module; average pooling and max pooling operations are applied along the channel axis to aggregate the channel information of the feature map, generating two two-dimensional feature maps. and After concatenating these two 2D feature maps to generate an effective feature descriptor, convolutional layers and sigmoid layers are applied to reduce the dimensionality and generate a 2D spatial attention map. : ,in This represents the Sigmoid activation function. Indicates the kernel size as The convolution operation.

Citation Information

Patent Citations

  • Coding and decoding structure semantic segmentation model based on position attention mechanism

    CN115908793A

  • Methods, apparatuses, and computer program products using a repeated convolution-based attention module for improved neural network implementations

    US20210064955A1