Compact enhanced deep source channel joint coding

By using the CEJSCC method and the CETL backbone module and SA module to process channel state information, the problems of high model complexity and redundancy in the Swin JSCC architecture are solved, achieving high-performance and low-latency image reconstruction, which is suitable for resource-constrained scenarios such as edge IoT.

CN121728253APending Publication Date: 2026-03-24CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-17
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

While maintaining high performance, the existing Swin JSCC architecture suffers from high model complexity and redundancy in channel adaptation modules, making it difficult to meet the deployment needs of resource-constrained scenarios such as edge IoT.

Method used

The Compact Enhanced Deep Source-Channel Joint Coding (CEJSCC) method is adopted. By constructing a CETL backbone module and an SA Module channel adaptation module, and combining feature sampling, compact self-attention and local enhancement mechanisms, the computational complexity is reduced and the model's anti-interference ability is enhanced. At the same time, structured SNR prior injection and cross-channel attention modulation are used to process channel state information.

Benefits of technology

While reducing model parameters and latency, it significantly improves image reconstruction performance, adapts to dynamic channel conditions, and meets the deployment requirements of resource-constrained scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121728253A_ABST
    Figure CN121728253A_ABST
Patent Text Reader

Abstract

The invention discloses a compact enhanced deep source channel joint coding (CEJSCC) method, and belongs to the field of source channel coding. The method comprises the following steps of: constructing an encoder-decoder architecture by taking a compact enhanced transform layer (CETL) as a core, and retaining global context and local details of an image while reducing the calculation complexity through three-stage operation of'feature sampling-compact self-attention-local enhancement 'of the CESA; a channel adaptive module (SA Module) based on sequence shuffling attention (SSA) is embedded in a deep stage of a codec, and dynamic channel adaptation is realized through double operations of structured SNR prior injection-cross-channel attention modulation. In order to evaluate the performance of the CEJSCC provided by the invention, Kodak24 and CLIC2021 data sets are adopted to carry out performance test on AWGN and Rayleigh channels, the result is transparent, the provided CEJSCC is obviously superior to Swin JSCC and ADJSCC in performance indexes of PSNR and MS-SSIM, meanwhile, the model parameter quantity is reduced by 56.70% compared with the Swin JSCC, the end-to-end delay is reduced by 32.66%, and the CEJSCC has the advantages of light weight and high robustness. The method is suitable for 6G low-delay communication, edge Internet of Things and other resource-limited scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of source-channel coding, specifically relating to a compact enhanced deep source-channel joint coding method (CEJSCC) for image sources. It is suitable for resource-constrained scenarios such as 6G low-latency high-reliability communication and edge IoT transmission, and can realize low-complexity encoding, decoding and transmission of high-fidelity images under dynamic channels. Background Technology

[0002] In traditional communication systems, source coding and channel coding are typically designed and implemented as two independent modules. Source coding (such as JPEG) is only responsible for compressing redundant information; channel coding (such as LDPC) is only responsible for adding anti-interference redundancy. Although this separate design performs well in many scenarios, this architecture has significant drawbacks in core 6G scenarios (such as low-latency, high-reliability communication and edge IoT transmission). First, there is a "cliff effect" (performance drops sharply when channel conditions are below a threshold); second, the cascading of multiple modules leads to high latency, making it difficult to adapt to dynamic channels; and third, the superposition of redundancy reduces transmission efficiency.

[0003] In recent years, Deep Learning-based Joint Source-Channel Coding (Deep JSCC) has broken the constraints of traditional separate design. This method leverages the powerful feature extraction and mapping capabilities of neural networks to integrate source coding, channel coding, and modulation / demodulation processes into an end-to-end trainable framework, achieving global joint optimization from source to channel. Deep JSCC not only adaptively learns the matching relationship between source semantic features and channel conditions, approaching the Shannon limit with finite code lengths, but also exhibits smooth performance degradation as the channel deteriorates, significantly outperforming traditional separate schemes.

[0004] Reference 1, “Xu J, Ai B, Chen W, et al. Wireless image transmission using deep source channel coding with attention modules[J]. IEEE Transactions on Circuits and Systems for Video Technology, 2021, 32(4): 2315-2328,” proposes a Deep Joint Source Channel Coding (ADJSCC) architecture based on an attention mechanism. Compared to Deep JSCC, ADJSCC introduces a channel-level soft attention mechanism to dynamically adjust feature weights based on real-time channel SNR, thereby enhancing the system's channel adaptability and robustness.

[0005] Reference 2, “Yang K, Wang S, Dai J, et al. Swinjscc: Taming swin transformer for deep joint source-channel coding[J]. IEEE Transactions on Cognitive Communications and Networking, 2024,” proposes the SwinJSCC architecture based on the Swin Transformer. Using the Swin Transformer as the backbone network, it leverages its hierarchical feature extraction capabilities and linear computational complexity with image size to address the problem of insufficient global information capture in CNNs, enhancing the model's ability to learn details in high-resolution images. Simultaneously, a Channel ModNet and a Rate ModNet are designed to scale the latent representation based on channel state information and the target transmission rate, improving the model's ability to adapt to different channel conditions and rate configurations.

[0006] While Swin JSCC achieves significant performance improvements, it still fails to resolve the challenge of balancing performance and complexity. On one hand, the number of parameters in its backbone network, Swin Transformer, is significantly larger than that of convolutional network-based joint coding methods, resulting in high computational overhead. On the other hand, the number of parameters in Channel ModNet is also significantly larger than that of the channel adaptation module (AF Module) in ADJSCC, highlighting module redundancy. This leads to high overall model storage overhead for Swin JSCC, making it unsuitable for deployment in resource-constrained scenarios such as edge IoT. Therefore, there is an urgent need for a deep joint source-channel coding method that can significantly reduce model complexity and redundancy while maintaining high performance. Summary of the Invention

[0007] The core objective of this invention is to address the dual problems of "high complexity of the backbone model" and "redundancy of the channel adaptation module" in the existing SwinJSCC architecture, and to provide a compact and enhanced deep source-channel joint coding method (CEJSCC) that achieves lightweight model and low latency while maintaining or even improving image reconstruction performance.

[0008] To achieve the above objectives, the technical solution of the present invention specifically includes the following steps:

[0009] Step 1: Build the encoder-decoder architecture of the overall CEJSCC model. The encoder and decoder each contain 3 progressive stages. Through the encoder's downsampling (to achieve feature compression and channel enhancement) and the decoder's upsampling (to achieve feature recovery and channel matching), multi-scale feature extraction and image reconstruction are completed.

[0010] Step 2: Construct an improved CETL backbone module and embed it between the various stages of the codec. Through the synergistic effect of the three mechanisms of "feature sampling - compact self-attention - local enhancement", the computational complexity is reduced while preserving the global context and local details of the image (such as edges and textures).

[0011] Step 3: Construct the SA Module channel adaptation module and embed it deep in the codec stage. Using SNR as the channel state information (CSI), the module achieves dynamic channel adaptation with extremely low parameter overhead through the dual operation of "structured SNR prior injection - cross-channel attention modulation", thereby enhancing the model's anti-interference capability.

[0012] Furthermore, in step 1, the encoder comprises three stages, with an input image size of (where is the image height, is the width, and 3 is the number of RGB channels); each stage contains a downsampling layer, and after the three stages, the image size is sequentially downsampled to . , , Finally, a 1×1 convolutional layer is used to adjust the number of channels so that the output feature dimension matches the target channel bandwidth ratio (CBR), generating channel symbols adapted for wireless transmission.

[0013] The decoder is structurally symmetrical to the encoder. After receiving symbols transmitted through the channel, it first adjusts the symbol dimensions using a 1×1 convolutional layer. Aligned with the encoder output dimension;

[0014] Then, similarly, the image is restored to its original size through three stages, each containing an upsampling layer. After these three stages, the image size is restored sequentially. , , The final output is a reconstructed image.

[0015] Furthermore, in step 2, CETL is a Transformer structure with "Compact Enhanced Self-Attention (CESA) + Gated Deep Feedforward Network (GDFN)" as its core.

[0016] CESA is the core attention unit of CETL, achieving efficient feature interaction through three steps:

[0017] Feature sampling: The input token sequence is sampled using an average pooling layer. Downsampling reduces spatial resolution to decrease the computational overhead of subsequent attention modules, while expanding the model's receptive field and enhancing its ability to capture global information.

[0018] Compact self-attention: First, the output features of the feature sampling layer are... Divided along the channel dimension and The system consists of two parts; subsequently, convolution and reshaping operations are applied to each part to generate query (Q), key (K), and value (V) vectors. The value vectors of the two sets are then swapped, and cross-attention is calculated. The calculation process can be represented as follows:

[0019] ;

[0020] Local enhancement: After the attention output, a 2×2 deconvolutional layer and a 1×1 convolutional layer are introduced to reconstruct and enhance high-frequency information (such as edges and textures), improving the quality of feature representation. The calculation process can be represented as follows:

[0021] ;

[0022] GDFN is a structural improvement on the traditional Feed-Forward Network (FFN) in the Transformer. It first extracts complementary features through two parallel convolutional paths:

[0023] Activation of Feature Path: First, a 1×1 convolutional layer is used to fuse multi-channel information from local pixels. Then, a 3×3 depthwise separable convolution is used to capture local structures (such as edges and textures) within a 3×3 area around the pixel channel by channel, overcoming the limitation of traditional FFN which only focuses on the channel dimension and ignores the spatial dimension. Finally, the GELU activation function is used to perform a non-linear transformation on the convolution output, enhancing the feature representation capability while avoiding the "dead neuron" problem of ReLU. Output:

[0024] ;

[0025] Basic Feature Path: The structure is basically the same as the activation feature path, except that GELU activation is disabled. Basic features are extracted through convolution, and the output is:

[0026] ;

[0027] The outputs of two parallel paths are fused through element-wise multiplication to generate a "gated signal" for filtering useful features. The process can be represented as:

[0028] ;

[0029] The gated signal is then passed through a 1×1 convolutional layer to adjust the channel dimension to match the input dimension of the subsequent residual connections.

[0030] Furthermore, in step 3, the design of an efficient channel adaptation module (SA Module) based on SSA mainly includes two stages:

[0031] Structured SNR Prior Injection: First, the input SNR (dB value) is injected according to the formula:

[0032] Convert to a linear scale SNR_linear;

[0033] Then take the reciprocal of SNR_linear to generate the SNR prior vector. ; will the Compared with the original input features Subtracting element by element yields the SNR perceptual features. ;

[0034] When the channel quality is good (high SNR), it approaches This can effectively avoid introducing redundant information;

[0035] When the channel quality is poor (low SNR). It can effectively characterize the attenuation effect of channel degradation on signal characteristics.

[0036] Cross-channel attention modulation: The original input features and the above-mentioned SNR-aware features are input into the SSA module.

[0037] The SSA module aggregates sequence information through a process of "global pooling - sequence shuffling - block processing - attention weight calculation - weighted summation," efficiently capturing complex dependencies across sequences and deeply fusing original input features with channel prior information derived from SNR. By fully leveraging the complementarity of these two elements, it significantly enhances the output features' ability to represent the channel state. The formula is as follows:

[0038] .

[0039] Next, the SA Module was embedded in the CEJSCC network. Two SA Modules were added to the encoder, one after the second stage and the other after the third stage. In the decoder, two SA Modules were used symmetrically with the encoder, one before the input of the first stage and the other before the input of the second stage.

[0040] The beneficial effects of this invention are:

[0041] This invention uses CETL as the main component of the joint codec, achieving a balance between "reducing complexity" and "preserving details." It obtains SNR-aware features through structured processing of Channel State Information (CSI), effectively characterizing the impact of channel state changes. Furthermore, it utilizes SSA to process the original input features and SNR-aware features, significantly enhancing the output features' ability to represent the channel state. Ultimately, compared to Swin JSCC, this invention significantly reduces model parameters, substantially decreases end-to-end latency, and improves encoding and decoding performance. Attached Figure Description

[0042] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0043] Figure 1 This is an overall structural diagram of the CEJSCC proposed in this invention;

[0044] Figure 2 This is a schematic diagram of the CETL structure used in this invention;

[0045] Figure 3 This is a schematic diagram of the SSA module used in this invention and the proposed SA Module.

[0046] Figure 4 This is a performance comparison chart of the method proposed in this invention with other methods at CBR=1 / 6;

[0047] Figure 5 This is a performance comparison chart of the method proposed in this invention with other methods at CBR=1 / 16; Detailed Implementation

[0048] To further clarify the purpose, technical solution, and implementation effects of the invention, the present invention will be described in further detail below with reference to the accompanying drawings and embodiments, but the scope of protection of the present invention is not limited thereto.

[0049] This invention relates to a compact and enhanced deep source-channel joint coding method. A specific implementation of this invention is given below, with the following parameter settings:

[0050] During the training of CEJSCC, the training dataset used was the DIV2K dataset, which contained 800 high-quality training images and 100 validation images. During training, the images were randomly cropped into pixel patches. The mean squared error (MSE) was selected as the training loss function, AdamW was selected as the training optimizer, the training learning rate was set to a certain value, and the batch size was set to 6.

[0051] The training used AWGN and Rayleigh channels, and training was conducted in scenarios with channel bandwidth ratios of CBR=1 / 6 and CBR=1 / 16.

[0052] The training process is divided into two phases:

[0053] Phase 1: Pre-train the model under a fixed signal-to-noise ratio (13 dB). At this time, turn off the model's SNR adaptive function so that the model can learn the basic representation of the source features.

[0054] Phase 2: Enable SNR adaptive function and randomly sample the signal-to-noise ratio in the range of 0~20 dB to simulate the dynamic channel environment in actual communication and further optimize the model's ability to adapt to channel changes.

[0055] Based on the above parameter settings, the specific steps of this method are as follows:

[0056] Step 1: Load training images from the DIV2K dataset. Each image needs to be standardized, normalizing pixel values ​​from the integer range [0, 255] to the floating-point range [0, 1] to reduce numerical instability during training. The data is then organized in batches of 6 images each, with the order randomly shuffled during training to enhance generalization ability.

[0057] Step 2: Convert the image vector Compared with the current signal-to-noise ratio (Uniform random sampling within the 0~20dB range; sampling is not required if training is under fixed channel ratio conditions) The encoder is input to the common input. The encoder performs the following operations in sequence:

[0058] After the first stage: the input image is adjusted through the downsampling layer to... The feature representation is then enhanced using two layers of CETL.

[0059] After the second stage: the image is adjusted through the downsampling layer... Then, the feature representation is enhanced through 4 layers of CETL, and subsequently compared with the current signal-to-noise ratio. Together, input the SA Module to explicitly learn and fuse channel state information;

[0060] After the third stage, the image is adjusted through the downsampling layer to Then, the feature representation is enhanced through 4 layers of CETL, and subsequently compared with the current signal-to-noise ratio. Together, input the SA Module to explicitly learn and fuse channel state information;

[0061] Finally, a 1×1 convolutional layer is used to adjust the number of channels, outputting the joint encoding vector. Where C is determined by the channel bandwidth ratio CBR. The structures of CEJSCC, CETL, and SA Module are shown in the attached figures. Figure 1 Appendix Figure 2 and attached Figure 3 As shown.

[0062] Step 3: Combine the joint encoding vector Simulate actual wireless communication processes using either the AWGN channel or the Rayleigh channel (depending on the current training channel settings).

[0063] Step 4: Convert the received channel output vector Compared with the current signal-to-noise ratio The input decoder adopts a structure symmetrical to the encoder, and its specific operation is as follows:

[0064] First, a 1×1 convolutional layer is used to adjust the number of channels, outputting the decoder input vector. C is determined by the channel bandwidth ratio CBR;

[0065] After the first stage: first comparing with the current signal-to-noise ratio The SA Module learns and fuses channel state information together; then the feature representation is enhanced through 4 layers of CETL; finally, it is adjusted through an upsampling layer. ;

[0066] After the second stage: the output sequentially passes through the SA Module, four CETL layers, and an upsampling layer. ;

[0067] After the third stage, through two CETL layers and an upsampling layer, the reconstructed image is output. ;

[0068] Step 5: Calculate the reconstructed image With the original image The mean squared error (MSE) between the two values ​​is used as the loss function, and the formula is as follows:

[0069] ,

[0070] Gradient descent is performed using the AdamW optimizer, with a learning rate set to... The gradient is calculated and the model parameters are updated for each batch, gradually optimizing the weights of the encoder and decoder.

[0071] Step 6: After each training round, calculate the Peak Signal-to-Noise Ratio (PSNR) on the validation set to quantitatively evaluate the reconstruction quality, monitor model performance, and save the weights of the best-performing model. PSNR reflects pixel-level distortion; its core idea is to calculate the ratio of the maximum possible power of the signal to the power of the distortion noise, as shown in the formula:

[0072]

[0073] Step 7: After training in the fixed signal-to-noise ratio stage, use the saved optimal model to train the dynamic channel (SNR random sampling), and repeat steps 1-6.

[0074] The final test used the Kodak24 dataset, and the test metrics used were PSNR and Multi-Scale Structural Similarity (MS-SSIM). MS-SSIM is an evaluation metric that considers the structural similarity of images at different scales. It extends the calculation of SSIM to multiple scales, thus providing a more comprehensive assessment of image quality. Its calculation formula is as follows:

[0075]

[0076] in Used to compare brightness; Used for comparing contrast; Used for comparison structures, indices , , Used to adjust the relative importance of different components.

[0077] The end-to-end latency test method is as follows: using the Kodak24 dataset, setting the batch size to 1, repeating the experiment ten times and taking the average value.

[0078] Appendix Figure 4 These are performance curves of PSNR and MS-SSIM for CEJSCC, Swin JSCC, and ADJSCC under AWGN and Rayleigh channels with CBR=1 / 6; (Attached) Figure 5 The graph shows the performance curves at CBR=1 / 16. As can be seen from the graph, whether it is PSNR or MS-SSIM, the CEJSCC proposed in this invention has a significant performance advantage over ADJSCC in Reference 1 and Swin JSCC in Reference 2, and the performance advantage of CEJSCC is more significant at higher CBR (CBR=1 / 6).

[0079] Furthermore, the Swin JSCC model has 18.34 + 9.86M parameters, while the CEJSCC model has only 11.63 + 0.58M parameters, representing a 56.70% reduction in model parameters. Specifically, the SA Module designed in this invention introduces only 5.88% of the parameters of the Channel ModeNet in SwinJSCC, significantly reducing additional overhead. This means that in practical applications, the storage space occupied by the CEJSCC model weight file will be significantly reduced compared to Swin JSCC, substantially lowering storage costs. Simultaneously, in end-to-end latency tests, the end-to-end latency of CEJSCC is 52.87 ms, while that of Swin JSCC is 78.51 ms, representing a 32.66% reduction in end-to-end latency for CEJSCC. This fully demonstrates the effectiveness of the method proposed in this invention.

Claims

1. A compact and enhanced deep source-channel joint coding method, characterized in that, Specifically, the following steps are included: Step 1: Build the CEJSCC encoder-decoder architecture. The encoder and decoder each contain 3 progressive stages. The encoder's downsampling operation realizes feature map scale compression and channel enhancement, and the decoder's upsampling operation realizes feature map scale restoration and channel adaptation. Step 2: Embed a compact enhancement transform layer (CETL) between each stage of the encoder and decoder. The CETL performs three stages of operation in sequence: feature sampling, compact self-attention, and local enhancement, which efficiently enhances the image feature representation at different scales. Step 3: Embed two channel adaptation modules (SA Modules) based on Sequence Shaking Attention (SSA) in the deep stages of the encoder and decoder; the SA Modules take SNR as the channel state information (CSI) input and perform a two-stage operation of "structured SNR prior injection - cross-channel attention modulation" in sequence to dynamically adjust the original input features; Step 4: After the encoder processes the input image, it maps it to a channel symbol adapted for wireless channel transmission; after the channel symbol is transmitted through the AWGN channel or Rayleigh fading channel, the decoder receives the channel symbol and processes it in reverse to reconstruct the source information of the input image.

2. The compact and enhanced deep source-channel joint coding method according to claim 1, characterized in that... In step 1: The encoder input image has a resolution of H×W×3 (3 being the number of RGB channels), and its dimensions after three stages of downsampling are as follows: Then adjust from a single 1×1 convolutional layer to The decoder receives The channel symbol is also first adjusted through a 1×1 convolutional layer to... Then, after three stages of upsampling, the size is restored sequentially. H×W×3, where C is determined by the channel bandwidth ratio CBR.

3. The compact and enhanced deep source-channel joint coding method according to claim 1, characterized in that... In step 2: The CETL described above is a Transformer structure centered around Compact-Enhanced Self-Attention (CESA) and a Gated-Dconv Feed-Forward Network (GDFN). CESA comprises three stages: (1). Feature sampling: The input token sequence is sampled using an average pooling layer. Downsampling is performed to reduce its spatial resolution and decrease the computational overhead of subsequent attention modules. (2). Compact self-attention: This involves sampling the features... Divided along the channel dimension and The system consists of two parts. Subsequently, convolution and reshaping operations are applied to both parts to generate query (Q), key (K), and value (V) vectors. The two sets of value vectors are then swapped, and cross-attention is performed to achieve efficient feature interaction and information complementarity. The computation process can be represented as follows: (3) Local Enhancement: After the attention output, a 2×2 deconvolution layer and a 1×1 convolution layer are introduced to reconstruct and enhance high-frequency information (such as edges and textures), thereby improving the quality of feature representation. The calculation process can be represented as follows: In GDFN, the activation feature path and activation feature path are first processed in parallel. The activation feature path first uses a 1×1 convolutional layer to fuse multi-channel information from local pixels; then a 3×3 depthwise separable convolution is used to capture local structure (such as edges and textures); finally, a GELU activation function is used to perform a non-linear transformation on the convolution output. The basic feature path and activation path are almost identical, except that the GELU activation function is removed. Finally, the outputs of the two parallel paths are fused together through element-wise multiplication to generate a "gated signal" to filter useful features. In addition, the SSA module was used to residual link the N-layer CETL of each stage, recalibrate the feature response by calculating the attention weights of the channel dimension, capture the dependencies between different scanning directions, and improve the robustness of the model to complex channel interference.

4. The compact and enhanced deep source-channel joint coding method according to claim 1, characterized in that... In step 3: The two-stage operation of the SA Module is as follows: (1). Structured SNR prior injection: First, the input SNR (dB value) is converted into a linear scale SNR_linear according to the formula SNR_linear=10^(SNR_dB÷10); the reciprocal of SNR_linear is taken to generate an SNR prior vector; the SNR prior vector is subtracted element by element from the original input feature to obtain the SNR perception feature (approaching the input feature when the SNR is high, and characterizing the channel attenuation effect when the SNR is low). (2). Cross-channel attention modulation: The original input features are concatenated with the SNR-sensing features and then input into the Sequence Shuffle Attention (SSA) module. The SSA module completes the aggregation of sequence information through the steps of "global pooling - sequence shuffle - block processing - attention weight calculation - weighted summation". On the one hand, it captures the complex feature dependencies across sequences, and on the other hand, it deeply integrates the input features with the channel prior information derived from SNR, which significantly enhances the ability of the output features to represent the channel state.