Semantic wireless image transmission method suitable for resource-constrained platform
By designing the MobileJSCC semantic wireless communication network architecture on a resource-constrained platform, combining the encoder structure of deep separable convolution and inverted residual connections, the problem of high computational complexity of existing deep learning models is solved, and efficient semantic wireless image transmission is achieved.
Patent Information
- Application Number
- CN202510002769.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-13
AI Technical Summary
Existing deep learning models are difficult to efficiently deploy on resource-constrained platforms, and cannot effectively reduce the computational complexity to support semantic wireless image transmission.
Using the MobileJSCC semantic wireless communication network architecture, the encoder is designed through a structural combination of deep-separable convolution and inverted residual connection, combining additive Gaussian white noise channel and Rayleigh fading channel model to simulate wireless channel transmission, and image features are restored by layer-by-layer decoding through deep-separable transposed convolution.
It significantly reduces the computational complexity and number of parameters of the neural network model in resource-constrained environments, while ensuring image transmission quality and achieving efficient semantic recovery.
Smart Images

Figure CN119996579A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of semantic communication, and in particular to a semantic wireless image transmission method applicable to a resource-constrained platform. Background Art
[0002] With the rapid advancement of information technology and artificial intelligence, the limitations of traditional communication systems have gradually become apparent. Traditional systems usually only focus on the accurate transmission of data, while ignoring the semantic information of the data, that is, the deeper meaning and intention carried by the data. This traditional method not only transmits redundant data, but also increases the communication burden and cannot fully support the needs of intelligent interaction. The introduction of semantic communication has changed this situation, allowing both parties in communication to not only transmit data, but also understand the semantic content of the data, thereby achieving more efficient and intelligent interaction.
[0003] As a key technology for semantic communication, Joint Source Channel Coding (JSCC) integrates semantic perception into the data transmission process through the feature extraction and pattern recognition capabilities of deep learning models. JSCC methods, such as DeepJSCC, Adaptive Channel JSCC, and JSCC based on SwinTransformer architecture, use end-to-end training strategies to optimize the synergy between encoders and decoders, thereby improving the overall efficiency of the system. JSCC not only transmits information at the data bit level, but also ensures the semantic accuracy of the transmitted data, including the expression of intent and visual effects.
[0004] The implementation of JSCC is highly dependent on deep learning models, and these complex neural networks usually require a lot of computing resources. This poses a huge challenge to resource-constrained platforms such as mobile devices and embedded systems, and traditional deep learning models are difficult to deploy efficiently on these devices. Summary of the invention
[0005] The purpose of the present invention is to provide a semantic wireless image transmission method suitable for resource-constrained platforms, aiming to optimize the neural network structure, significantly reduce the computational complexity in a resource-constrained environment, and ensure the image transmission quality.
[0006] To achieve the above object, the present invention provides a semantic wireless image transmission method applicable to a resource-constrained platform, comprising the following steps:
[0007] Step 1: Construct a MobileJSCC semantic wireless communication network architecture, wherein the MobileJSCC semantic wireless communication network architecture includes a MobileJSCC encoder, a wireless channel and a MobileJSCC decoder;
[0008] Step 2: Convert the image into feature representation using a lightweight MobileJSCC encoder;
[0009] Step 3: Transmit the feature representation through the wireless channel;
[0010] Step 4: Use the lightweight MobileJSCC decoder to decode the received feature representation and restore the original image.
[0011] Optionally, the lightweight MobileJSCC encoder in step 2 includes a standard convolutional layer, two depth-wise separable convolutional layers, and two inverted residual depth-wise separable convolutional layers.
[0012] Optionally, in step 3, the wireless channel uses an additive white Gaussian noise channel and a Rayleigh fading channel model to simulate wireless channel transmission.
[0013] Optionally, the lightweight MobileJSCC decoder in step 4 includes five depth-wise separable transposed convolutional layers.
[0014] Optionally, both the MobileJSCC encoder and the MobileJSCC decoder adopt end-to-end training to optimize model parameters by minimizing the mean square error between the original image and the reconstructed image.
[0015] The present invention provides a semantic wireless image transmission method suitable for resource-constrained platforms. By establishing a MobileJSCC semantic wireless communication network architecture, a structure combining deep separable convolution and inverted residual connection is used in the encoder part to reduce the computational complexity. At the same time, an additive Gaussian white noise channel and a Rayleigh fading channel model are used to simulate wireless channel transmission. In the decoder part, deep separable transposed convolution is used to decode and restore image features layer by layer to achieve reconstruction of the original image. The encoder and decoder of the present invention adopt a joint optimization training strategy to optimize the weights and bias parameters of the encoding and decoding modules by minimizing the mean square error between the input image and the reconstructed image to achieve efficient semantic restoration. It has been verified that the present invention greatly reduces the amount of calculation and parameters required for the neural network model under different image resolutions, signal-to-noise ratios and channel conditions. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0017] Figure 1 It is a schematic flow chart of the steps of a semantic wireless image transmission method applicable to a resource-constrained platform of the present invention.
[0018] Figure 2 It is a schematic diagram of the MobileJSCC semantic wireless communication network architecture of a specific embodiment of the present invention.
[0019] Figure 3 It is a structural diagram of a MobileJSCC encoder and a MobileJSCC decoder according to a specific embodiment of the present invention.
[0020] Figure 3 (a) is a schematic diagram of the neural network structure of the MobileJSCC encoder of a specific embodiment of the present invention.
[0021] Figure 3 (b) is a schematic diagram of the neural network structure of the MobileJSCC decoder of a specific embodiment of the present invention.
[0022] Figure 4 It is a schematic diagram of the structure of a specific convolutional layer used in MobileJSCC of a specific embodiment of the present invention.
[0023] Figure 4 (a) is a schematic diagram of a specific convolutional layer DSC structure of a specific embodiment of the present invention.
[0024] Figure 4 (b) is a schematic diagram of the specific convolutional layer ResDSC structure of a specific embodiment of the present invention.
[0025] Figure 4 (c) is a schematic diagram of the DSTC structure of a specific convolutional layer in a specific embodiment of the present invention.
[0026] Figure 5 The figure is a schematic diagram of the peak signal-to-noise ratio (PSNR) performance comparison between MobileJSCC and DeepJSCC in AWGN and Rayleigh fading channels in a specific embodiment of the present invention.
[0027] Figure 6 This is a schematic diagram of the peak signal-to-noise ratio (PSNR) performance comparison between MobileJSCC and DeepJSCC at different bandwidth compression ratios in a specific embodiment of the present invention. DETAILED DESCRIPTION
[0028] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0029] See also Figure 1 The present invention provides a semantic wireless image transmission method applicable to a resource-constrained platform, comprising the following steps:
[0030] S1: constructing a MobileJSCC semantic wireless communication network architecture, wherein the MobileJSCC semantic wireless communication network architecture includes a MobileJSCC encoder, a wireless channel and a MobileJSCC decoder;
[0031] S2: Convert the image into feature representation using a lightweight MobileJSCC encoder;
[0032] S3: Transmit feature representation via wireless channel;
[0033] S4: Use the lightweight MobileJSCC decoder to decode the received feature representation and restore the original image.
[0034] Specifically, the MobileJSCC encoder part adopts a structure that combines depthwise separable convolution (DSC) and DSC with inverted residual connection (Residual DSC, ResDSC) to reduce computational complexity. By decomposing the standard convolution into spatial convolution and channel convolution, the DSC structure can reduce the amount of computation while maintaining the image feature expression capability. The ResDSC structure can learn deep and complex feature information and enhance the model's expression capability without increasing the difficulty of training. Through the combination of multiple DSCs and ResDSCs, the MobileJSCC encoder can extract semantically rich image features with limited computing resources and perform effective compression, thereby adapting to wireless communication environments with limited bandwidth.
[0035] At the same time, in view of the characteristics of wireless channels in practical applications, the wireless channels of the present invention include additive white Gaussian noise (AWGN) channels and Rayleigh fading channels. By simulating the effects of multipath and noise interference in the wireless transmission process, the model has strong robustness under real channel conditions.
[0036] Furthermore, the MobileJSCC decoder uses Depthwise Separable Transposed Convolution (DSTC) to decode and restore image features layer by layer to achieve reconstruction of the original image. The design of DSTC not only effectively reduces the computational complexity, but also ensures the accuracy of image reconstruction under wireless transmission conditions.
[0037] In addition, to ensure the end-to-end image transmission quality, the MobileJSCC encoder and decoder in the present invention adopt a joint optimization training strategy. By minimizing the mean square error (MSE) between the input image and the reconstructed image, the weights and bias parameters of the encoding and decoding modules are optimized to achieve efficient semantic recovery. This joint optimization mechanism enhances the synergy between the encoding and decoding modules, reduces the loss of information at the bit level, ensures that the semantic information of the image is retained during the transmission process, and improves the transmission efficiency.
[0038] The present invention is further described in conjunction with specific embodiments. Figures 2 to 6 :
[0039] like Figure 2 The figure shows the MobileJSCC semantic wireless communication network architecture constructed in this embodiment. The system includes a MobileJSCC encoder, a wireless channel and a MobileJSCC decoder.
[0040] The MobileJSCC encoder receives the input image x, extracts features from it and encodes it through the designed deep learning model. The encoder’s task is to convert the input image into compact image features y to reduce the transmission bandwidth requirements while retaining the key semantic information of the image.
[0041] The wireless channel is used to transmit the encoded image features y. The channel model can include various channel conditions such as AWGN or Rayleigh fading. The features may be affected by noise or signal attenuation in the channel. The affected image features obtained by the receiver are expressed as
[0042] The MobileJSCC decoder receives the It is restored by the decoder and the restored image is represented as It is called image estimation. The task of the decoder is to restore (reconstruct) an image that is semantically consistent with the original input x as much as possible.
[0043] Combination Figure 3 As shown in (a), the image input to the MobileJSCC encoder is Where H and W represent the height and width of the image, respectively, and 3 represents the number of color channels (usually red, green, and blue). represents the set of all real numbers, which means x is a real tensor. x first passes through a standard convolutional layer (Conv) to capture the initial features. Then, it passes through a depth-wise separable convolutional layer, which separates the spatial and channel operations, making it more efficient than a standard convolutional layer. Next, the DSC layer is followed by two DSC layers with inverted residual connections, called ResDSC layers, which help learn deeper and more complex features without increasing the difficulty of training. Finally, it ends with a DSC layer to produce a compressed feature map in and represent the height and width of the compressed image, respectively, and Represents the number of channels in the compressed image representation. It is worth noting that each convolutional layer in the MobileJSCC encoder is followed by a nonlinear ReLU activation function. The role of these convolutional layers is to extract meaningful features from the input image, while the nonlinear activation function ReLU promotes the learning of nonlinear mapping from the source signal space to the encoded signal space.
[0044] set up represents the bandwidth compression ratio of the semantic wireless communication system, where and DIM x They are the compressed feature maps and the dimension of the input image x. And DIM x =H×W×3. and The size of is usually fixed after convolution. The present invention can be achieved by adjusting the size of the compressed image. To control the bandwidth compression ratio α. Then, the compressed feature map Through a normalization layer, the output image features of the MobileJSCC encoder are obtained in express The conjugate transpose of represents the square root operation, P is the average transmission power, and the image feature y satisfies the average power constraint where y * represents the conjugate transpose of y, and E represents the expectation operation.
[0045] like Figure 3 As shown in (b), the MobileJSCC decoder uses five depth-wise separable transposed convolutional layers (DSTC) to reverse the operations performed by the MobileJSCC encoder and transform the affected image features into The image estimate is mapped back to the original image x
[0046] The MobileJSCC encoder and decoder are trained together as a whole, rather than separately. This design allows for better synergy between the various parts, thereby optimizing the overall performance of the task. The average distortion between and is used to obtain the optimal parameters of MobileJSCC:
[0047]
[0048] Where θ and ω represent the weights and biases of the proposed MobileJSCC network, MSE (Mean Squared Error) is the mean square error, and Where M is the number of samples and m is the mth sample.
[0049] The DSC layer in the MobileJSCC encoder is composed of Figure 4 (a) shows the decomposition of Conv in the DeepJSCC encoder, which consists of two main parts: deep convolutional layer and point convolutional layer. The deep convolutional layer applies an independent L×L convolution kernel to each input channel, while the point convolutional layer uses a 1×1 convolution kernel to mix features between different channels. Note that BN (BatchNormalization) stands for batch normalization, which is used to speed up the training process of neural networks and reduce the impact of internal covariate shift. Since each input channel passes through a separate L×L convolution kernel, the computational complexity of the deep convolutional layer is H in ×W in ×C in ×L 2 , where H in and W in Respectively represent the height and width of the input feature map, C in is the number of input channels. Obviously, the computational complexity of the depthwise convolutional layer is significantly lower than that of the Conv layer due to the reduced interaction between channels. On the other hand, the computational complexity of the pointwise convolutional layer is H out ×W out ×C in ×C out , where H out and W out Respectively represent the height and width of the output feature map, C out is the number of output channels. Therefore, the computational complexity of the entire DSC layer is:
[0050] Complexity DSC =H in ×W in ×C in ×L2 +H out ×W out ×C in ×C out (2)
[0051] When H out ×W out ≈H in ×W in Compared with the computational complexity of the Conv layer in the DeepJSCC encoder, Conv =H in ×W in ×C in ×C out ×L 2 , the percentage reduction of the computational complexity of the DSC layer in the MobileJSCC encoder is:
[0052]
[0053] In the MobileJSCC encoder, since the DSC layer uses a 5×5 convolution kernel, if the number of output channels is 20, the computational complexity of DSC can be reduced by up to 91% compared to Conv.
[0054] The ResDSC layer in the MobileJSCC encoder is composed of Figure 4 As shown in (b), it consists of three main parts: 1) Expansion layer: a 1×1 point convolution to increase the number of channels. 2) Depth convolution layer: an L×L depth convolution to perform spatial filtering. 3) Reduction layer: another 1×1 point convolution to reduce the number of channels to the original or desired number.
[0055] Assuming β is the expansion factor of ResDSC, the computational complexity of the expansion layer, depth convolution layer and reduction layer of ResDSC are H in ×W in ×C in ×β、H in ×W in ×β×C in ×L 2 and H in ×W in ×β×C in ×C out Therefore, the total computational complexity of the ResDSC layer is:
[0056] Complexity ResDSC =H in ×W in ×β×C in ×(1+L 2 +Cout ) (4)
[0057] Compared with Conv, the percentage of computational complexity reduction of ResDSC is:
[0058]
[0059] In the MobileJSCC encoder, ResDSC uses a 5×5 convolution kernel and an expansion factor C out =32, if β=6, then the computational complexity of ResDSC is reduced by about 56.5% compared to Conv. If β=1, the number of output channels after the expansion layer is the same as the number of input channels, that is, the expansion layer does not actually perform any operation and does not contribute any computational complexity to the overall calculation. In this case, the computational complexity of ResDSC can be reduced by 92.875% compared to Conv.
[0060] The DSTC layer in the MobileJSCC decoder is composed of Figure 4 As shown in (c), it consists of two parts: a depth-wise transposed convolution layer that performs spatial upsampling on each input channel separately, and a point-wise transposed convolution layer that merges the upsampled features across channels. The computational complexity of the depth-wise transposed convolution and point-wise transposed convolution of DSTC is δ 2 ×H in ×W in ×C in ×L 2 and δ 2 ×H in ×W in ×C in ×C out , where δ represents the step size of the transposed convolution. Therefore, the total computational complexity of the DSTC layer is:
[0061] Complexity DSTC =δ 2 ×H in ×W in ×C in ×(L 2 +C out ) (6)
[0062] Compared with the computational complexity of the Trans-Conv layer used in the DeepJSCC decoder, Trans-Conv =δ 2 ×H in ×W in ×C in ×C out ×L 2, the percentage reduction in computational complexity of the DSTC layer used in the MobileJSCC decoder can be expressed as
[0063]
[0064] This is the same percentage reduction as DSC compared to Conv.
[0065] In summary, the special convolutional layers used in MobileJSCC provide a significant reduction in computational complexity, which can lead to faster and more efficient processing on mobile devices. The specific percentage of computational complexity reduction depends on specific parameters, such as the convolution kernel size L, the number of output channels C, and the number of kernels. out , expansion factor β, and the step size δ of the transposed convolution.
[0066] In this embodiment, the present invention uses the CIFAR-10 dataset to train and evaluate the proposed MobileJSCC model. In order to evaluate the quality of the restored image in the context of semantic wireless communication, the present invention uses the peak signal-to-noise ratio (PSNR) as a key indicator. In addition, in order to verify the computational efficiency of MobileJSCC, the present invention compares it with two state-of-the-art JSCC schemes, including DeepJSCC and SwinJSCC, in terms of computational complexity and parameter complexity.
[0067] As shown in Table 1, the parameter complexity (evaluated by Params, i.e. the number of parameters) and computational complexity (evaluated by FLOPs, i.e. the number of floating-point operations per second) of MobileJSCC, DeepJSCC, and SwinJSCC models on the CIFAR-10 dataset are compared. In terms of parameter complexity, MobileJSCC reduces 52.6% and 99.5% of Params compared to DeepJSCC and SwinJSCC, respectively. The Params of SwinJSCC are significantly higher than those of MobileJSCC and DeepJSCC. This is because SwinJSCC adopts the Transformer architecture, which usually involves a large number of parameters to capture complex patterns and dependencies in the data. In terms of computational complexity, MobileJSCC always exhibits the lowest FLOPs, followed by DeepJSCC, while SwinJSCC has the highest FLOPs. This is because MobileJSCC adopts depthwise separable convolutions and inverted residual connections, which significantly reduces computational complexity.
[0068] Table 1 Comparison of MobileJSCC, DeepJSCC, and SwinJSCC model parameters
[0069]
[0070] Figure 5 The PSNR performance comparison of MobileJSCC and DeepJSCC under AWGN and Rayleigh fading channels is shown. For the AWGN channel, MobileJSCC and DeepJSCC show similar PSNR performance at lower signal-to-noise ratios (SNRs), indicating comparable image recovery quality after wireless transmission. However, when the SNR exceeds 20dB, the performance gap between the two widens slightly, with DeepJSCC being approximately 1dB higher than MobileJSCC. Despite this small gap, it is quite acceptable given the significant savings in computational resources achieved by MobileJSCC. For the Rayleigh fading channel, the PSNR of both MobileJSCC and DeepJSCC models decreases compared to that of the AWGN channel. However, at an SNR of 10dB, the PSNR of MobileJSCC is slightly higher than that of DeepJSCC, indicating that MobileJSCC exhibits better robustness under certain channel conditions.
[0071] Figure 6 The PSNR performance comparison of MobileJSCC and DeepJSCC at different bandwidth compression ratios is shown. As the bandwidth compression ratio increases, the PSNR of DeepJSCC is slightly higher than that of MobileJSCC in the AWGN channel with a signal-to-noise ratio of 0dB and the Rayleigh fading channel with a signal-to-noise ratio of 10dB, but the difference is less than 1dB. For the AWGN channel with a signal-to-noise ratio of 10dB, when the bandwidth compression ratio is within a certain range, the PSNR of MobileJSCC is higher than that of DeepJSCC. This shows that within this bandwidth compression ratio range, MobileJSCC is superior to DeepJSCC in terms of image restoration quality.
[0072] It should be noted that in order to address the computational challenges faced by resource-constrained platforms such as mobile and embedded devices, this embodiment introduces a MobileJSCC semantic wireless communication network architecture. Specifically, the MobileJSCC encoder consists of 1 convolution layer Conv, 2 depth-separable convolution layers DSC, and 2 inverted residual depth-separable convolution layers ResDSC, while the MobileJSCC decoder consists of 5 depth-separable transposed convolution layers DSTC. Experiments show that although DeepJSCC has a slightly higher PSNR under certain conditions, MobileJSCC is still an ideal choice for resource-constrained environments due to its lightweight design and significant computational resource savings. Especially in applications that require efficient processing and transmission of high-quality images, MobileJSCC provides a good balance by ensuring image quality while reducing the demand for computing resources.
[0073] What is disclosed above is only a preferred embodiment of the present invention, and it certainly cannot be used to limit the scope of rights of the present invention. Ordinary technicians in this field can understand that all or part of the processes of the above embodiment and equivalent changes made according to the claims of the present invention still fall within the scope of the invention.
Claims
1. A semantic wireless image transmission method suitable for resource-constrained platforms, characterized in that: The following steps are involved: Step 1: Construct a MobileJSCC semantic wireless communication network architecture, wherein the MobileJSCC semantic wireless communication network architecture includes a MobileJSCC encoder, a wireless channel and a MobileJSCC decoder; Step 2: Convert the image into feature representation using a lightweight MobileJSCC encoder; Step 3: Transmit the feature representation through the wireless channel; Step 4: Use the lightweight MobileJSCC decoder to decode the received feature representation and restore the original image.
2. The semantic wireless image transmission method applicable to a resource-constrained platform as claimed in claim 1, characterized in that: The lightweight MobileJSCC encoder in step 2 consists of one standard convolutional layer, two depthwise separable convolutional layers, and two inverted residual depthwise separable convolutional layers.
3. The semantic wireless image transmission method applicable to a resource-constrained platform as claimed in claim 2, characterized in that: In step 3, the wireless channel uses an additive white Gaussian noise channel and a Rayleigh fading channel model to simulate wireless channel transmission.
4. The semantic wireless image transmission method applicable to a resource-constrained platform as claimed in claim 3, characterized in that: The lightweight MobileJSCC decoder in step 4 consists of five depth-wise separable transposed convolutional layers.
5. The semantic wireless image transmission method applicable to a resource-constrained platform as claimed in claim 4, characterized in that: The MobileJSCC encoder and the MobileJSCC decoder are both trained end-to-end to optimize model parameters by minimizing the mean square error between the original image and the reconstructed image.
Citation Information
Cited By
Multi-mode face semantic communication method and device and medium
CN120318893A