Lightweight semantic image communication method and system based on deep learning
By introducing the ConvNeXt-T backbone network and lightweight signal-to-noise ratio adaptive module in the wireless image transmission system, the problem of lightweight and insufficient feature expression capabilities of the system model is solved, bandwidth and signal-to-noise ratio adaptive is achieved, and excellent performance is shown in high-resolution image transmission.
Patent Information
- Application Number
- CN202510200397.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-27
AI Technical Summary
While existing deep learning-based wireless image transmission systems are both channel adaptation and bandwidth ratio adaptation, it is difficult to achieve lightweight system models, and their feature expression capabilities are limited, which fails to effectively solve the model storage problem.
A lightweight semantic image communication method based on deep learning is proposed. By introducing an efficient and lightweight ConvNeXt-T backbone network and a lightweight signal-to-noise ratio adaptive module based on large convolution kernels, the system's parameter amount and storage amount are reduced, and bandwidth and signal-to-noise ratio adaptive is realized.
While maintaining performance competitiveness, the model's parameter volume and storage requirements are significantly reduced, especially in high-resolution image transmission, and the performance is significantly improved, and it has good signal-to-noise ratio and bandwidth adaptive characteristics.
Smart Images

Figure CN120050420A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and more specifically, to a lightweight semantic image communication method and system based on deep learning. Background Art
[0002] In traditional wireless communication systems, source coding and channel coding are often designed separately as independent modules, aiming to improve communication efficiency and reliability. However, this separate design approach has relatively low efficiency under harsh channel conditions and is prone to the "cliff effect", resulting in a sharp decline in communication performance.
[0003] With the rapid development of deep learning, researchers have begun to apply deep learning to joint source-channel coding (JSCC, Joint Source-Channel Coding) for wireless image transmission. Although these methods are trained under specific channel conditions and bandwidth ratios, they show significant advantages compared with traditional separate methods. They not only improve the transmission performance but also show stronger stability to changes in channel conditions, and also overcome the "cliff effect" caused by harsh channel conditions in traditional communication.
[0004] Currently, many researchers have turned their attention to this emerging communication paradigm of JSCC methods based on deep learning, and have not only made progress in semantic and task-oriented communication, but also many research results have emerged in image transmission. For example, channel output feedback is used to improve the reconstruction quality at the image receiving end and enhance the robustness of the system to channel changes; for the problem of signal-to-noise ratio adaptation, a signal-to-noise ratio adaptive image transmission system based on an attention module is proposed, enabling communication transmission to adapt to more diverse channel conditions. The JSCC method based on deep learning not only shows exciting capabilities in wireless image transmission, but also performs well in text and voice transmission.
[0005] Although this communication method of JSCC methods based on deep learning has shown great potential and made some progress, there are still some problems to be solved. In the existing technologies, the design of wireless image transmission tasks focuses more on the investigation of reconstruction quality, and most of them cannot take into account channel adaptability and bandwidth ratio adaptability; the proposed wireless image transmission communication systems can meet the signal-to-noise ratio adaptation requirements, but fail to effectively solve the problem of lightweight system models, and at the same time, their feature expression capabilities are limited due to the small receptive fields of the models; although some of the proposed methods maintain considerable competitiveness in performance, they do not consider the model storage problem. At the same time, in the next-generation mobile communication system, the interconnection of all things has become a trend of the times, and people's demand for intelligent wireless communication services is growing continuously, and many emerging intelligent applications such as holographic communication, mixed reality, and smart cities are emerging one after another.
[0006] Therefore, to meet the new requirements, we not only need to focus on improving the transmission performance but also pay attention to the system storage capacity, especially considering that some users' mobile devices do not have enough storage space to accommodate a large number of model parameters. Summary of the Invention
[0007] In view of the above problems, the purpose of the present invention is to provide a lightweight semantic image communication method and system based on deep learning, which is used to reduce the number of parameters and storage capacity of the wireless image transmission semantic communication system, making it lightweight and capable of achieving bandwidth and signal-to-noise ratio adaptation.
[0008] The first aspect of the present invention provides a lightweight semantic image communication method based on deep learning, including:
[0009] The image to be transmitted is input into the joint source-channel encoder, and at the same time, the channel signal-to-noise ratio information is embedded into the encoder for encoding to obtain a complex-domain semantic code;
[0010] Before the complex-domain semantic code is transmitted to the channel, its code length is adjusted through semantic masking operation to achieve bandwidth compression, and the masked semantic code to be transmitted is obtained;
[0011] The masked semantic code to be transmitted is transmitted through the channel, and a noisy semantic code is obtained at the receiving end;
[0012] The masked positions of the noisy semantic code are filled with zero values, and the received semantic code is obtained after zero-padding;
[0013] The received semantic code enters the joint source-channel decoder, and at the same time, the channel signal-to-noise ratio information is embedded into the decoder for decoding and reconstruction to obtain the output image.
[0014] In this solution, the formula for obtaining the complex-domain semantic code is specifically:
[0015]
[0016] where represents the complex domain, represents the input image, represents the real domain, the number of dimensions N = H × W × C, and the parameters H, W, and C respectively represent the height, width, and number of channels of the input image; μ is the channel signal-to-noise ratio, which is estimated according to the channel conditions and fed back to the joint source-channel encoder is the parameter of the encoder.
[0017] In this solution, the formula for obtaining the masked semantic code to be transmitted is specifically:
[0018] z′ = z · γ;
[0019] where γ ∈ {0, 1} N denotes a binary semantic code mask vector for masking the complex domain semantic code z, · denotes the dot product operation, V = ρN represents the transmission bandwidth, ρ is the bandwidth ratio, and the mask vector γ is set to γ i = 1 indicates that the i-th symbol will be selected for transmission; γ i = 0 indicates that the i-th symbol will be masked; after the masking operation, an input power constraint is imposed on each masked semantic code z′, that is, it satisfies
[0020] In this scheme, the formula for obtaining the noisy semantic code is specifically:
[0021]
[0022] where n is complex Gaussian distributed noise with mean zero and variance σ 2 , that is, n ~ η(0, σ 2 I).
[0023] In this scheme, the obtained received semantic code is obtained by zero-padding the masked positions of the noisy semantic code with zero values.
[0024] In this scheme, the formula for obtaining the output image is specifically:
[0025]
[0026] where μ is the channel signal-to-noise ratio, which is estimated according to the channel conditions and fed back to the joint source-channel decoder D θ , and θ is the parameter of the decoder.
[0027] This scheme also includes:
[0028] Before training, load the training dataset from the given dataset, and set the batch size, learning rate, and total number of training epochs to B, α, and T respectively; during training, set the patience parameter and minimum improvement threshold of the early stopping mechanism to ε and δ respectively, which are used to control early stopping of training and adjust the learning rate;
[0029] For each epoch during training, extract a batch of data A = {x 1 , x 2 , …, x B} from the training dataset, where A contains B samples;
[0030] For each data sample x in the batch i , the DeepJSCC-T model proposed in the present invention is used to encode and decode the sample to obtain a reconstructed output The formula is:
[0031]
[0032] Calculate the reconstructed image and the original image x i to obtain the mean squared error (MSE) between them, and obtain the loss value of the corresponding data sample The formula is:
[0033]
[0034] For the samples in each batch, further calculate the average loss of the batch The formula is:
[0035]
[0036] After calculating , update the model parameters and θ by the backpropagation algorithm to minimize the loss function. The formula is:
[0037]
[0038] Increment the number of rounds for updating the model parameters by 1;
[0039] If the number of rounds is less than the preset total number of rounds, the corresponding model continues to be trained; before the number of rounds reaches the preset total number of rounds, if the loss function does not decrease during the rounds, record the number of rounds when the loss function does not decrease. When the number of rounds when the loss function does not decrease is greater than the patience parameter, the model stops training according to the early stopping mechanism. And during the training process, if the number of rounds when the loss function does not decrease reaches the set minimum improvement threshold, adjust the learning rate and then train the model; if the number of rounds is greater than the preset total number of rounds, the corresponding model stops training and outputs the optimal model parameters and θ * , and the minimum improvement threshold is less than the patience parameter.
[0040] The second aspect of the present invention provides a lightweight semantic image communication system based on deep learning, including a memory and a processor. A program of a lightweight semantic image communication method based on deep learning is stored in the memory. When the program of the lightweight semantic image communication method based on deep learning is executed by the processor, the following steps are implemented:
[0041] The image to be transmitted is input into the joint source-channel encoder, and at the same time, the channel signal-to-noise ratio information is embedded into the encoder for encoding to obtain a complex-domain semantic code;
[0042] Before the complex-domain semantic code is transmitted to the channel, bandwidth compression is achieved by adjusting its code length through a semantic masking operation to obtain the masked semantic code to be transmitted;
[0043] The masked semantic code to be transmitted is transmitted through the channel, and a noisy semantic code is obtained at the receiving end;
[0044] The masked positions of the noisy semantic code are filled with zeros, and the received semantic code is obtained after zero-padding;
[0045] The received semantic code enters the joint source-channel decoder, and at the same time, the channel signal-to-noise ratio information is embedded into the decoder for decoding and reconstruction to obtain the output image.
[0046] In this scheme, the formula for obtaining the complex-domain semantic code is specifically:
[0047]
[0048] where represents the complex domain, represents the input image, represents the real domain, the number of dimensions N = H × W × C, and the parameters H, W, and C represent the height, width, and number of channels of the input image respectively; μ is the channel signal-to-noise ratio, which is estimated according to the channel conditions and fed back to the joint source-channel encoder is the parameter of the encoder.
[0049] In this scheme, the formula for obtaining the masked semantic code to be transmitted is specifically:
[0050] z′ = z·γ;
[0051] where γ ∈ {0, 1} B represents a binary semantic code masking vector for masking the complex-domain semantic code z, · represents the dot product operation, V = ρN represents the transmission bandwidth, ρ is the bandwidth ratio, and the masking vector γ is set to γ i = 1 indicates that the i-th symbol will be selected for transmission; γ i = 0 indicates that the i-th symbol will be masked; after the masking operation, an input power constraint is imposed on each masked semantic code z′, that is, it satisfies
[0052] The present invention discloses a lightweight semantic image communication method and system based on deep learning, which is used to reduce the parameter quantity and storage capacity of a wireless image transmission semantic communication system, making it have lightweight characteristics and enabling bandwidth and signal-to-noise ratio adaptation. The method provided by the present invention introduces an efficient and lightweight ConvNeXt-T as a new backbone network for joint source-channel coding (JSCC) in the system architecture design, and designs a lightweight signal-to-noise ratio adaptive module based on large convolutional kernels. The present invention demonstrates significant lightweight advantages in key indicators such as model parameter quantity and storage requirements, maintains quite competitiveness in performance such as peak signal-to-noise ratio (PSNR), especially for high-resolution images, the performance improvement is more significant, and it has good signal-to-noise ratio and bandwidth adaptation characteristics. The present invention provides an effective solution for high-quality image transmission, especially in resource-constrained scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 The flowchart of a lightweight semantic image communication method based on deep learning according to the present invention is shown;
[0054] Figure 2 The overall architecture diagram of the Deep JSCC-T system proposed by the present invention is shown;
[0055] Figure 3 The structural diagram of the pure convolutional module ConvNeXt Block is shown;
[0056] Figure 4 The block diagram of the lightweight signal-to-noise ratio adaptive module based on large convolutional kernels is shown;
[0057] Figure 5 The structural diagram of the lightweight signal-to-noise ratio adaptive module based on large convolutional kernels is shown;
[0058] Figure 6 The structural diagrams of the downsampling module and the upsampling module are shown;
[0059] Figure 7 The comparison of PSNR performance between the Deep JSCC-T model proposed by the present invention and other models under the conditions of dataset CI FAR-10 and bandwidth ratio ρ = 1 / 12 is shown;
[0060] Figure 8 The comparison of PSNR performance between the Deep JSCC-T model proposed by the present invention and other models under the conditions of dataset CI FAR-10 and bandwidth ratio ρ = 1 / 6 is shown;
[0061] Figure 9Shows the comparison of the structural similarity index between the Deep JSCC-T model proposed by the present invention and other models under the conditions of the dataset CIFAR-10 and the bandwidth ratio ρ = 1 / 12;
[0062] Figure 10 Shows the comparison of the structural similarity index between the Deep JSCC-T model proposed by the present invention and other models under the conditions of the dataset CIFAR-10 and the bandwidth ratio ρ = 1 / 6;
[0063] Figure 11 Shows the comparison of the PSNR performance between the Deep JSCC-T model proposed by the present invention and other models under the conditions of the dataset Kodak and the bandwidth ratio ρ = 1 / 12;
[0064] Figure 12 Shows the comparison of the PSNR performance between the Deep JSCC-T model proposed by the present invention and other models under the conditions of the dataset Kodak and the bandwidth ratio ρ = 1 / 6;
[0065] Figure 13 Shows the block diagram of a lightweight semantic image communication system based on deep learning according to the present invention;
[0066] Figure 14 Shows the comparison data between the Deep JSCC-T model of the present invention and other models. Detailed implementation manners
[0067] In order to more clearly understand the above objects, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.
[0068] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0069] Figure 1 Shows the flowchart of a lightweight semantic image communication method based on deep learning according to the present invention.
[0070] S101, The image to be transmitted is input into the joint source-channel encoder, and at the same time, the channel signal-to-noise ratio information is embedded into the encoder for encoding to obtain a complex-domain semantic code;
[0071] S102, Before the complex-domain semantic code is transmitted to the channel, its code length is adjusted through semantic masking operation to achieve bandwidth compression, and a masked semantic code to be transmitted is obtained;
[0072] S103. The masked semantic code to be transmitted is transmitted through the channel, and the noisy semantic code is obtained at the receiving end;
[0073] S104. The masked positions of the noisy semantic code are filled with zeros, and the received semantic code is obtained after zero-padding;
[0074] S105. The received semantic code enters the joint source-channel decoder, and at the same time, the channel signal-to-noise ratio information is embedded into the decoder for decoding and reconstruction to obtain the output image.
[0075] According to the embodiments of the present invention, for a wireless image transmission semantic communication system, a lightweight joint source-channel coding (DeepJSCC) scheme based on a convolutional neural network (ConvNeXt), called DeepJSCC-T, is proposed; wherein, the lightweight wireless image transmission system model consists of a trainable joint source-channel encoder semantic mask bandwidth compression module, an untrainable physical channel, zero-padding, and a trainable decoder D θ and so on, where and θ respectively represent the parameters of the encoder and the decoder.
[0076] According to the embodiments of the present invention, the formula for obtaining the complex-domain semantic code is specifically:
[0077]
[0078] where represents the complex domain, represents the input image, represents the real domain, the number of dimensions N = H×W×C, and the parameters H, W, and C respectively represent the height, width, and number of channels of the input image; μ is the channel signal-to-noise ratio, which is estimated according to the channel conditions and fed back to the joint source-channel encoder is the parameter of the encoder.
[0079] According to the embodiments of the present invention, the formula for obtaining the masked semantic code to be transmitted is specifically:
[0080] z′ = z·γ;
[0081] where γ ∈ {0, 1} N represents a binary semantic code mask vector for masking the complex-domain semantic code z, · represents the dot product operation, V = ρN represents the transmission bandwidth, ρ is the bandwidth ratio, and the mask vector γ is set to γ i = 1 indicates that the i-th symbol will be selected for transmission; γ i= 0 indicates that the i-th symbol will be masked; after the masking operation, an input power constraint is imposed on each masked semantic code z′, that is, it satisfies
[0082] According to an embodiment of the present invention, the formula for obtaining the noisy semantic code is specifically:
[0083]
[0084] where n is a complex Gaussian distributed noise with a mean of zero and a variance of σ 2 i.e., n ∼ η(0, σ 2 I).
[0085] According to an embodiment of the present invention, the obtained received semantic code is obtained by zero-padding the masked positions of the noisy semantic code with zero values.
[0086] According to an embodiment of the present invention, the formula for obtaining the output image is specifically:
[0087]
[0088] where μ is the channel signal-to-noise ratio, which is estimated according to the channel conditions and fed back to the joint source-channel decoder D θ , and θ is the parameter of the decoder.
[0089] According to an embodiment of the present invention, it further includes:
[0090] Before training, load the training dataset from the given dataset, and set the batch size, learning rate, and total number of training epochs to B, α, and T respectively; during training, set the patience parameter and minimum improvement threshold of the early stopping mechanism to ε and δ respectively, which are used to control early stopping of training and adjust the learning rate.
[0091] For each epoch during training, extract a batch of data A = {x 1 , x 2 , …, x B} from the training dataset, where A contains B samples;
[0092] For each data sample x i in the batch, encode and decode the sample based on the DeepJSCC-T model proposed by the present invention to obtain the reconstructed output whose formula is:
[0093]
[0094] Calculate the reconstructed image and the original image x i to obtain the loss value of the corresponding data sample by calculating the Mean Squared Error (MSE) between them Its formula is as follows:
[0095]
[0096] For each batch of samples, further calculate the batch average loss Its formula is as follows:
[0097] where \(i\in B\);
[0098] After calculating , update the model parameters \(\omega\) and \(\theta\) by minimizing the loss function through the backpropagation algorithm, and its formula is: and \(\theta\), and its formula is:
[0099]
[0100] Increment the number of rounds for updating the model parameters by 1;
[0101] If the number of rounds is less than the preset total number of rounds, the corresponding model continues to be trained; before the number of rounds reaches the preset total number of rounds, if the loss function does not decrease during the rounds, record the number of rounds when the loss function does not decrease. When the number of rounds when the loss function does not decrease is greater than the patience parameter, then according to the early stopping mechanism, the model stops training. And during the training process, if the number of rounds when the loss function does not decrease reaches the set minimum improvement threshold, then adjust the learning rate and train the model again; if the number of rounds is greater than the preset total number of rounds, the corresponding model stops training and outputs the optimal model parameters \(\omega\) and \(\theta\). and \(\theta\) * , and the minimum improvement threshold is less than the patience parameter.
[0102] It should be noted that during the training process, the Adam optimizer is used to update the model parameters, and the learning rate is dynamically adjusted through the MultiplicativeLR scheduler to avoid falling into local optima; at the same time, when monitoring the change of the model performance on the validation set, the early stopping mechanism is adopted to prevent the model from overfitting.
[0103] Figure 2 is the overall architecture diagram of the Deep JSCC-T system.
[0104] As Figure 2As shown, the input image is fed into the encoder for processing to extract multi-level semantic features. The encoder adopts a two-stage design architecture, which is responsible for extracting preliminary and deep semantic features respectively. By gradually compressing the spatial dimension of the input image, each stage not only completes the extraction of semantic features at different levels but also effectively reduces the computational complexity. The first stage (Stage 1) mainly completes the preliminary extraction of semantic features, and the dimension of its output features is compressed to This module consists of two-dimensional convolution (Conv2d), pure convolutional module (ConvNeXt Block), and LASMod (Lightweight Adaptive SNR Module Based on Large Convolutional Kernels), etc. Among them, Conv2d is used to capture the spatial features of the image, ConvNeXt Block provides more efficient local and global feature modeling, and the LASMod module adaptively adjusts the model parameters to different signal-to-noise ratio conditions and reduces the number of parameters and storage requirements during this process. The second stage (Stage 2) further extracts deep semantic features and optimizes the semantic feature representation to adapt to channel transmission. This stage includes a downsampling module (Downsample), ConvNeXt Block, and LASMod module. The downsampling module further compresses the image dimension to to create conditions for subsequent deep semantic feature extraction. The ConvNeXt Block and LASMod modules enhance the representation ability of the image semantic information in this stage and further improve the robustness and adaptability of transmission. The decoder is designed symmetrically to the encoder and is used to gradually restore the semantic feature information during the transmission process to generate a high-quality reconstructed image. Corresponding to the downsampling module in the encoder, the decoder uses an upsampling module (Upsample) to restore the spatial resolution of the image and combines transposed convolution (convTranspose2d) to restore the semantic features to the original image. The decoder also includes ConvNeXt Block and LASMod modules at each stage to restore high-level semantic features and optimize the restoration quality.
[0105] Figure 3 It is the structural diagram of the pure convolutional module ConvNeXt Block.
[0106] As Figure 3As shown, a pure convolutional module (ConvNeXt Block) is adopted in each stage of the joint source-channel encoder and decoder to provide more efficient local and global feature modeling; this module uses depthwise separable convolutions, which can effectively reduce the number of model parameters, thereby reducing the model complexity; in the structure diagram of the pure convolutional module, k represents the convolutional kernel, s represents the stride, and p represents the padding; for example, k7 means the convolutional kernel size is 7, s1 means the stride is 1, and p3 means the padding size is 3; during data processing, the input data (with size h×w×c) enters the pure convolutional module (ConvNeXt Block), first passing through a depthwise separable convolutional layer (Depthwise Conv2d), which uses a large 7x7 convolutional kernel and p3 padding to keep the height and width of the input unchanged while changing the number of channels; then the data passes through layer normalization (Layer Norm) and a 1x1 convolutional layer (Conv2d) to adjust the number of channels without changing the spatial dimensions; subsequently, the data passes through the GELU activation function to introduce non-linearity, increasing the model's expressive power, and passes through a 1x1 convolutional layer and layer normalization again to further adjust the number of channels; then a possible dropout operation is performed through the regularization method Drop Path to reduce overfitting; finally, the entire module is added to the input through a residual connection (skip connection), allowing direct propagation of gradient information; after being processed by the ConvNeXt Block, its output size is the same as the input size, and the output size is still The stacking numbers of the ConvNeXt Blocks in Stage 1 and Stage 2 are set to M1 and M2 respectively.
[0107] Figure 4 and Figure 5 are the block diagram and structure diagram of the lightweight signal-to-noise ratio adaptive module based on large convolutional kernels.
[0108] As Figure 4 and Figure 5 shown, the lightweight signal-to-noise ratio adaptive module based on large convolutional kernels (Lightweight Adaptive SNR Module Based on Large Convolutional Kernels, LASMod) consists of a lightweight module based on large convolutional kernels (Lightweight Module Based on Large Convolutional Kernels, LLK), a signal-to-noise ratio information extraction module (SNR Information Extraction, SIE), and a 1×1 convolutional layer Module composition; By introducing a decoupled large convolutional kernel design, LASMod can not only significantly increase the receptive field of the model, but also effectively control the computational complexity and parameter scale; while maintaining the lightweight of the model, it can significantly improve its adaptability to complex and changing channel conditions, making it particularly suitable for communication scenarios with limited computing resources; the fusion of signal-to-noise ratio information improves the adaptive ability of the model and maintains high-efficiency transmission performance under different signal-to-noise ratio conditions; the LLK sub-module extracts local spatial features g from the input information y using the decoupled large convolutional kernel network 1 and global spatial features g 2 The formulas for are and where is a depthwise convolution with a convolutional kernel of k i and a dilation rate of d i (i = 1, 2), select k 1 = 5, d 1 = 1, k 2 = 7, d 2 = 3; Two depthwise convolutions with gradually increasing kernel sizes and increasing dilation rates cooperate with each other to construct a larger convolutional kernel; after extracting local and global features, first pass through a 1×1 convolutional layer respectively, and then perform channel fusion of the spatial feature vectors, and splice the features obtained by the kernels with different receptive field ranges to obtain the feature information g, and its calculation formula is where represents a 1×1 convolutional layer; after splicing, perform average pooling and max pooling on the feature information g respectively to capture the distribution and significance of the features, which are used to calculate the adaptive attention weights and perform weighted fusion on the multi-scale features to highlight the key information, thereby enhancing the adaptability of the model under complex input conditions; at the same time, in order to realize the information interaction of different spatial features, after average pooling and max pooling, use a 2×2 convolutional layer to integrate spatial and channel information, and apply the Sigmoid activation function to the spatial attention feature map to obtain the spatial selection mask λ, and its calculation formula is where and represent channel-based average pooling and max pooling for extracting spatial relationships, represents a 2×2 convolutional layer, and Sigmoid(.) represents the Sigmoid activation function; after obtaining the spatial selection mask, weight the features obtained by the decoupled large convolutional kernel with the corresponding spatial selection mask to obtain the fused feature g' of the LLK sub-module, and its calculation formula is where λ 1 and λ 2Separate masks are selected for the independent spaces corresponding to the large convolution kernels decomposed in λ. To avoid excessive increase in network complexity, the extraction of signal-to-noise ratio information is implemented using a simple neural network consisting of two fully connected (FC) layers and two activation layers. After each FC layer, there is a ReLu activation function, which maps the signal-to-noise ratio information μ into high-level features u suitable for fusion with the output features of the sub-module LLK, so that the entire network can adapt to complex and changing channel conditions. The mapping process of u is u = ReLU(W 2 (ReLU(μ·W 1 +b 1 ))·+b 2 ), where W1 and b 1 represent the weights and biases of the first FC layer, and W2 and b 2 represent the weights and biases of the second FC layer, and ReLU represents the activation function. To enhance the adaptive ability of the model to the dynamic changes of the signal-to-noise ratio, LASMod fuses the features output by the large convolution kernel network and the extracted signal-to-noise ratio information. The mapped signal-to-noise ratio information u is concatenated with the fused spatial features g' obtained through LLK, and is fused through a 1×1 convolution layer to obtain the new attention feature q. Its calculation formula is After obtaining the attention feature q, an element-wise multiplication operation is performed between it and the input y of LASMod to obtain the output y' of LASMod. Its calculation formula is y' = y·q.
[0109] Figure 6 Structural diagrams of the downsampling module (Downsample) and the upsampling module (Upsample).
[0110] As Figure 6 shown, the downsampling module (Downsample) and the upsampling module (Upsample) are respectively used in Stage 2 of the encoder and Stage 1 of the decoder; the downsampling and upsampling modules are used to adjust the spatial dimensions of the input, which helps the model extract and recover higher-level features; for the downsampling module, the input feature tensor is first normalized by Layer Norm, and then the spatial resolution is halved through a convolutional layer (Conv2d) with a kernel size of 2 and a stride of 2; the Upsample module is the inverse operation of Downsample. The input feature tensor is first magnified in spatial resolution through a ConvTranspose2d layer (transposed convolution) with a kernel size of 2 and a stride of 2, and then passed through Layer Norm (layer normalization) to optimize the feature distribution and stabilize the training.
[0111] Figures 7 to 12This is the comparison of the transmission performance between the Deep JSCC-T model proposed in the present invention and other models.
[0112] First, the proposed scheme was trained on the CIFAR-10 image dataset. After the training phase was completed, the performance of each scheme was tested on 10,000 test images in the CIFAR-10 dataset that were different from the training images. For high-resolution images, the DIV2K dataset was selected as the training set. During the training phase, in order to enhance the generalization ability of the model, the method of random cropping was adopted to adjust the image size to the format of 3×256×256. In the testing phase, the Kodak image dataset was selected as the test set. When evaluating the transmission performance of high-resolution images, in order to reduce the random influence brought by random channel noise, each image was transmitted 10 times. By transmitting the images multiple times and taking the average, a more real and accurate evaluation of the transmission performance can be obtained.
[0113] During the experiment, the initial learning rate was set to 10 -4 ; during the experiment process, a learning rate strategy of dynamic adjustment was adopted. If the validation loss did not decrease significantly after 20 training epochs, the learning rate was reduced to 0.9 times the original. For the entire training process, the total number of training epochs was set to 1000. For low-resolution images, if the validation loss did not decrease after the first 100 training epochs, the training would automatically terminate; while for high-resolution images, if the validation loss did not decrease after the first 150 training epochs, the training would also automatically stop. In the system architecture proposed in the present invention, the number of ConvNeXt Blocks in each stage was set to M 1 = 3, M2 = 6; for low-resolution images, the number of channels was set to [C 1 , C 2 = [96, 192]. For high-resolution images, the number of channels was set to [C 1 , C 2 = [128, 256]. In the bandwidth ratio setting, the bandwidth ratio was selected as ρ ∈ {1 / 24, 1 / 12, 1 / 8, 1 / 6, 5 / 24, 1 / 4}. The training covered the signal-to-noise ratio range from 0 dB to 25 dB to ensure the performance of the model under different signal-to-noise ratio conditions. All experiments were carried out on a Linux server using a single NVIDIA RTX 3090 GPU.
[0114] Figure 7 and Figure 8This is a comparison of the PSNR performance of the Deep JSCC-T model proposed in the present invention with other models under the conditions of the CIFAR-10 dataset, AWGN channel, and different bandwidth ratios (ρ = 1 / 12 and ρ = 1 / 6). For all Deep JSCC models, they are trained at specific signal-to-noise ratio values, specifically SNR train = 1dB, 4dB, 7dB, 13dB, 19dB; while for the ADJSCC model and the Deep JSCC-V model, they are trained under a wider signal-to-noise ratio variation range ([0, 25]dB). For the Deep JSCC-T proposed in the present invention, the same signal-to-noise ratio test range ([0, 25]dB) is selected for testing and comparison. As Figure 7 shown, in the case of bandwidth ratio ρ = 1 / 12, the Deep JSCC-T model proposed in the present invention outperforms the Deep JSCC model trained at a specific signal-to-noise ratio (SNRtrain) at all tested signal-to-noise ratio (SNRtest) values; and as SNRtest increases, the performance of Deep JSCC-T continuously improves, and its best performance exceeds that of the Deep JSCC model (under the condition of SNRtrain = 1dB) by 6.34dB. Thus, it can be seen that the proposed Deep JSCC-T model shows robust adaptability to different SNRtest, further indicating its good robustness under different signal-to-noise ratio conditions; compared with the latest signal-to-noise ratio adaptive and bandwidth adaptive Deep JSCC-V method, Deep JSCC-T maintains considerable competitiveness throughout the tested signal-to-noise ratio range and has good signal-to-noise ratio and bandwidth adaptive characteristics; compared with the ADJSCC method trained separately under a specific bandwidth ratio condition, although there is a gap of about 0.36dB in peak signal-to-noise ratio (PSNR) for Deep JSCC-T, this is because the use of large convolutional kernels with a large receptive field may cause some detailed feature information to be ignored when processing low-resolution images. However, thanks to the powerful feature extraction ability of the ConvNeXt Block structure, Deep JSCC-T still performs well in feature extraction and processing. As Figure 8 shown, in the case of bandwidth ratio ρ = 1 / 6, it can be observed that compared with Figure 7The conclusion of similar trends. It should be emphasized that as the bandwidth ratio increases, the performance advantage of Deep JSCC-T compared to Deep JSCC (under the condition of SNRtrain = 1dB) in the high SNR region becomes more obvious, and its best performance advantage is up to 9.07dB; while the performance gap between Deep JSCC-T and the ADJSCC and Deep JSCC-V models will gradually narrow with the increase of SNRtest, and even surpass these two methods at SNRtest = 16dB. Thus, it can be seen that the model proposed in the present invention does not cause excessive losses in bandwidth compression or SNR adaptability while improving performance.
[0115] Figure 9 and Figure 10 shows the comparison of the Structural Similarity Index (SSIM) between the Deep JSCC-T model proposed in the present invention and other models under different bandwidth ratios (ρ = 1 / 12 and ρ = 1 / 6) on the CIFAR-10 dataset. SSIM measures the difference between the original image and the reconstructed image from the perspective of structural similarity. As Figure 9 and Figure 10 shown, under different bandwidth ratios, whether it is low bandwidth or high bandwidth, Deep JSCC-T can outperform other comparison models in terms of SSIM. Similar to PSNR, the SSIM value increases with the increase of the test SNR, and Deep JSCC-T shows better retention of the original image in terms of structural similarity, especially prominent at high SNR.
[0116] Figure 11 and Figure 12 shows the PSNR performance comparison between the Deep JSCC-T model proposed in the present invention and other models on the higher-resolution Kodak dataset, under the AWGN channel and different bandwidth ratios (ρ = 1 / 12 and ρ = 1 / 6). For high-resolution images, the DIV2K dataset and the Kodak dataset are selected as the training set and the test set respectively. During training, the image size is adjusted by the method of random cropping to enhance the generalization ability of the model; during testing, the method of multiple transmissions is used to reduce the randomness impact brought by random channel noise. As Figure 11 and Figure 12As shown, under the training conditions based on the high-resolution DIV2K dataset, compared with several other comparison models, the proposed Deep JSCC-T model shows obvious performance advantages both under the bandwidth ratio ρ = 1 / 12 and ρ = 1 / 6; as the test signal-to-noise ratio (SNRtest) increases, the performance advantage of the proposed Deep JSCC-T model in terms of peak signal-to-noise ratio (PSNR) becomes more obvious. In particular, compared with the ADJSCC method trained separately under a specific bandwidth ratio, the maximum performance advantage of the peak signal-to-noise ratio (PSNR) of Deep JSCC-T can reach more than 5 dB within the test signal-to-noise ratio range.
[0117] Table 1 shows the comparison of the inference time, FLOPs, number of model parameters (#param), and storage overhead between the proposed Deep JSCC-T model and other models on the Kodak dataset (batch size is 1). The inference time is related to the transmission size, and an image with a transmission resolution of 512×512 is selected to verify the transmission delay. FLOPs, number of model parameters (#param), and storage overhead are independent of the image size and only related to the model itself.
[0118] Under the condition that the bandwidth ratio R = 1 / 6, the comparison cases are shown in Table 1. It can be seen from Table 1 that under the condition that the bandwidth ratio R = 1 / 6, the inference time of the Deep JSCC-T model proposed by the present invention is the least, and it has an obvious advantage in terms of latency. However, in terms of FLOPs, the number of model parameters, and storage requirements, the Deep JSCC-T proposed by the present invention shows obvious advantages. The FLOPs value of the Deep JSCC-T proposed by the present invention is reduced by 30.75 times compared with ADJSCC and 31.12 times compared with Deep JSCC-V. Therefore, the Deep JSCC-T proposed by the present invention can significantly reduce its computational overhead and resource requirements. The main reason lies in the efficient operation mechanism of depthwise separable convolution in the designed backbone network and signal-to-noise ratio adaptive module. In terms of the number of model parameters, the Deep JSCC-T proposed by the present invention is reduced by 24.65% compared with ADJSCC and 27.52% compared with Deep JSCC-V; and the significant reduction in the number of parameters also directly reduces the storage requirements; the storage overhead of the Deep JSCC-T proposed by the present invention is reduced by 30.98% compared with ADJSCC and 33.47% compared with Deep JSCC-V. In addition, by evaluating the average PSNR performance of the three methods on the Kodak dataset under the condition that the signal-to-noise ratio range is [1, 25] dB, it can be found that the PSNR of Deep JSCC-T is 3.68 dB higher than that of ADJSCC and 4.65 dB higher than that of Deep JSCC-V. Therefore, the Deep JSCC-T proposed by the present invention performs similarly to ADJSCC and Deep JSCC-V in terms of inference time, but it shows significant lightweight advantages in key indicators such as FLOPs, the number of model parameters, and storage requirements, and shows excellent improvement in performance. This shows that the Deep JSCC-T proposed by the present invention is not only more efficient in computing resource utilization but also provides great potential and scalability for applications in actual communication scenarios.
[0119] Figure 13 A block diagram of a lightweight semantic image communication system based on deep learning according to the present invention is given.
[0120] As Figure 13 shown, the second aspect of the present invention provides a lightweight semantic image communication system 13 based on deep learning, including a memory 131 and a processor 132. A program of a lightweight semantic image communication method based on deep learning is stored in the memory. When the program of the lightweight semantic image communication method based on deep learning is executed by the processor, the following steps are implemented:
[0121] The image to be transmitted is input into the joint source-channel encoder, and at the same time, the channel signal-to-noise ratio information is embedded into the encoder for encoding to obtain a complex-domain semantic code;
[0122] Before the complex-domain semantic code is transmitted to the channel, its code length is adjusted through semantic masking operation to achieve bandwidth compression, and the masked semantic code to be transmitted is obtained;
[0123] The masked semantic code to be transmitted is transmitted through the channel, and a noisy semantic code is obtained at the receiving end;
[0124] The masked positions of the noisy semantic code are filled with zero values, and the received semantic code is obtained after zero-padding;
[0125] The received semantic code enters the joint source-channel decoder, and at the same time, the channel signal-to-noise ratio information is embedded into the decoder for decoding and reconstruction to obtain the output image.
[0126] In this scheme, the formula for obtaining the complex-domain semantic code is specifically:
[0127]
[0128] where represents the complex domain, represents the input image, represents the real domain, the number of dimensions N = H × W × C, and the parameters H, W, and C represent the height, width, and number of channels of the input image respectively; μ is the channel signal-to-noise ratio, which is estimated according to the channel conditions and fed back to the joint source-channel encoder is the parameter of the encoder.
[0129] In this scheme, the formula for obtaining the masked semantic code to be transmitted is specifically:
[0130] z′ = z·γ;
[0131] where γ ∈ {0, 1} N represents a binary semantic code masking vector for masking the complex-domain semantic code z, · represents the dot product operation, V = ρN represents the transmission bandwidth, ρ is the bandwidth ratio, and the masking vector γ is set to γ i = 1 means that the i-th symbol will be selected for transmission; γ i = 0 means that the i-th symbol will be masked; after the masking operation, an input power constraint is imposed on each masked semantic code z′, that is, it satisfies
[0132] The present invention discloses a lightweight semantic image communication method and system based on deep learning, which is used to reduce the parameter quantity and storage capacity of a wireless image transmission semantic communication system, endow it with lightweight characteristics, and enable bandwidth and signal-to-noise ratio adaptability. The method provided by the present invention introduces an efficient and lightweight ConvNeXt-T as a new backbone network for joint source-channel coding (JSCC, Joint Source-Channel Coding) in the system architecture design, and designs a lightweight signal-to-noise ratio adaptability module based on large convolutional kernels. The present invention demonstrates significant lightweight advantages in key indicators such as model parameter quantity and storage requirements, maintains quite competitiveness in performance such as peak signal-to-noise ratio (PSNR), especially for high-resolution images, the performance improvement is more significant, and it has good signal-to-noise ratio and bandwidth adaptability characteristics. The present invention provides an effective solution for high-quality image transmission, especially in resource-constrained scenarios.
[0133] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.
[0134] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units; they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0135] In addition, each functional unit in the embodiments of the present invention can be all integrated in one processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit; the above integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.
[0136] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the above method embodiments. The aforementioned storage medium includes various media that can store program codes, such as removable storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0137] Alternatively, if the above integrated units of the present invention are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media that can store program codes, such as removable storage devices, ROM, RAM, magnetic disks, or optical discs.
Claims
1. A lightweight semantic image communication method based on deep learning, characterized in that: include: The image to be transmitted is input into the joint source channel encoder, and the channel signal-to-noise ratio information is embedded into the encoder for encoding to obtain a complex domain semantic code; Before the complex domain semantic code is transmitted to the channel, the code length is adjusted through the semantic mask operation to achieve bandwidth compression, and the masked semantic code to be transmitted is obtained; The masked semantic code to be transmitted is transmitted through the channel, and the noisy semantic code is obtained at the receiving end; The mask position of the noisy semantic code is filled with zero values, and the received semantic code is obtained after zero filling; The received semantic code enters the joint source channel decoder, and the channel signal-to-noise ratio information is embedded in the decoder for decoding and reconstruction to obtain the output image.
2. According to claim 1, a lightweight semantic image communication method based on deep learning is characterized in that: The obtained plural domain semantic code The formula is: in, represents the complex domain, represents the input image, represents the real number domain, with dimension N = H × W × C, where the parameters H, W, and C represent the height, width, and number of channels of the input image, respectively; μ is the channel signal-to-noise ratio, which is estimated based on the channel conditions and fed back to the joint source-channel encoder are the parameters of the encoder.
3. The method for lightweight semantic image communication based on deep learning according to claim 1, characterized in that: The mask semantic code to be transmitted is obtained The formula is: z′=z·γ; Among them, γ∈{0,1} N represents a binary semantic code mask vector for masking the complex domain semantic code z, · represents the dot product operation, V = ρN represents the transmission bandwidth, ρ is the bandwidth ratio, and the mask vector γ is set to γ i =1 means the i-th symbol will be selected for transmission; γ i = 0 means that the i-th symbol will be masked; after the masking operation is completed, an input power constraint is imposed on each mask semantic code z′, that is, 4. The method of lightweight semantic image communication based on deep learning according to claim 1, characterized in that: The noisy semantic code is obtained The formula is: Among them, n is a model with mean zero and variance σ 2 Complex Gaussian distribution noise, that is, n~η(0,σ 2 I).
5. The method of lightweight semantic image communication based on deep learning according to claim 1, characterized in that: The received semantic code is through the noisy semantic code The mask positions of are zero-filled with zero values.
6. The method of lightweight semantic image communication based on deep learning according to claim 1, characterized in that: The output image is obtained The formula is: Where μ is the channel signal-to-noise ratio, which is estimated based on the channel conditions and fed back to the joint source-channel decoder D θ , θ is the parameter of the decoder.
7. The method of lightweight semantic image communication based on deep learning according to claim 1, characterized in that: Also includes: Before training, load the training dataset from the given dataset, and set the batch size, learning rate, and total number of training rounds to B, α, and T respectively; During the training process, the patience parameter and minimum improvement threshold of the Early Stopping mechanism are set to ε and δ, which are used to control early stopping of training and adjust the learning rate, respectively. For each epoch in the training process, a batch of data A = {x1, x2, ..., x B }, where A contains B samples; For each data sample x in the batch i , based on the DeepJSCC-T model proposed in this invention, the sample is encoded and decoded to obtain the reconstructed output The formula is: Computational reconstruction of the image With the original image x i The mean square error (MSE) between them is used to obtain the loss value of the corresponding data sample. The formula is: For each batch of samples, the batch average loss is further calculated The formula is: where i∈B; Calculate Then, the model parameters are updated by minimizing the loss function through the back propagation algorithm. and θ, the formula is: The number of rounds for updating model parameters is increased by 1; If the number of rounds is less than the preset total number of rounds, the corresponding model continues to be trained; before the number of rounds reaches the preset total number of rounds, if the model loss function does not decrease during the rounds, the number of rounds in which the corresponding loss function does not decrease is recorded. When the number of rounds in which the loss function does not decrease is greater than the patience parameter, the model stops training according to the early stopping mechanism, and during the training process, if the number of rounds in which the loss function does not decrease reaches the set minimum improvement threshold, the learning rate is adjusted and the model is trained again; if the number of rounds is greater than the preset total number of rounds, the corresponding model stops training and outputs the optimal model parameters. and θ * , the minimum improvement threshold is less than the patient parameter.
8. A lightweight semantic image communication system based on deep learning, characterized in that: The invention comprises a memory and a processor, wherein a lightweight semantic image communication method program based on deep learning is stored in the memory, and the lightweight semantic image communication method program based on deep learning is executed by the processor to implement the following steps: The image to be transmitted is input into the joint source channel encoder, and the channel signal-to-noise ratio information is embedded into the encoder for encoding to obtain a complex domain semantic code; Before the complex domain semantic code is transmitted to the channel, the code length is adjusted through the semantic mask operation to achieve bandwidth compression, and the masked semantic code to be transmitted is obtained; The masked semantic code to be transmitted is transmitted through the channel, and the noisy semantic code is obtained at the receiving end; The mask position of the noisy semantic code is filled with zero values, and the received semantic code is obtained after zero filling; The received semantic code enters the joint source channel decoder, and the channel signal-to-noise ratio information is embedded in the decoder for decoding and reconstruction to obtain the output image.
9. The light-weight semantic image communication system based on deep learning according to claim 8, characterized in that: The obtained plural domain semantic code The formula is: in, represents the complex domain, represents the input image, represents the real number domain, with dimension N = H × W × C, where the parameters H, W, and C represent the height, width, and number of channels of the input image, respectively; μ is the channel signal-to-noise ratio, which is estimated based on the channel conditions and fed back to the joint source-channel encoder are the parameters of the encoder.
10. A light-weight semantic image communication system based on deep learning according to claim 8, characterized in that: The mask semantic code to be transmitted is obtained The formula is: z′=z·γ; Among them, γ∈{0,1} N represents a binary semantic code mask vector for masking the complex domain semantic code z, · represents the dot product operation, V = ρN represents the transmission bandwidth, ρ is the bandwidth ratio, and the mask vector γ is set to γ i =1 means the i-th symbol will be selected for transmission; γ i = 0 means that the i-th symbol will be masked; after the masking operation is completed, an input power constraint is imposed on each mask semantic code z′, that is,