Self-adaptive rate image semantic communication system fusing potential diffusion model
By integrating the adaptive rate control and visual effect optimization of the potential diffusion model, the visual effect distortion problem of the image semantic communication system under low signal-to-noise ratio conditions is solved, and adaptive rate control and image quality improvement are achieved.
Patent Information
- Application Number
- CN202510379938.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
The existing image semantic communication system based on joint encoding of deep source channels lacks an image content feedback mechanism, resulting in distortion of image transmission quality and visual effects under low signal-to-noise ratio conditions, and low system adaptability and flexibility.
Adaptive rate image semantic communication system that fuses potential diffusion model is adopted to realize adaptive rate control of combined feedback of image content and channel signal-to-noise ratio through a policy network and a feature dynamic encoder, and combines the potential diffusion model optimization module to optimize the visual effect of the transmitted image.
Adaptive rate control under low signal-to-noise ratio conditions is realized, transmission resource consumption is reduced, image visual effect distortion is improved through potential diffusion models, and the adaptability and practicality of image semantic communication is improved.
Smart Images

Figure CN120263985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image semantic communication, and in particular, to an adaptive rate image semantic communication system integrating a latent diffusion model (Latent Diffusion Model, i.e., LDM, latent diffusion model). Background Art
[0002] Most current communication technologies are based on Shannon's information theory. As the culmination of traditional mobile communication technologies, 5G has gradually approached the Shannon limit in both source coding technology and channel coding technology. Therefore, it is necessary to break away from the basic framework of probability information and make technological changes to modern communication technologies to better meet the development needs of future transmission communication. With the development of the 6G era and the Internet of Things, the utilization of relevant communication resources needs to be more efficient, and more intelligent data needs to be transmitted with fewer communication resources. As one of the key technologies of 6G, semantic communication focuses on understanding, extracting, and transmitting the semantic features of the source, and ensuring that the destination can understand the received source semantic features, so as to successfully recover the source information based on semantics. Image semantic communication is one of the important research contents of semantic communication. Image semantic communication technology mainly relies on computer vision, and its main focus is on the learning, understanding, and creation of image information by computers. With the development of computer technology, the core technology of computer image processing has shifted to deep learning technology represented by convolutional neural networks.
[0003] In recent years, image semantic communication systems based on deep joint source-channel coding (Deep-JSCC) have developed rapidly. Since the content transmitted by the image semantic communication system is no longer a bit stream but a feature vector corresponding to the image, the key to achieving transmission rate control in image semantic communication lies in controlling the size of the image feature vector. Existing image semantic communication systems based on deep joint source-channel coding mainly achieve adaptive transmission rate through channel signal-to-noise ratio (SNR) feedback, and its rate index is CPP (Channel Usage per Pixel). The CPP of image semantic communication is related to the size of the transmitted image features, and CPP can be expressed as:
[0004] However, in the context of limited, expensive communication transmission resources in the Internet of Things and the increasingly rich content and quality of transmitted images, the adaptive rate image semantic communication based on channel SNR feedback lacks an image content feedback mechanism, resulting in low adaptability, flexibility, and practicality of the system. Therefore, how to effectively implement an image content feedback mechanism has become one of the key issues in the field of image semantic communication.
[0005] Currently, the commonly used performance evaluation metrics for deep source-channel joint coding image semantic communication systems are Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity Index (SSIM). Generally, the higher these two performance metrics are, the higher the quality of the transmitted image. However, when the channel signal-to-noise ratio is poor, even if the values of the peak signal-to-noise ratio and the structural similarity index are within the normal range, the transmitted image will still have visual distortion. How to improve the visual effect of the transmitted image under low signal-to-noise ratio conditions is an important challenge facing the current development of semantic communication. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide an adaptive rate image semantic communication system integrating a latent diffusion model in view of the deficiencies of the above-mentioned prior art. An image content feedback mechanism is realized through the proposed policy network and Features Dynamic Encoder\Decoder, and an adaptive rate control scheme based on joint feedback of image content and channel signal-to-noise ratio is realized. At the same time, a latent diffusion model (LDM) optimization module is developed and integrated to optimize the visual effect of the transmitted image, so as to solve the problem of visual distortion of the transmitted image in the case of poor channel signal-to-noise ratio in image semantic communication.
[0007] To solve the above technical problems, the technical solutions adopted by the present invention are as follows:
[0008] An adaptive rate image semantic communication system integrating a latent diffusion model, comprising a transmitter module, a receiver module, a channel module, and an image visual effect optimization module based on the latent diffusion model;
[0009] The transmitter module includes a source-channel joint encoder, a policy network, and a feature dynamic encoder; the source-channel joint encoder processes the input original image X using a convolutional neural network and converts it into the corresponding feature F of the image; the feature dynamic encoder processes the image feature F and divides the feature F into two parts of features G s and G n where s and n are preset values; the policy network generates a binary mask W with the same two-dimensional size s as G s and controls and adjusts the image feature F transmitted in the channel through the binary mask W s ; the image features to be transmitted in G p are controlled and screened by multiplying G s by the binary mask W to achieve adaptive rate control; the finally transmitted image feature F s is expressed as F p = G p =Gs ×W + G n ;
[0010] The channel module includes an additive white Gaussian noise channel and a Rayleigh channel; Feature F p After being transmitted through the channel module, that is, in Feature F p random noise is added, and the feature after passing through the channel module is expressed as Then where θ is a noise vector with the same size as Feature F p and the values follow a standard Gaussian distribution;
[0011] Feature F p needs to be restored to an image after being processed by the receiver module. The receiver module includes a feature dynamic decoder and a source-channel joint decoder; The image feature after passing through the channel needs the feature dynamic decoder to process its feature size, and after processing, the feature size is the same as that of the original image feature F; The image feature is finally converted into the transmitted image after being processed by the source-channel joint decoder The source-channel joint decoder performs the opposite operation to the source-channel joint encoder;
[0012] The image visual effect optimization module based on the latent diffusion model is used to optimize the transmitted and restored image when the channel signal-to-noise ratio is poor, avoiding the phenomenon of visual effect distortion. The image visual effect optimization module based on the latent diffusion model includes a conditional generator module and a latent diffusion processing module;
[0013] The conditional generator module includes BLIPv2, a CLIP encoder, and an LDM encoder; BLIPv2 is a pre-trained model for image-to-text, extracting the text information T of the transmitted image The text information T is encoded by the CLIP encoder to obtain its corresponding text feature f txt , which is a multi-dimensional feature vector; The LDM encoder encodes the transmitted picture into the latent space, that is, obtaining the latent space feature f of the transmitted image spa ; The finally obtained conditional set is represented by c, that is, c = {f spa , f txt , λ}, where λ is the known channel signal-to-noise ratio;
[0014] The latent diffusion processing module performs a reverse diffusion denoising operation with the conditional set c as a parameter. The latent diffusion processing module includes several denoisers and an LDM decoder. Among them, the denoiser is a UNet structure based on the stable diffusion model, and the LDM decoder decodes the latent space representation of the image processed by the denoiser to obtain the visually optimized image X final ;
[0015] The training of the latent diffusion model includes a forward noise addition process and a reverse denoising process. The original image passes through the VAE encoder to generate the latent space feature z0; the final latent space feature z is inferred based on the latent space feature z0 final , z final is converted into an image X with optimized vision through the LDM decoder final .
[0016] Furthermore, in the source-channel joint encoder, the original image X first undergoes a reflection padding operation and then undergoes a cyclic operation with the structural sequence of a two-dimensional convolutional layer, a normalization layer, and a ReLU activation function for several times, and finally outputs the feature F corresponding to the image
[0017] Furthermore, the specific processing process of the feature dynamic encoder is as follows: the image feature F output by the source-channel joint encoder is processed through two cyclic operations with the structural sequence of a residual network and SNR adaption, and then undergoes convolutional processing by a two-dimensional convolutional layer. Finally, through a reconstruction grouping operation, the second dimension of the input feature is divided into two groups, s and n, and the corresponding two-part features G s and G n ;
[0018] The SNR adaption is implemented based on a multi-layer perceptron design, and the input of the module is adjusted by the multiplication and addition coefficients generated based on the channel signal-to-noise ratio; the SNR adaption first performs an average pooling operation on the input and then concatenates it with the channel SNR, and after being processed by the MLP, the multiplication and addition factors obtained are respectively multiplied and added to the original input
[0019] Furthermore, the input of the policy network is the image feature F generated by the source-channel joint encoder. The feature F first undergoes an average pooling operation and then is concatenated with the channel signal-to-noise ratio, and then undergoes multi-layer perceptron processing to generate a continuous value tensor, where each value represents the probability that the corresponding feature group is activated. The output of the multi-layer perceptron is processed by Gumbel-Softmax and then converted into a corresponding probability distribution, and finally, after one-hot to Thermal processing, the final corresponding binary mask W is obtained
[0020] Furthermore, the feature dynamic decoder performs zero-padding and reshaping operations on the input feature and then adopts the reverse operation in the feature dynamic encoder, that is, a two-dimensional convolutional layer and two cyclic operations with the structural sequence of SNR adaption and a residual network, and finally obtains an image feature with the same size as the original feature F
[0021] Furthermore, the image feature In the source-channel joint decoder, after performing deconvolution operations with the same number of convolutional layers as the source-channel joint encoder, followed by ReLU and sigmoid activations, and removing the pixels added by the reflection padding of the source-channel joint encoder, the transmitted and restored image is finally obtained.
[0022] Furthermore, the forward noise addition process is specifically as follows: z0 undergoes t rounds of forward noise addition operations to generate pure Gaussian noise z. t , where t is a preset value, and the forward noise addition operation is to add Gaussian white noise; the forward noise addition process is a Markov chain, and the mean of the noise in each round of forward noise addition is 0, with variance β t ∈ (0, 1), and finally is obtained, where ε represents the noise sampled from the standard Gaussian distribution, and α t = 1 - β t , represents the cumulative product of α i from the 1st round to the t-th round, α i is an intermediate variable, and i ∈ [1, t].
[0023] Furthermore, the reverse denoising process is specifically as follows: Under the guidance of z0, z t is estimated from the pure Gaussian noise z t-1 , and its posterior distribution is denoted as p(z t-1 | z t , z0), and it follows a Gaussian distribution where is the covariance matrix, and μ t (z t , z0) represents the mean, as shown in the following equation:
[0024]
[0025] The reverse denoising process uses the conditional set c, the pre-set denoising step size t d , the neural network parameters of the denoiser for predicting the noise , and the scaling factor γ as the parameters in the denoising process of the denoiser to denoise the pure noise z t . The training loss function of the Latent Diffusion Model is expressed as:
[0026]
[0027] where, represents the expected value of , represents the noise predicted by the denoiser, represents calculating the mean square error MSE.
[0028] Further, the final latent space feature z is inferred based on the latent space feature z0 final In this case, the image restored by the JSCC decoder is used as an intermediate guide to generate the intermediate state during the denoising process Then there is:
[0029]
[0030] Through Calculated to get Expressed as:
[0031]
[0032] where φ represents The alignment level between and f spa C l 、H l 、W l respectively represent the number of channels, height, and width of the features in the latent space. Finally Replacing z0, there is a posterior distribution Finally, the final latent space feature z is inferred final .
[0033] The beneficial effects of adopting the above technical solutions are as follows: The adaptive rate image semantic communication system integrating the latent diffusion model provided by the present invention realizes adaptive rate control and image visual effect optimization on the basis of the paradigm of the semantic image communication system; the adaptive rate control can adjust the image feature size transmitted by the semantic communication according to the image content and the channel signal-to-noise ratio to control the transmission rate CPP, which can effectively reduce the transmission resource consumption caused by simple picture content or poor channel signal-to-noise ratio conditions; the image visual effect optimization performs diffusion denoising processing on the transmitted picture based on the latent diffusion model, which can effectively improve the problem of visual effect distortion in image semantic communication. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a block diagram of a deep source-channel joint coding adaptive rate control image semantic communication system provided by an embodiment of the present invention;
[0035] Figure 2 It is a block diagram of a deep source-channel joint coding adaptive rate control image semantic communication system based on an LDM image visual optimization module provided by an embodiment of the present invention;
[0036] Figure 3 It is a block diagram of a source-channel joint decoder module provided by an embodiment of the present invention;
[0037] Figure 4 It is a block diagram of a feature dynamic encoder module provided by an embodiment of the present invention;
[0038] Figure 5 Block diagram of the signal-to-noise ratio adaptive module provided by the embodiment of the present invention;
[0039] Figure 6 Block diagram of the policy network structure provided by the embodiment of the present invention;
[0040] Figure 7 Block diagram of the feature dynamic decoder structure provided by the embodiment of the present invention;
[0041] Figure 8 Block diagram of the source-channel joint decoder structure provided by the embodiment of the present invention;
[0042] Figure 9 Block diagram of the conditional generator structure provided by the embodiment of the present invention;
[0043] Figure 10 Block diagram of the latent diffusion processing module structure provided by the embodiment of the present invention. Detailed implementation manners
[0044] The following combines the accompanying drawings and embodiments to further describe in detail the specific implementation manners of the present invention. The following embodiments are used to illustrate the present invention, but are not used to limit the scope of the present invention.
[0045] This embodiment proposes a deep source-channel joint coding image semantic communication system capable of realizing adaptive rate control, as Figure 1 shown. The system includes a transmitter module, a receiver module, and a channel module. The transmitter module includes a source-channel joint encoder, a policy network, and a feature dynamic encoder. The receiver module includes a source-channel joint decoder and a feature dynamic decoder. The channel module includes an additive white Gaussian noise (AWGN) channel and a Rayleigh channel. This embodiment also proposes an image visual effect optimization module based on LDM. This module includes a conditional generator and a latent diffusion processing module. The deep source-channel joint coding adaptive rate control image semantic communication system integrating the LDM image visual optimization module is as Figure 2 shown.
[0046] The block diagram of the transmitter module is as shown in the transmitter part of Figure 1 . The source-channel joint encoder mainly uses a convolutional neural network to process the input original image X and convert it into the corresponding feature F of the image. The corresponding block diagram of the source-channel joint encoder is as Figure 3 shown. The original image X is first processed through a reflection padding operation, and then processed through a cycle of operations with the structural order of a two-dimensional convolutional layer, a normalization layer, and a ReLU activation function, and finally outputs the corresponding feature F of the image.
[0047] The source-channel joint encoder processes the image feature F and divides the feature F into G s and G n into two parts. The structural block diagram of the source-channel joint encoder is as shown in Figure 4 . The image feature F output by the source-channel joint encoder is processed through two cyclic operations with the structural order of: residual network, signal-to-noise ratio adaptation, and then through two-dimensional convolutional layer convolution processing. Finally, through the reconstruction grouping operation, the second dimension of the input feature is divided into two groups s and n, and the corresponding two-part features G s and G n are obtained, where s and n are preset values. The signal-to-noise ratio adaptation module in the block diagram is designed and implemented based on the Multilayer Perceptron (MLP). The structural block diagram of the signal-to-noise ratio adaptation is as shown in Figure 5 . The signal-to-noise ratio adaptation first performs an average pooling operation on the input and then concatenates it with the channel signal-to-noise ratio. After being processed by the MLP, the multiplication and addition factors obtained are respectively multiplied and added to the original input. The role of the signal-to-noise ratio adaptation is to adjust the module input based on the multiplication and addition coefficients generated by the channel signal-to-noise ratio.
[0048] The policy network generates a binary mask W with the same two-dimensional size s as G s . The proposed adaptive rate control scheme is to control and adjust the image feature F transmitted in the channel through the binary mask W and G s generated by the above source-channel joint encoder. The structure of the policy network is as shown in p Figure 6 .
[0049] The input of the policy network is the image feature F generated by the source-channel joint encoder. The feature F first undergoes an average pooling operation and then is concatenated with the channel signal-to-noise ratio. After being processed by the MLP, a continuous value tensor is generated, and each value represents the probability that the corresponding feature group is activated. The output of the MLP is processed by Gumbel-Softmax and converted into a corresponding probability distribution. Finally, through one-hot to Thermal processing, the final corresponding binary mask W is obtained.
[0050] The binary mask W generated by the existing policy network, G s and G n are multiplied by G s and the binary mask W to control and filter the image features to be transmitted in G s to achieve adaptive rate control. The finally transmitted image feature F p can be expressed as F p =G s ×W + G n .
[0051] Feature F p is transmitted through the channel module, that is, random noise is added to Feature F p , and the feature after passing through the channel module is represented as Then where θ is a noise vector with the same size as Feature F p and the values follow a standard Gaussian distribution.
[0052] Feature F p needs to be restored to an image after being processed by the receiver module. The receiver module includes: a feature dynamic decoder and a source-channel joint decoder. The structure of the receiver module is as shown in the receiver part of Figure 1 .
[0053] The image feature after passing through the channel requires the feature dynamic decoder to process its feature size. After processing, the feature size is the same as that of the original image feature F. The structural block diagram of the feature dynamic decoder is as shown in Figure 7 . The feature dynamic decoder performs zero-padding and reshaping operations on the input feature , and then adopts the reverse operations in the source-channel joint encoder, that is, two-dimensional convolutional layers and two cyclic operations with the structure order of signal-to-noise ratio adaptation and residual network. Finally, an image feature with the same size as the original feature size F is obtained
[0054] Image feature is finally converted into the transmitted image after being processed by the source-channel joint decoder The source-channel joint decoder performs the opposite operations to the source-channel joint encoder. The structural block diagram of the source-channel joint decoder is as shown in Figure 8 . Image feature In the source-channel joint decoder, after performing the same number of transposed convolution operations as in the source-channel joint encoder, ReLU and sigmoid activations, and removing the pixels added by the reflection padding of the source-channel joint encoder, the finally transmitted and restored image is obtained
[0055] The above adaptive rate control image semantic communication scheme based on deep source-channel joint coding is regarded as a whole. During training, an end-to-end training method is adopted, and the representation of the loss function during the training process is as follows:
[0056]
[0057] where the first term represents the reconstruction loss between the original image X and the restored image after transmission , and the second term represents the transmission rate CPP, where μ is a set weight coefficient.
[0058] When the transmitted and restored image is in a channel with a relatively poor signal-to-noise ratio, the visual effect will be distorted. Therefore, an image visual effect optimization module based on the latent diffusion model is added on the basis of the adaptive rate control of the deep source-channel joint coding. The module consists of two parts: a conditional generator module and a latent diffusion processing module.
[0059] The conditional generator module consists of three parts: BLIPv2, CLIP encoder, and LDM encoder. The structure of the conditional generator is as Figure 9 shown. BLIPv2 is a pre-trained model for image-to-text, which extracts the text information T of the transmitted image . The text information T is encoded by the CLIP encoder to obtain its corresponding text feature f txt (multi-dimensional feature vector). The LDM encoder encodes the transmitted picture into the latent space, that is, the latent space feature f of the transmitted image is obtained spa . The finally obtained condition not only includes the latent space feature f spa , the text feature f txt , but also includes the channel signal-to-noise ratio (known). The condition set is represented by c, that is, c = {f spa , f txt , λ}, where λ is the known channel signal-to-noise ratio.
[0060] The latent diffusion processing module performs reverse diffusion denoising operations with the condition set c as a parameter. The structure of the latent diffusion processing module includes: several denoisers and an LDM decoder, and its structural block diagram is as Figure 10 shown.
[0061] The training of the latent diffusion model (Latent Diffusion Model) includes a forward noise addition process and a reverse denoising process. The original image generates a latent space feature z0 through the VAE encoder.
[0062] Forward noise addition process: z0 undergoes t (preset value) rounds of forward noise addition operations (adding Gaussian white noise) to generate pure Gaussian noise z t . The forward noise addition process is a Markov chain. The noise mean of each round of forward noise addition is 0, and the variance β t ∈ (0, 1). Finally, is obtained, where ε represents the noise sampled from the standard Gaussian distribution, and α t = 1 - β t , represents the cumulative product of α i from the first round to the t-th round, and α i is an intermediate variable, i ∈ [1, t].
[0063] Reverse denoising process: Under the guidance of z0, estimate z from pure Gaussian noise z t among which t-1 , and its posterior distribution can be expressed as p(z t-1 |z t , z0), and follows a Gaussian distribution where is the covariance matrix, and μ t (z t , z0) represents the mean:
[0064]
[0065] The reverse denoising process uses the conditional set c, the pre-set denoising step t d , the neural network parameters for the denoiser to predict noise and the scale factor γ as parameters in the denoising process of the denoiser to denoise the pure noise z t . The training loss function of the latent diffusion model can be expressed as:
[0066]
[0067] where represents the expected value of , represents the noise predicted by the denoiser, represents calculating the mean square error MSE.
[0068] To ensure that when the trained latent diffusion model is optimized, its optimization effect does not deviate from the original image X, the image restored by the source-channel joint decoder is used as an intermediate guide to generate the intermediate state during the denoising process. Then there is:
[0069]
[0070] Through calculation, is expressed as:
[0071]
[0072] where φ represents the alignment level between spa and f l , C l , H l , W Finally, replace z0, then there is the posterior distribution Finally, through inference, the final latent space feature z final is obtained, and z finalConvert it into a visually optimized image X through the LDM decoder final 。
[0073] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. An adaptive rate image semantic communication system integrating a latent diffusion model, characterized in that: The system includes a transmitter module, a receiver module, a channel module, and an image visual effect optimization module based on the latent diffusion model; The transmitter module includes a source-channel joint encoder, a policy network, and a feature dynamic encoder; the source-channel joint encoder processes the input original image X using a convolutional neural network and converts it into the corresponding feature F of the image; the feature dynamic encoder processes the image feature F and divides the feature F into two parts of features G, where s and n are preset values; the policy network generates a binary mask W with the same two-dimensional size s as G, and controls and adjusts the image feature F transmitted in the channel through the binary mask W and the feature G s and G n ; the image features to be transmitted in G are controlled and screened by multiplying G by the binary mask W to achieve adaptive rate control; the finally transmitted image feature F is expressed as F s = G s ×W + G p ; s ; s p p s n The channel module includes an additive white Gaussian noise channel and a Rayleigh channel; Feature F p After transmission through the channel module, that is, in Feature F p random noise is added, and the feature after passing through the channel module is expressed as Then where θ is a noise vector with the same size as Feature F p and the values follow a standard Gaussian distribution; Feature F p It needs to be restored to an image after being processed by the receiver module, which includes a feature dynamic decoder and a source-channel joint decoder; the image features after passing through the channel The feature dynamic decoder needs to process its feature size, and the processed feature size is the same as the size of the original image feature F; the image features Finally, it is converted into the transmitted image after being processed by the source-channel joint decoder The source-channel joint decoder performs the opposite operation to the source-channel joint encoder; The image visual effect optimization module based on the latent diffusion model is used to optimize the image after transmission recovery when the channel signal-to-noise ratio is poor, avoiding the phenomenon of visual effect distortion. The image visual effect optimization module based on the latent diffusion model includes a conditional generator module and a latent diffusion processing module; The conditional generator module includes BLIPv2, a CLIP encoder, and an LDM encoder; BLIPv2 is a pre-trained model for image-to-text, which extracts the text information T of the transmitted image The text information T is encoded by the CLIP encoder to obtain its corresponding text feature f txt , which is a multi-dimensional feature vector; the LDM encoder encodes the transmitted picture into the latent space, that is, the latent space feature f of the transmitted image is obtained spa ; the finally obtained conditional set is represented by c, that is, c = {f spa , f txt , λ}, where λ is the known channel signal-to-noise ratio; The latent diffusion processing module performs reverse diffusion denoising operations with the condition set c as a parameter. The latent diffusion processing module includes several denoisers and an LDM decoder. Among them, the denoiser is a UNet structure based on the stable diffusion model, and the LDM decoder decodes the image latent space representation processed by the denoiser to obtain the image X with optimized visual effects. final ; Latent diffusion model training includes a forward noise addition process and a reverse denoising process. The original image generates latent space feature z0 through the VAE encoder; the final latent space feature z is inferred based on the latent space feature z0 final , z final is converted into an image X with optimized vision through the LDM decoder final .
2. An adaptive rate image semantic communication system integrating a potential diffusion model according to claim 1, characterized in that: In the source-channel joint encoder, the original image X is first processed by a reflection padding operation, and then processed through a cyclic operation with the structural sequence of a two-dimensional convolutional layer, a normalization layer, and a ReLU activation function for several times, and finally the feature F corresponding to the image is output.
3. The adaptive rate image semantic communication system integrating a potential diffusion model according to claim 2, wherein: The specific processing procedure of the feature dynamic encoder is as follows: the image feature F output by the source-channel joint encoder is processed through two cyclic operations with the structural order of residual network and signal-to-noise ratio adaptation, and then through two-dimensional convolutional layer convolution processing. Finally, through the reconstruction grouping operation, the second dimension of the input feature is divided into two groups, s and n, and the corresponding two-part features G are obtained. s and G n ; The signal-to-noise ratio adaptation is realized based on the multi-layer perception design, and the input of the module is adjusted by the multiplication and addition coefficients generated based on the channel signal-to-noise ratio; For the signal-to-noise ratio adaptation, the input is first subjected to an average pooling operation and then concatenated with the channel SNR. After being processed by the MLP, the multiplication and addition factors obtained are respectively multiplied and added to the original input.
4. An adaptive rate image semantic communication system integrating a potential diffusion model according to claim 3, characterized in that: The input of the policy network is the image feature F generated by the source-channel joint encoder. The feature F is first subjected to an average pooling operation and then concatenated with the channel signal-to-noise ratio, and then processed through a multi-layer perception to generate a continuous value tensor, where each value represents the probability that the corresponding feature group is activated. The output of the multi-layer perception is processed by Gumbel-Softmax and converted into a corresponding probability distribution, and finally the final corresponding binary mask W is obtained through one-hot to Thermal processing.
5. An adaptive rate image semantic communication system integrating a potential diffusion model, characterized in that: The feature dynamic decoder performs zero-padding and reshaping operations on the input features After that, it adopts the reverse operations in the feature dynamic encoder, namely the cyclic operations of a two-dimensional convolutional layer and two structures in the order of signal-to-noise ratio adaptation and residual network, and finally obtains image features with the same size F as the original feature size 6. The adaptive rate image semantic communication system integrating a latent diffusion model according to claim 5, characterized in that: The described image features In the source-channel joint decoder, after performing deconvolution operations with the same number of convolutional operations as the source-channel joint encoder, along with ReLU and sigmoid activations, and removing the pixels added by the reflection padding of the source-channel joint encoder, the transmitted and restored image is finally obtained 7. An adaptive rate image semantic communication system integrating a potential diffusion model, characterized in that: The specific process of forward noise addition is as follows: z0 undergoes a pre-round forward noise addition operation to generate pure Gaussian noise z t , where t is a preset value, and the forward noise addition operation is to add white Gaussian noise; the forward noise addition process is a Markov chain. The mean of the noise added in each round of forward noise addition is 0, and the variance is β t ∈(0,1), and finally we get where ε represents the noise sampled from the standard Gaussian distribution, and α t =1-β t , represents the cumulative product of α from the 1st round to the t-th round i , α i is an intermediate variable, and i ∈ [1, t].
8. An adaptive rate image semantic communication system integrating a potential diffusion model, characterized in that: The reverse denoising process is specifically as follows: under the guidance of z0, from pure Gaussian noise z t Medium estimate z t-1 , whose posterior distribution is expressed as p(z t-1 |z t ,z0), and obeys Gaussian distribution in is the covariance matrix, μ t (z t ,z0) represents the mean, as follows: The reverse denoising process uses the conditional set c, the pre-set denoising step size t d , the neural network parameters of the denoiser for predicting noise and the scaling factor γ as parameters in the denoising process of the denoiser to denoise the pure noise z t The training loss function of the Latent Diffusion Model is expressed as: Among them, denotes the expected value of denotes the noise predicted by the denoiser, denotes the calculation of the mean square error MSE.
9. An adaptive rate image semantic communication system integrating a potential diffusion model, characterized in that: The final latent space feature z is inferred based on the latent space feature z0 final Among them, the image restored by the JSCC decoder is used as an intermediate guidance to generate the intermediate state during the denoising process Then there is: By Calculated to obtain Expressed as: where φ represents the alignment level with f spa , and C l , H l , and W l represent the number of channels, height, and width of the features in the latent space respectively; finally replaces z0, and then there is the posterior distribution Finally, the final latent space feature z is inferred final .
Citation Information
Cited By
Robust semantic communication method based on antagonism purification
CN120930520A
Communication device and method in communication device
JP7874810B1