Unmanned aerial vehicle semantic communication system based on diffusion model

By introducing image enhancement technology based on diffusion model in the drone semantic communication system, combined with channel gain calculation of the channel estimation module, the problem of poor signal reconstruction effect in the existing system in a low signal-to-noise environment is solved, and more stable and adaptable image transmission is achieved.

CN120017758APending Publication Date: 2025-05-16BEIHANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510010976.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing drone semantic communication systems have low signal-to-noise and severe signal scattering, and the IoT receiver's poor reconstruction of data is mainly due to the overfitting of the channel model in the training stage and lack of generalization ability to real-life scenarios.

Method used

Using a drone semantic communication system based on diffusion model, abstract semantic information is extracted through a joint encoding module and converted into signal data that can be directly transmitted. The channel estimation module calculates the gain value of the current channel, and the diffusion module enhances image based on the pre-trained diffusion model and channel estimation value.

Benefits of technology

It improves the stability and adaptability of the semantic communication system in the signal transmission process, enhances the quality of image transmission, and can better cope with complex and variable channel conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017758A_ABST
    Figure CN120017758A_ABST
Patent Text Reader

Abstract

The invention relates to the field of unmanned aerial vehicle semantic communication, in particular to an unmanned aerial vehicle semantic communication system based on a diffusion model, which is characterized in that a joint coding module is arranged to extract abstract semantic information and convert the abstract semantic information into signal data capable of being directly transmitted for transmission, and semantic decoding is performed through a joint decoder to reconstruct image data; in a signal data transmission period, a channel estimation module is adopted, a current channel state is taken as a basis, a channel estimation value is calculated, an image is enhanced adaptively through a diffusion module in combination with the channel estimation value, and then the image is transmitted to an Internet of Things receiver, so that the channel adaptive capability is realized on the basis of a semantic communication system, and the signal transmission efficiency is improved. And the signal transmission stability of the semantic communication system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of unmanned aerial vehicle semantic communication, and in particular to a unmanned aerial vehicle semantic communication system based on a diffusion model. Background Art

[0002] As a new generation of IoT devices, drones can effectively build line-of-sight channels and improve the quality of wireless communications in urban scenarios by leveraging their advantages in coverage and mobility. By equipping drones with semantic communication systems, they can further provide high-quality, low-latency intelligent services to ground users.

[0003] For example, Chinese patent publication number: CN118470581A discloses a cross-modal semantic communication method and system for drones for target detection tasks, which involves the field of communication technology, including collecting training data, building a data set, and annotating the category and location information of the data set; building and pre-training a target detection model under channel-free conditions, and obtaining the network parameters of the pre-trained model; building a channel environment, and designing semantic communication-related channel encoding and decoding, and building an overall model after adding the channel environment; migrating the network parameters of the pre-trained model, and performing migration training on the obtained parameters, and training the overall model after adding the channel environment; judging whether the network training is completed, inputting the test set data into the network, and outputting the detection results. The invention uses semantic information of different scales, comprehensively utilizes multi-modal information for complementarity, and improves detection accuracy. It is also an effective method for drones and other Internet devices to offload operations to edge servers when facing complex channel environments for target detection tasks.

[0004] However, although drones provide a good line-of-sight channel, the reconstruction of data by the IoT receiver is not satisfactory when the signal-to-noise ratio is low and the signal scattering is severe.

[0005] The root cause is that semantic communication systems usually overfit the channel model in the training phase and lack the ability to generalize to real-world scenarios. Image enhancement in existing semantic communication systems mainly relies on generative artificial intelligence, which uses the model's image generation ability to compensate for image distortion caused by communication noise. Generative adversarial networks (GANs) have performed well in the field of image generation, and many studies have integrated GANs as image enhancement modules into the semantic communication framework. However, the adversarial loss of GANs makes its training process highly uncertain, and improper hyperparameter selection may cause model collapse and difficulty in convergence. To solve this problem, variational autoencoders (VAEs) optimize the shortcomings of GANs in terms of training stability. VAEs introduce implicit probability spaces based on the autoencoder architecture and achieve controllable image generation by optimizing the divergence. However, VAEs have limitations in image generalization capabilities, and the limited model capacity makes them mainly suitable for specific types of images, making it difficult to cope with complex and diverse image enhancement requirements. In addition, VAE image generation relies on the learned probability space, making it difficult to make adaptive image enhancements based on current channel conditions.

[0006] In summary, in the semantic communication architecture, it is particularly important to implement an image enhancement unit that is stable, controllable, highly generalizable, and can make corresponding adjustments according to channel conditions. Summary of the invention

[0007] To this end, the present invention provides a UAV semantic communication system based on a diffusion model to overcome the problem that the semantic communication system in the prior art is not adaptable to wireless channels. It usually relies on the channel model assumptions used by the model in the training phase and is difficult to cope with the complex and changeable channel conditions in real scenarios.

[0008] To achieve the above object, the present invention provides a UAV semantic communication system based on a diffusion model, which comprises:

[0009] A joint encoding module, which is used to receive the original image data input by the UAV sender to extract abstract semantic information and convert it into signal data that can be directly transmitted, wherein the conversion includes mapping the abstract semantic information into a complex domain to adapt to signal transmission conditions;

[0010] A channel estimation module, which is used to determine the channel gain in the current communication environment based on the pilot signal to determine a channel estimation value;

[0011] A joint decoding module, which is used to receive the signal data to perform semantic decoding and obtain reconstructed image data, wherein the semantic decoding includes mapping the signal data from a complex domain to abstract semantic information and decoding the abstract semantic information;

[0012] The diffusion module is used to receive the channel estimation value and the reconstructed image data as input, and perform image enhancement on the reconstructed image data based on a pre-trained diffusion model.

[0013] Furthermore, the joint encoding module is used to extract abstract semantic information including:

[0014] For converting the original image data into an image embedding block through a convolutional neural network;

[0015] Used to extract abstract semantic information from each of the image embedding blocks through a plurality of Transformer blocks.

[0016] Furthermore, the channel estimation module is used to construct a linear relationship between the channel estimation value and the signal received by the IoT receiver, and set a minimum root mean square error term to determine an optimal linear coefficient matrix;

[0017] The optimal linear coefficient matrix is ​​calculated based on the partial derivative of the minimized root mean square error term.

[0018] Furthermore, the channel estimation module is used to determine the channel estimation value, including:

[0019] for determining a preliminary representation of a channel estimation matrix according to the optimal coefficient matrix;

[0020] replacing corresponding matrix items in the preliminary representation according to the mean matrix to obtain the channel estimation matrix, and determining the channel estimation matrix as the channel estimation value;

[0021] The obtained channel estimation matrix includes constants based on the constellation diagram and the current channel signal-to-noise ratio.

[0022] Furthermore, the diffusion model predicts the noise of the reconstructed image data according to the learnable parameters during the diffusion stage.

[0023] The learnable parameters are related to the number of diffusions and the current image, and the current image is an image obtained under the current number of diffusions.

[0024] The diffusion module determines a description of signal transmission for a wireless channel according to a channel gain matrix, including:

[0025] to obtain reconstructed image data output by the joint decoding module;

[0026] The description is determined according to the reconstructed image data, a channel gain matrix and additive Gaussian noise.

[0027] Furthermore, the diffusion module is used to use the likelihood probability gradient of the current image of the diffusion model in the sampling stage as a guide item.

[0028] Furthermore, the diffusion module is used to introduce a classifier network to predict the category of the current image, and to construct a guide term function based on the prediction result to determine a complete noise prediction for the sampling stage.

[0029] Furthermore, the diffusion module receives the channel estimation value and the reconstructed image data as input, and determines the description of the noise term including:

[0030] To determine the projection of the original data in the column space and the null space of the channel gain matrix, and to determine the estimated value for the original image in combination with the channel estimated value;

[0031] It is used to introduce the influence of channel noise during the transmission process, optimize the estimated value for the original image, and obtain the optimized estimated value;

[0032] A description of the noise term is determined according to the optimized estimate.

[0033] Furthermore, the joint encoding module is used to extract abstract semantic information based on a pre-trained semantic communication model.

[0034] Compared with the prior art, the beneficial effect of the present invention lies in that a joint coding module is provided to extract abstract semantic information, and convert it into signal data that can be directly transmitted for transmission, and semantic decoding is performed through a joint decoder to reconstruct image data. During the transmission of signal data, a channel estimation module is used to calculate the channel estimation value based on the current channel state, and the image is enhanced by the diffusion module combined with the channel estimation value adaptively, and then transmitted to the Internet of Things receiver. Furthermore, channel adaptability is achieved on the basis of the semantic communication system, thereby improving the stability of the semantic communication system for signal transmission.

[0035] In particular, the present invention calculates the channel estimation value through the channel estimation module, and subsequently uses the channel estimation value and the reconstructed image data as the input of the diffusion module. The channel estimation value is determined according to the channel gain of the current channel, and can reflect the relevant characteristics of the channel, so that the subsequent diffusion module can adapt to the current channel when performing image enhancement, thereby improving the stability of the semantic communication system during the signal transmission process. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a structural schematic diagram of a UAV semantic communication system based on a diffusion model according to an embodiment of the invention;

[0037] Figure 2 A schematic diagram of a diffusion model theoretical framework of an embodiment of the invention;

[0038] Figure 3 A schematic diagram of a diffusion model network architecture of an embodiment of the invention;

[0039] Figure 4 A schematic diagram of a residual unit network architecture according to an embodiment of the invention;

[0040] Figure 5 A schematic diagram of a residual downsampling unit network architecture according to an embodiment of the invention;

[0041] Figure 6 A schematic diagram showing a comparison of peak signal-to-noise ratio performances of the present invention and a conventional algorithm under different signal-to-noise ratios under a Rayleigh channel of an embodiment of the invention;

[0042] Figure 7 Schematic diagram of comparing the multi-scale structural similarity performance of the present invention and the traditional algorithm under different signal-to-noise ratios under the Rayleigh channel of an embodiment of the present invention. DETAILED DESCRIPTION

[0043] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0044] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0045] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the drawings. This is merely for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0046] In addition, it should be noted that in the description of the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0047] See also Figure 1 As shown, it is a structural schematic diagram of a UAV semantic communication system based on a diffusion model according to an embodiment of the present invention. The UAV semantic communication system based on a diffusion model according to an embodiment of the present invention includes:

[0048] A joint encoding module, which is used to receive the original image data input by the UAV sender to extract abstract semantic information and convert it into signal data that can be directly transmitted, wherein the conversion includes mapping the abstract semantic information into a complex domain to adapt to signal transmission conditions;

[0049] A channel estimation module, which is used to determine the channel gain in the current communication environment based on the pilot signal to determine a channel estimation value;

[0050] A joint decoding module, which is used to receive the signal data to perform semantic decoding to obtain reconstructed image data, wherein the semantic decoding includes mapping the signal data from a complex domain to a real domain;

[0051] The diffusion module is used to receive the channel estimation value and the reconstructed image data as input, and perform image enhancement on the reconstructed image data based on a pre-trained diffusion model.

[0052] Specifically, the joint encoding module is used to extract abstract semantic information including:

[0053] For converting the original image data into an image embedding block through a convolutional neural network;

[0054] Used to extract abstract semantic information from each of the image embedding blocks through a plurality of Transformer blocks.

[0055] It can be understood that, for the signal data sent by the joint coding module, the signal is transmitted through the wireless channel. Generally, the process can be described as follows:

[0056]

[0057] In formula (1), H represents the channel gain matrix, z represents the additive Gaussian noise, represents the signal received by the receiver, and e represents the signal data sent by the joint coding module.

[0058] Specifically, benefiting from the excellent channel conditions provided by the UAV, the channel modeling can be simplified to a certain extent and implemented using the linear minimum mean square error. In the implementation, under the influence of channel gain and additive Gaussian noise, it can be obtained.

[0059] y=Hu+z. (2)

[0060] In formula (2), u represents the pilot signal matrix, and y represents the signal received by the receiving end under the influence of channel noise.

[0061] Specifically, the channel estimation module is used to construct a linear relationship between the channel estimation value and the signal received by the IoT receiver, and set the minimum root mean square error term to determine the optimal linear coefficient matrix;

[0062] The optimal linear coefficient matrix is ​​calculated based on the partial derivative of the minimized root mean square error term.

[0063] Furthermore, in one possible implementation, the channel estimation module constructs a channel estimation value The linear relationship between the signal y received by the receiving end under the influence of channel noise and the pilot signal is: Then the optimization objective can be described as minimizing the following mean square error term,

[0064]

[0065] In formula (3), superscript (·) H represents the conjugate transpose of a matrix, stands for the probability expectation operator.

[0066] The optimal solution can be calculated by taking the partial derivative of the mean square error term L(A), as shown below:

[0067]

[0068] In formula (4),

[0069] Based on this, the optimal linear coefficient matrix is ​​calculated based on the partial derivative of the minimized root mean square error term, and the obtained linear coefficient matrix is,

[0070]

[0071] Specifically, the channel estimation module is used to determine the channel estimation value including:

[0072] for determining a preliminary representation of a channel estimation matrix according to the optimal coefficient matrix;

[0073] replacing corresponding matrix items in the preliminary representation according to the mean matrix to obtain the channel estimation matrix, and determining the channel estimation matrix as the channel estimation value;

[0074] The obtained channel estimation matrix includes constants based on the constellation diagram and the current channel signal-to-noise ratio.

[0075] In one possible implementation, the preliminary representation of the channel estimation matrix determined by the channel estimation module according to the optimal coefficient matrix is:

[0076]

[0077] In formula (6), σ z represents the standard deviation of Gaussian noise z, It can be obtained through historical statistics or predefined channel models.

[0078] It can be understood by those skilled in the art that H u is a square matrix, in which the diagonal elements represent the power of the pilot signal, and the other elements are the correlation values ​​of the pilot signal and the delayed signal, which can approximately satisfy the standard Gaussian distribution. Therefore,

[0079] Using the mean matrix Replace u H u, reducing the computational complexity of matrix inversion at the expense of accuracy, and finally obtaining the channel estimation matrix

[0080]

[0081] In formula (7), κ is a constant that depends on the constellation diagram, and SNR represents the current channel signal-to-noise ratio.

[0082] It is understood that in implementation, the channel estimation matrix As the channel estimation value, it can be used as one of the inputs of the diffusion model to achieve image enhancement.

[0083] It can be understood by those skilled in the art that the diffusion model includes a diffusion stage and a sampling stage, wherein the diffusion stage includes that the diffusion model gradually converts the complete image into randomly distributed noise by adding noise to the image, and in the diffusion stage, the diffusion model will learn to predict the added noise according to the state of the current image and the iteration round. The sampling stage includes that the diffusion model gradually denoises the noisy image to restore it to the original image. After the diffusion stage is over, the diffusion model has the ability to predict noise. The present invention will use the ability of the diffusion model to optimize the irreversible impact of the wireless channel on the image transmission process through the sampling stage to achieve image enhancement.

[0084] It should be noted that the noise added during the diffusion model training process is only added to allow the model to learn how to predict noise and restore the image. This noise is not channel noise.

[0085] However, in actual communication systems, the diffusion model is used to predict the noise caused by channel transmission.

[0086] Therefore, in the subsequent explanation of the diffusion model principle, noise refers to the noise added by model training.

[0087] See also Figure 2 As shown, it is a schematic diagram of the theoretical framework of the diffusion model of an embodiment of the present invention. In implementation,

[0088] The original image input into the diffusion model can be expressed as x0. It can be understood that the original image is the reconstructed image data output by the joint decoding module. If the subscript represents the number of diffusion times, the image obtained by the tth noise diffusion can be expressed as x t , assuming that the diffusion process contains T steps, the final diffusion result can be expressed as x T .

[0089] Since the diffusion process is to add Gaussian noise to the original image, this process can be described in the form of probability, and q is defined to represent x t-1 Under the condition x t The probability distribution of

[0090]

[0091] In formula (8), N represents Gaussian distribution, Represents x t The change of the mean is The variance is β t Gaussian distribution of I. β t is a constant related to the diffusion number t, satisfying β t <β t+1 ∈(0,1), I represents the identity matrix.

[0092] Please continue reading Figure 2 As shown, assuming that the diffusion process satisfies the Markov assumption, then q(x t |x t-1 ,x0)=q(x t |x t-1 ), the diffusion process can be expressed as

[0093]

[0094] in ε represents the random noise added by the diffusion process, satisfying ε~N(0,I).

[0095] It is understandable that since the diffusion process satisfies the Gaussian distribution, it can also be assumed that the sampling stage satisfies the Gaussian distribution, that is, according to x t Infer x t- 1. Define p to represent x t Under the condition x t-1 The probability distribution of , then we have,

[0096] p θ (x t-1 |x t )~N(x t-1 ;μ θ (x t ,t),σ θ (x t,t)),(10)

[0097] In formula (10), θ represents the learnable parameters of the diffusion model, μ θ (x t ,t) and σ θ (x t , t) represent the mean and variance of the Gaussian distribution in the sampling stage, which are learnable parameters and are related to the current image x t , is related to the number of diffusion times t.

[0098] Specifically, it can be understood that the diffusion model can predict the current image x with the help of parameters θ t The noise ε, and from x t Subtract ε from the image to achieve image denoising.

[0099] The sampling phase is to obtain T The original image x0 is obtained, and the generation of x0 can be described by the maximum likelihood probability, expressed as P(x0).

[0100] Since the diffusion model optimizes the confidence lower bound of the likelihood probability, the optimization objective can be expressed as

[0101]

[0102] In formula (11), x T :x0 represents from x T The sampling process to x0, x1:x T Represents from x1 to x T The diffusion process. P(x T :x0) represents the likelihood probability of the sampling process, which can be described as follows because it satisfies the Markov distribution:

[0103]

[0104] In formula (12), x T can be regarded as sampling from Gaussian noise, can be regarded as deterministic probability, and has nothing to do with the parameter θ. Therefore, the optimization objective can be simplified as follows:

[0105]

[0106] In formula (13), KL(·||·) represents the Kullback-Leibler divergence, which describes the distance between probability distributions.

[0107] It is understandable that the simplified optimization objective p θ (x0|x1) is actually independent of the parameter θ, because x0 is known and the first diffusion step is deterministic, and x0 can be obtained directly from x1.

[0108] Therefore, the loss function can be described as

[0109]

[0110] Using the Bayesian formula, we can get:

[0111]

[0112] The loss function can be simplified as,

[0113]

[0114] Ignore variance and constant term coefficient, the diffusion process can be expressed as

[0115]

[0116] In formula 16, ε θ represents the model's prediction of noise, θ is a learnable parameter, and the model receives the current image and diffusion times t as input, predicting the noise ε θ After calculating the mean square error of the original noise ε, the gradient of the parameter θ is calculated to update the model. Accordingly, the sampling process can be expressed as

[0117]

[0118] where ε θ is the noise in the model prediction.

[0119] Since the sampling process satisfies the Markov assumption, T rounds of sampling are required. Now consider that the sampling process conforms to a non-Markov process. Also assume that the sampling process conforms to a Gaussian distribution and define ξ to represent the initial state x0 and x t Under the condition x t-1 The probability distribution of , then we have

[0120]

[0121] in and ψ represent unknown parameters, the sampling process can be expressed as Since the diffusion process remains unchanged, there is still So we can substitute it into

[0122]

[0123] By using the undetermined coefficient method, the first parameter can be obtained by comparing the diffusion process and the value of the second parameter ψ includes,

[0124]

[0125] Therefore, the sampling process can be modified as follows:

[0126]

[0127] Since the diffusion process remains unchanged, x0 can be obtained by the following formula:

[0128]

[0129] Under the non-Markov assumption, the sampling process can be a subset of the diffusion process. The sampling process is defined as τ, then

[0130]

[0131] Specifically, in order to improve the enhancement effect of the diffusion model on the image, the diffusion module of the present invention uses the likelihood probability gradient of the current image of the diffusion model in the sampling stage as a guide item.

[0132] In some possible implementations, the boot item is described as,

[0133]

[0134] Please read further Figure 2 As shown, the diffusion module is used to introduce a classifier network to predict the category of the current image, and to construct a guide term function based on the prediction result to determine a complete noise prediction for the sampling stage.

[0135] In some possible implementations, an additional classifier network is introduced Predict the current image x t If the predicted category label is l, the process can be expressed as

[0136] in Represents the parameters of the classifier network. Combined with x t The image generation process.

[0137] Then, the complete guide term function can be expressed as,

[0138]

[0139] Therefore, the complete noise prediction at the sampling stage can be expressed as,

[0140]

[0141] It is understandable that the noise prediction here It means that during the model training phase, the model is trained based on the current image x t And the noise in the image is predicted at iteration round t.

[0142] The diffusion model achieves image restoration by predicting the noise of the current image and removing it. Therefore, by using this feature, the diffusion model can predict the impact of the channel on the image and achieve image enhancement.

[0143] Specifically, for the specific use of the diffusion model, that is, we apply the diffusion model to the semantic communication system by combining the current channel estimate to achieve image enhancement.

[0144] Specifically, the diffusion module receives the channel estimation value and reconstructing image data s as input to achieve image enhancement;

[0145] The diffusion module determines a description of signal transmission for a wireless channel according to a channel gain matrix, including:

[0146] to obtain reconstructed image data output by the joint decoding module;

[0147] The description is determined based on the reconstructed image data, the channel gain matrix and the additive Gaussian noise. In some possible implementations, under the condition that only the channel effect is considered, the image transmission can be described as:

[0148] s=Hx+z.(29)

[0149] For the channel gain matrix H, if H * represents the pseudo-inverse matrix of H, then HH * H=H,

[0150] The diffusion module receives the channel estimate and the reconstructed image data as input, and determines a description for the noise term including:

[0151] To determine the projection of the original data in the column space and the null space of the channel gain matrix, and to determine the estimated value for the original image in combination with the channel estimated value;

[0152] It is used to introduce the influence of channel noise during the transmission process, optimize the estimated value for the original image, and obtain the optimized estimated value;

[0153] A description of the noise term is determined according to the optimized estimate.

[0154] The channel transmission process can be further expressed as,

[0155]

[0156] In formula (30), represents the projection of the original data x in the column space of the channel gain matrix H, that is, the observable data, and represents the projection of x onto the null space of H, which can be viewed as the image distortion caused by data transmission.

[0157] In the channel estimation part, the channel estimation value can be used Instead of H, the estimated value of the original image can be expressed as,

[0158]

[0159] In formula (31), Ω t is the coefficient matrix associated with the current iteration number t.

[0160] Therefore, the sampling process can be further optimized as follows:

[0161]

[0162] Considering the influence of channel noise z during transmission, can be further transformed into,

[0163]

[0164] Where Ω t =VΛ t V H Represents the pair matrix Ω t The singular value decomposition of Ω t It has the good properties of a diagonal matrix. In the result of singular value decomposition, Ω t The left and right singular vector matrices of are equal and can be represented by V, while Λ t Is included Ω t The diagonal matrix of eigenvalues. Also consider The singular value decomposition result of Where V and U represent The left singular vector matrix and the right singular vector matrix of , Σ is a diagonal matrix.

[0165] Therefore the noise term can be described as,

[0166]

[0167] Specifically, the present invention is specifically manifested in two points for the image enhancement step. The first is the consideration of channel gain, that is, the channel gain matrix H, which is reflected in formula 29 and formula 30. Based on the decomposition result of formula 30, formula 33 can be further obtained. The second point is the consideration of channel noise, that is, the channel noise z in formula 29. Formula 33 is obtained while considering the channel gain matrix H.

[0168] It can be understood that for Formula 33, the last item after the plus sign is a description of the noise z, so Formula 34 solves the problem of z. The specific method is to solve the form of the noise distribution of z and sample from it to realize the noise estimation, which will not be repeated here.

[0169] Specifically, the semantic communication model needs to be trained during implementation, and the physical channel can be used as a simulation of the wireless communication environment without participating in the back-propagation process of the model.

[0170] In one implementation, the training process for the semantic communication model includes:

[0171] Step S01, setting input, including data set I, batch size B1, and learning rate η1;

[0172] Step S02, setting output, including, a pre-trained semantic communication model;

[0173] Step S03, repeating the semantic communication model training step S030 until the model converges;

[0174] Wherein, the model training step S030 includes:

[0175] Step S031, select a batch of data containing B1 images from data set I

[0176] Step S032: one by one, the i-th image x i Extracting abstract semantic information through joint encoding module processing i , the abstract semantic information is transmitted through the channel and reaches the joint decoding module to obtain The joint decoding module performs semantic decoding to restore the image

[0177] Step S033, calculate the loss using formula (35),

[0178]

[0179] Step S034, using the gradient descent method to solve the gradient value of the loss function, and using this as the descent direction to update the model parameters.

[0180] Specifically, see Figure 3 , Figure 4 as well as Figure 5 , which are respectively a schematic diagram of a diffusion model network architecture, a schematic diagram of a residual unit network architecture, and a schematic diagram of a residual downsampling unit network architecture of an embodiment of the invention,

[0181] In implementation, the diffusion model architecture is as follows Figure 3 As shown, a symmetrical Unet architecture is adopted.

[0182] Among them, the first data processing module is composed of two residual modules and one residual downsampling module. The module is repeated three times, and the result is input into the second data processing module. The second data processing module is composed of two residual modules, two convolutional attention modules and a residual downsampling module. The result obtained by the second data processing module enters the residual module and then the convolutional attention module.

[0183] Please continue reading Figure 3 As shown, the diffusion model connects the left-right symmetrical network architecture through the convolutional attention module, and each data processing module performs a direct connection operation on the sampled data.

[0184] The difference between the residual module and the residual downsampling module is that an additional downsampling module is provided in the residual downsampling module to modify the size of the image.

[0185] At the same time, each data processing includes a time embedding operation. The time embedding layer contains a SiLU activation layer and a linear layer, and the current iteration number t is added to the output of the first convolutional layer.

[0186] In implementation, the diffusion model needs to be trained;

[0187] In one implementation, the training process for the diffusion model includes:

[0188] Step S11, set the input, including data set I, diffusion sequence {0,…,T}, parameter sequence {β0,…,β T}, learning rate η2, where β T is a constant related to the diffusion number T.

[0189] Step S12, setting output, including, a pre-trained diffusion model;

[0190] Step S13, repeating the diffusion model training step S130,

[0191] Wherein, the diffusion model training step S130 includes:

[0192] S131, initializing network parameters;

[0193] S132, randomly sampling image x0 from data set I as training data;

[0194] S133, randomly sampling diffusion times from the diffusion sequence {0,…,T};

[0195] S134, calculate the loss function according to formula (17), and find the gradient of the model parameter θ to update the model.

[0196] Specifically, it is necessary to combine the trained semantic communication model and the diffusion model to train the UAV semantic communication system model based on the diffusion model, including:

[0197] Step S21, set the input, including batch size B2, sampling sequence τ, pre-trained semantic communication model, pre-trained diffusion model, classifier network Learning rate η3;

[0198] Step S22, setting output, including a UAV semantic communication system model based on a diffusion model;

[0199] Step S23, repeating the UAV semantic communication system model training step S230 based on the diffusion model until the model converges;

[0200] Among them, the process of training the UAV semantic communication system model based on the diffusion model includes:

[0201] Step S231, select a batch of data containing B2 images from data set I As training data;

[0202] Step S232: the i-th image x is i The abstract semantic information is extracted by the joint encoding module to obtain i , the abstract semantic information is transmitted through the channel and reaches the joint decoding module to obtain The joint decoding module performs semantic decoding to restore the image The channel estimation module obtains the current channel estimation value Diffusion module combination and i Enhance the image to get the final image

[0203] Step S233, the loss mean of all images obtained by solving formula 35 for the B2 final images is used to update the parameters of the semantic communication model using gradients.

[0204] Specifically, see Figure 6 as well as Figure 7 As shown, Figure 6 Schematic diagram of peak signal-to-noise ratio performance comparison between the present invention and the traditional algorithm under different signal-to-noise ratios under the Rayleigh channel of an embodiment of the present invention; Figure 7 Schematic diagram of multi-scale structural similarity performance comparison between the present invention and the traditional algorithm under different signal-to-noise ratios under the Rayleigh channel of an embodiment of the present invention;

[0205] A comparative example of the present invention and the conventional method is provided, wherein:

[0206] The traditional method uses JPEG for source coding and LDPC for channel coding, and selects the DIV2K dataset as the test standard.

[0207] In typical urban scenarios, UAV-based air-to-ground communications usually adopt Rayleigh channel modeling. Peak signal-to-noise ratio and multi-scale structural similarity are selected as evaluation indicators to test the image transmission performance under different signal-to-noise ratios, with the unit of decibel (dB).

[0208] Table 1. Comparison of peak signal-to-noise ratio performance of the present invention and the traditional algorithm under different signal-to-noise ratios in Rayleigh channel (unit: dB)

[0209]

[0210] Table 2. Comparison of multi-scale structural similarity performance between the present invention and the traditional algorithm under different signal-to-noise ratios in Rayleigh channel (unit: dB)

[0211]

[0212] visible,

[0213] Figure 6 Table 1 shows the peak signal-to-noise ratio performance of this algorithm. Figure 7 Table 2 shows the multi-scale structural similarity performance of the present invention. It can be seen that the performance of the present invention is better than that of the traditional method when the peak signal-to-noise ratio and the multi-scale structural similarity are used as evaluation indicators.

[0214] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A UAV semantic communication system based on diffusion model, characterized in that: include: A joint encoding module, which is used to receive the original image data input by the UAV sender to extract abstract semantic information and convert it into signal data that can be directly transmitted, wherein the conversion includes mapping the abstract semantic information into a complex domain to adapt to signal transmission conditions; A channel estimation module, which is used to determine the channel gain in the current communication environment based on the pilot signal to determine a channel estimation value; A joint decoding module, which is used to receive the signal data to perform semantic decoding and obtain reconstructed image data, wherein the semantic decoding includes mapping the signal data from a complex domain to abstract semantic information and decoding the abstract semantic information; The diffusion module is used to receive the channel estimation value and the reconstructed image data as input, and perform image enhancement on the reconstructed image data based on a pre-trained diffusion model.

2. The UAV semantic communication system based on diffusion model according to claim 1 is characterized in that: The joint encoding module is used to extract abstract semantic information including: For converting the original image data into an image embedding block through a convolutional neural network; Used to extract abstract semantic information from each of the image embedding blocks through a plurality of Transformer blocks.

3. The UAV semantic communication system based on diffusion model according to claim 2 is characterized in that: The channel estimation module is used to construct a linear relationship between the channel estimation value and the signal received by the IoT receiver, and set the minimum root mean square error term to determine the optimal linear coefficient matrix; The optimal linear coefficient matrix is ​​calculated based on the partial derivative of the minimized root mean square error term.

4. The UAV semantic communication system based on diffusion model according to claim 3 is characterized in that: The channel estimation module is used to determine the channel estimation value, including: for determining a preliminary representation of a channel estimation matrix according to the optimal coefficient matrix; replacing corresponding matrix items in the preliminary representation according to the mean matrix to obtain the channel estimation matrix, and determining the channel estimation matrix as the channel estimation value; The obtained channel estimation matrix includes constants based on the constellation diagram and the current channel signal-to-noise ratio.

5. The UAV semantic communication system based on diffusion model according to claim 1 is characterized in that: The diffusion model predicts the noise of the reconstructed image data according to the learnable parameters during the diffusion phase, The learnable parameters are related to the number of diffusions and the current image, and the current image is an image obtained under the current number of diffusions.

6. The UAV semantic communication system based on diffusion model according to claim 5 is characterized in that: The diffusion module is used to use the likelihood probability gradient of the current image of the diffusion model in the sampling stage as a guide item.

7. The UAV semantic communication system based on diffusion model according to claim 6 is characterized in that: The diffusion module is used to introduce a classifier network to predict the category of the current image, and to construct a guide term function based on the prediction result to determine a complete noise prediction for the sampling stage.

8. The UAV semantic communication system based on diffusion model according to claim 1 is characterized in that: The diffusion module determines a description of signal transmission for a wireless channel according to a channel gain matrix, including: to obtain reconstructed image data output by the joint decoding module; The description is determined according to the reconstructed image data, a channel gain matrix and additive Gaussian noise.

9. The UAV semantic communication system based on diffusion model according to claim 1 is characterized in that: The diffusion module receives the channel estimate and the reconstructed image data as input, and determines a description for the noise term including: To determine the projection of the original data in the column space and the null space of the channel gain matrix, and to determine the estimated value for the original image in combination with the channel estimated value; It is used to introduce the influence of channel noise during the transmission process, optimize the estimated value for the original image, and obtain the optimized estimated value; A description of the noise term is determined according to the optimized estimate.

10. The UAV semantic communication system based on diffusion model according to claim 1 is characterized in that: The joint encoding module is used to extract abstract semantic information based on a pre-trained semantic communication model.

Citation Information

Patent Citations

  • Target detection task-oriented unmanned aerial vehicle cross-modal semantic communication method and system

    CN118470581A