End-to-end communication method based on joint source channel coding
By applying joint source channel coding technology and neural network code rate adaptation module in end-to-end communication scenarios, the problem that DJSCC in the prior art is difficult to adapt to end-to-end communication and core network transmission rate constraints is solved, and efficient wired data transmission and dynamic code rate adaptation are achieved.
Patent Information
- Application Number
- CN202510190599.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-20
AI Technical Summary
The existing DJSCC technology is mainly aimed at point-to-point communication scenarios, and lacks a transmission method or framework suitable for end-to-end communication scenarios. Especially under the constraints of the transmission rate of the core network, it is difficult to achieve efficient wired data transmission.
An end-to-end communication method based on joint source channel encoding is proposed. By performing channel estimation between the source node and the destination node, joint encoding and decoding is performed using the DJSCC encoder and the decoder, and the neural network code rate adaptation module is used in the core network to adapt to different transmission rates.
It effectively reduces the data load inside the core network, dynamically adapts to different core network data transmission rate constraints, and improves end-to-end communication performance.
Smart Images

Figure CN120185769A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of wireless communication technologies, and in particular, to an end-to-end communication method based on joint source-channel coding. Background Art
[0002] DJSCC (Deep Joint Source-Channel Coding) technology is an emerging wireless communication technology that directly maps source data to channel symbols by jointly using a deep neural network and source coding and channel coding modules in a traditional communication system. Compared with traditional communication systems based on separate source-channel coding, DJSCC can achieve better source reconstruction performance.
[0003] Currently, the existing DJSCC technology in the prior art mainly targets point-to-point communication scenarios, that is, direct communication between two communication nodes is considered. However, actual end-to-end communication scenarios involve wireless communication between the sending and receiving ends and the core network, as well as wired communication with transmission rate constraints between communication nodes within the core network.
[0004] With the rapid increase in the number of communication users and real-time application requirements in recent years, the data load within the core network has increased rapidly and has caused a communication bottleneck to a certain extent. When extending DJSCC technology to end-to-end communication scenarios, since channel symbols contain channel redundancy information added to resist wireless channels, directly encoding channel symbols into bits and transmitting them in the core network will significantly increase the data load within the core network, resulting in inefficient wired data transmission. In addition, in end-to-end communication scenarios, DJSCC technology needs to adapt to different core network transmission rate constraints to achieve dynamic encoding and decoding.
[0005] Currently, the existing DJSCC technology in the prior art mainly targets point-to-point communication scenarios. Patent CN 116599628 A proposes a digital DJSCC transmission framework for point-to-point communication; Patent CN118944812 A proposes a data quantization and inverse quantization method based on DJSCC, which realizes the quantization of the analog quantity output by the encoder for further encoding into a bit stream. In addition, Patent CN 118714624 A, Patent 118606883A, and Patent 116388829A consider DJSCC communication scenarios assisted by relay nodes, where data processing and forwarding are implemented at the relay nodes to improve communication performance. In terms of code rate control, Patent CN118433129 A uses reinforcement learning to obtain a code rate control strategy that can maximize semantic efficiency; Patent CN 107623560A targets LDPC systems under separate source-channel coding and proposes a code rate allocation method to allocate suitable source coding rates and channel coding rates for each frame of the source to optimize image transmission quality and transmission efficiency.
[0006] The disadvantages of the DJSCC technology in the above prior art include: The existing DJSCC technology mainly targets point-to-point communication scenarios with 2 communication nodes or communication scenarios with additional relay nodes, lacking a DJSCC transmission method or framework for actual end-to-end communication scenarios that simultaneously considers wireless and wired links. In addition, there is no DJSCC adaptation method for core network transmission rate constraints. Summary of the Invention
[0007] An embodiment of the present invention provides an end-to-end communication method based on joint source-channel coding to effectively improve end-to-end communication performance.
[0008] To achieve the above object, the present invention adopts the following technical solutions.
[0009] An end-to-end communication method based on joint source-channel coding, comprising:
[0010] The source access node performs channel estimation on the first wireless channel between it and the source node to obtain the first signal-to-noise ratio, and feeds back the first signal-to-noise ratio to the source node. The destination node performs channel estimation on the second wireless channel between it and the destination access node to obtain the second signal-to-noise ratio, and feeds back the second signal-to-noise ratio to the destination access node;
[0011] The source node sends a certain modality of data to be transmitted and the first signal-to-noise ratio to the joint source-channel coding DJSCC encoder 1, obtains the source-channel input signal according to the joint coding result returned by the DJSCC encoder 1, and sends the source-channel input signal after normalization;
[0012] The source access node converts the received source-channel input signal into a tensor, sends the first signal-to-noise ratio and the tensor to the DJSCC decoder 1, and then processes the reconstructed features returned by the DJSCC decoder 1 through a prediction network, a transformation network, and a hyperprior encoder to obtain the total bitstream. The source access node sends the total bitstream to the core network;
[0013] The destination access node uses an entropy decoding algorithm to convert the received total bitstream into decoded features, sends the second signal-to-noise ratio and the decoded features to the DJSCC encoder 2, and sends the destination-channel input signal returned by the DJSCC encoder 2 after normalization;
[0014] The destination node converts the received destination-channel input signal into a tensor, sends the second signal-to-noise ratio and the tensor to the DJSCC decoder 2, and restores the tensor output by the DJSCC decoder 2 to pixel values to obtain a reconstructed image.
[0015] Preferably, the source node sends a certain modality of data to be transmitted and the first signal-to-noise ratio to the joint source-channel coding DJSCC encoder 1, obtains the source-channel input signal according to the joint coding result returned by the DJSCC encoder 1, and sends the source-channel input signal after normalization, including:
[0016] The source node obtains a certain modality of data to be transmitted and obtains the input tensor x;
[0017] The source node sends the first signal-to-noise ratio in floating-point form and the input tensor x to the DJSCC encoder 1. The DJSCC encoder 1 completes the joint coding process adapted to the first wireless channel, obtains the joint coding result in real value, combines the real-valued joint coding results in pairs into complex symbols to obtain the source-channel input signal s1, and normalizes the power of the source-channel input signal s1 so that the average power of the source-channel input signal s1 is 1;
[0018] The source node sends the normalized source-channel input signal s1.
[0019] Preferably, the source access node converts the received source-channel input signal into a tensor, sends the first signal-to-noise ratio and the tensor to the DJSCC decoder 1, and then processes the reconstructed features returned by the DJSCC decoder through the prediction network, transformation network, and hyperprior encoder to obtain the total bitstream. The source access node sends the total bitstream to the core network, including:
[0020] The source access node disassembles each complex symbol of the received source-channel input signal to obtain real symbols, and combines all real symbols into a tensor;
[0021] The source access node sends the first signal-to-noise ratio in floating-point form and the combined tensor to the DJSCC decoder 1. The DJSCC decoder 1 completes the joint decoding process adapted to the first wireless channel, obtains the reconstructed feature y, and the DJSCC decoder 1 returns the reconstructed feature y to the source access node;
[0022] The source access node inputs the target code rate in floating-point form and the reconstructed feature y into the prediction network to obtain the Lagrange multiplier λ, and the prediction network returns the Lagrange multiplier λ to the source access node;
[0023] The source access node inputs the Lagrange multiplier λ in floating-point form and the reconstructed feature y into the transformation network to obtain the transformed feature y λ and quantizes the transformed feature y λ to obtain the quantized version The transformation network returns the quantized version to the source access node;
[0024] The source access node inputs the Lagrange multiplier λ in floating-point form and the transformed feature y λ into the hyperprior encoder to obtain the hyperprior z, and quantizes the hyperprior z to obtain the quantized version Input the quantized version and the Lagrange multiplier λ into the hyperprior decoder to obtain the probability distribution parameters of the transformed feature y λ and The hyperprior decoder inputs the probability distribution parameters of the transformed feature y λ and and returns them to the source access node;
[0025] The source access node uses the entropy coding algorithm to encode the quantized transformed feature and into the bitstream b , encodes the quantized hyperprior y into the bitstream b , encodes the Lagrange multiplier λ into the bitstream b z , combines the bitstreams b λ , b y , and b z and b λ into the total bitstream b;
[0026] The source access node sends the total bitstream b into the core network.
[0027] Preferably, the destination access node uses the entropy decoding algorithm to convert the received total bitstream into decoded features, sends the second signal-to-noise ratio and the decoded features to the DJSCC encoder 2, and sends out the normalized destination channel input signal returned by the DJSCC encoder 2, including:
[0028] The total bitstream b is transmitted between communication nodes in the core network via a lossless wired link until it leaves the core network and reaches the destination access node. The destination access node divides the total bitstream b into bitstreams b y , b z and b λ ;
[0029] The destination access node decodes the bitstream b λ into the Lagrange multiplier λ, uses the entropy decoding algorithm to decode the bitstream b z into the hyperprior Input the hyperprior and the Lagrange multiplier λ into the hyperprior decoder to obtain the probability distribution parameters and
[0030] The destination access node uses the entropy decoding algorithm according to the probability distribution parameters and decode the bitstream b y into decoded features
[0031] The destination access node sends the second signal-to-noise ratio in floating-point form and the decoded features to the DJSCC encoder 2. The DJSCC encoder 2 completes the joint encoding process for the second wireless channel adaptation, combines the real-valued joint encoding results in pairs into complex symbols, and obtains the channel input signal s2;
[0032] The destination access node performs power normalization on the channel input signal s2 so that its average power is 1, and sends the normalized channel input signal s2.
[0033] Preferably, the destination node converts the received destination channel input signal into a tensor, sends the second signal-to-noise ratio and the tensor to the DJSCC decoder 2, and restores the tensor output by the DJSCC decoder 2 to pixel values to obtain a reconstructed image, including:
[0034] The destination node disassembles each complex symbol of the received destination channel input signal to obtain real symbols, and combines all the real symbols into a tensor;
[0035] The destination node sends the second signal-to-noise ratio in floating-point form and the combined tensor to the DJSCC decoder 2. The DJSCC decoder 2 completes the joint decoding process for the second wireless channel adaptation and obtains an output tensor The DJSCC decoder 2 returns the output tensor x to the destination node;
[0036] The destination node limits each value of the output tensor x to the range (0, 1), sets the value exceeding 1 to 1, and sets the value less than 0 to 0, and restores the output tensor x to pixel values to obtain a reconstructed image.
[0037] Preferably, the source access node inputs the target code rate in floating-point form and the reconstruction feature y into the prediction network to obtain the Lagrange multiplier λ, including: The prediction network first processes the reconstruction feature y by using 3 residual blocks, pooling, and a multi-layer perceptron in sequence, and outputs the feature vector g y , and the normalization parameters a, b. At the same time, it processes the target code rate R by using a multi-layer perceptron and outputs the code rate vector g R , and processes the code rate vector g by using the normalization parameters a, b R to obtain g R ;
[0038]
[0039] Combine With g y , the multi-layer perceptron is used for processing and the Sigmoid activation function is used to limit the output value r between 0 and 1. r is scaled to obtain the predicted Lagrange multiplier λ:
[0040] λ = r × (λ max ―λ min ) + λ min
[0041] where λ max and λ min are the maximum and minimum values within the value range of λ, respectively.
[0042] Preferably, the source access node inputs the Lagrange multiplier λ in floating-point form and the reconstructed feature y into the transformation network to obtain the transformed feature y λ , including:
[0043] The transformation network first uses grouped convolution to process the input Lagrange multiplier λ and the reconstructed feature y, outputs intermediate features, then uses the Lagrange multiplier λ to generate weights and biases, multiplies the weights and the intermediate features element-wise, adds the biases and the intermediate features element-wise, outputs fused features, and finally processes the fused features with a convolution kernel of size 1x1 and a stride of 1, GeLU activation, and a convolution kernel of size 1x1 and a stride of 1 in sequence and adds them to the residuals to obtain the output transformed feature y λ .
[0044] It can be seen from the technical solutions provided by the embodiments of the present invention described above that the embodiments of the present invention propose a bit rate adaptation module based on a neural network to adapt to different core network transmission rate constraints. The bit rate adaptation module uses a neural network to transform the reconstructed source features. The bitstream obtained by entropy coding based on the transformed features can adapt to any given core network rate constraint, further improving the end-to-end communication performance.
[0045] Additional aspects and advantages of the present invention will be given in part in the following description, and these will become apparent from the following description or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 FIG. is a schematic diagram of an end-to-end communication framework based on DJSCC provided by an embodiment of the present invention;
[0048] Figure 2 A residual block and a prediction network structure diagram provided by an embodiment of the present invention;
[0049] Figure 3 A code rate residual block structure, a transformation network structure, a hyperprior encoder, and a hyperprior decoder structure diagram provided by an embodiment of the present invention;
[0050] Figure 4 A training data processing flowchart of a DJSCC encoder / decoder 1 and a DJSCC encoder / decoder 2 provided by an embodiment of the present invention;
[0051] Figure 5 A data processing flowchart for end-to-end training of the entire framework considering a transformation network, a compression module, and a decompression module in a code rate adaptation module provided by an embodiment of the present invention;
[0052] Figure 6 A training data processing flowchart of a prediction network in a code rate adaptation module provided by an embodiment of the present invention. Detailed implementation manners
[0053] The following details the implementation manners of the present invention. Examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The implementation manners described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation to the present invention.
[0054] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "including" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or coupling. The term "and / or" used herein includes any unit and all combinations of one or more of the associated listed items.
[0055] Those skilled in the art can understand that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by those of ordinary skill in the art to which this invention belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted in an idealized or overly formal sense unless defined as such here.
[0056] For ease of understanding the embodiments of the present invention, the following will further explain with several specific embodiments in conjunction with the accompanying drawings, and each embodiment does not constitute a limitation to the embodiments of the present invention.
[0057] The present invention is directed to an end-to-end communication scenario, and proposes a transmission framework based on DJSCC technology and a corresponding training strategy to maximize the end-to-end communication performance under the constraint of the core network transmission rate. The present invention proposes that before core network transmission, the source features are reconstructed using the received channel symbols, and referring to the idea of data compression based on neural networks, the probability distribution of the reconstructed source features is estimated using a neural network, and the reconstructed source features are encoded into a bitstream using entropy coding methods such as arithmetic coding, so as to achieve efficient wired communication.
[0058] The present invention proposes a schematic diagram of an end-to-end communication framework based on DJSCC as Figure 1 shown. The framework includes a source node, a source access node, multiple communication nodes in the core network, a destination access node, and a destination node. Among them, the source node and the source access node realize wireless data transmission through the first wireless channel, and the destination access node and the destination node realize wireless data transmission through the second wireless channel. The communication nodes in the core network realize wired data transmission through a lossless wired link with a transmission rate constraint. The DJSCC encoder / decoder 1, the code rate adaptation module, the compression / decompression module, and the DJSCC encoder / decoder 2 of the framework are all composed of neural networks. The code rate adaptation module includes a prediction network and a transformation network.
[0059] The present invention is applicable to the transmission of various modal data (such as images, videos, etc.), and is not limited to specific modal data.
[0060] The processing flow of an end-to-end communication method based on joint source-channel coding proposed by the present invention includes the following processing steps:
[0061] Step 1. The source access node performs channel estimation on the first wireless channel to obtain the first signal-to-noise ratio, and feeds back the first signal-to-noise ratio to the source node through a feedback mechanism. At the same time, the destination node performs channel estimation on the second wireless channel to obtain the second signal-to-noise ratio, and feeds back the second signal-to-noise ratio to the destination access node through a feedback mechanism.
[0062] The source node performs the following steps:
[0063] Step 2. Given a certain modality data to be transmitted, obtain the input tensor x. Taking an image as an example, normalize the pixel values of the image so that each normalized pixel value is within the range (0, 1), and represent the normalized image as the input tensor x.
[0064] Step 3. Input the first signal-to-noise ratio in floating-point form and the input tensor x into the DJSCC encoder 1. The DJSCC encoder 1 completes the joint coding process adaptive to the first wireless channel and obtains a real-valued joint coding result. Then, combine the real-valued joint coding results in pairs into complex symbols to obtain the source channel input signal s1. Normalize the power of the source channel input signal s1 so that the average power of the source channel input signal s1 is 1. Transmit the normalized source channel input signal s1.
[0065] The steps of the source node are completed.
[0066] Step 4. The normalized source channel input signal s1 is transmitted through the first wireless channel to the source access node.
[0067] The source access node performs the following steps:
[0068] Step 5. The source access node disassembles each complex symbol of the received source channel input signal to obtain real symbols. Combine all the real symbols into a tensor.
[0069] Step 6. Input the first signal-to-noise ratio in floating-point form and the combined tensor into the DJSCC decoder 1. The DJSCC decoder 1 completes the joint decoding process adaptive to the first wireless channel and obtains the reconstructed feature y. The DJSCC decoder 1 returns the reconstructed feature y to the source access node.
[0070] Step 7. The source access node inputs the target code rate in floating-point form and the reconstructed feature y into the prediction network to obtain the Lagrange multiplier λ. The prediction network returns the Lagrange multiplier λ to the source access node.
[0071] Figure 2 This is a residual block and a prediction network structure diagram provided by an embodiment of the present invention. As Figure 2 (b) shows, the basic processing unit of the prediction network structure is a residual block, as Figure 2 (a) shows. First, use grouped convolution to process the input features and output intermediate features. Then, sequentially use a convolution with a kernel size of 1x1 and a stride of 1, a GeLU activation, and a convolution with a kernel size of 1x1 and a stride of 1 to process the intermediate features and add them to the residuals to obtain the output features. The prediction network first sequentially uses 3 residual blocks, pooling, and a multi-layer perceptron to process the reconstructed feature y and outputs the feature vector gy , and the normalization parameters a and b. At the same time, the multi-layer perceptron is used to process the target bitrate R, and the bitrate vector g is output R . The bitrate vector g is processed using the normalization parameters a and b R to obtain g R .
[0072]
[0073] Then, g is merged R with g y . The multi-layer perceptron is used for processing and the Sigmoid activation function is used to limit the output value r between 0 and 1. Finally, r is scaled to obtain the predicted Lagrange multiplier λ:
[0074] λ = r × (λ max ―λ min ) + λ min
[0075] where λ max and λ min are the maximum and minimum values within the range of λ respectively.
[0076] Step 8. The source access node inputs the Lagrange multiplier λ in floating-point form and the reconstructed feature y into the transformation network to obtain the transformed feature y λ . The transformed feature y is quantized λ to obtain the quantized version The transformation network returns the quantized version to the source access node.
[0077] The basic processing unit of the transformation network is the bitrate residual block, and its structure is as shown in Figure 3 (a). First, the input feature is processed using grouped convolution to output the intermediate feature. Then, the Lagrange multiplier λ is used to generate the weights and biases. The weights are multiplied element-wise with the intermediate feature, and the biases are added element-wise to the intermediate feature to output the fused feature. Finally, the fused feature is processed sequentially with a convolution kernel of size 1x1 and stride 1, GeLU activation, and a convolution kernel of size 1x1 and stride 1, and added to the residual to obtain the output feature. The structure of the transformation network is as shown in Figure 3 (b), which contains 3 bitrate residual blocks for processing the reconstructed feature y into the transformed feature y λ .
[0078] Step 9. The source access node inputs the Lagrange multiplier λ in floating-point form and the transformed feature y λ into the hyperprior encoder to obtain the hyperprior z. The hyperprior z is quantized to obtain the quantized version The quantized version is input into the hyperprior decoder to obtain the transformed feature yλ Probability distribution parameters and The hyperprior decoder transforms the feature y λ Probability distribution parameters and And returns them to the source access node.
[0079] Figure 3 This is a structure diagram of a rate residual block, a transformation network structure, a hyperprior encoder, and a hyperprior decoder provided by an embodiment of the present invention. The basic processing unit of the hyperprior encoder / decoder is a rate residual block, and its structure is as shown in Figure 3 (a). The structure of the hyperprior encoder is as shown in Figure 3 (c), which includes 1 downsampling convolution and 3 rate residual blocks for processing the transformed feature y λ Into the hyperprior z. The structure of the hyperprior decoder is as shown in Figure 3 (d), which includes 3 rate residual blocks and 1 upsampling convolution for processing the quantized hyperprior Into probability distribution parameters and
[0080] Step 10. The source access node uses the entropy coding algorithm to encode the quantized transformed feature and Into the bitstream b , encodes the quantized hyperprior y Into the bitstream b , encodes the Lagrange multiplier λ into the bitstream b z , combines the bitstreams b λ , b y , and b z And b λ Into the total bitstream b.
[0081] Step 11. The source access node sends the total bitstream b into the core network.
[0082] The core network contains multiple communication nodes, and data communication between nodes is completed through lossless wired links with rate constraints. After the total bitstream enters the core network, it is sequentially transmitted to these communication nodes until it leaves the core network.
[0083] The steps of the source access node are completed.
[0084] Step 12. The total bitstream b is transmitted between communication nodes in the core network through lossless wired links until it leaves the core network and reaches the destination access node.
[0085] The destination access node performs the following steps:
[0086] Step 13. The destination access node divides the total bitstream b into bitstreams b y , b z and b λ .
[0087] Step 14. The destination access node decodes the bitstream b λ into the Lagrange multiplier λ, and uses the entropy decoding algorithm to decode the bitstream b z into the hyperprior Input the hyperprior into the hyperprior decoder to obtain the probability distribution parameters and
[0088] The destination access node uses the entropy decoding algorithm to decode the bitstream b and into the decoded features y
[0089] Step 15. The destination access node inputs the second signal-to-noise ratio in floating-point form and the decoded features into the DJSCC encoder 2 to complete the joint coding process for the second wireless channel adaptation. Combine the joint coding results of real values in pairs into complex symbols to obtain the destination channel input signal s2
[0090] The destination access node normalizes the power of the destination channel input signal s2 so that its average power is 1. Transmit the normalized destination channel input signal s2
[0091] The steps of the destination access node are completed
[0092] Step 16. The normalized destination channel input signal s2 is transmitted to the destination node through the second wireless channel
[0093] The destination node performs the following steps
[0094] Step 17. The destination node disassembles each complex symbol of the received destination channel input signal to obtain real symbols. Combine all real symbols into a tensor
[0095] Step 18. The destination node inputs the second signal-to-noise ratio in floating-point form and the combined tensor into the DJSCC decoder 2. The DJSCC decoder 2 completes the joint decoding process for the second wireless channel adaptation to obtain the output tensor The DJSCC decoder 2 returns the output tensor to the destination node
[0096] Step 19. The destination node processes the output tensor Each value of is normalized, and the output tensor values exceeding 1 are set to 1, and the output tensor values less than 0 are set to 0. The output tensor is restored to pixel values to obtain the reconstructed image.
[0097] The execution of the destination node step is completed.
[0098] The embodiment of the present invention provides a progressive training strategy for the above-mentioned end-to-end communication framework based on DJSCC, and the training of the overall framework is completed in multiple steps. The training of each step requires loading the network parameters trained in the previous step. The detailed training steps are as follows:
[0099] 1. Without considering the code rate adaptation module, compression module, and decompression module, train DJSCC encoder 1 and DJSCC encoder 2. Among them, the reconstructed features are directly input into DJSCC encoder 2 for encoding. A training data processing flow of DJSCC encoder 1 and DJSCC encoder 2 provided by the embodiment of the present invention is as Figure 4 shown. In each iteration process, fix the signal-to-noise ratio of the second wireless channel (20 dB - 30 dB), and randomly sample the signal-to-noise ratio of the first wireless channel within the training signal-to-noise ratio range for training until convergence. The loss function of this step is the distortion between the input tensor x and the output tensor as shown in Equation 1. This distortion includes objective distortion (measured by metrics such as Peak Signal to Noise Ratio, PSNR) and perceptual distortion (measured by metrics such as Learned Perceptual Image Patch Similarity, LPIPS).
[0100]
[0101] where d1(·,·) is the distortion function.
[0102] 2. On the basis of step 1, train DJSCC encoder 1 and DJSCC encoder 2. The data processing flow is as Figure 4 shown. In each iteration process, randomly sample the signal-to-noise ratio of the first wireless channel and the signal-to-noise ratio of the second wireless channel within the training signal-to-noise ratio range for training until convergence. The loss function of this step is the same as that of step 1.
[0103] 3. Based on step 2, for the transformation network, compression module, and decompression module in the rate adaptation module provided by the embodiments of the present invention, the data processing flow for end-to-end training of the entire framework is as follows Figure 5 shown. During each iteration, within the training signal-to-noise ratio range, the signal-to-noise ratio of the first wireless channel and the signal-to-noise ratio of the second wireless channel are randomly sampled respectively, and at the same time, λ is randomly sampled within the training Lagrange multiplier λ range for training until convergence. The loss function for this step is the distortion between the input tensor x and the output tensor and the trade-off between the rate of the transformed feature and the hyperprior , as shown in Equation 2.
[0104]
[0105] where λ controls the degree of trade-off, R y represents the rate of the transformed feature , and R z represents the rate of the hyperprior .
[0106] 4. Based on step 3, considering the prediction network in the rate adaptation module, while fixing the remaining neural network parameters, the data processing flow for training only this prediction network is as follows Figure 6 shown. During each iteration, within the training signal-to-noise ratio range, the signal-to-noise ratio of the first wireless channel is randomly sampled, and at the same time, the rate R is randomly sampled within the training rate range for training until convergence. The loss function for this step is the deviation between the target rate R and the actual rate (R y +R z ), as shown in Equation 3. This deviation can be measured by the Mean Square Error (MSE) or the Mean Absolute Error (MAE).
[0107] Loss = d2(R, (R y +R z )) #(3)
[0108] where d2(·,·) is the deviation function.
[0109] In summary, the method of the present invention is applicable to the end-to-end communication scenario based on DJSCC, taking into account both wireless and wired data transmission, which can effectively reduce the data load within the core network and dynamically adapt to different core network data transmission rate constraints, thereby improving the end-to-end communication performance.
[0110] Those of ordinary skill in the art can understand that: The attached drawings are only schematic diagrams of an embodiment, and the modules or processes in the attached drawings are not necessarily essential for implementing the present invention.
[0111] As can be understood from the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present invention, in essence or the part that makes contributions to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0112] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device or system embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments. The device and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0113] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for end-to-end communication based on joint source-channel coding, characterized in that: include: The source access node performs channel estimation on a first wireless channel between itself and the source node, obtains a first signal-to-noise ratio, and feeds the first signal-to-noise ratio back to the source node; the destination node performs channel estimation on a second wireless channel between itself and the destination access node, obtains a second signal-to-noise ratio, and feeds the second signal-to-noise ratio back to the destination access node; The source node sends a certain modal data to be transmitted and a first signal-to-noise ratio to the joint source channel coding DJSCC encoder 1, obtains a source channel input signal according to the joint coding result returned by the DJSCC encoder 1, and sends the source channel input signal after normalization; The source access node converts the received source channel input signal into a tensor, sends the first signal-to-noise ratio and the tensor to the DJSCC decoder 1, and then processes the reconstructed features returned by the DJSCC decoder 1 through the prediction network, the transformation network and the super-a priori encoder to obtain a total bit stream, and the source access node sends the total bit stream to the core network; The destination access node converts the received total bit stream into a decoding feature using an entropy decoding algorithm, sends the second signal-to-noise ratio and the decoding feature to the DJSCC encoder 2, and normalizes the destination channel input signal returned by the DJSCC encoder 2 and sends it out; The destination node converts the received destination channel input signal into a tensor, sends the second signal-to-noise ratio and the tensor to the DJSCC decoder 2, and restores the tensor output by the DJSCC decoder 2 to a pixel value to obtain a reconstructed image.
2. The method according to claim 1, characterized in that The source node sends a certain modal data to be transmitted and a first signal-to-noise ratio to a joint source channel coding DJSCC encoder 1, obtains a source channel input signal according to a joint coding result returned by the DJSCC encoder 1, and sends the source channel input signal after normalization, including: The source node obtains a certain modality data to be transmitted and obtains the input tensor x; The source node sends a first signal-to-noise ratio and an input tensor x in floating-point form to the DJSCC encoder 1. The DJSCC encoder 1 completes a joint coding process adaptive to the first wireless channel to obtain a joint coding result of a real value. The joint coding results of the real value are combined into complex symbols in pairs to obtain a source channel input signal s1. The source channel input signal s1 is power normalized so that the average power of the source channel input signal s1 is 1. The source node sends a normalized source channel input signal s1.
3. The method according to claim 2, characterized in that The source access node converts the received source channel input signal into a tensor, sends the first signal-to-noise ratio and the tensor to the DJSCC decoder 1, and then processes the reconstructed features returned by the DJSCC decoder through a prediction network, a transformation network and a super-a priori encoder to obtain a total bit stream, and the source access node sends the total bit stream to the core network, including: The source access node receives the source channel input signal Decompose each complex number symbol to obtain the real number symbol, and merge all the real number symbols into a tensor; The source access node sends the first signal-to-noise ratio in floating point form and the combined tensor to the DJSCC decoder 1. The DJSCC decoder 1 completes the joint decoding process for the first wireless channel adaptation to obtain the reconstructed feature y. The DJSCC decoder 1 returns the reconstructed feature y to the source access node. The source access node inputs the target bit rate and the reconstructed feature y in floating point form into the prediction network to obtain the Lagrange multiplier λ, and the prediction network returns the Lagrange multiplier λ to the source access node; The source access node inputs the floating-point Lagrange multiplier λ and the reconstructed feature y into the transformation network to obtain the transformed feature y λ , quantized transformation feature y λ Get the quantized version The transformation network quantizes the Return to the source access node; The source access node converts the floating-point Lagrange multiplier λ and the transformation feature y λ Input to the super prior encoder to obtain the super prior z, and quantize the super prior z to obtain the quantized version The quantized version And the Lagrange multiplier λ is input to the super prior decoder to obtain the transformed feature y λ The probability distribution parameters of and The super prior decoder transforms the feature y λ The probability distribution parameters of and Return to the source access node; The source access node uses an entropy coding algorithm to calculate the probability distribution parameters and The quantized transformation features Encoded as bit stream b y , the quantized hyper-prior Encoded as bit stream b z , the Lagrange multiplier λ is encoded into a bit stream b λ , merge bitstream b y , b z and b λ is the total bit stream b; The source access node sends the total bit stream b into the core network.
4. The method according to claim 3, characterized in that The destination access node converts the received total bit stream into a decoding feature using an entropy decoding algorithm, sends the second signal-to-noise ratio and the decoding feature to the DJSCC encoder 2, and normalizes the destination channel input signal returned by the DJSCC encoder 2 and sends it out, including: The total bit stream b is transmitted between communication nodes in the core network via lossless wired links until it leaves the core network and reaches the destination access node. The destination access node divides the total bit stream b into bit streams b y , b z and b λ ; The destination access node sends the bit stream b λ Decoded into Lagrange multiplier λ, the bit stream b is decoded using entropy decoding algorithm z Decoding as a Super Prior Super Prior and the Lagrange multiplier λ are input to the super-prior decoder to obtain the probability distribution parameters and The destination access node uses an entropy decoding algorithm based on the probability distribution parameters and The bit stream b y Decode to decode features The destination access node sends the second signal-to-noise ratio and decoding characteristics in floating point form To DJSCC encoder 2, DJSCC encoder 2 completes the joint coding process for the second wireless channel adaptation, combines the joint coding results of the real values into complex symbols in pairs, and obtains the channel input signal s2; The destination access node normalizes the power of the channel input signal s2 so that its average power is 1, and sends the normalized channel input signal s2.
5. The method according to claim 4, characterized in that The destination node converts the received destination channel input signal into a tensor, sends the second signal-to-noise ratio and the tensor to the DJSCC decoder 2, and restores the tensor output by the DJSCC decoder 2 to a pixel value to obtain a reconstructed image, including: The destination node receives the input signal of the destination channel Decompose each complex number symbol to obtain the real number symbol, and merge all the real number symbols into a tensor; The destination node sends the second signal-to-noise ratio in floating point form and the combined tensor to DJSCC decoder 2, and DJSCC decoder 2 completes the joint decoding process for the second wireless channel adaptation and obtains the output tensor DJSCC decoder 2 will output a tensor Return to the destination node; The destination node will output a tensor Each value of is limited to the range (0, 1), values exceeding 1 are set to 1, and values less than 0 are set to 0, and the output tensor Restore to pixel values and get the reconstructed image.
6. The method according to claim 4, characterized in that The source access node inputs the target bit rate in floating point form and the reconstructed feature y into the prediction network to obtain the Lagrange multiplier λ, including: the prediction network first uses three residual blocks, pooling and multi-layer perceptron to process the reconstructed feature y in sequence, and outputs the feature vector g y , and standardized parameters a, b, and use the multilayer perceptron to process the target bit rate R and output the bit rate vector g R , use the standardized parameters a, b to process the bit rate vector g R get merge With g y , use the multi-layer perceptron to process and use the Sigmoid activation function to limit the output value r between 0 and 1, scale r, and get the predicted Lagrange multiplier λ: λ=r×(λ max ―l min )+λ min Among them, λ max and λ min are the maximum and minimum values within the range of λ respectively.
7. The method according to claim 4, characterized in that The source access node inputs the floating-point Lagrange multiplier λ and the reconstruction feature y into the transformation network to obtain the transformation feature y λ ,include: The transformation network first uses grouped convolution to process the input Lagrange multiplier λ and the reconstructed feature y, outputs the intermediate feature, then uses the Lagrange multiplier λ to generate weights and biases, multiplies the weights by the intermediate feature element by element, adds the bias to the intermediate feature element by element, outputs the fused feature, and finally uses convolution with a kernel size of 1x1 and a step size of 1, GeLU activation, convolution with a kernel size of 1x1 and a step size of 1 to process the fused feature and add it to the residual to obtain the output transformed feature y λ .
Citation Information
Patent Citations
Image transmission rate self-adaptive distribution method based on joint source channel coding
CN107623560A
Semantic communication code rate control method
CN118433129A