Satellite network semantic communication method, system and device based on reliable joint source channel coding and storage medium

Through a semantic communication method of satellite network based on reliable joint source channel encoding, the Vision Transformer model and NOMA technology are used to solve the problems of scarcity of spectrum resources and degradation of signal quality in satellite Internet of Things, and efficient and reliable satellite-ground communication is achieved.

CN120263347APending Publication Date: 2025-07-04HARBIN INST OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510337177.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Satellite Internet of Things faces the problems of scarce spectrum resources, declining signal quality, decreasing communication system robustness and poor data transmission stability, especially in the satellite-ground links, and existing semantic communication research lacks specific scenario applications.

Method used

The semantic communication method of satellite network based on reliable joint source channel encoding is adopted, and the image semantic representation is extracted using the Vision Transformer model, and the channel estimation and joint loss function are encoded, and power distribution is performed through NOMA technology to realize signal transmission and recovery.

Benefits of technology

Significantly reduce the amount of data transmitted, improve communication efficiency, enhance the robustness and reliability of the satellite-ground link, improve channel capacity, and improve user experience, especially under low signal-to-noise ratio conditions, image quality is improved by 14.8% to 58.8%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005322078830000031
    Figure BDA0005322078830000031
  • Figure BDA0005322078830000035
    Figure BDA0005322078830000035
  • Figure BDA0005322078830000036
    Figure BDA0005322078830000036
Patent Text Reader

Abstract

The invention relates to a satellite network semantic communication method based on reliable joint source channel coding. The satellite network semantic communication method comprises a signal transmitting method and a signal receiving method. The method comprises the following steps of: 1, extracting features of an image by utilizing a Vision Transform model to obtain semantic representation of the image; 2, constructing a channel estimation and joint loss function; 3, encoding the semantic representation in the step 1 by using a joint source channel encoder and combining the channel estimation in the step 2 and the joint loss function; 4, NOMA power distribution is carried out on the coding result in the step 3, and the coding result is transmitted through a wireless link; the signal receiving method comprises the following steps: step 5, a receiving end receives a signal transmitted in step 4 from a link; and 6, recovering the original user data from the noise by using a joint source channel decoder. According to the method, the semantic information is extracted from the image by using the pre-trained Vision Transform model, so that the amount of data to be transmitted is remarkably reduced, and meanwhile, the communication efficiency is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information and communication technology, and in particular to a satellite network semantic communication method, system, device and storage medium based on reliable joint source channel coding. Background Art

[0002] With the explosive growth of information technology and intelligent applications, satellite Internet of Things faces many severe challenges. Spectrum resources are extremely scarce, and commonly used frequency bands are largely occupied, resulting in spectrum congestion and increasing interference, which seriously limits the performance improvement and expansion of communication systems. Satellite-to-ground links are affected by ground wireless equipment and complex physical environments, resulting in reduced signal quality, impaired robustness of communication systems, and reduced stability and reliability of data transmission. During large-scale access, the channel capacity approaches the Shannon limit, making it difficult to increase the transmission rate. The use of advanced technologies will lead to problems such as increased bit error rate, increased complexity, and increased latency and power consumption, hindering the widespread application of satellite Internet of Things.

[0003] As an innovative communication paradigm, semantic communication has brought new hope for solving many problems faced by satellite IoT. It is in a unique semantic layer in the communication system, and mainly pre-processes information through advanced artificial intelligence technology. In terms of spectrum resource utilization, semantic communication can accurately extract key semantic content from a large amount of information, discard redundant information, and achieve efficient information compression. This efficient information processing method significantly improves the channel capacity under limited spectrum resource conditions, thereby effectively alleviating the current situation of spectrum congestion. In terms of coping with complex transmission environments, semantic communication enhances the anti-interference ability of information itself by intelligently encoding and extracting features of information. Even in an environment such as the satellite-to-ground link that is easily interfered by multiple factors, information pre-processed by semantic communication can be transmitted with higher stability, reducing the impact of signal quality degradation on communication, thereby improving the robustness of the communication system. In addition, semantic communication can also perform personalized processing for different types of information, adopt different transmission strategies according to the importance and sensitivity of the information, further optimize the communication process, and ensure the reliable transmission of key information. However, the existing research on semantic communication is some conceptual models, lacking research on specific scenarios, especially for future satellite-to-ground integrated networks.

[0004] Therefore, finding a new semantic communication method for satellite networks is a key issue that needs to be urgently solved in the current technical field. Summary of the invention

[0005] In view of the above problems, the present invention proposes a satellite network semantic communication method and system based on reliable joint source channel coding.

[0006] A satellite network semantic communication method based on reliable joint source-channel coding of the present invention includes a signal transmission method and a signal reception method;

[0007] The signal transmission method includes the following steps:

[0008] Step 1: Use the Vision Transformer model to extract the features of the image to obtain the semantic representation of the image;

[0009] Step 2: Construct a channel estimation and joint loss function;

[0010] Step 3: Use a joint source-channel encoder to encode the semantic representation in Step 1 in combination with the channel estimation and joint loss function in Step 2;

[0011] Step 4: Perform NOMA power allocation on the encoding result in Step 3 and transmit it through a wireless link;

[0012] The signal reception method includes the following steps:

[0013] Step 5: The receiving end receives the signal transmitted in Step 4 from the link;

[0014] Step 6: Use a joint source-channel decoder to recover the original user data from the noise;

[0015] Step 7: Use the joint source-channel decoder to map the output of the decoder in Step 6 back to a high-dimensional space to prepare for image reconstruction;

[0016] Step 8: Use the Vision Transformer model to recover the source information according to the semantic representation in Step 7.

[0017] Further, the specific method of Step 1 is:

[0018] The input image is input into a pre-trained Vision Transformer model; the Vision Transformer (ViT) model divides the image into small blocks of 16×16 pixels and converts each small block into a feature vector; ViT captures the global dependencies between different small blocks in the image through a multi-head attention mechanism, generates a highly abstract semantic representation, and outputs a feature vector Z0. The specific formula is as follows:

[0019] z i =W patch ·x i +b patch ,i=1,2,…,N,

[0020] Z0=Z+E pos ,

[0021] Among them, W patch is a trainable weight matrix, b patch is a bias vector, E pos is a position encoding, z i and Z are embedding vectors in the intermediate process.

[0022] Furthermore, the specific method of step two is as follows:

[0023] The channel state information includes the signal-to-noise ratio and the channel type; through the channel state information embedding layer, the signal-to-noise ratio and the channel type are converted into feature vectors of a fixed dimension, and the CSI feature vector is fused with the image feature vector Z0 to generate a joint feature representation. The specific formula is as follows:

[0024] x fused = Z0 + CSI

[0025] where CSI is the feature vector obtained by converting the signal-to-noise ratio and the channel type through the embedding layer;

[0026] The joint loss function combines the image reconstruction loss and the CSI prediction loss. The image reconstruction loss uses the mean squared error (MSE) to measure the difference between the predicted CSI and the true CSI. The CSI prediction loss uses the mean squared error as the image reconstruction loss to measure the difference between the predicted CSI and the true CSI. The formula for the joint loss function is:

[0027] Loss Total = α·Loss Recon + β·Loss CSI

[0028] where Loss Recon and Loss CSI are the image reconstruction loss and the channel prediction loss respectively, and α and β are the corresponding weight coefficients used to balance image reconstruction and CSI prediction;

[0029] Loss Recon and Loss CSI are calculated as follows:

[0030]

[0031] where and x (i) represent the reconstructed image and the original image respectively, and represent the predicted channel state information and the true channel state information respectively.

[0032] Furthermore, the specific method of step three is as follows:

[0033] The joint source-channel encoder combines the semantic information x emb with the channel state information to generate a low-dimensional coded signal suitable for transmission over the wireless channel: The CSI features are directly added to the image features using element-wise addition to complete the fusion of the two; the fused features are input into a Transformer-based encoder, which consists of multiple Transformer layers, each layer containing a multi-head attention mechanism and a feed-forward neural network. The specific formula is as follows:

[0034] C = x fused W enc + b enc

[0035] where W enc is the coding weight matrix and b enc is the coding bias vector; to enhance the non-linear expression ability of the encoder, an activation function is also introduced:

[0036] C' = ReLU(C) ∈ C B×M

[0037] After that, to meet the power constraint of channel transmission, the coded signal needs to be normalized:

[0038]

[0039] where P is the transmission power and ||C|| is the Frobenius norm of matrix C', and the calculation method is:

[0040]

[0041] Through the self-attention mechanism, the encoder can capture the global dependencies between features, thereby generating a more robust signal C encoded , and the specific formula is as follows:

[0042] C encoded = C” ∈ C B×M

[0043] C encoded contains the semantic information of the image and the channel state information.

[0044] Furthermore, the specific method of step four is:

[0045] The encoded signal x encoded is assigned to multiple users, and each user is weighted according to its power allocation ratio; the system has n users, and the power allocation ratio of each user is p i , where it is required to satisfy:

[0046]

[0047] Then the signal of each user is:

[0048]

[0049] The signals of all users are superimposed together to form the final transmitted signal x transmit ; The transmitted signal is transmitted through the wireless channel, and the receiving end receives the signal y received ;

[0050] The specific method of Step 6 is:

[0051] The joint source-channel decoder first preprocesses the received signal, including denoising and normalization operations, to eliminate noise and interference in the channel; the signal is input into the decoder based on Transformer, and the decoder consists of multiple Transformer linear layers, each layer containing a multi-head attention mechanism and a feed-forward neural network. The first linear layer maps the received signal to the hidden space, and then increases the non-linearity through the activation function:

[0052] H1 = ReLU(W1C′ k + b1)

[0053] where W1 is the weight matrix of the first layer, b1 is the bias vector of the first layer, and H1 is the number of hidden units in the first layer; then feature extraction is performed through several hidden layers:

[0054] H l = ReLU(W l H l-1 + b l ), l = 2, 3, …, L - 1,

[0055] where W l is the weight matrix of the l-th layer, b l is the bias vector of the l-th layer, H l is the number of hidden units in the l-th layer, and the last layer maps the hidden features back to the dimension of the semantic feature vector:

[0056]

[0057] where W L is the weight matrix of the output layer, b L is the bias vector of the output layer, and D is the dimension of the semantic feature vector;

[0058] Through the self-attention mechanism, the decoder captures the global dependencies in the signal, gradually reconstructs the semantic features of the image, and finally the decoder outputs the reconstructed feature vector

[0059] Further, the specific method of step seven is as follows:

[0060] The decoded features are input into the reconstruction head, which includes a fully connected layer and an upsampling layer. The fully connected layer converts the feature vector into image patches, which represent a low-resolution version of the image. The upsampling layer enlarges these patches to the resolution of the original image, and the reconstructed image x reconstructed is output.

[0061] Further, the specific method of step eight is as follows:

[0062] After obtaining the reconstructed image x reconstructed , semantic optimization of the image is performed through the masked self-attention mechanism of the Transformer; the masked self-attention mechanism captures the global dependencies between different regions in the image, thereby restoring the global structure of the image; the cross-attention module is used to query the correlation degree between the target sequence and the feature vector; the cross-attention mechanism helps the model understand the relationship between different parts of the image to improve the reconstruction quality; the mapping from the feature vector to the output sequence is completed through a feed-forward neural network to achieve the restoration of semantic information and obtain the final optimized image.

[0063] The present invention also relates to a satellite network semantic communication system based on reliable joint source-channel coding, and the system includes a computer module that utilizes the above-mentioned multi-objective spatio-temporal switching decision method for the satellite-ground integrated network.

[0064] The present invention also relates to a computer device, including a memory and a processor. The memory stores a computer program, and it is characterized in that when the processor executes the computer program, the steps of the satellite network semantic communication method based on reliable joint source-channel coding are realized.

[0065] The present invention also relates to a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the satellite network semantic communication method based on reliable joint source-channel coding are realized.

[0066] Beneficial effects

[0067] The present invention is directed to a satellite-ground integrated network, and proposes a satellite network semantic communication method and system based on reliable joint source-channel coding. The present invention uses a pre-trained Vision Transformer (ViT) model to extract semantic information from images, significantly reducing the amount of data to be transmitted while maintaining communication efficiency.

[0068] In addition, to improve the communication performance of the space-ground integrated network under low signal-to-noise ratio (SNR) conditions, the present invention proposes a reliable joint source channel coding (Reliable Joint Source Channel Coding, JSCC, RJSCC) method for semantic communication systems. This method significantly improves system reliability through accurate channel estimation and a carefully designed joint loss function, and jointly optimizes channel adaptation and semantic reconstruction to obtain robust performance in complex space-ground links. To cope with the situation of large-scale access, the present invention also introduces NOMA technology to allocate different communication powers to users to ensure reliable communication under different channel conditions.

[0069] The simulation results show that the RJSCC algorithm proposed by the present invention can improve the communication quality of the system under low SNR. The image quality is improved by 14.8% and 58.8% in the space-ground channel and the Rician channel respectively. Compared with orthogonal multiple access, the NOMA scheme has higher spectral efficiency and user service quality. Brief Description of the Drawings

[0070] Figure 1 It is a semantic communication scenario diagram of the space-ground integrated network of the present invention.

[0071] Figure 2 It is a technical roadmap of the semantic communication method for satellite networks based on reliable joint source channel coding of the present invention.

[0072] Figure 3 It is a comparison diagram of the performance of semantic communication and traditional communication of the present invention.

[0073] Figure 4a It is a comparison diagram of the PSNR performance of the reliable joint source channel coding scheme and the ordinary joint source channel coding scheme of the present invention on the AWGN channel.

[0074] Figure 4b It is a comparison diagram of the PSNR performance of the reliable joint source channel coding scheme and the ordinary joint source channel coding scheme of the present invention on the Rician channel.

[0075] Figure 4c It is a comparison diagram of the PSNR performance of the reliable joint source channel coding scheme and the ordinary joint source channel coding scheme of the present invention on the space-ground channel.

[0076] Figure 4d It is a comparison diagram of the SSIM performance of the reliable joint source channel coding scheme and the ordinary joint source channel coding scheme of the present invention on the AWGN channel.

[0077] Figure 4eThis is a comparison chart of the SSIM performance between the reliable joint source-channel coding scheme of the present invention and the ordinary joint source-channel coding scheme over a Rician channel.

[0078] Figure 4f This is a comparison chart of the SSIM performance between the reliable joint source-channel coding scheme of the present invention and the ordinary joint source-channel coding scheme over a satellite-ground channel.

[0079] Figure 5a This is a comparison chart of the performance of the PSNR difference between users under the NOMA and OMA frameworks for the satellite network semantic communication method based on reliable joint source-channel coding of the present invention.

[0080] Figure 5b This is a comparison chart of the performance of the SSIM difference between users under the NOMA and OMA frameworks for the satellite network semantic communication method based on reliable joint source-channel coding of the present invention. Detailed implementation manners

[0081] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but it is not intended to limit the present invention.

[0082] In the present invention, a satellite network semantic communication method based on reliable joint source-channel coding is proposed for satellite-ground networks. The present invention uses a pre-trained ViT model to perform semantic extraction on the image to be transmitted, and uses a Transformer model to fuse channel state information and a joint loss function to propose a reliable joint source-channel coding method, and uses NOMA technology for non-orthogonal multiple access in channel transmission, realizing semantic communication of a reliable joint source-channel coding scheme under satellite-ground networks.

[0083] The satellite network semantic communication method based on reliable joint source-channel coding of the present invention includes a signal transmission method and a signal reception method;

[0084] The signal transmission method includes the following steps:

[0085] Step 1: Use a Vision Transformer model to extract the features of the image to obtain the semantic representation of the image;

[0086] The input image is input into a pre-trained Vision Transformer model; the Vision Transformer model divides the image into small blocks of 16×16 pixels and converts each small block into a feature vector; the Vision Transformer model captures the global dependencies between different small blocks in the image through a multi-head attention mechanism, generates a highly abstract semantic representation, and outputs a feature vector Z0. The specific formula is as follows:

[0087] zi = W patch ·x i + b patch , i = 1, 2, …, N,

[0088] Z0 = Z + E pos ,

[0089] where W patch is a trainable weight matrix, b patch is a bias vector, E pos is a position encoding, z i and Z are embedding vectors of intermediate processes.

[0090] Step 2: Construct the channel estimation and joint loss function;

[0091] The channel state information includes the signal-to-noise ratio and the channel type; through the channel state information embedding layer, the signal-to-noise ratio and the channel type are converted into feature vectors of a fixed dimension, and the CSI feature vector is fused with the image feature vector Z0 to generate a joint feature representation. The specific formula is as follows:

[0092] x fused = Z0 + CSI

[0093] where CSI is the feature vector converted from the signal-to-noise ratio and the channel type through the embedding layer;

[0094] The joint loss function combines the image reconstruction loss and the CSI prediction loss. The image reconstruction loss uses the mean squared error (MSE) to measure the difference between the predicted CSI and the true CSI, and the CSI prediction loss uses the mean squared error of the image reconstruction loss to measure the difference between the predicted CSI and the true CSI; the formula of the joint loss function is:

[0095] Loss Total = α·Loss Recon + β·Loss Recon

[0096] where Loss Recon and Loss CSI are the image reconstruction loss and the channel prediction loss respectively, and α and β are the corresponding weight coefficients used to balance the image reconstruction and the CSI prediction;

[0097] Loss Recon and Loss CSI are calculated as follows:

[0098]

[0099] where and x(i) respectively represent the reconstructed image and the original image, and respectively represent the predicted channel state information and the true channel state information.

[0100] Step 3: Use the joint source-channel encoder to encode the semantic representation in Step 1 by combining the channel estimation and the joint loss function in Step 2;

[0101] The joint source-channel encoder combines the semantic information x emb with the channel state information to generate a low-dimensional encoded signal suitable for transmission in the wireless channel: directly add the CSI features to the image features element-wise to complete the fusion of the two; the fused features are input into an encoder based on Transformer, which consists of multiple Transformer layers, and each layer contains a multi-head attention mechanism and a feed-forward neural network. The specific formula is as follows:

[0102] C = x fused W enc + b enc

[0103] where, W enc is the encoding weight matrix, and b enc is the encoding bias vector; to enhance the non-linear expression ability of the encoder, an activation function also needs to be introduced:

[0104] C' = ReLU(C) ∈ C B×M

[0105] After that, to meet the power constraint of channel transmission, the encoded signal also needs to be normalized:

[0106]

[0107] where, P is the transmission power, ||C|| is the Frobenius norm of matrix C', and the calculation method is:

[0108]

[0109] Through the self-attention mechanism, the encoder can capture the global dependence relationship between features, thereby generating a more robust signal C encoded as follows:

[0110] C encoded = C” ∈ C B×M

[0111] C encoded contains the semantic information of the image and the channel state information.

[0112] Step 4: Perform NOMA power allocation on the encoding result of Step 3 and transmit it through the wireless link;

[0113] The encoded signal x encoded is allocated to multiple users, and each user is weighted according to its power allocation ratio; the system has n users, and the power allocation ratio of each user is p i , where it is necessary to satisfy:

[0114]

[0115] Then the signal of each user is:

[0116]

[0117] The signals of all users are superimposed together to form the final transmitted signal x transmit ; The transmitted signal is transmitted through the wireless channel, and the receiving end receives the signal y received ;

[0118] The signal receiving method includes the following steps:

[0119] Step 5: The receiving end receives the signal transmitted in Step 4 from the link;

[0120] Step 6: Use the joint source-channel decoder to recover the original user data from the noise;

[0121] The joint source-channel decoder first preprocesses the received signal, including denoising and normalization operations, to eliminate noise and interference in the channel; the signal is input into a decoder based on Transformer, and the decoder consists of multiple Transformer linear layers, each layer containing a multi-head attention mechanism and a feed-forward neural network. The first linear layer maps the received signal to the hidden space, and then the activation function is used to increase the non-linearity:

[0122] H1 = ReLU(W1C′ k + b1)

[0123] where, W1 is the weight matrix of the first layer, b1 is the bias vector of the first layer, and H1 is the number of hidden units of the first layer; subsequently, feature extraction is performed through several hidden layers:

[0124] H l = ReLU(W l H l-1 + b l ), l = 2, 3, …, L - 1,

[0125] where, W l is the weight matrix of the l-th layer, b l is the bias vector of the l-th layer, Hl is the number of hidden units in the l-th layer, and the dimension of the last layer that maps the hidden features back to the semantic feature vector:

[0126]

[0127] where W L is the weight matrix of the output layer, b L is the bias vector of the output layer, and D is the dimension of the semantic feature vector;

[0128] Through the self-attention mechanism, the decoder captures the global dependencies in the signal and gradually reconstructs the semantic features of the image. Finally, the decoder outputs the reconstructed feature vector

[0129] Step 7: Use the joint source-channel decoder to map the output of the decoder in Step 6 back to the high-dimensional space to prepare for image reconstruction;

[0130] The decoded features are input into the reconstruction head, which includes a fully connected layer and an upsampling layer. The fully connected layer converts the feature vector into image patches, which represent a low-resolution version of the image. The upsampling layer enlarges these patches to the resolution of the original image, and the reconstructed image x reconstructed is output.

[0131] Step 8: Use the Vision Transformer model to recover the source information according to the semantic representation in Step 7.

[0132] After obtaining the reconstructed image x reconstructed the semantic optimization of the image is performed through the masked self-attention mechanism of the Transformer; the masked self-attention mechanism captures the global dependencies between different regions in the image, thereby restoring the global structure of the image; the cross-attention module is used to query the correlation degree between the target sequence and the feature vector; the cross-attention mechanism helps the model understand the relationship between different parts of the image to improve the reconstruction quality; the mapping between the feature vector and the output sequence is completed through the feed-forward neural network to achieve the restoration of semantic information and obtain the final optimized image.

[0133] Multi-user communication model based on semantic communication

[0134] In this embodiment, the present invention relates to a semantic communication architecture in a space-ground integrated network based on non-orthogonal multiple access technology.

[0135] The first step of transmission is to use the ViT model to perform semantic extraction on the original image. This model first divides the input image into several small patches, and then flattens each patch x i into a vector X i, and then it is mapped to dimension D through a linear transformation to obtain the embedding vector z i , and the specific process is as follows:

[0136] z i = W patch ·x i + b patch , i = 1, 2, …, N,

[0137] Among them, W is the linear embedding weight matrix, and b is the bias vector. In order to retain the position information of each small block in the image, a position encoding E is introduced, and the final input embedding vector Z0 can be expressed as:

[0138] Z0 = Z + E pos ,

[0139] Among them, Z is the combination of the embedding vectors of all small blocks. In each layer of the Transformer, the multi-head attention mechanism module captures the global dependencies through the query (Query), key (Key), and value (Value) matrices, and their calculation formulas are as follows in turn:

[0140] Q = Z l W Q ,

[0141] K = Z l W K ,

[0142] V = Z l W V ,

[0143] Among them, Z l is the input of the l-th layer, and W Q , W K and W V are the weight matrices of the query, key, and value. Then, according to the number of heads of the multi-head attention, Q, K, and V are partitioned, and the attention weight matrix is calculated for each head respectively. After weighted summation, the outputs of all heads are concatenated and passed through a linear transformation to obtain the final output of the multi-head attention. Then, residual connection and layer normalization are also required. Just like this, after passing through the L-layer Transformer module, the final semantic feature vector is obtained.

[0144] After the semantic feature vector passes through the joint source-channel coding, it will be sent into the channel for transmission. In order to achieve the simultaneous transmission of multiple users, the present invention adopts the NOMA technology to allocate different transmission powers to different users, and uses superposition coding and successive interference cancellation for decoding. There are 4 users accessing simultaneously in the system of the present invention, and the power allocated to each user is P k , and it satisfies the total power constraint:

[0145]

[0146] Among them, P k is the allocated power of the k-th user, and P total is the total transmission power. After superposition coding, the coded signal C of each user k after power allocation is superimposed to obtain:

[0147]

[0148] Then it is sent into the wireless channel for transmission. Based on the power allocation strategy, the receiving end decodes the received signals for each user in the order of decreasing power to perform effective successive interference cancellation. First, decode the signal of the user with k = 1 having a higher power:

[0149]

[0150] After successful decoding, eliminate the signal interference of this user from the received signal:

[0151]

[0152] Then repeat the above process to decode the signals of the remaining users in turn.

[0153] Reliable joint source-channel coding scheme

[0154] In this embodiment, the present invention relates to a reliable joint source-channel coding method that integrates channel state information and a joint loss function. The channel state information includes signal-to-noise ratio and channel type, which will be fused with the image feature vector to generate a joint feature representation. The specific method is as follows:

[0155] x fused = x emb + CSI

[0156] Among them, x emb represents the semantic feature vector, and CSI represents the channel state information. The joint loss function combines the image reconstruction loss and the CSI prediction loss. The image reconstruction loss uses the mean squared error (MSE) to measure the difference between the predicted CSI and the true CSI, and the CSI prediction loss also uses MSE to measure the difference between the predicted CSI and the true CSI. The formula for the joint loss function is:

[0157] Loss Total = α·Loss Recon + β·Loss Recon ,

[0158] Among them, Loss Recon and Loss CSIThey are the image reconstruction loss and the channel prediction loss respectively, and α and β are the corresponding weight coefficients used to balance image reconstruction and CSI prediction. Loss Recon and Loss CSI are calculated as follows:

[0159]

[0160] where, and x (i) represent the reconstructed image and the original image respectively, and represent the predicted channel state information and the true channel state information respectively. In addition, according to different task requirements, metrics such as peak signal-to-noise ratio and structural similarity of the image can be added to the joint loss function to adjust the focus of the system.

[0161] Joint source-channel codec based on Transformer

[0162] In this embodiment, the present invention relates to a joint source-channel encoder that encodes the extracted semantic feature vectors into signals suitable for channel transmission, enhancing the anti-interference ability and data reconstruction ability.

[0163] First, after obtaining the semantic vector S, the encoder will perform encoding weights and biases, and then perform an offline transformation on S. The specific method is as follows:

[0164] C = SW enc + b enc ,

[0165] where, W enc is the encoding weight matrix, and b enc is the encoding bias vector. To enhance the non-linear expression ability of the encoder, an activation function also needs to be introduced:

[0166] C' = ReLU(C) ∈ C B×M .

[0167] After that, to meet the power constraint of channel transmission, the encoded signal also needs to be normalized:

[0168]

[0169] where, P is the transmission power, and ||C|| is the Frobenius norm of matrix C', and the calculation method is:

[0170]

[0171] Finally, the encoded signal is obtained:

[0172] C encoded = C” ∈ CB×M 。

[0173] The main task of the joint source-channel decoder is to restore the received signal C' k to the original semantic feature vector. First, the first linear layer of the decoder maps the received signal to the hidden space, and then adds non-linearity through the activation function:

[0174] H1 = ReLU(W1C′ K +b1)

[0175] where W1 is the weight matrix of the first layer, b1 is the bias vector of the first layer, and H1 is the number of hidden units in the first layer. Subsequently, feature extraction is performed through several hidden layers:

[0176] H l = ReLU(W l H l-1 +b l ), l = 2, 3, …, L-1,

[0177] where W l is the weight matrix of the l-th layer, b l is the bias vector of the l-th layer, H l is the number of hidden units in the l-th layer, and the last layer maps the hidden features back to the dimension of the semantic feature vector:

[0178]

[0179] where W L is the weight matrix of the output layer, b L is the bias vector of the output layer, and D is the dimension of the semantic feature vector.

[0180] In this way, the design of the joint source-channel codec is completed.

[0181] Channel simulation

[0182] In this embodiment, a satellite channel model proposed by the present invention is presented. The input signal first calculates the Doppler frequency shift based on the carrier frequency and velocity of the satellite to simulate the frequency change caused by satellite movement. The signal sequence is randomly divided into multiple segments, and each segment may be subject to different types of interference, including the Doppler effect, multipath fading, and shadow fading. These interference types are implemented through random selection to simulate the variable channel conditions in the actual environment. The Doppler effect modulates the signal with a cosine function to reflect the frequency offset caused by the satellite velocity; multipath fading simulates the superposition effect of the signal after propagating through multiple paths by adding random noise, enhancing the uncertainty of the signal; shadow fading adjusts the signal strength through a randomly generated scaling factor to simulate the attenuation effect of obstacles on the signal. These processing steps ensure that the signal can truly reflect various complex interferences in the satellite channel during transmission.

[0183] After processing each interference segment, the model also adds additive white Gaussian noise (AWGN) according to the set signal-to-noise ratio (SNR) to further simulate the influence of channel noise. Specifically, the average power of the signal is first calculated, and then the standard deviation of the noise is determined according to the given SNR value. Random noise with the same dimension as the signal is generated and scaled proportionally before being superimposed on the signal. This process ensures that the final output signal not only contains various interferences but also has a controllable noise level, enabling the testing and evaluation of the robustness and performance of communication systems under different channel conditions. In addition, the model is made flexible by parameterizing the carrier frequency, satellite velocity, and standard deviation of shadow fading, enabling it to adapt to different satellite communication scenarios and requirements.

[0184] System Training and Performance Simulation Analysis

[0185] In this embodiment, the PSNR and SSIM performance analysis of reliable joint channel coding for semantic communication in a space-ground integrated network is given. In the simulation experiment, the semantic communication model integrates the ViT and Transformer models to respectively achieve semantic information extraction and recovery and joint source-channel coding and decoding. The system first generates image embeddings through the model and encodes them with CSI, then performs NOMA power allocation to generate signals for each user, and finally transmits the signals through a simulated channel. The images of each user are constructed through the decoder, thus realizing semantic communication in a multi-user scenario in the space-ground integrated network. Figure 2 The specific relevant parameters are shown. Table 1 shows the simulation parameters of the reliable joint source-channel coding method and system for semantic communication in the space-ground integrated network.

[0186] Table 1

[0187]

[0188]

[0189] As Figure 3 shown, the performance of four users in the system is averaged, and the performance of traditional communication and semantic communication is compared. The results show that semantic communication has significant advantages under low signal-to-noise ratio conditions. For example, when the SNR is 0, the PSNR of the traditional communication method in the satellite-ground channel is 9.72 dB, and the SSIM is 0.0528, with very poor image quality. While for semantic communication in this channel, the PSNR reaches 20.66 dB and the SSIM is 0.7513, with significantly improved image quality. Similarly, under AWGN conditions, although the performance difference between the two communication methods decreases at moderate SNR, overall, semantic communication still outperforms traditional communication at medium and low levels of SNR, demonstrating stronger transmission capabilities. Under the Rician channel, although the performance of the traditional communication method has improved, semantic communication still maintains an obvious advantage. That is to say, semantic communication can transmit high-quality image information under limited bandwidth and harsh conditions, ensuring the effectiveness and reliability of communication.

[0190] In Figure 4, the performance differences between the improved JSCC algorithm and the original JSCC algorithm under different conditions are further compared, and it is found that the improved algorithm indeed performs better in the satellite-ground channel. As shown in Figure 4, when the SNR is 0 dB, the PSNR of the improved JSCC algorithm increases from 20.6658 to 21.2958, and the SSIM increases from 0.7513 to 0.7697, with improvements of 0.63 dB and 0.0184 respectively. According to the calculation formula, the quality of the reconstructed image is improved by approximately 14.8%. When the SNR is 5, the improved algorithm still maintains a strong advantage, indicating the wide applicability of the improved algorithm in practical applications. However, at high SNR, although the PSNR of the improved algorithm is slightly lower than that of the original algorithm, the difference is small. At the same time, the SSIM value of the improved algorithm is always higher than that of the original algorithm, indicating that it has more advantages in terms of the fidelity of the image structure. Under other channel conditions, the improved algorithm also shows excellent performance. Especially under the Rician channel, the PSNR of the improved algorithm is increased by approximately 3.79 dB compared to the original algorithm, and the SSIM is increased by approximately 0.069. According to the calculation formula, the quality of the reconstructed image is improved by approximately 58.8%, indicating its adaptability and robustness under complex channel conditions, further enhancing the practicality and application prospects of the algorithm.

[0191] As shown in Figure 5, in terms of user fairness, NOMA also performs excellently in the satellite-terrestrial channel. At an SNR of 0 dB, the PSNR between NOMA users is only 0.0003 dB, and the SSIM difference is only 0.000018, indicating that there is almost no difference in the signal quality among users, ensuring the balance of the user experience. In contrast, OMA performs poorly in the satellite-terrestrial channel, with a PSNR difference reaching 0.0589 and an SSIM difference of 0.0013, showing a relatively significant difference in the user experience. Such unequal signal differences may lead to poor communication for some users, affecting the overall user perception. In the AGWN channel, although the signal differences between NOMA users also perform well, they are still relatively inferior compared to the satellite-terrestrial channel. The user fairness in the NOMA system also remains good in the Rician channel. For the satellite-terrestrial communication environment, users are widely distributed and the channel conditions are complex. Good user fairness ensures that all users can obtain stable communication services regardless of the environment.

[0192] Although the present invention has been described with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the present invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed, as long as they do not depart from the spirit and scope of the present invention as defined by the appended claims. It should be understood that the features described in the different dependent claims and in the present invention can be combined in a different manner than that described in the original claims. It should also be understood that the features described in connection with a single embodiment can be used in other described embodiments.

Claims

1. A satellite network semantic communication method based on reliable joint source-channel coding, characterized in that, Including a signal transmission method and a signal reception method; The signal transmission method includes the following steps: Step 1: Use the Vision Transformer model to extract the features of the image to obtain the semantic representation of the image; Step 2: Construct a channel estimation and joint loss function; Step 3: Use a joint source-channel encoder to encode the semantic representation in Step 1 in combination with the channel estimation and joint loss function in Step 2; Step 4: Perform NOMA power allocation on the encoding result in Step 3 and transmit it through a wireless link; The signal reception method includes the following steps: Step 5: The receiving end receives the signal transmitted in Step 4 from the link; Step 6: Use a joint source-channel decoder to recover the original user data from the noise; Step 7: Use the joint source-channel decoder to map the output of the decoder in Step 6 back to a high-dimensional space to prepare for image reconstruction; Step 8: Use the Vision Transformer model to recover the source information according to the semantic representation in Step 7.

2. The satellite network semantic communication method based on reliable joint source-channel coding according to claim 1, characterized in that, The specific method of Step 1 is: The input image is input into a pre-trained Vision Transformer model; the Vision Transformer model divides the image into small blocks of 16×16 pixels and converts each small block into a feature vector; the Vision Transformer model captures the global dependencies between different small blocks in the image through the multi-head attention mechanism, generates a highly abstract semantic representation, and outputs the feature vector Z0. The specific formula is as follows: z i = W patch · x i + b patch , i = 1, 2, …, N, Z0 = Z + E pos , Among them, W patch is a trainable weight matrix, b patch is a bias vector, E pos is a positional encoding, z i and Z are embedding vectors in the intermediate process.

3. The satellite network semantic communication method based on reliable joint source-channel coding according to claim 1, characterized in that The specific method of Step 2 is: The channel state information includes the signal-to-noise ratio and the channel type; through the channel state information embedding layer, the signal-to-noise ratio and the channel type are converted into feature vectors of a fixed dimension, and the CSI feature vector is fused with the image feature vector Z0 to generate a joint feature representation. The specific formula is as follows: x fused = Z0 + CSI Among them, CSI is the feature vector obtained by converting the signal-to-noise ratio and the channel type through the embedding layer; The joint loss function combines the image reconstruction loss and the CSI prediction loss. The image reconstruction loss uses the mean squared error (MSE) to measure the difference between the predicted CSI and the true CSI, and the CSI prediction loss uses the mean squared error of the image reconstruction loss to measure the difference between the predicted CSI and the true CSI; the formula of the joint loss function is: Loss Total = α·loss Recon + β·Loss CSI where Loss Recon and Loss CSI are the image reconstruction loss and the channel prediction loss, respectively, and α and β are the corresponding weight coefficients used to balance image reconstruction and CSI prediction; Loss Recon and Loss CSI The calculation formula is as follows: Among them, and x (i) represent the reconstructed image and the original image respectively, and represent the predicted channel state information and the true channel state information respectively.

4. The satellite network semantic communication method based on reliable joint source-channel coding according to claim 1, wherein The specific method of Step 3 is: The joint source-channel encoder combines the semantic information x emb with the channel state information to generate a low-dimensional coded signal suitable for transmission in the wireless channel: the CSI features are directly added to the image features by element-wise addition to complete the fusion of the two; the fused features are input into a Transformer-based encoder, which consists of multiple Transformer layers, each layer containing a multi-head attention mechanism and a feed-forward neural network, and the specific formula is as follows: C = x fused W enc + b enc Among them, W enc is the encoding weight matrix, and b enc is the encoding bias vector; in order to enhance the non-linear expression ability of the encoder, an activation function also needs to be introduced: C' = ReLU(C) ∈ C B×M After that, in order to meet the power constraint of channel transmission, the encoded signal also needs to be normalized: Among them, P is the transmission power, ||C|| is the Frobenius norm of matrix C', and the calculation method is: Through the self-attention mechanism, the encoder can capture the global dependencies between features, thereby generating a more robust signal C encoded , and the specific formula is as follows: C encoded = C” ∈ C B×M C encoded Contains semantic information of the image and channel state information.

5. The reliable joint source-channel coding method and system based on semantic communication under a satellite-ground network according to claim 1, characterized in that The specific method of Step 4 is: Encoded signal x encoded is assigned to multiple users, and each user is weighted according to its power allocation ratio; there are n users in the system, and the power allocation ratio of each user is p i , where it is necessary to satisfy: Then the signal of each user is: The signals of all users are superimposed together to form the final transmitted signal x transmit ; The transmitted signal is transmitted through a wireless channel, and the receiving end receives the signal y received ; The specific method of Step 6 is: The joint source-channel decoder first preprocesses the received signal, including denoising and normalization operations, to eliminate noise and interference in the channel; the signal is input into a Transformer-based decoder, which consists of multiple Transformer linear layers, and each layer contains a multi-head attention mechanism and a feed-forward neural network. The first linear layer maps the received signal to a hidden space, and then the activation function is used to increase non-linearity: H1 = ReLU(W1C′ k + b1) Among them, W1 is the weight matrix of the first layer, b1 is the bias vector of the first layer, and H1 is the number of hidden units in the first layer; then feature extraction is performed through several hidden layers: H l = ReLU(WlHl -1 + bl), l = 2, 3, …, L - 1, Among them, W l is the weight matrix of the l-th layer, b l is the bias vector of the l-th layer, H l is the number of hidden units in the l-th layer. The last layer maps the hidden features back to the dimension of the semantic feature vector: Among them, W L is the weight matrix of the output layer, b L is the bias vector of the output layer, and D is the dimension of the semantic feature vector; Through the self-attention mechanism, the decoder captures the global dependencies in the signal, gradually reconstructs the semantic features of the image, and finally the decoder outputs the reconstructed feature vector 6. The satellite network semantic communication method based on reliable joint source-channel coding according to claim 1, wherein, The specific method of step seven is: Decoded features are input into a reconstruction head, which includes a fully connected layer and an upsampling layer. The fully connected layer transforms the feature vectors into image patches that represent a low-resolution version of the image. The upsampling layer enlarges these patches to the resolution of the original image, and the reconstructed image x reconstructed is output.

7. The satellite network semantic communication method based on reliable joint source-channel coding according to claim 1, wherein The specific method of step eight is: After obtaining the reconstructed image x reconstructed the semantic optimization of the image is performed through the masked self-attention mechanism of the Transformer; the masked self-attention mechanism captures the global dependencies between different regions in the image, thereby restoring the global structure of the image; the cross-attention module is used to query the correlation degree between the target sequence and the feature vector; the cross-attention mechanism helps the model understand the relationship between different parts in the image to improve the reconstruction quality; the mapping from the feature vector to the output sequence is completed through the feed-forward neural network to realize the restoration of semantic information and obtain the final optimized image.

8. A satellite network semantic communication system based on reliable joint source-channel coding, characterized in that The system includes a computer module that uses the above multi-objective space-time switching decision method for satellite-ground integrated networks.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the satellite network semantic communication method based on reliable joint source-channel coding.

10. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the satellite network semantic communication method based on reliable joint source-channel coding.

Citation Information

Cited By

  • Joint source channel coding channel state information feedback method based on dynamic resource selection

    CN122293269A

  • Joint source-channel coding channel state information feedback method based on dynamic resource selection

    CN122293269B