A method for adaptive semantic communication based on reinforcement learning
By adopting an adaptive semantic communication method based on reinforcement learning, the resource allocation imbalance problem in semantic communication under time-varying fading channels in existing technologies is solved. Adaptive scheduling and joint decision-making across different fading scenarios are realized, improving the performance and efficiency of communication tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- UESTC (SHENZHEN) ADVANCED RES INST
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-08
AI Technical Summary
Existing digital semantic communication schemes struggle to achieve end-to-end adaptive optimal transmission under time-varying fading channels and multi-objective trade-offs. Furthermore, the joint decision-making of parameters such as feature selection, quantization bit allocation, and coding rate is difficult to resolve efficiently, leading to resource allocation imbalances.
An adaptive semantic communication method based on reinforcement learning is adopted. By constructing an adaptive semantic transmission architecture, the method utilizes a policy network to make adaptive joint decisions on physical layer parameters such as coding rate, modulation order, and transmit power. By combining the importance and relevance of semantic features, adaptive scheduling is achieved across different fading scenarios.
It achieves better task performance and transmission efficiency under the constraints of latency and energy budget, and improves the overall efficiency and performance of semantic communication.
Smart Images

Figure CN121619066B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of semantic communication, and in particular to an adaptive semantic communication method based on reinforcement learning. Background Technology
[0002] Existing digital semantic communication schemes largely rely on fixed or empirical transmission and resource allocation strategies, making it difficult to achieve end-to-end adaptive optimal transmission under time-varying fading channels and multi-objective trade-offs. Furthermore, semantic information representation typically consists of multiple features, with significant differences in their contribution to the final task, and features often exhibit redundancy. The lack of a mechanism to simultaneously characterize "task importance" and "relevance" at the feature level leads to unbalanced resource allocation, limiting overall efficiency and performance improvement. Moreover, the joint decision-making of feature selection, quantization bit allocation, and physical layer parameters such as coding rate, modulation order, and transmit power constitutes a high-dimensional hybrid optimization problem, which is difficult to solve using traditional analytical or heuristic methods. Therefore, there is an urgent need for a semantic feature-oriented scheduling and physical layer joint configuration scheme that can dynamically adapt to channel and constraints to meet communication requirements and improve task performance. Summary of the Invention
[0003] The purpose of this invention is to overcome the shortcomings of the prior art and provide an adaptive semantic communication method based on reinforcement learning, which can realize the selection of semantic feature transmission across different fading scenarios and make adaptive joint decisions on physical layer parameters such as coding rate, modulation order, and transmit power to obtain better task performance and transmission efficiency.
[0004] The objective of this invention is achieved through the following technical solution: an adaptive semantic communication method based on reinforcement learning, comprising the following steps:
[0005] Step S1: Construct an adaptive semantic transmission architecture based on reinforcement learning:
[0006] The adaptive semantic transmission architecture includes a transmitter, a receiver, and a control module; the transmitter includes a semantic encoder, a quantization module, a channel coding module, and a modulation module; the receiver includes a demodulation module, a channel decoding module, an inverse quantization module, and a semantic decoder; the wireless channel between the transmitter and receiver adopts a block fading model: the channel conditions remain quasi-static during a semantic feature transmission, and the channel conditions will change between different feature transmissions;
[0007] The control module includes a policy network based on reinforcement learning, used to control the module parameters of the transmitting end;
[0008] Step S2: Pre-train several sets of semantic encoder-decoder models and construct a lookup table to record different operating points and their corresponding source rates and distortion metrics;
[0009] Step S3: Train the policy network based on the Markov decision process and reward function. The training adopts a domain randomization mechanism: for each round, randomly select or combine different fading models and statistical parameters and generate block fading sequences, and iteratively update the policy parameters on the multi-domain channel distribution.
[0010] Step S4: Based on the trained policy network, perform adaptive semantic sample transmission.
[0011] The beneficial effects of this invention are: by introducing a semantic importance metric that takes into account both the differences in task contribution and the correlation between features, this invention enables the selection of semantic feature transmission across different fading scenarios under the constraints of latency and energy budget, and performs adaptive joint decision-making on physical layer parameters such as coding rate, modulation order, and transmit power to obtain better task performance and transmission efficiency. Attached Figure Description
[0012] Figure 1 This is a schematic diagram illustrating the principle of the present invention;
[0013] Figure 2 This is a schematic diagram of an adaptive semantic transmission architecture based on reinforcement learning. Detailed Implementation
[0014] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.
[0015] like Figure 1 As shown, an adaptive semantic communication method based on reinforcement learning includes the following steps:
[0016] Step S1: Construct an adaptive semantic transmission architecture based on reinforcement learning, such as... Figure 2 As shown:
[0017] The adaptive semantic transmission architecture includes a transmitter, a receiver, and a control module; the transmitter includes a semantic encoder, a quantization module, a channel coding module, and a modulation module; the receiver includes a demodulation module, a channel decoding module, an inverse quantization module, and a semantic decoder; the wireless channel between the transmitter and receiver adopts a block fading model: the channel conditions remain quasi-static during a semantic feature transmission, and the channel conditions will change between different feature transmissions;
[0018] The control module includes a policy network based on reinforcement learning, used to control the module parameters of the transmitting end;
[0019] Step S2: Pre-train several sets of semantic encoder-decoder models and construct a lookup table to record different operating points and their corresponding source rates and distortion metrics;
[0020] S201. Construct a semantic encoder-semantic decoder model, which includes a semantic encoder and a semantic decoder, both of which are implemented using neural networks;
[0021] S202. The source rate is processed by a semantic encoder. The semantic data (e.g., an image) is encoded, and then the encoded result is decoded by a semantic decoder. The loss between the original semantic data and the decoded data is then calculated and denoted as source distortion. The calculation methods include mean squared error or cross-entropy loss;
[0022] Through parameters and ,calculate The semantic encoder-semantic decoder model is trained using the loss function of the semantic encoder-semantic decoder model.
[0023] S203. At different source rates Next, repeat step S202 to obtain multiple sets of semantic encoder-semantic decoder model training results, and then obtain the source rate. Distortion from source The Lagrange objective function is denoted as:
[0024]
[0025] The Lagrange objective function refers to: taking, from the multiple sets of semantic encoder-semantic decoder model training results obtained by repeatedly executing step S202, the optimal value... Minimum training results As a working point model pair; where, , express Minimize the semantic encoder and semantic decoder parameters;
[0026] S204. Change parameters Obtain multiple working point model pairs Where V represents the number of working point model pairs obtained, v represents the v-th working point model pair, v=1,2,…,V; and the source rate and source distortion of each working point model pair are recorded and stored in a lookup table (LUT).
[0027] Step S3: Train the policy network based on the Markov decision process and reward function. The training adopts a domain randomization mechanism: for each round, randomly select or combine different fading models and statistical parameters and generate block fading sequences, and iteratively update the policy parameters on the multi-domain channel distribution.
[0028] To guide the allocation of resources for semantic features, at the beginning of each sample transfer, the importance weight of each extracted semantic feature is first calculated. This is to simultaneously reflect the differences in task contributions and the dependencies between features. Specifically:
[0029] S301. Suppose that the sample set contains multiple samples, and each sample is a semantic data to be transmitted;
[0030] S302. The transmitter selects a sample from the sample set and provides the source rate; it then obtains the operating point model pair at the source rate from the LUT.
[0031] The semantic encoder in the working point model is used to semantically encode the selected samples to obtain the latent semantic representation. ,Will The data is divided into blocks to obtain K semantic feature blocks. The semantic feature block is a matrix with W rows and H columns, where W and H are the height and width of the semantic feature block, respectively.
[0032] S303. At the start of each sample transmission, first calculate the importance weight of each extracted semantic feature block. This is to simultaneously reflect the differences in task contributions and the dependencies between features:
[0033] Calculate task relevance factors The selected sample is forward-propagated through the semantic encoder and semantic decoder of the selected working point, that is, after being processed according to step S302, it is decoded by the decoder to obtain the decoding result.
[0034] Let the selected sample be denoted as The decoding result is ;
[0035] Calculate reconstruction loss for features The gradient magnitude is calculated, and a global average pooling is performed in the spatial dimension to obtain the factor: ,in For the first The value in the w-th row and h-th column;
[0036] Calculate the correlation factor between features :calculate Compared with any other feature cosine similarity The average of their absolute values is obtained , used to characterize the feature-related redundancy strength;
[0037] Calculating importance weights: Importance weights are defined as the product of the two factors. Then, normalize to obtain the weight vector. ;
[0038] S304. Features to be sent The sending end first allocates the number of quantization bits. The quantization module converts it into a bitstream. Then, the channel coding rate is selected sequentially. Modulation order (e.g., {4,16,64,128}-QAM) and transmit power The signal is encoded in the channel coding module, modulated in the modulation module, and then transmitted according to the transmit power.
[0039] For the features to be sent The minimum number of modulation symbols required is Transmission time is ,in, Symbol rate;
[0040] Set a transmission time budget for the sample With total energy budget This ensures that the sample transfer process satisfies And energy consumption meets .
[0041] S305. Treat the transmission of each sample as a round, model the stepwise transmission process of the corresponding semantic features as a Markov decision process, and achieve online adaptive joint decision-making through reinforcement learning.
[0042] MDP modeling
[0043] If the transmission of each sample image is regarded as a round, the stepwise transmission process of the corresponding semantic features can be modeled as a Markov decision process (MDP), and online adaptive joint decision-making can be achieved through reinforcement learning.
[0044] 1) Status: During transmission, the sending end at each decision step... Observation status: in This represents the current block fading channel gain. This is the semantic importance vector of the current sample; For feature indicator vectors, , This indicates whether the feature has been sent; 0 indicates that it has been sent, and 1 indicates that it has not been sent. and These are the remaining symbol budget and the power budget, respectively.
[0045] 2) Action: In the i-th decision step, the policy network is based on... Output action: That is: first, select the feature index to be sent in this step. Then determine the physical layer parameters corresponding to this feature, including quantization bits. coding rate Modulation order With transmission power ;
[0046] 3) State transition: After the action is executed, the symbol and power budget are deterministically updated according to the selected action cost, and the transmitted features are removed from the available set; the channel gain is implemented using a block fading first-order Markov model. The evolution proceeds to the next decision-making step;
[0047] in, This represents the channel gain transition probability at the (i+1)th decision step relative to the previous decision steps. This represents the channel gain transition probability from the i-th decision step to the (i+1)-th decision step, meaning that the channel gain transition probability of the (i+1)-th decision step relative to the previous decision steps is only related to the i-th decision step.
[0048] 4) Reward Function: The reward function is based on the features actually recovered at the receiving end. The calculation, along with latency, forms the reward:
[0049] (1)
[0050] in As a weighting factor, Transmission delay of the i-th decision step;
[0051] By training a policy network to maximize the long-term average reward, the weighted sum of long-term average semantic distortion and latency is minimized. Maximizing the long-term average reward is achieved by summing all the reward functions of each feature block for each sample and then dividing by the average of the feature block data of all samples.
[0052] 5) Domain Randomization: To improve the generalization ability of the strategy under different fading types and channel statistical conditions, this invention introduces a domain randomization mechanism for the fading channel during the training phase: at the beginning of each training round, the channel model and its statistical parameters are randomly sampled to generate the block fading process for that round, thereby obtaining the corresponding current block fading channel gain. The random sampling includes at least one or a combination of the following: (1) fading distribution type (e.g., Rayleigh / Rician / Nakagami); (2) multipath intensity or K-factor, Nakagami-m parameter, etc.; (3) average SNR / path loss, shadow fading variance; (4) time correlation coefficient or Doppler parameter; (5) if Markov channel modeling is used, the channel state transition probability matrix should also be included. By covering multiple channel domains during training, the policy network learns a robust decision mapping to channel uncertainty, thereby maintaining stable performance even when deployed to unknown or changing fading channels.
[0053] In the embodiments of this application, a small amount of target deployment scenario data can be introduced at the end of training for fine-tuning to further improve the performance stability of a specific environment: after the domain randomization training converges, the channel domain parameters initialized in the round are switched from global random sampling to statistical parameters of the target deployment scenario (or a small range of perturbation thereof), and a small amount of target scenario data is used to continue to update the policy for several rounds according to the original reinforcement learning update rules, so as to achieve fine-tuning at the end of training and improve the performance stability of a specific environment.
[0054] Step S4: Based on the trained policy network, perform adaptive semantic sample transmission.
[0055] (1) The system receives input samples and selects the target operating point model pair from the LUT. ;
[0056] (2) Semantic encoder extracts latent semantic representation And obtain K semantic feature blocks ;
[0057] (3) Calculate the semantic importance weight of each feature to obtain the importance weight vector. ;
[0058] (4) Initialize the sample's delay, symbol budget, and energy budget, and observe the current block fading channel state;
[0059] (5) At each decision step The policy network is based on the state First, select which feature to send, and then determine its physical layer parameters: And perform the corresponding quantization, encoding, modulation, and transmission processes;
[0060] (6) The receiving end performs demodulation, decoding, and inverse quantization to obtain The semantic decoder outputs the task results, and the corresponding reward is obtained according to formula (1); the system updates the remaining budget and proceeds to the next decision step;
[0061] (7) The transmission process of the sample ends when all features have been sent or the budget is exhausted. Since the strategy has been trained on multiple fading domains, it still has good adaptability and robustness when the channel statistics change.
[0062] In summary, this invention proposes an adaptive joint optimization scheme for digital semantic communication in time-varying fading channels. By constructing a decision mechanism that takes into account both task and semantic features, reinforcement learning is used to achieve online joint configuration of semantic feature scheduling and physical layer parameters such as coding rate, modulation order, and transmit power. Furthermore, a robust and real-time deployable transmission strategy is obtained through generalization design across fading scenarios, which significantly improves transmission efficiency.
[0063] The foregoing description illustrates and describes a preferred embodiment of the present invention. However, as previously stated, it should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the inventive concept described herein through the foregoing teachings or techniques or knowledge in related fields. Any modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. An adaptive semantic communication method based on reinforcement learning, characterized in that: Includes the following steps: Step S1: Construct an adaptive semantic transmission architecture based on reinforcement learning: The adaptive semantic transmission architecture includes a transmitter, a receiver, and a control module; the transmitter includes a semantic encoder, a quantization module, a channel coding module, and a modulation module; the receiver includes a demodulation module, a channel decoding module, an inverse quantization module, and a semantic decoder; the wireless channel between the transmitter and receiver adopts a block fading model: the channel conditions remain quasi-static during a semantic feature transmission, and the channel conditions will change between different feature transmissions; The control module includes a policy network based on reinforcement learning, used to control the module parameters of the transmitting end; Step S2: Pre-train several sets of semantic encoder-decoder models and construct a lookup table to record different operating points and their corresponding source rates and distortion metrics; Step S2 includes: S201. Construct a semantic encoder-semantic decoder model, which includes a semantic encoder and a semantic decoder, both of which are implemented using neural networks; S202. The source rate is processed by a semantic encoder. The semantic data is encoded, and then the encoded result is decoded by a semantic decoder. The loss between the original semantic data and the decoded data is then calculated and denoted as source distortion. The calculation methods include mean squared error or cross-entropy loss; Through parameters and ,calculate The semantic encoder-semantic decoder model is trained using the loss function of the semantic encoder-semantic decoder model. S203. At different source rates Next, repeat step S202 to obtain multiple sets of semantic encoder-semantic decoder model training results, and then obtain the source rate. Distortion from source The Lagrange objective function is denoted as: ; The Lagrange objective function refers to: taking, from the multiple sets of semantic encoder-semantic decoder model training results obtained by repeatedly executing step S202, the optimal value... Minimum training results As a working point model pair; where, , express Minimize the semantic encoder and semantic decoder parameters; S204. Change parameters Obtain multiple working point model pairs Where V represents the number of working point model pairs obtained, v represents the v-th working point model pair, v=1,2,…,V; and the source rate and source distortion of each working point model pair are recorded and stored in a lookup table (LUT). Step S3: Train the policy network based on the Markov decision process and reward function. The training adopts a domain randomization mechanism: for each round, randomly select or combine different fading models and statistical parameters and generate block fading sequences, and iteratively update the policy parameters on the multi-domain channel distribution. Step S4: Based on the trained policy network, perform adaptive semantic sample transmission.
2. The adaptive semantic communication method based on reinforcement learning according to claim 1, characterized in that: Step S3 includes: S301. Suppose that the sample set contains multiple samples, and each sample is a semantic data to be transmitted; S302. The transmitter selects a sample from the sample set and provides the source rate; it then obtains the operating point model pair at the source rate from the LUT. The semantic encoder in the working point model is used to semantically encode the selected samples to obtain the latent semantic representation. ,Will The data is divided into blocks to obtain K semantic feature blocks. The semantic feature block is a matrix with W rows and H columns, where W and H are the height and width of the semantic feature block, respectively. S303. At the start of each sample transmission, first calculate the importance weight of each extracted semantic feature block. This is to simultaneously reflect the differences in task contributions and the dependencies between features: Calculate task relevance factors The selected sample is forward-propagated through the semantic encoder and semantic decoder of the selected working point, that is, after being processed according to step S302, it is decoded by the decoder to obtain the decoding result. Let the selected sample be denoted as The decoding result is ; Calculate reconstruction loss for features The gradient magnitude is calculated, and a global average pooling is performed in the spatial dimension to obtain the factor: ,in For the first The value in the w-th row and h-th column; Calculate the correlation factor between features :calculate Compared with any other feature cosine similarity The average of their absolute values is obtained , used to characterize the feature-related redundancy strength; Calculating importance weights: Importance weights are defined as the product of the two factors. Then, normalize to obtain the weight vector. ; S304. Features to be sent The sending end first allocates the number of quantization bits. The quantization module converts it into a bitstream. Then select the channel coding rate in sequence. Modulation order and transmission power The signal is encoded in the channel coding module, modulated in the modulation module, and then transmitted according to the transmit power. For the features to be sent The minimum number of modulation symbols required is Transmission time is ,in, Symbol rate; Set a transmission time budget for the sample With total energy budget This ensures that the sample transfer process satisfies And energy consumption meets ; S305. Treat the transmission of each sample as a round, model the stepwise transmission process of the corresponding semantic features as a Markov decision process, and achieve online adaptive joint decision-making through reinforcement learning.
3. The adaptive semantic communication method based on reinforcement learning according to claim 2, characterized in that: Step S305 includes: 1) Status: During transmission, the sending end at each decision step... Observation status: in This represents the current block fading channel gain. This is the semantic importance vector of the current sample; For feature indicator vectors, , This indicates whether the feature has been sent; 0 indicates that it has been sent, and 1 indicates that it has not been sent. and These are the remaining symbol budget and the power budget, respectively. 2) Action: In the i-th decision step, the policy network is based on... Output action: That is: first, select the feature index to be sent in the i-th decision step. Then determine the physical layer parameters corresponding to this feature, including the quantization bits of the i-th decision step. coding rate Modulation order With transmission power ; 3) State transition: After the action is executed, the symbol and power budget are deterministically updated according to the selected action cost, and the transmitted features are removed from the available set; the channel gain is implemented using a block fading first-order Markov model. The evolution proceeds to the next decision-making step; in, This represents the channel gain transition probability at the (i+1)th decision step relative to the previous decision steps. This represents the channel gain transition probability from the i-th decision step to the (i+1)-th decision step, meaning that the channel gain transition probability of the (i+1)-th decision step relative to the previous decision steps is only related to the i-th decision step. 4) Reward Function: The reward function is based on the features actually recovered at the receiving end. The calculation, along with latency, forms the reward: in As a weighting factor, Transmission delay of the i-th decision step; By training a policy network to maximize the long-term average reward, the weighted sum of long-term average semantic distortion and latency is minimized. Maximizing the long-term average reward is achieved by summing all the reward functions of each feature block for each sample and then dividing by the average of the feature block data of all samples. 5) Domain randomization: At the beginning of each training round, the channel model and its statistical parameters are randomly sampled to generate the block fading process for that round, thereby obtaining the corresponding current block fading channel gain. ; The random sampling includes at least one or a combination of the following: (1) Types of fading distribution; (2) Multipath intensity or K-factor, Nakagami-m parameter; (3) Average SNR / Path Loss, Shadow Fading Variance; (4) Time correlation coefficient or Doppler parameters; (5) Channel state transition probability matrix; By covering multiple channel domains during training, the policy network learns a robust decision mapping for channel uncertainty.
4. The adaptive semantic communication method based on reinforcement learning according to claim 3, characterized in that: Step S4 includes: S401. The system receives input samples and selects the target operating point model pair from the LUT. ; S402. Semantic encoder extracts latent semantic representation. And obtain K semantic feature blocks ; S403. Calculate the semantic importance weight of each feature to obtain the importance weight vector. ; S404. Initialize the sample's delay, symbol budget, and energy budget, and observe the current block fading channel state; S405. At each decision step The policy network is based on the state First, select which feature to send, and then determine its physical layer parameters: And perform the corresponding quantization, encoding, modulation, and transmission processes; S406. The receiver performs demodulation, decoding, and inverse quantization to obtain... The semantic decoder outputs the task results, and the system updates the remaining budget and proceeds to the next decision step based on the corresponding reward. S407. The transmission process of this sample ends when all features have been sent or the budget has been exhausted.
Citation Information
Patent Citations
Semantic communication method for semantic source channel adaptive coding under parallel channels
CN119766400A
Task-oriented semantic communication method and related device
CN120656463A