Intelligent resource scheduling method for non-orthogonal multiple access backscattering system
By constructing a non-orthogonal multiple access backscattering system, employing a centralized training distributed execution framework and a federated averaging algorithm, and jointly optimizing channel correlation, SIC decoding order, and reader transmit power, the problems of limited spectrum resources and difficulty in obtaining CSI in multi-cell backscattering communication systems are solved, achieving maximum system performance and enhanced robustness.
Patent Information
- Application Number
- CN202511556062.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-29
- Publication Date
- 2026-01-09
AI Technical Summary
In multi-cell backscatter communication systems, existing technologies have failed to effectively solve the interference problem caused by limited spectrum resources, and the channel state information (CSI) is difficult to obtain accurately, resulting in resource scheduling strategies deviating from the optimal solution and insufficient system robustness and communication reliability.
A nonorthogonal multiple access backscattering system is constructed, employing a centralized training distributed execution framework and a federated averaging algorithm to jointly optimize channel correlation, SIC decoding order, tag transmission coefficients, and reader transmission power. Resource scheduling is achieved through a Markov decision process, and policy optimization is performed by combining a deep reinforcement learning model.
In imperfect CSI scenarios, the system and speed are improved, robustness and training efficiency are enhanced, dynamic network environments are adapted, and the system's spectrum utilization and communication reliability are increased.
Smart Images

Figure CN121310291A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of wireless communication technology, and in particular to an intelligent resource scheduling method for non-orthogonal multiple access backscattering systems. Background Technology
[0002] Backscatter communication technology, with its advantages of ultra-low power consumption and low cost, has become a key technology supporting the ultra-low power requirements of large-scale Internet of Things (IoT), showing broad application prospects in scenarios such as smart grids, smart agriculture, logistics tracking, and environmental monitoring. This technology achieves data transmission by reflecting existing radio frequency signals in the environment, without the need for actively transmitting radio frequency signals, significantly reducing terminal energy consumption and network deployment costs.
[0003] With the surge in the number of IoT devices, dense deployment of readers and tags in multi-cell backscatter communication systems has become a trend. However, it also faces two major challenges: First, severe inter-cell interference caused by limited spectrum resources restricts system and rate performance; second, the two-hop cascaded structure (carrier source-tag-reader) and noise effects of the backscatter link make it difficult to accurately obtain channel state information (CSI). Imperfect CSI can cause resource scheduling strategies to deviate from the optimal solution, significantly reducing system robustness and communication reliability.
[0004] To improve spectral efficiency, Non-Orthogonal Multiple Access (NOMA) technology has been introduced into backscatter communication systems, allowing multiple tags to share the same time-frequency resources, thus significantly improving spectral utilization. However, existing research on multi-cell backscatter resource scheduling based on NOMA has significant limitations: most studies are based on the perfect CSI assumption, failing to consider the impact of CSI estimation errors in real-world scenarios; the serial interference cancellation (SIC) decoding order is not sufficiently optimized, resulting in poor interference suppression; traditional optimization algorithms (such as convex optimization and alternating optimization) suffer from high computational complexity, slow convergence speed, and difficulty adapting to dynamic network environments when facing joint optimization problems involving channel correlation, decoding order, reflection coefficient, and transmit power; the "environment instability" problem easily occurs during multi-agent training, and experience cannot be effectively shared among agents, leading to low training efficiency and poor policy generalization ability. Therefore, there is an urgent need for a resource scheduling method that can jointly optimize multi-dimensional resource parameters in imperfect CSI scenarios, balancing system and rate improvement, robustness enhancement, and training efficiency, to meet the practical application needs of multi-cell NOMA backscatter communication systems. Summary of the Invention
[0005] The main objective of this invention is to propose an intelligent resource scheduling method for non-orthogonal multiple access backscattering systems, which can maximize system performance and rate under imperfect CSI scenarios, thereby enhancing robustness and training efficiency.
[0006] This invention is achieved through the following technical solution: A smart resource scheduling method for non-orthogonal multiple access backscattering systems includes the following steps: Step S1: Construct a non-orthogonal multiple access backscatter system. All cells in the system share the same spectrum resources. All tags in the cell communicate uplink via NOMA. The receiver of the system uses SIC for signal decoding. Each tag receives excitation signals from all readers in the system. Step S2: Construct an optimization problem that maximizes the sum rate of the system by jointly optimizing channel correlation, SIC decoding order, tag transmission coefficients, and reader transmission power. Step S3: Transform the optimization problem into a Markov decision process, adopt a centralized training distributed execution framework, and introduce a federated averaging algorithm based on this framework to obtain a resource scheduling strategy.
[0007] Furthermore, in step S1, in the non-orthogonal multiple access backscattering system, within each cell, the frequency band is planned to be divided into N / 2 sub-bands, and each sub-band allows two devices to perform NOMA pairing, using the formula... The rule stipulates that only one pair of tags can be accessed in each sub-band f. The rule stipulates that each pair of tags can only access one sub-frequency band, where, This indicates the association between the tag and the access sub-band. If the tag In sub-band If a reader is connected to the c-th cell, then... ,otherwise, C represents the number of cells. Let N be the set of available sub-bands, N be the number of tags in the system, and n be the nth tag.
[0008] Furthermore, in step S1, the receiving end of the system uses a decoding order set. To indicate the decoding order of SIC, if the tag If the reflected signal is decoded first, then during its decoding process, the tag paired with it using NOMA will be... The reflected signal is considered interference, and at this time, let's assume... Conversely, if the label If the reflected signal is not decoded first, then assume .
[0009] Furthermore, in step S1, the label The power of the received excitation signal is expressed as ,Label After receiving the excitation signal, it is modulated and reflected, and then the tag... The reflected signal power is expressed as ,in, For the first The transmission power of the reader in each community, For tags The reflection coefficient, This indicates that the reader c and the tag The imperfect channel gain between them.
[0010] Furthermore, in step S2, the optimization problem is expressed as: ,in, For the set of channel correlation indicator coefficients, For the decoding order set, For the set of reflectance coefficients of the label, This refers to the set of transmit powers of the reader. This represents the sum rate across the entire frequency band in the c-th cell. This represents the sum rate of the reader in the c-th cell on sub-frequency band f. This indicates that the reader in cell c received the tag. SINR of the reflected signal, where W is the available bandwidth. This represents the power of additive white Gaussian noise. This represents the transmit power of the reader in the c-th cell. This represents the maximum power limit of the reader in the c-th cell. This indicates the minimum SINR constraint for the label. .
[0011] Furthermore, in step S3, during the Markov process, the state space of the agent in the c-th cell is represented as follows: The action space of the intelligent agent is represented as The reward function of this intelligent agent Represented as ,in, , This is the penalty value.
[0012] Furthermore, in step S3, the centralized training distributed execution framework specifically includes: during the training phase, each agent shares global environmental information to achieve joint optimization based on a global perspective; during the execution phase, each agent makes independent decisions based only on its own observable local information.
[0013] Furthermore, in step S3, the federated averaging algorithm introduced based on the centralized training distributed execution framework specifically includes: modeling each agent as an independent deep reinforcement learning model; at the start of training, the central server of the non-orthogonal multiple access backscattering system uniformly initializes the global model parameters and distributes the global model parameters to the agents in each cell; each agent in each cell performs deep reinforcement training based on its own network environment and observation information to continuously optimize its local deep reinforcement learning model; during the training iteration process, each agent periodically uploads the updated local deep reinforcement learning model parameters to the central server; the central server uses the federated averaging algorithm to weight and aggregate the model parameters uploaded by each agent to obtain a new global model, and then distributes the new global model to all agents.
[0014] Furthermore, in step S3, during training, a combination of simple reward centralization and federated averaging algorithm is introduced, and an actor-critic architecture is adopted. Specifically, this includes setting up two sub-actor networks to handle discrete actions and continuous actions respectively, and a global critic network to guide the updating of network parameters.
[0015] Furthermore, artificially scrambled CSI is introduced during the training phase to assist learning.
[0016] As can be seen from the above description of the present invention, compared with the prior art, the present invention has the following beneficial effects: This invention first constructs a non-orthogonal multiple access backscattering system. All cells in this system share the same spectrum resources, and all tags within a cell communicate uplink via NOMA. The receiver uses SiC for signal decoding, and each tag receives excitation signals from all readers within the system. Then, an optimization problem is constructed based on this system. This optimization problem maximizes the system's sum rate by jointly optimizing channel correlation, SiC decoding order, tag reflection coefficients, and reader transmission power. Finally, the optimization problem is transformed into a Markov decision process, employing a centralized training distributed execution framework. Based on this framework, a federated averaging algorithm is introduced to obtain a resource scheduling strategy. During the process, by jointly optimizing channel correlation, SiC decoding order, tag reflection coefficients, and reader transmission power, the system's sum rate is improved under different reader coverage radii. Combining the experience-sharing mechanism of federated learning with a stability enhancement strategy based on reward centralization, a resource scheduling strategy that maximizes the system's sum rate in imperfect CSI scenarios is obtained. During training, the federated averaging algorithm, based on centralized training distributed execution, effectively improves the agent's policy coordination and training stability. Attached Figure Description
[0017] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] Figure 1 This is a flowchart of the present invention.
[0019] Figure 2 This is a schematic diagram of the non-orthogonal multiple access backscattering system of the present invention.
[0020] Figure 3 This is a performance comparison chart of the present invention and the comparative scheme under different reader coverage radii. Detailed Implementation
[0021] The present invention will be further described below through specific embodiments.
[0022] like Figure 1 As shown, the intelligent resource scheduling method for non-orthogonal multiple access backscattering systems includes the following steps: Step S1: Construct a non-orthogonal multiple access backscatter system. All cells in the system share the same spectrum resources. All tags in the cell communicate uplink via NOMA. The receiver of the system uses SIC for signal decoding. Each tag receives excitation signals from all readers in the system. More specifically, nonorthogonal multiple access backscattering systems such as Figure 2 As shown, each cell shares the same spectrum resources, with an available frequency band bandwidth of W. Within each cell, the frequency band is divided into N / 2 sub-bands, and the set of available sub-bands is denoted as . Each sub-band allows two devices to perform NOMA pairing, through the formula The rule stipulates that only one pair of tags can be accessed in each sub-band f. The rule stipulates that each pair of tags can only access one sub-frequency band, where, This indicates the association between the tag and the access sub-band. If the tag In sub-band If a reader is connected to the c-th cell, then... ,otherwise, C represents the number of cells, N represents the number of tags in the system, and n represents the nth tag.
[0023] Since the uplink communication of tags within the cell uses NOMA, the receiver needs to use SIC for signal decoding accordingly. The received signal power is affected by several optimizable variables, such as the backscatter coefficient, channel gain, and reader transmit power. Therefore, the decoding order must also be considered during the optimization phase to further improve system performance. A decoding order set is adopted. To indicate the decoding order of SIC, if the tag If the reflected signal is decoded first, then during its decoding process, the tag paired with it using NOMA will be... The reflected signal is considered interference, and at this time, let's assume... Conversely, if the label If the reflected signal is not decoded first, then assume .
[0024] In multi-community scenarios, tags The power of the received excitation signal is expressed as ,Label After receiving the excitation signal, it is modulated and reflected, and then the tag... The reflected signal power is expressed as ,in, For the first The transmission power of the reader in each community, For tags The reflection coefficient.
[0025] This system considers a standard block fading channel model, with reader c and tag c. The channel between them can be represented as ,in This indicates the path loss at a reference distance of one meter. For reader c and tags The distance between them This is the path loss index. This refers to a fading channel. However, in practical wireless communication systems, due to factors such as channel estimation errors and environmental interference, it is difficult to obtain a perfect CSI. Therefore, this invention adopts the following imperfect CSI model. ,in For channel estimation error, Let V be the channel noise variance, and let the imperfect channel gain be expressed as: .
[0026] Step S2: Construct an optimization problem that maximizes the sum rate of the system by jointly optimizing channel correlation, SIC decoding order, tag transmission coefficients, and reader transmission power. Specifically, the optimization problem is represented as ,in, For the set of channel correlation indicator coefficients, For the decoding order set, For the set of reflectance coefficients of the label, This refers to the set of transmit powers of the reader. This represents the sum rate across the entire frequency band in the c-th cell. This represents the sum rate of the reader in the c-th cell on sub-frequency band f. This indicates that the reader in cell c received the tag. The SINR of the reflected signal, due to noise and inter-tag interference in the sub-band, can be expressed as follows: The received signal of the reader in the c-th cell at sub-band f can be expressed as... W represents the available bandwidth. This represents the power of additive white Gaussian noise. This represents the transmit power of the reader in the c-th cell. This represents the maximum power limit of the reader in the c-th cell. This indicates the minimum SINR constraint for the label. .
[0027] Step S3: The optimization problem is transformed into a Markov decision process. A centralized training distributed execution framework is adopted, and a federated averaging algorithm is introduced based on this framework to obtain a resource scheduling strategy. Because the optimization variables are mutually coupled and exhibit non-convex characteristics during the joint optimization of channel correlation, SIC decoding order, tag reflection coefficient, and reader transmit power, traditional mathematical optimization methods are prone to getting trapped in local optima. To address this issue, the joint optimization problem can be transformed into a Markov decision process. By defining a structured space to address the coupling relationships between channel correlation, decoding order, reflection coefficient, and transmit power, the dispersed optimization variables can be integrated into a co-determined whole. State space: Sets the state of the agent in the c-th cell. The sum rate of its cell Then the state space of the agent can be represented as In this context, the intelligent agent is the reader of the base station; Action Space: Given that the actions of an agent in cell c are composed of four parts: channel correlation indicator coefficient, SIC decoding order coefficient, backscattering coefficient of tags within cell c, and transmit power of reader c, the action space of this agent can be expressed as follows: Each agent needs to allocate access channel sub-bands for tags within the cell for NOMA pairing, set the SIC decoding order, set the backscatter coefficient for the tags, and set the transmit power for the reader; Reward Function: The objective of the joint optimization is to maximize the sum rate of the entire backscatter communication system. Simultaneously, to ensure that various constraints are satisfied during the optimization process, corresponding penalty terms need to be designed for different constraints. Therefore, the reward function consists of two parts: the main reward and the penalty terms. The main reward is set as the sum rate of the reader. This is done to maximize the overall system speed; while the penalty term is used to ensure that the signal-to-noise ratio (SINR) requirement of each tag is met, specifically defined as: , The penalty value is a positive number that adapts to the objective optimization; therefore, the reward function of the agent in the c-th cell is... Represented as .
[0028] In the following description, Let represent the time step of the agent in the c-th cell. t The state, actions, and rewards under these conditions.
[0029] Because the system of this invention is a multi-agent model with a discrete-continuous hybrid action space, centralized or distributed training methods have the following limitations: while centralized training can fully utilize global information, it is difficult to scale to large-scale systems during the execution phase; and while fully distributed training can reduce execution overhead, it often lacks global coordination, leading to a decrease in learning efficiency and system performance. Therefore, this invention adopts a Centralized Training Distributed Execution (CTDE) framework to balance the cooperation between agents and training complexity. Specifically, this framework includes: during the training phase, each agent shares global environmental information to achieve joint optimization based on a global perspective, which not only improves the stability and convergence speed of learning but also helps alleviate the non-stationarity problem common in multi-agent environments; during the execution phase, each agent makes independent decisions based only on its own observable local information, without needing to acquire the global state. This effectively reduces the communication and computational overhead of the system in actual deployment and enhances the scalability and practicality of the method in large-scale dynamic environments.
[0030] The federated averaging algorithm introduced based on the centralized training distributed execution framework specifically includes: modeling each agent as an independent deep reinforcement learning model; at the start of training, the central server of the non-orthogonal multiple access backscattering system uniformly initializes the global model parameters and distributes them to the agents in each cell; each agent in each cell performs deep reinforcement learning training based on its own network environment and observation information to continuously optimize its local deep reinforcement learning model; during training iterations, each agent periodically uploads the updated local deep reinforcement learning model parameters to the central server; the central server uses the federated averaging algorithm to weight and aggregate the model parameters uploaded by each agent to obtain a new global model, which is then distributed to all agents. Artificially scrambled CSI is introduced during the training phase to assist model learning, allowing the model to maintain stable decision-making ability and communication optimization performance even when there is interference or information bias in the actual channel environment, thereby improving the model's robustness.
[0031] More specifically, the federated averaging algorithm proceeds as follows: The local model parameters of the agents in the c-th cell are set to... The aggregated global model parameters are , This represents the number of tags in cell c. This represents the total number of labels in the system. During model training, the agent periodically uploads the updated local model to the central server. The central server then uses a federated averaging algorithm to weight and aggregate the models from each cell, forming a new global model. Specifically, the central server uses the following federated averaging formula to globally aggregate the network parameters: , , , Subsequently, the central server broadcasts the aggregated model to each cell, and the agent loads the model to continue the next round of local training, iterating until convergence.
[0032] In another embodiment, during training, a combination of simple reward centralization and federated averaging is introduced to further ensure the stability of the training process. More specifically, the algorithm based on federated averaging and simple reward centralization employs an actor-critic architecture, comprising two sub-actor networks to handle discrete and continuous actions respectively, and a global critic network to guide the updating of network parameters. The stochastic strategy for generating discrete actions is described below. A discrete actor network will output k values for k discrete actions. These values are then fed into a softmax function to obtain a probability distribution, from which discrete actions are randomly sampled to generate discrete actions. Continuous actor networks generate stochastic policies with continuous parameters by outputting the mean and variance of a Gaussian distribution for each parameter. .
[0033] During the data acquisition phase, to efficiently utilize data generated by old strategies, the algorithm leverages the importance sampling principle, reusing historical experience by calculating the ratio of new to old strategies. Specifically, the agent interacts with the environment for T steps, utilizing the old strategies... Sampling collects a set of empirical samples And store it in the experience pool To enhance the learning efficiency and stability of the algorithm, Simple Reward Centralization (SRC) is employed for the reward sequence. Perform the following processing ,in .
[0034] During the training phase, from A small batch of samples of size B is randomly selected from the sample, and artificially scrambled CSI is used. Rewards after SRC processing To calculate the loss function of each network. The loss function for the discrete actor network is: ,in, This represents the empirical mean of a finite batch of samples. It is the generalized advantage estimation (GAE). Importance sampling techniques were considered to measure generalized advantage. By limiting The ratio of the hyperparameters to the gradient limits prevents the policy from exceeding a reasonable range, resulting in smoother policy updates and ensuring training stability. The hyperparameters are used to constrain gradient updates. Generalized advantage estimation... The calculation is based on an estimate of future returns, and the specific formula is as follows: Where T represents the total number of time steps in which the agent interacts with the environment under a given policy. It is a discount factor. It's a hyperparameter. It is a state-value function. It is in state The actual cumulative reward. The loss functions for the continuous actor network and the critic network are respectively... and Finally, the Adam optimizer is used to update the network parameters, and the old policy parameters of the actor network are periodically updated. .
[0035] like Figure 3 The figure shows a performance comparison between the present invention and the comparative scheme under different reader coverage radii. The present invention is FedHPPO-SRC, while the comparative schemes include Multi-Agent Deep Q-Network (MADQN), Multi-Agent Value Decomposition Algorithm (MAPDQN), Multi-Agent Hierarchical Proximal Policy Optimization (MAHPPO), and FedHPPO without the SRC strategy. Figure 3 As can be seen, the system and rate decrease with the increase of the coverage radius, and the present invention achieves the best performance in terms of both the sum and rate.
[0036] In this invention, the terms "first," "second," and "third," etc., are used only to distinguish similar objects and are not necessarily used to describe a specific order or sequence, nor should they be construed as indicating or implying relative importance. The use of terms such as "upper," "lower," "left," "right," "front," and "rear" to indicate orientation or positional relationships is based on the orientation or positional relationships shown in the accompanying drawings and is only for the convenience of describing the invention, not to indicate or imply that the device referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation on the scope of protection of this invention. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0037] Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0038] The above are merely specific embodiments of the present invention, but the design concept of the present invention is not limited thereto. Any non-substantial modifications made to the present invention using this concept shall be considered as infringing upon the protection scope of the present invention.
Claims
1. A smart resource scheduling method for non-orthogonal multiple access backscattering systems, characterized in that: Includes the following steps: Step S1: Construct a non-orthogonal multiple access backscatter system. All cells in the system share the same spectrum resources. All tags in the cell communicate uplink via NOMA. The receiver of the system uses SIC for signal decoding. Each tag receives excitation signals from all readers in the system. Step S2: Construct an optimization problem that maximizes the sum rate of the system by jointly optimizing channel correlation, SIC decoding order, tag transmission coefficients, and reader transmission power. Step S3: Transform the optimization problem into a Markov decision process, adopt a centralized training distributed execution framework, and introduce a federated averaging algorithm based on this framework to obtain a resource scheduling strategy.
2. The intelligent resource scheduling method for non-orthogonal multiple access backscattering systems according to claim 1, characterized in that: In step S1, in the non-orthogonal multiple access backscattering system, within each cell, the frequency band is planned into N / 2 sub-bands, and each sub-band allows two devices to perform NOMA pairing, as specified by the formula. The rule stipulates that each sub-band f can only access one pair of tags. The rule stipulates that each pair of tags can only access one sub-frequency band, where, This indicates the association between the tag and the access sub-band. If the tag In sub-band If a reader is connected to the c-th cell, then... ,otherwise, C represents the number of cells. Let N be the set of available sub-bands, N be the number of tags in the system, and n be the nth tag.
3. The intelligent resource scheduling method for non-orthogonal multiple access backscattering systems according to claim 2, characterized in that: In step S1, the receiving end of the system uses a decoding order set. To indicate the decoding order of SIC, if the tag The reflected signal is decoded first, and during its decoding process, the tag paired with it using NOMA is... The reflected signal is considered interference, and at this time, let's assume... Conversely, if the label If the reflected signal is not decoded first, then assume .
4. The intelligent resource scheduling method for non-orthogonal multiple access backscattering systems according to claim 3, characterized in that: In step S1, the label The power of the received excitation signal is expressed as ,Label After receiving the excitation signal, it is modulated and reflected, and the tag... The reflected signal power is expressed as ,in, For the first The transmission power of the reader in each community, For tags The reflection coefficient, For reader c and tags The imperfect channel gain between them.
5. The intelligent resource scheduling method for non-orthogonal multiple access backscattering systems according to claim 4, characterized in that: In step S2, the optimization problem is expressed as: ,in, For the set of channel correlation indicator coefficients, For the decoding order set, For the set of reflectance coefficients of the label, This refers to the set of transmit powers of the reader. This represents the sum rate across the entire frequency band in the c-th cell. This represents the sum rate of the reader in the c-th cell on sub-frequency band f. This indicates that the reader in cell c received the tag. SINR of the reflected signal, where W is the available bandwidth. This represents the power of additive white Gaussian noise. This represents the transmit power of the reader in the c-th cell. This represents the maximum power limit of the reader in the c-th cell. This indicates the minimum SINR constraint for the label. .
6. The intelligent resource scheduling method for non-orthogonal multiple access backscattering systems according to claim 5, characterized in that: In step S3, during the Markov process, the state space of the agent in the c-th cell is represented as follows: The action space of the intelligent agent is represented as The reward function of the intelligent agent Represented as ,in, , This is the penalty value.
7. The intelligent resource scheduling method for non-orthogonal multiple access backscattering systems according to claim 6, characterized in that: In step S3, the centralized training distributed execution framework specifically includes: during the training phase, each agent shares global environmental information to achieve joint optimization based on a global perspective; during the execution phase, each agent makes independent decisions based only on its own observable local information.
8. The intelligent resource scheduling method for non-orthogonal multiple access backscattering systems according to claim 7, characterized in that: In step S3, a federated averaging algorithm is introduced based on the centralized training distributed execution framework. Specifically, this includes: modeling each agent as an independent deep reinforcement learning model; at the start of training, the central server of the non-orthogonal multiple access backscattering system initializes the global model parameters uniformly and distributes the global model parameters to the agents in each cell; each agent in each cell performs deep reinforcement training based on its own network environment and observation information to continuously optimize its local deep reinforcement learning model; during the training iteration process, each agent periodically uploads the updated local deep reinforcement learning model parameters to the central server; the central server uses the federated averaging algorithm to weight and aggregate the model parameters uploaded by each agent to obtain a new global model, and then distributes the new global model to all agents.
9. A smart resource scheduling method for a non-orthogonal multiple access backscattering system according to claim 8, characterized in that: In step S3, during training, a combination of simple reward centralization and federated averaging algorithm is introduced, and an actor-critic architecture is adopted. Specifically, this includes setting up two sub-actor networks to handle discrete actions and continuous actions respectively, and a global critic network to guide the update of network parameters.
10. A smart resource scheduling method for a non-orthogonal multiple access backscattering system according to claim 9, characterized in that: Artificially scrambled CSI is introduced during the training phase to aid learning.
Citation Information
Patent Citations
Uplink NOMA power distribution method based on dynamic decoding SIC receiver
CN106385300A
Backscatter communication system and energy beam forming optimization method thereof
CN110430148A
A non-orthogonal multiple access backscatter communication method
CN119766319A
Radio base station, user terminal, radio communication method and radio communication system
US20160142193A1