RIS-Enhanced Cooperative Transmission Method for Uplink OFDMA System

Through the MADDQN and MADDPG algorithm combined with the BDTL framework, the resource allocation of RIS enhanced OFDMA system is optimized, which solves the problem of inefficient training when network configuration changes, and realizes a fast adaptability and efficient resource allocation strategy, which improves system performance and user speed.

CN115866611BActive Publication Date: 2025-07-08NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211481575.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-24
Publication Date
2025-07-08
Estimated Expiration
2042-11-24

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and efficiently train new networks when network configuration changes in RIS enhanced OFDMA systems, resulting in low resource allocation efficiency and deep Q networks have quantization errors and overly optimistic action value estimation problems in continuous action tasks.

Method used

Multi-agent dual-deep Q network (MADDQN) algorithm and multi-agent deep deterministic strategy gradient (MADDPG) algorithm are used, combined with the bidirectional transfer learning (BDTL) framework, subcarrier allocation, power allocation and RIS control strategies are optimized, and the adaptability of neural networks in different environments is improved through transfer learning.

Benefits of technology

It realizes the rapid and effective adjustment of resource allocation strategies when network configuration changes, maximizes users and speed, improves system performance and accelerates the convergence speed of neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FDA0005432148930000013
    Figure FDA0005432148930000013
  • Figure FDA0005432148930000021
    Figure FDA0005432148930000021
  • Figure FDA0005432148930000022
    Figure FDA0005432148930000022
Patent Text Reader

Abstract

The present invention belongs to the field of wireless communication. Specifically, it is a cooperative transmission method for RIS-enhanced uplink OFDMA systems, which specifically includes: defining the key parameters of the uplink OFDMA system; designing a multi-agent deep reinforcement learning algorithm to solve this optimization problem; changing the current environment with different schemes, and proposing a bidirectional transfer learning framework to enhance the adaptability of the neural network in different environments by utilizing the common configurations in different environments, so that the agent can more efficiently dynamically adjust the resource allocation and RIS control strategy to maximize the sum rate of all users. The present invention can effectively improve the system performance, accelerate the convergence speed of the neural network in different environments, and enable the agent to adapt to different environments faster and more effectively.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of wireless communication. Specifically, it is a cooperative transmission method for RIS-enhanced uplink OFDMA system, which is based on two-way migration reinforcement learning. Background Art

[0002] In recent years, due to the continuous growth of the requirements for massive access and low-latency communication, the resource allocation problem in the fifth-generation (5G) technology has attracted wide attention. As the main access method of wireless communication systems, orthogonal frequency division multiple access (OFDMA) divides the transmission bandwidth into a series of orthogonal and non-overlapping subcarrier sets, and allocates different subcarrier sets to different users within the same time slot to achieve multiple access. OFDMA technology adaptively allocates resources according to the channel gain, greatly improving the system performance. In addition, as a revolutionary technology, the reconfigurable intelligent surface (RIS) can significantly improve the performance of wireless communication systems. It consists of multiple reflecting elements and can be used to reconfigure the received signal by simultaneously controlling its amplitude and phase shift.

[0003] Currently, a large amount of research work has focused on resource allocation and reflecting element control in RIS-enhanced OFDMA systems. By optimizing the beamformers of all users and the phase shifts of RIS, the throughput of multi-cell RIS-enhanced OFDMA systems is maximized, and the sum rate of OFDMA systems is improved by utilizing the significant beamforming gain brought by RIS. However, due to the dynamics and uncertainties of RIS-aided wireless systems, it is usually difficult to establish the accurate mathematical models required by the above model-driven methods. As a model-free method, the deep Q-network (DQN) greatly reduces the difficulty of mathematical modeling by introducing a trial-and-error mechanism and optimizing the output strategy through interaction with the environment. Some work has used the DQN method to solve the resource allocation problem in OFDMA systems. However, the main drawback of DQN is that the output actions can only be discrete, causing quantization errors for continuous action tasks (such as power allocation, amplitude, and phase shift control). On the other hand, DQN uses the same value to select and evaluate actions, resulting in overly optimistic value estimates for certain actions. Fortunately, based on the actor-critic architecture, the deep deterministic policy gradient (DDPG) has been applied to tasks with continuous action spaces. In addition, the double deep Q-network (DDQN) uses different networks to select and evaluate actions, addressing the problem of overly optimistic value estimates. Some studies have used DDQN or DDPG in wireless systems to maximize system capacity or other objectives. However, these studies have not yet solved the following important problems commonly existing in wireless communication systems: when the network configuration changes, how to quickly and effectively train a new network in the new network configuration. Summary of the Invention

[0004] The object of the present invention is to overcome the deficiencies of the prior art and provide a cooperative transmission method for RIS-enhanced uplink OFDMA systems based on bidirectional transfer reinforcement learning. The present invention designs a multi-agent double deep Q-network (MADDQN) algorithm to optimize the subcarrier allocation strategy, proposes a multi-agent deep deterministic policy gradient (MADDPG) algorithm to jointly optimize the transmit power of users and the amplitude and phase offsets of RIS reflection elements, and proposes a bidirectional transfer learning (BDTL) framework to improve the adaptability of neural networks in different environments by utilizing the same configurations of different environments, thereby dynamically adjusting the resource allocation strategy to maximize the sum rate of all users.

[0005] The specific technical solution adopted by the present invention is as follows:

[0006] A cooperative transmission method for RIS-enhanced uplink OFDMA systems, which is based on bidirectional transfer reinforcement learning, includes the following steps:

[0007] Step 1: Define the key parameters of the uplink OFDMA system;

[0008] Step 2: Design a multi-agent double deep Q-network (MADDQN) algorithm to solve the subcarrier allocation problem, design a multi-agent deep deterministic policy gradient (MADDPG) algorithm to solve the power allocation and RIS control problems, and use the corresponding algorithms to train the corresponding agents in the current environment;

[0009] Step 3: Change the current environment with different schemes, and propose a bidirectional transfer learning (BDTL) framework to enhance the adaptability of neural networks in different environments by utilizing the common configurations of different environments, so that the agents can more efficiently dynamically adjust the resource allocation (including subcarrier allocation and power allocation) and RIS control strategies to maximize the sum rate of all users.

[0010] For a further improvement of the present invention, the specific process of Step 1 is as follows: Set the RIS-enhanced uplink multi-cell OFDMA system to be composed of N cells, each cell is equipped with a base station (BS) with L antennas, and M single-antenna users are randomly distributed in each cell; Represent the base station set and the user set related to base station n as α = {1, 2,..., N} and β = {1, 2,..., M}; Set that there are K subcarriers in each cell, and the subcarrier set is defined as γ = {1, 2,..., K}; Set a RIS to be composed of E reflection elements, and the set of reflection elements is represented as ε = {1, 2,..., E}; Define the subcarrier allocation parameter as where indicates that user m in cell n is allocated to subcarrier k at time slot t; The data rate of user m in base station n at time slot t on subcarrier k is defined as Finally, the sum of the transmission rates of all users is maximized.

[0011] For a further improvement of the present invention, in step two, the specific method of training the current agent in the current environment using the MADDQN algorithm and the MADDPG algorithm is as follows: The resource allocation framework and RIS control are divided into a subcarrier allocation module, a power allocation module, and an RIS control module; for the subcarrier allocation module and the power allocation module, each user is regarded as an agent; for the RIS control module, the RIS is regarded as an agent; in addition, the OFDMA system is regarded as the environment; for the subcarrier allocation module, each agent is equipped with a DDQN unit composed of a training Q-network and a target Q-network; for the power allocation module and the RIS control module, each agent is equipped with a DDPG unit composed of a training actor network, a target actor network, a training critic network, and a target critic network. Since there are multiple agents, a multi-agent network is formed, namely MADDQN and MADDPG. In the three modules, the training process of each agent is as follows: In time slot t, the OFDMA system feeds back its state to each agent. Then, each Q-network in the subcarrier allocation module randomly selects an action from the action space of this module with probability ε, or selects the action that maximizes the Q function value according to Equation (1) with probability 1 - ε

[0012]

[0013] Where, is the action generated by each agent in the subcarrier allocation module, is the state fed back by the environment to each agent in the subcarrier allocation module, is the Q-network parameter of each agent in the subcarrier allocation module, and Α is the action space of the subcarrier allocation module. Each training actor network in the power allocation module and the RIS control module respectively generates a deterministic action according to the state fed back by the OFDMA system to the agent and the amplitude and phase of the RIS (as the state of the agent RIS). At the same time, a random noise is used together with the output action for network exploration. According to the deterministic policy gradient theory, each training actor network updates its parameter μ, and each target critic network evaluates this action according to this deterministic action.

[0014] After each agent in the subcarrier allocation module and the power allocation module executes the selected action, it obtains the returned real-time reward from the OFDMA system. Since the objective of the present invention is to maximize the sum rate of all users, therefore, in the present invention, the real-time rewards of the subcarrier allocation module and the power allocation module are uniformly defined as Equations (2) and (3):

[0015]

[0016] Among them

[0017]

[0018] Among them, represents the data rate of user m assigned to subcarrier k in cell n, represents the penalty term assigned to user m in cell n, ρ is the penalty coefficient, and the reward function of the agent RIS in the RIS control module is Finally, the OFDMA system correspondingly switches to a new state in the next time slot t + 1. The agents in the subcarrier allocation module, power allocation module, and RIS control module continuously interact with the OFDMA system to continuously obtain real-time samples and and store the real-time samples correspondingly in the experience pool of each module. In addition, an experience replay method is introduced to eliminate data correlation. Part of the samples are randomly drawn from the experience pools of the subcarrier allocation module, power allocation module, and RIS control module and Therefore, the loss functions of each DDQN and DDPG unit in the subcarrier allocation module, power allocation module, and RIS control module are respectively defined as

[0019]

[0020] Among them

[0021]

[0022] Among them, τ and τ - are the parameters of the training Q network and the target Q network respectively, μ and μ - are the parameters of the training actor network and the target actor network respectively, θ and θ - are the parameters of the training critic network and the target critic network respectively, γ is the discount rate of DDQN, and η is the discount rate of DDPG.

[0023] During the training process, for each DDQN unit and DDPG unit of each agent, the RMSProp optimizer is used to update the parameters of the training Q network and the training critic network by minimizing the loss functions (4) and (5). In addition, every T s time slots, the parameters τ of the training Q network and the parameters θ of the training critic network in the subcarrier allocation module, power allocation module, and RIS control module are respectively copied to update the parameters τ - of the target Q network and the parameters θ - of the target critic network in the subcarrier allocation module, power allocation module, and RIS control module.

[0024] For a further improvement of the present invention, in step three: the specific method for training a new agent in a new environment using a bidirectional transfer learning framework based on MADDQN and MADDPG is as follows: First, three different environments are set: Environment 1, keeping the default environment parameters unchanged; Environment 2, since there may be a line-of-sight (LoS) path in the RIS channel, changing the small-scale fading between the RIS and the BS (i.e., Q RB (t)) to a Rice distribution while keeping other simulation parameters unchanged; (3) reducing the correlation coefficient φ to 0.64 to simulate an environment where channel parameters change rapidly over time, while keeping other simulation parameters unchanged; Then, in the new environment, the proposed bidirectional transfer learning framework is used to train the new agent; during the training process, when calculating the target Q values of each new DDQN unit in the subcarrier allocation module and the target Q values of each new DDPG unit in the power allocation module and the RIS control module, the knowledge extracted from the old agent and the experience collected from the new agent are considered simultaneously. Therefore, the loss functions of each new DDQN unit in the subcarrier allocation module and each new DDPG unit in the power allocation module and the RIS control module are respectively expressed as:

[0025]

[0026] where

[0027]

[0028] where, and Q f respectively represent the network trained in environment g and the new network in environment f, and π(·) respectively represent the trained actor network and the new actor network. ψ represents a scaling factor that takes values in the range (0, 1] and gradually decreases according to the rule ψ←ψ / (1 + Θ) at each time slot t, where Θ is the decay factor. This indicates that as time goes by, each new agent in the subcarrier allocation module, the power allocation module, and the RIS control module will increasingly use its own experience for training.

[0029] The beneficial effects of the present invention are as follows: The present invention is applicable to the uplink OFDMA system. By using a bidirectional transfer learning framework based on MADDQN and MADDPG to complete resource allocation with the goal of maximizing the sum rate of all users, the system performance can be effectively improved, the convergence speed of the neural network can be accelerated, and the new agent can adapt to the new network environment faster and more effectively. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a graph showing the sum rate performance of MADDQN&MADDPG using default simulation parameters and the baseline algorithm in an embodiment of the present invention.

[0031] Figure 2 Sum rate performance graphs of MADDQN&MADDPG and the baseline algorithm for changing the distribution of QRB(t) in the embodiments of the present invention.

[0032] Figure 3 Sum rate performance graphs of MADDQN&MADDPG and the baseline algorithm for reducing the channel correlation coefficient in the embodiments of the present invention.

[0033] Figure 4 Sum rate performance graphs of BDTL and the baseline algorithm using default simulation parameters in the embodiments of the present invention.

[0034] Figure 5 Sum rate performance graphs of BDTL and the baseline algorithm for changing the distribution of QRB(t) in the embodiments of the present invention.

[0035] Figure 6 Sum rate performance graphs of BDTL and the baseline algorithm for reducing the channel correlation coefficient in the embodiments of the present invention.

[0036] Specific implementation

[0037] To deepen the understanding of the present invention, the following will further describe the present invention in detail with reference to the drawings and embodiments. These embodiments are only used to explain the present invention and do not limit the protection scope of the present invention. Embodiment

[0038] A cooperative transmission method for a RIS-enhanced uplink OFDMA system, which is based on two-way transfer reinforcement learning and includes the following steps:

[0039] Step 1: The uplink multi-cell OFDMA system consists of N cells, each cell is equipped with a base station (BS), the base station has L (L>1) antennas, and M single-antenna users are randomly distributed in each cell. We respectively represent the base station set and the user set related to base station n as α = {1, 2,..., N} and β = {1, 2,..., M}. The cells reuse the same frequency resources, and the total bandwidth is equally divided into K orthogonal subcarriers, which we represent as γ = {1, 2,..., K}.

[0040] In addition, we assume that a RIS composed of E low-cost reflection elements is deployed inside the system. The set of reflection elements is represented as ε = {1, 2,..., E}. It is assumed that each BS only decodes the messages from its related users, and the signals from users in other cells are regarded as interference. In addition, the subcarriers in the same cell are orthogonal to each other, thus avoiding intra-cell interference.

[0041] Step 2: An algorithm based on MADDQN and MADDPG is designed to solve the resource allocation (including power allocation and subcarrier allocation) and RIS control problems, and this algorithm is used to train the current agent in the current environment. In the present invention, DRL (including DDQN and DDPG) can be described as a stochastic game problem, which consists of two entities (i.e., the agent and the environment) and three elements (i.e., state, action, and reward). We regard the RIS-enhanced multi-cell OFDMA system as the environment and each user as an agent. In addition, the entire RIS is regarded as an agent. For subcarrier allocation, we give the state, action, and reward of each user as follows.

[0042] The subcarrier allocation process based on MADDQN is as follows. First, we equip each agent with a DDQN unit, which consists of a training Q-network and a target Q-network. Then, the training process of each agent starts. In time slot t, the environment feeds back its state to each agent. Then, the training Q-network of each agent selects an ε-greedy strategy. Specifically, each training Q-network randomly selects an action in the action space with probability ε, or selects the action that maximizes the following Q-estimate value with probability 1 - ε:

[0043]

[0044] where A represents the action space of subcarrier allocation, represents the parameters of the training Q-network. The value of ε gradually decreases over time, which means that the randomness gradually decreases and the network tends to select the best action. Each agent executes the selected action and obtains the returned reward from the environment. Finally, the environment switches to a new state in the next time slot t + 1. During consecutive interactions, we continuously obtain samples

[0045] Then the samples are stored in the experience replay buffer. When the experience replay buffer is full of samples, the latest samples will overwrite the earliest samples in time. In addition, we also introduce the experience replay method to eliminate data correlation. Each time we train, we randomly draw samples with a capacity of D1 from the experience pool Therefore, the loss function of each DDQN unit is defined as

[0046]

[0047] where

[0048]

[0049] where τ and τ - are the parameters of the training Q-network and the target Q-network respectively, and γ is the discount rate of DDQN.

[0050] MADDPG is used to find the optimal power allocation strategy for all users and design the RIS reflection coefficient composed of amplitude and phase offsets. Each agent has three DDPG units for power allocation, amplitude control, and phase shift design respectively. Each DDPG unit consists of four networks: a training actor network, a target actor network, a training critic network, and a target critic network. Different from the ε-greedy algorithm used in DDQN, the training actor network generates deterministic actions based on the current state. In addition, random noise is combined with the output actions to balance the exploration and exploitation of actions. The action of agent j (j = 1, …, M×N+1) is defined as:

[0051]

[0052] where μ is the parameter of the training actor network.

[0053] Similar to DDQN, in DDPG, we use experience replay to train the network. After determining the actions, each target critic network generates the target Q value for training to evaluate the selected actions by using the sampled samples with a capacity of D2 The target Q value is defined as follows:

[0054]

[0055] In addition, the loss function of each DDPG unit is defined as

[0056]

[0057] where μ - is the parameter of the target actor network, θ and θ - are the parameters of the training critic network and the target critic network respectively, and η is the discount rate of DDPG.

[0058] Finally, we introduce the BDTL framework to improve the adaptability of new agents in different environments by leveraging the common configurations of these environments. The process of BDTL can be divided into two steps. Step 1: Set different environments, and then use the MADDQN and MADDPG algorithms to train agents in these environments. Finally, save the trained agents after the training process ends. Step 2: Train new agents in the environments set in Step 1. Suppose we have set a total of F environments. In the f-th environment (f = 1, …, F), when calculating the target Q value of each new agent, we simultaneously consider the knowledge extracted from the agents trained in the environment g (g ≠ f) and the experience collected from the new agents. Then, we define the training process of each agent under bidirectional transfer learning as follows. For each DDQN unit and DDPG unit, we rewrite the loss function and the target Q value as:

[0059]

[0060] wherein

[0061]

[0062] wherein and Q f respectively represent the network trained in environment g and the new network in environment f, and π(·) respectively represent the trained actor network and the new actor network. ψ represents a scaling factor that takes values in the range (0, 1] and gradually decreases according to the rule of ψ←ψ / (1 + Θ) at each time slot t, where Θ is the decay factor. This indicates that as time goes by, the DDQN unit and DDPG unit of each agent will increasingly use the experience of the new network for training.

[0063] Under the steps of the above embodiments, simulations are carried out in different scenarios to illustrate the beneficial effects of the present invention. In our simulation, we set F = 3 environments as follows: Environment 1: Keep the default environment parameters unchanged. Environment 2: Since there may be a line-of-sight (LoS) path in the RIS channel, we change the small-scale fading between the RIS and the BS (i.e., QRB(t)) to a Rician distribution, and keep other simulation parameters unchanged. Environment 3: Reduce the correlation coefficient φ to 0.64 to simulate an environment where the channel parameters change rapidly over time, and keep other simulation parameters unchanged. The sum-rate performance simulation results of MADDQN&MADDPG and the baseline algorithms under Environment 1 are as Figure 1 shown; the sum-rate performance simulation results of MADDQN&MADDPG and the baseline algorithms under Environment 2 are as Figure 2 shown; the sum-rate performance simulation results of MADDQN&MADDPG and the baseline algorithms under Environment 3 are as Figure 3 shown. As Figures 1-3 shown, our MADDQN&MADDPG algorithm is superior to the other three baseline algorithms in terms of convergence speed and sum rate. The reasons are as follows: 1) The action space of the single-agent algorithm grows exponentially with the increase in the number of actions. This means that the single-agent method requires more action selections to obtain the same result as the multi-agent algorithm, which leads to higher computational complexity and performance degradation. The sum-rate performance simulation results of BDTL and the baseline algorithms under Environment 1 are as Figure 4 shown; the sum-rate performance simulation results of BDTL and the baseline algorithms under Environment 2 are as Figure 5 shown; the sum-rate performance simulation results of BDTL and the baseline algorithms under Environment 3 are as Figure 6 shown. From Figures 4-6It can be seen that, compared with the baseline algorithm, BDTL has a faster convergence rate, especially in the early stage of training. This shows that under the premise of dynamically adjusting the ratio of the knowledge of the old agent and the experience of the new agent, bidirectional transfer learning does help the new agent better adapt to different environments.

[0064] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. A cooperative transmission method for RIS-enhanced uplink OFDMA system, which is based on two-way migration reinforcement learning, is characterized in that It includes the following steps: Step 1: Define the key parameters of the uplink OFDMA system; Step 2: Design a multi-agent double deep Q-network algorithm to solve the subcarrier allocation problem, design a multi-agent deep deterministic policy gradient algorithm to solve the power allocation and RIS control problems, and use the corresponding algorithms to train the corresponding agents in the current environment; Step 3: Change the current environment with different schemes, set up a two-way transfer learning framework, and enhance the adaptability of the neural network in different environments by using the common configurations of different environments, so that the agents can more efficiently dynamically adjust the resource allocation and RIS control strategies to maximize the sum rate of all users: Step 1: Set different environments, then use the MADDQN and MADDPG algorithms to train the agents in these environments, and finally, save the trained agents after the training process ends. Step 2: Train new agents in the environments set in Step 1; The specific process of Step 1 is as follows: Process 1: Set that the RIS-enhanced uplink multi-cell OFDMA system consists of N cells, each cell is equipped with a base station with L antennas, and M single-antenna users are randomly distributed in each cell; Process 2: Represent the set of base stations and the set of users associated with base station n as α = {1, 2,..., N} and β = {1, 2,..., M}; Process 3: Set that there are K subcarriers in each cell, and the subcarrier set is defined as γ = {1, 2,..., K}; Process 4: Set that a RIS consists of E reflection elements, and the set of reflection elements is represented as ε = {1, 2,..., E}; Process Five: Define the subcarrier allocation parameter as wherein indicates that user m in cell n is allocated to subcarrier k in time slot t; Step 6. The data rate of user m in time slot t on subcarrier k in base station n is defined as Process 7: Obtain the maximum sum of the transmission rates of all users; The specific process of Step 2 is as follows: Process 1: Divide the resource allocation framework and RIS control into a subcarrier allocation module, a power allocation module, and a RIS control module; Process 2: Regard the users in each subcarrier allocation module and power allocation module as an agent, regard the RIS control module as an agent, and regard the OFDMA system as the environment; Process 3: For the subcarrier allocation module, equip each agent with a DDQN unit composed of a training network and a target network; Process 4: Equip each agent set in Process 2 with a DDPG unit composed of a training actor network, a target actor network, a training critic network, and a target critic network. Since there are multiple agents, a multi-agent network is formed, namely MADDQN and MADDPG.

2. The RIS-enhanced uplink OFDMA system cooperative transmission method according to claim 1, wherein In the fourth process described above, in the three modules, the training process of each agent is as follows: In time slot t, the OFDMA system feeds back its state to each agent. Then, each Q network in the subcarrier allocation module randomly selects an action from the action space of this module with probability ε, or selects the action that maximizes the Q function value according to Equation (1) with probability 1 - ε: Among them, is the action generated by each agent in the subcarrier allocation module, is the state feedback from the environment to each agent in the subcarrier allocation module, is the Q-network parameter of each agent in the subcarrier allocation module, and Α is the action space of the subcarrier allocation module.

3. The RIS-enhanced uplink OFDMA system cooperative transmission method according to claim 2, wherein After each agent in the subcarrier allocation module and the power allocation module executes the selected action, it obtains the real-time reward returned from the OFDMA system, and the real-time rewards of the subcarrier allocation module and the power allocation module are uniformly defined as Equations (2) and (3): Among them Among them, represents the data rate of user m assigned to subcarrier k in cell n, represents the penalty term assigned to user m in cell n, ρ is the penalty coefficient, and the reward function of the agent RIS in the RIS control module is 4. The RIS-enhanced uplink OFDMA system cooperative transmission method according to claim 3, characterized in that The agents in the subcarrier allocation module, power allocation module, and RIS control module continuously interact with the OFDMA system to continuously obtain real-time samples and and store the real-time samples in the experience pool of each module accordingly. The experience replay method is introduced to eliminate data correlation, and part of the samples are randomly drawn from the experience pools of the subcarrier allocation module, power allocation module, and RIS control module respectively and The loss functions of each DDQN and DDPG unit in the subcarrier allocation module, power allocation module, and RIS control module are respectively defined as Among them Among them, τ and τ - are the parameters of the training Q-network and the target Q-network respectively, μ and μ - are the parameters of the training actor network and the target actor network respectively, θ and θ - are the parameters of the training critic network and the target critic network respectively, γ is the discount rate of DDQN, η is the discount rate of DDPG; for the DDQN unit and the DDPG unit of each agent, the RMSProp optimizer is used to update the parameters of the training Q-network and the training critic network by minimizing the loss functions (4) and (5).

5. The RIS-enhanced uplink OFDMA system cooperative transmission method according to claim 4, characterized in that In step three: The specific method of training a new agent in a new environment using a bidirectional transfer learning framework based on MADDQN and MADDPG is as follows: First, three different environments are set: Environment 1, keeping the default environment parameters unchanged; Environment 2, since there may be a line-of-sight path in the RIS channel, changing the small-scale fading between the RIS and the BS to a Rice distribution, and keeping other simulation parameters unchanged; Environment 3, reducing the correlation coefficient φ to 0.64 to simulate an environment where channel parameters change rapidly over time, and keeping other simulation parameters unchanged; Then, in the new environment, the proposed bidirectional transfer learning framework is used to train the new agent. During the training process, the loss functions of each new DDQN unit in the subcarrier allocation module and each new DDPG unit in the power allocation module and the RIS control module are respectively expressed as: Among them Among them, and Ω f respectively represent the network trained in environment g and the new network in environment f. and π(·) respectively represent the trained actor network and the new actor network. ψ represents a scaling factor that takes values in the range (0, 1] and gradually decreases according to the rule ψ←ψ / (1 + Θ) at each time slot t, where Θ is the decay factor.