A model training and uplink resource occupation method, device, equipment and medium
By training the Actor and Critic networks of user equipment using a multi-agent reinforcement learning strategy, the problem of insufficient uplink transmission channel state feedback in active open-loop networks is solved, achieving high-reliability and low-latency network transmission.
Patent Information
- Application Number
- CN202310166096.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2026-01-20
- Estimated Expiration
- 2043-02-22
AI Technical Summary
In an active open-loop network, the uplink transmitter cannot obtain channel status feedback in real time and accurately, which leads to the inability to make reasonable resource allocation and reduces transmission reliability.
A multi-agent reinforcement learning strategy is adopted to model the active network architecture, train the Actor network and Critic network of the user equipment, and solve the problem of channel state feedback in uplink transmission by centralized training and distributed execution.
It improves network reliability, ensures high reliability with extremely low latency, and avoids control latency overhead and resource conflicts caused by real-time data interaction.
Smart Images

Figure CN116321434B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wireless communication technology, and more specifically, to a model training, uplink resource allocation method, apparatus, device, and medium. Background Technology
[0002] Currently, mobile communications based on closed-loop communication, despite employing numerous technologies to compress latency, still primarily suffer from latency due to complex closed-loop control and signaling interactions. Therefore, a fundamental change to the closed-loop network architecture is necessary. Active open-loop networks eliminate all direct control and signaling interactions in their architecture, achieving extremely low-latency communication to support more 5G and even 6G vertical applications.
[0003] However, because active open-loop networks eliminate all control signaling interactions, the sender cannot obtain channel status feedback and other information in real time during uplink transmission, thus failing to make reasonable resource allocation, resulting in reduced transmission reliability. Summary of the Invention
[0004] The problem addressed in this application is that in active open-loop networks, the uplink transmitter cannot accurately obtain information such as channel state feedback in real time, thus making it impossible to make reasonable resource allocation.
[0005] To address the aforementioned problems, the first aspect of this application provides a model training method, comprising:
[0006] Modeling of proactive network architecture based on Markov decision-making determines the state model, action model, and reward model of user equipment in the proactive network architecture.
[0007] The network models constructed from the state model, action model, and reward model of the user device in the active network architecture are trained using a multi-agent reinforcement learning strategy, resulting in the trained Actor network and Critic network of the user device.
[0008] A second aspect of this application provides an uplink resource occupancy method, which includes:
[0009] Get the current state of the user device;
[0010] The Actor network and Critic network of the user device are obtained, and the Actor network and Critic network are trained according to the model training method described above;
[0011] The current state is input into the Actor network and the Critic network to obtain the output action of the user equipment.
[0012] The uplink resources occupied by the user equipment are determined based on the output action.
[0013] A third aspect of this application provides a model training apparatus, comprising:
[0014] The modeling module is used to model the proactive network architecture based on Markov decision-making, and to determine the state model, action model and reward model of user equipment in the proactive network architecture.
[0015] The training module is used to train the network model constructed from the state model, action model, and reward model of the user device in the active network architecture through a multi-agent reinforcement learning strategy, so as to obtain the trained Actor network and Critic network of the user device.
[0016] A fourth aspect of this application provides an uplink resource occupancy device, comprising:
[0017] The status acquisition module is used to acquire the current status of the user equipment;
[0018] The model acquisition module is used to acquire the Actor network and Critic network of the user device, which are trained according to the model training method described above.
[0019] The output acquisition module is used to input the current state into the Actor network and the Critic network to obtain the output action of the user equipment.
[0020] The resource determination module is used to determine the uplink resources occupied by the user equipment based on the output action.
[0021] A fifth aspect of this application provides a terminal device, which includes: a memory and a processor;
[0022] The memory is used to store programs;
[0023] The processor, coupled to the memory, is used to execute the program for performing the aforementioned model training method, or for performing the aforementioned uplink resource occupancy method.
[0024] The sixth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the model training method described above, or to implement the uplink resource occupancy method described above.
[0025] In this application, a distributed execution method with centralized training using a multi-agent reinforcement learning strategy can be adopted to train the resource occupancy model of user equipment based on the global information of the active network. This solves the problem that the uplink transmission sender in the active network cannot obtain global information such as channel state feedback in real time and thus cannot make reasonable resource occupancy.
[0026] In this application, compared with the traditional resource allocation method of closed-loop network, the focus is no longer on optimizing system energy consumption, but on network reliability, thereby improving network reliability and ensuring high reliability under extremely low latency.
[0027] In this application, a comprehensive scheme for training a multi-agent reinforcement learning strategy is adopted. Each user device executes the optimal resource occupancy strategy based on the trained resource occupancy model. This can guide the proactive occupancy of online resources in application scenarios with extremely high real-time requirements, thereby avoiding the control latency overhead caused by real-time data interaction and reducing resource conflicts to ensure network reliability. Attached Figure Description
[0028] Figure 1 This is a schematic diagram of an active open-loop network;
[0029] Figure 2 This is a flowchart of a model training method according to an embodiment of this application;
[0030] Figure 3 This is a flowchart of a model training method S200 according to another embodiment of this application;
[0031] Figure 4 This is a schematic diagram of a multi-agent reinforcement learning model according to an embodiment of this application;
[0032] Figure 5 This is a schematic diagram of the wireless resource model in this application;
[0033] Figure 6 This is a flowchart of an uplink resource occupancy method according to an embodiment of this application;
[0034] Figure 7 This is a structural block diagram of a model training apparatus according to an embodiment of this application;
[0035] Figure 8 This is a structural block diagram of an uplink resource occupancy device according to an embodiment of this application;
[0036] Figure 9 This is a structural block diagram of a terminal device according to an embodiment of this application. Detailed Implementation
[0037] To make the above-mentioned objects, features, and advantages of this application more apparent and understandable, specific embodiments of this application will be described in detail below with reference to the accompanying drawings. Although exemplary embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of this application and to fully convey the scope of this application to those skilled in the art.
[0038] It should be noted that, unless otherwise stated, the technical or scientific terms used in this application shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains.
[0039] URLLC (Ultra-Reliable Low-Latency Communication), as one of the three major service categories of 5G, has enabled the rapid development of vertical fields such as factory automation, autonomous driving, and augmented / virtual reality (smart factories, smart assembly, and vehicle-to-everything). However, existing 5G URLLC technologies are all based on closed-loop communication processes. Although many technologies have been adopted to compress data (re)transmission latency, the huge latency overhead caused by the complex control signaling interaction in closed-loop communication is the main source of 5G air interface latency. To effectively reduce latency, the complex control signaling interaction process must be eliminated. Therefore, a network architecture based on active open-loop communication has emerged.
[0040] In an active open-loop network architecture, there is no complex control information exchange. A specific active open-loop network deployment diagram is shown below. Figure 1 As shown, each User Equipment (UE) is constantly moving and actively associating with nearby Access Points (APs) to form a virtual cell (VC) centered on itself. It autonomously selects appropriate radio resources for unlicensed transmission and uses spatial diversity and other technologies to ensure reliable data transmission, thereby eliminating the complex control feedback process and bringing low latency requirements. (Note that in traditional closed-loop communication, the network and radio resources required for UE uplink transmission need to be uniformly allocated by the base station (BS) and other control centers, and the allocation results are notified to each UE through broadcast. We call this process resource allocation; while in the active open-loop network, the UE autonomously selects network and radio resources. This process is called resource occupancy.)
[0041] Because the notification and feedback process of closed-loop communication is eliminated, the UE needs to make autonomous decisions to select appropriate radio and network resources during uplink transmission without any auxiliary information such as channel state information. This inevitably leads to resource conflicts (conflict is defined as when different UEs select the same radio resource block for transmission, for example, in...). Figure 2 In this scenario, UE1 and UE2 select the same AP2 for data transmission (leading to a conflict due to selecting the same radio resource block), which reduces the data transmission success rate and degrades network reliability. Therefore, avoiding blind uplink decision-making and ensuring high reliability while maintaining extremely low latency in open-loop transmission is of paramount importance.
[0042] To address the aforementioned issues, this application provides a novel model training scheme that, by introducing a multi-agent reinforcement learning algorithm and employing a centralized training and distributed execution approach, solves the problem in active open-loop networks where the uplink transmitter cannot accurately obtain channel state feedback information in real time, thus hindering reasonable resource utilization.
[0043] For ease of understanding, the following terms may be used and are explained below:
[0044] UE (User Equipment): User equipment, user terminal, refers to the equipment or terminal that performs user functions in mobile communications.
[0045] APs (Access Points): Wireless access points are the access points for mobile devices / user devices (including mobile robots, vehicles, mobile phones, etc.) to enter the communication network. They are mainly used in broadband homes, inside buildings, and inside parks, and can cover tens to hundreds of meters.
[0046] VC (Virtual Cell): A virtual cell is a virtual area consisting of user equipment and associated wireless access points.
[0047] This application provides a model training method, which can be executed by a model training device that can be integrated into electronic devices such as anchor nodes, base stations, pads, computers, servers, computer clusters, and data centers. Figure 2 The diagram shown is a flowchart of a model training method according to an embodiment of this application; wherein the model training method includes:
[0048] S100, based on Markov decision-making, models the active network architecture and determines the state model, action model and reward model of user equipment in the active network architecture;
[0049] By modeling the active network architecture, the uplink resource occupancy problem in the active network is transformed into a Markov decision process, thereby converting the uplink resource occupancy problem into a decision problem.
[0050] S200 trains the network model constructed from the state model, action model, and reward model of the user device in the active network architecture through a multi-agent reinforcement learning strategy, resulting in the trained Actor network and Critic network of the user device.
[0051] Among them, multi-agent reinforcement learning strategies can include the MADDPG (Multi-agent deep deterministic policy gradient) algorithm, the COMA (counterfactual multi-agent policy gradients) algorithm, the MAPPO (Multi-Agent PPO) algorithm, and the VDN (Value Decomposition Networks) algorithm, etc.
[0052] In the multi-agent reinforcement learning strategy, each user device has an Actor network and a Critic network, as well as a target Actor network and a target Critic network. In this application, the model training method can train the Actor network and Critic network of the user device, as well as the target Actor network and target Critic network. However, in the uplink resource occupancy method, only the Actor network and Critic network of the user device need to be used.
[0053] It should be noted that, as mentioned above, the network models constructed from the user device's state model, action model, and reward model include an Actor network and a Critic network, as well as a target Actor network and a target Critic network. After training, the trained Actor network and Critic network, as well as the target Actor network and target Critic network, can be obtained. However, since the target Actor network and target Critic network are not required in the subsequent acquisition method, this application only needs to train the Actor network and Critic network for acquiring the user device, and there is no restriction on whether to acquire the target Actor network and target Critic network.
[0054] In this application, the network model constructed during training of the network model built from the user device's state model, action model, and reward model in an active network architecture means that the user device's state model, action model, and reward model construct the network model, and this network model is trained in the active network architecture. This network model refers to the user device's Actor network and Critic network, as well as the target Actor network and target Critic network (the target Actor network and target Critic network are used in the training process but not in subsequent resource consumption processes). Alternatively, it can be described as training the user device's network model in an active network architecture, which is constructed from the user device's state model, action model, and reward model.
[0055] In this application, a distributed execution method with centralized training using a multi-agent reinforcement learning strategy can be adopted to train the resource occupancy model of user equipment based on the global information of the active network. This solves the problem that the uplink transmission sender in the active network cannot obtain global information such as channel state feedback in real time and thus cannot make reasonable resource occupancy.
[0056] In this application, compared with the traditional resource allocation method of closed-loop network, the focus is no longer on optimizing system energy consumption, but on network reliability, thereby improving network reliability and ensuring high reliability under extremely low latency.
[0057] In this application, a comprehensive scheme for training a multi-agent reinforcement learning strategy is adopted. Each user device executes the optimal resource occupancy strategy based on the trained resource occupancy model. This can guide the proactive occupancy of online resources in application scenarios with extremely high real-time requirements, thereby avoiding the control latency overhead caused by real-time data interaction and reducing resource conflicts to ensure network reliability.
[0058] In one implementation, such as Figure 3 Combination Figure 4 As shown, training the network model constructed from the state model, action model, and reward model of the user device in the active network architecture using a multi-agent reinforcement learning strategy includes:
[0059] S201, construct the Actor network, Critic network, target Actor network, and target Critic network of the user equipment;
[0060] Among them, there are multiple multi-agent reinforcement learning strategies. In this application, the MADDPG strategy is used as an example for model training.
[0061] In this application, the user equipment is an intelligent agent, which can be represented as an intelligent agent (Agent). i .
[0062] In this application, for the intelligent agent... i Its Actor network is represented as:
[0063] μ i (S i ;θ i )
[0064] Among them, S i Indicates Agent i Observational information, θ i These are the parameters of its Actor network;
[0065] In this application, for the intelligent agent... i Its Critic network is represented as:
[0066] q i (S1,S2,...,S N A1, A2, ..., A N ;ω i )
[0067] Where, ω i These are the parameters of its Critic network;
[0068] In this application, for the intelligent agent... i Its target Actor network is represented as:
[0069] μ′ i (S i ;θ′ i )
[0070] Where, θ′ i These are the parameters of its Actor network;
[0071] In this application, for the intelligent agent... i Its target Critic network is represented as:
[0072] q′ i (S1,S2,...,S N A1, A2, ..., A N ;ω′ i )
[0073] Where, ω′ i These are the parameters of its Critic network.
[0074] S202, initialize the network parameters of the Actor network, Critic network, target Actor network, and target Critic network, the experience replay pool, and the maximum number of training iterations;
[0075] S203, randomly determine the initial state of each user equipment;
[0076] S204, In each time slot, for each user equipment, perform an action in the current state to determine the acquired reward value and the next state;
[0077] S205, store the current state, the action, the reward value, and the next state into the experience replay pool; and update the current state to the next state;
[0078] S206, For each user equipment, randomly sample multiple samples from the experience replay pool, and determine the target network estimate for each user equipment based on the samples;
[0079] S207, Update the Actor network and Critic network of the user equipment based on the target network estimate;
[0080] Among them, Agent i The expected return in the long term is:
[0081]
[0082] Where r i For the discount factor to range from 0 to 1, for θ i The gradient is used for updating the Actor network, i.e.:
[0083]
[0084] Where D is the experience replay pool used to store the sampling state action information (S, A1, A2, ..., A) for each sampling. N (S1', S2', ..., S2') is sampled and updated when updating the Actor network parameters; N ') represents all agents in state (S1, S2, ..., S...). N The action to be taken when (A1, A2, ..., A) N The observations of the new state that has been transferred.
[0085] The modeling loss function is used for updating the Critic network, as shown below:
[0086] L i =E[(q(S1,S2,...,S N A1, A2, ..., A N ;ω i )-y i ) 2 ]
[0087] Among them, y i This is an estimate of the target network, calculated as follows:
[0088] y i =R i +γq′ i (S1,S2,...,S N A1, A2, ..., A N ;ω′ i )
[0089] S208, after all user equipment Actor and Critic networks have been updated, soft update the target Actor and target Critic networks of the user equipment;
[0090] Repeat step S204. In each time slot, for each user device, perform the current action in the current state, determine the obtained reward value and the next state, until the maximum number of training iterations is reached.
[0091] In each training round, the Agent i The target Actor network parameters are given by θ′ i =αθ i +(1-α)θ′ i Perform soft updates; similarly, the target Critic network parameters are determined by ω′. i =αω i +(1-α)ω′ i Perform a soft update.
[0092] If the maximum number of training iterations has not been reached, steps S204-S208 are executed repeatedly.
[0093] In this application, each user equipment continuously collects information in a dynamic environment through a multi-agent reinforcement learning strategy and transmits it uplink to the base station for storage and computation. The base station then feeds back the computation results and useful training information to each user equipment via the downlink. As time goes by, the base station and user equipment collect more and more global information, which enables the continuous iterative updates of the Actor network and Critic network parameters, allowing the algorithm to converge to a stable state and complete the training.
[0094] In this application, the fully trained Actor network will enable user equipment to autonomously decide on proactive open-loop network uplink resource allocation strategies in dynamically changing environments to achieve high network reliability.
[0095] In one implementation, step S100 involves modeling the proactive network architecture based on Markov decision principles to determine the state model, action model, and reward model of user equipment within the proactive network architecture.
[0096] The state model of the user equipment is as follows:
[0097]
[0098] Among them, L t For user equipment locations, RB stands for Radio Resource Block, and N stands for... b S represents the number of radio resource blocks that all channels can provide within a time slot. i (t) represents the state of the system within time slot t. Let i be the number of radio resource blocks required by the i-th user equipment. Let be the matrix of all RBs in the entire system for the i-th user equipment within time slot t.
[0099] In one implementation, the action model of the user equipment is as follows:
[0100] A i (t)={a1(t),a2(t),...,a N (t)}
[0101] a i (t)={p i,c U i ,s t i,c,m}
[0102] Among them, U i Let s be the set of access points associated with the i-th user equipment. t i,c,m To determine whether the m-th radio resource block of the c-th sub-channel is occupied by the i-th user equipment, p i,c It is the power allocated by the i-th user equipment to the c-th sub-channel, a i (t) represents the i-th user equipment.
[0103] In one implementation, the reward model obtained by the user equipment is as follows:
[0104]
[0105] Among them, s t i,c,m To determine whether the m-th radio resource block of the c-th sub-channel is occupied by the i-th user equipment, N b This represents the number of radio resource blocks that all channels can provide within a time slot. For user equipment data within time slot t, The l-th independent sub-part a in the user equipment data l The signal-to-noise ratio, N r S represents the number of radio resource elements that all channels can provide within a time slot, where s = S t a=A t, R i For UE i The reward.
[0106] Thus, within time slot t, if UE i If the uplink transmission SNR meets the minimum signal-to-noise ratio requirement and the number of RBs occupied by the UE per unit time slot, then the data transmission is successful. At this time, the UE... i The reward is 1 if you receive 1 otherwise, it is 0.
[0107] In one implementation, using the same reward function for all agents ensures consistency in their objectives. This maximizes the system's transmission success rate rather than the success rate of any individual agent, resulting in a total system reward that is the sum of the rewards for all users (UEs).
[0108] In one implementation, before S100, modeling the proactive network architecture based on Markov decision and determining the state model, action model, and reward model of user equipment in the proactive network architecture, the method further includes:
[0109] Construct an active open-loop network system model;
[0110] Construct a wireless resource model;
[0111] Construct a wireless channel communication model between the UE and the AP;
[0112] Construct a data SIMO multipath transmission model;
[0113] Establish an optimization problem model for uplink resource utilization in an active open-loop network.
[0114] The steps for constructing the active open-loop network system model are as follows:
[0115] It should be noted that in the active open-loop network architecture of this application, the network is divided into two parts: a radio access network and a core network. The radio access network consists of K access points (APs) and one base station (BS) managing the APs. The base station (BS) and each AP are connected via wired connections. The base station (BS) has certain storage and computing capabilities to manage the radio access network and is directly connected to the core network.
[0116] In constructing an active open-loop network system model, it is assumed that there are N real-time mobile UEs in the network, where UEs can be devices such as cars, industrial intelligent robots, and drones;
[0117] In constructing the active open-loop network system model, the set of UEs is defined as follows:
[0118] N = {1, 2, ..., N}
[0119] In constructing the active open-loop network system model, the set of UE locations is defined as follows:
[0120] L t ={(x1) t ,y1 t ),(x2 t ,y2 t ),...,(x N t ,y N t )}
[0121] In this process, each User Equipment (UE) selects the k nearest Access Points (APs) to form a virtual cell. The k APs in the virtual cell provide communication services to the UE in the form of Radio Resource Blocks (RBs) through spatial diversity, and finally transmit the data to the Base Station (BS) for processing.
[0122] In constructing an active open-loop network system model, the UE is defined. i The set of k associated APs is as follows:
[0123] U i ={AP i,1 AP i,2 ,...,AP i,k}
[0124] In constructing an active open-loop network system model, the entire network system time T = {1, 2, ..., T} is divided into T equal time slots, and the length of each time slot is τ.
[0125] The steps for constructing the wireless resource model are as follows:
[0126] It should be noted that, at the physical layer, data transmission for each UE is accomplished using Orthogonal Frequency Division Multiple Access (OFDMA). Radio resources are divided in the time and frequency domains in the form of Radio Resource Units (RUs), where each RU contains one subcarrier and one symbol.
[0127] Among them, the radio resource block (RB) is a radio resource composed of a fixed number of adjacent RUs and mapped to the link layer, with a transmission capacity of l.
[0128] like Figure 5 As shown, each small square is a RU, and four adjacent RUs form an RB.
[0129] Among them, different adjacent fixed numbers of RU can be selected to form RB.
[0130] Wherein, if all channels in each time slot can provide N r One RU can be mapped to N b There are RBs, therefore each RB contains N. m One RU can determine N r =N b ×N m .
[0131] Among them, if RU m Within a certain time slot, by a certain AP j Data occurring within time slot t is denoted as s. j,m =1; otherwise sj,m =0 indicates that it was not used by AP j Occupancy. The occupancy matrix of all RUs in the entire system within time slot t can be represented as:
[0132]
[0133] The steps for constructing the wireless channel communication model between the UE and the AP are as follows:
[0134] When UE i AP within the VC associated with itself in time slot t j When communicating, the channel gain coefficient between the two is expressed as:
[0135]
[0136] in, It is UE i and AP j The small-scale fading factors between them follow a Rayleigh distribution. For UE i and AP j The Euclidean distance between them, where α is the path loss exponent, with a value ranging from 2 to 5.
[0137] It should be noted that the wireless channel communication model between the UE and APs can also use any other channel fading model such as the Nakagami-m distribution, and no specific model is restricted in this application.
[0138] The steps for constructing the data SIMO multipath transmission model are as follows:
[0139] It should be noted that in order to meet the minimum latency requirement in the uplink transmission of an open-loop network, each data packet is transmitted only once in the channel. If the data cannot be decoded by the base station, the transmission fails or the data packet is lost.
[0140] Due to the multipath transmission formed by the UE and k APs, the uplink transmission is a single-input multiple-output (SIMO) system.
[0141] The transmission process is configured such that, within time slot t, the UE... i Data needs to be Divide into n independent sub-parts, and select L resource blocks RB through the associated AP. j Uploaded to the base station (BS), the resource block is represented as follows:
[0142]
[0143] Among them, UE i Different resource blocks (RBs) are selected for transmission for each independent sub-part, and a binary variable s is defined.t i,c,m Indicates whether the m-th RB of the c-th sub-channel is controlled by the UE. i Occupy, s t i,c,m =1 indicates that the RB is controlled by the UE. i Occupied, or not occupied by UE i Occupied.
[0144] At this point, the SIMO system can be further interpreted as the UE. i Data When transmission is performed through k associated APs, each of which is associated with the UE i Each associated AP independently completes the data transmission process described above. The transmission.
[0145] Therefore, it can be determined that the UE within time slot t... i AP usage j When the m-th RB of the c-th sub-channel transmits the l-th sub-part of the data, the signal-to-noise ratio (SNR) of this communication link is:
[0146]
[0147] in, It is UE i The power allocated to the c-th sub-channel.
[0148] In this process, the base station (BS) side adopts a selective merging method to determine the final SNR of each sub-part of the link with the highest SNR, where the l-th sub-part a l The SNR is:
[0149]
[0150] The base station (BS) was able to successfully receive the entire data from the user (UE). i Data The following signal-to-noise ratio needs to be met:
[0151]
[0152] Where ε is the minimum signal-to-noise ratio threshold that the signal can be accurately received and decoded by the base station.
[0153] It can be seen that UE i Data Whether the data can be successfully received by the base station (BS) depends on the smallest sub-data portion of the SIN. Therefore, as long as the data... If the above formula is satisfied during transmission, then the data is considered valid. Transmission successful; otherwise, it is not.
[0154] The steps involved in establishing the optimization problem model for uplink resource utilization in an active open-loop network are as follows:
[0155] In this application, the success rate ρ of data packets for all UEs during the long-term time T→∞ is defined as a reliability index of the uplink, as follows:
[0156]
[0157] Where O(t) is the total number of all UE transmission tasks in time slot t, δ i (t) is UE i The number of tasks successfully transmitted at time t.
[0158] Therefore, the uplink transmission reliability optimization problem based on open-loop transmission is formulated as follows:
[0159]
[0160] C3:||U i ||≤K
[0161]
[0162] Among them, P max C1 represents the UE's maximum transmit power, C2 represents the UE's total data transmission power limit, and C2 represents the AP's maximum transmit power. j The maximum number of RBs occupied does not exceed N c C3 is the limit on the maximum number of APs that can be associated with each UE, and C4 is the requirement for successful data transmission in uplink communication.
[0163] In this application, the ultimate goal of the uplink resource occupancy strategy is to find the optimal strategy π*:
[0164]
[0165] Where ρ is the reliability index of the uplink.
[0166] This application provides an uplink resource occupancy method, which can be executed by an uplink resource occupancy device that can be integrated into electronic devices such as anchor nodes, base stations, pads, computers, servers, computer clusters, and data centers. Figure 6 The diagram shows a flowchart of an uplink resource occupancy method according to an embodiment of this application; wherein the uplink resource occupancy method includes:
[0167] S10, Obtain the current status of the user equipment;
[0168] The method for obtaining the state of the user equipment in this step can refer to the method for obtaining the state of the user equipment in the aforementioned model training method. The specific acquisition process will not be repeated here.
[0169] S20, Obtain the Actor network and Critic network of the user device, wherein the Actor network and Critic network are trained according to the model training method described above;
[0170] The training process for the Actor network and Critic network of the user device in this step can refer to the aforementioned model training method; the specific acquisition process will not be repeated here.
[0171] S30, input the current state into the Actor network and Critic network to obtain the output action of the user equipment;
[0172] In one implementation, multiple output actions of the user are determined by an Actor network, and the user's preferences (or probabilities of performing) of the multiple output actions are determined by a Critic network. The final output action is then determined based on the preferences.
[0173] In one implementation, the Actor network and Critic network are systems with preset preference selection strategies. The current state is input into the system to obtain the output action of the user device.
[0174] S40, determine the uplink resources occupied by the user equipment based on the output action.
[0175] Among them, the uplink resources occupied by user equipment include resource allocation policies, power allocation, and associated APs.
[0176] In this application, through an offline-trained multi-agent reinforcement learning network, each user device can determine uplink resource occupancy information based on its current state, thereby ensuring ultra-low latency and high reliability in executing uplink data transmission tasks.
[0177] In this application, the training process of the multi-agent reinforcement learning strategy is to actively select appropriate associated APs and available RB resources for data transmission in a real-time dynamic environment through coordination and cooperation among UEs and by utilizing the global non-real-time information obtained by the AN, thereby avoiding resource conflicts and completing high-reliability, low-latency data transmission in application scenarios with extremely high real-time requirements.
[0178] In this application, the global non-real-time information obtained through the anchor node (AN) (which originates from the previous uplink transmission of the user equipment UE) is used to train the resource occupancy scheme of each user equipment UE offline in a centralized manner using the MADDPG algorithm. Then, based on the fully trained scheme, each user equipment UE executes the optimal occupancy strategy online in a distributed manner. This can guide the active online resource occupancy in application scenarios with extremely high real-time requirements, thereby avoiding the control delay overhead caused by real-time data interaction and reducing resource conflicts to ensure network reliability.
[0179] This application provides a model training apparatus for executing the model training method described above. The model training apparatus will be described in detail below.
[0180] like Figure 7 As shown, the model training device includes:
[0181] Modeling module 101 is used to model the active network architecture based on Markov decision and determine the state model, action model and reward model of user equipment in the active network architecture.
[0182] Training module 102 is used to train the network model constructed by the state model, action model and reward model of the user device in the active network architecture through a multi-agent reinforcement learning strategy, so as to obtain the trained Actor network and Critic network of the user device.
[0183] In one embodiment, the training module 102 is further configured to:
[0184] Construct the Actor network, Critic network, target Actor network, and target Critic network for the user equipment;
[0185] Initialize the network parameters of the Actor network, Critic network, target Actor network, and target Critic network, as well as the experience replay pool and the maximum number of training iterations.
[0186] Randomly determine the initial state of each user device;
[0187] Within each time slot, for each user device, perform an action in the current state, determine the acquired reward value and the next state;
[0188] The current state, the action, the reward value, and the next state are stored in the experience replay pool; and the current state is updated to the next state.
[0189] For each user equipment, multiple samples are randomly sampled from the experience replay pool, and the target network estimate for each user equipment is determined based on the samples;
[0190] Update the Actor network and Critic network of the user equipment based on the target network estimate;
[0191] After all user equipment's Actor and Critic networks are updated, the target Actor and target Critic networks of the user equipment are soft-updated.
[0192] Repeat the process of performing the current action in the current state for each user device within each time slot, determining the obtained reward value and the next state, until the maximum number of training iterations is reached.
[0193] In one implementation, the state model of the user equipment is:
[0194]
[0195] Among them, L t For user equipment locations, RB stands for Radio Resource Block, and N stands for... b This represents the number of radio resource blocks that all channels can provide within a time slot.
[0196] In one implementation, the action model of the user equipment is as follows:
[0197] A i (t)={a1(t),a2(t),...,a N (t)}
[0198] a i (t)={p i,c U i ,s t i,c,m}
[0199] Among them, U i Let s be the set of access points associated with the i-th user equipment. t i,c,m To determine whether the m-th radio resource block of the c-th sub-channel is occupied by the i-th user equipment, p i,c It is the power allocated by the i-th user equipment to the c-th sub-channel.
[0200] In one implementation, the reward model obtained by the user equipment is as follows:
[0201]
[0202] Among them, s t i,c,m To determine whether the m-th radio resource block of the c-th sub-channel is occupied by the i-th user equipment, N b This represents the number of radio resource blocks that all channels can provide within a time slot. For user equipment data within time slot t, The l-th independent sub-part a in the user equipment data l The signal-to-noise ratio, N r This represents the number of radio resource units that all channels can provide within a time slot.
[0203] The model training apparatus provided in the above embodiments of this application and the model training method provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.
[0204] This application provides an uplink resource occupancy device for executing the uplink resource occupancy method described above. The uplink resource occupancy device will be described in detail below.
[0205] like Figure 8 As shown, the uplink resource occupancy device includes:
[0206] Status acquisition module 201 is used to acquire the current status of the user equipment;
[0207] The model acquisition module 202 is used to acquire the Actor network and Critic network of the user device, wherein the Actor network and Critic network are trained according to the model training method described above.
[0208] The output acquisition module 203 is used to input the current state into the Actor network and the Critic network to obtain the output action of the user equipment.
[0209] Resource determination module 204 is used to determine the uplink resources occupied by the user equipment based on the output action.
[0210] The uplink resource occupancy device and the uplink resource occupancy method provided in the above embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.
[0211] The above describes the internal functions and structure of the model training device / uplink resource occupancy device, such as... Figure 9 As shown, in practice, the model training device / uplink resource occupancy device can be implemented as a terminal device, including: memory 301 and processor 303.
[0212] Memory 301 can be configured to store a program.
[0213] Additionally, memory 301 can also be configured to store various other data to support operation on the terminal device. Examples of this data include instructions for any application or method used on the terminal device, contact data, phonebook data, messages, pictures, videos, etc.
[0214] The memory 301 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0215] Processor 303, coupled to memory 301, is used to execute programs in memory 301 for:
[0216] Modeling of proactive network architecture based on Markov decision-making determines the state model, action model, and reward model of user equipment in the proactive network architecture.
[0217] The network models constructed from the state model, action model, and reward model of the user device in the active network architecture are trained using a multi-agent reinforcement learning strategy, resulting in the trained Actor network and Critic network of the user device.
[0218] In one implementation, processor 303 is specifically used for:
[0219] Construct the Actor network, Critic network, target Actor network, and target Critic network for the user equipment;
[0220] Initialize the network parameters of the Actor network, Critic network, target Actor network, and target Critic network, as well as the experience replay pool and the maximum number of training iterations.
[0221] Randomly determine the initial state of each user device;
[0222] Within each time slot, for each user device, perform an action in the current state, determine the acquired reward value and the next state;
[0223] The current state, the action, the reward value, and the next state are stored in the experience replay pool; and the current state is updated to the next state.
[0224] For each user equipment, multiple samples are randomly sampled from the experience replay pool, and the target network estimate for each user equipment is determined based on the samples;
[0225] Update the Actor network and Critic network of the user equipment based on the target network estimate;
[0226] After all user equipment's Actor and Critic networks are updated, the target Actor and target Critic networks of the user equipment are soft-updated.
[0227] Repeat the process of performing the current action in the current state for each user device within each time slot, determining the obtained reward value and the next state, until the maximum number of training iterations is reached.
[0228] In one implementation, the state model of the user equipment is:
[0229]
[0230] Among them, L t For user equipment locations, RB stands for Radio Resource Block, and N stands for... b This represents the number of radio resource blocks that all channels can provide within a time slot.
[0231] In one implementation, the action model of the user equipment is as follows:
[0232] A i (t)={a1(t),a2(t),...,a N (t)}
[0233] a i (t)={p i,c U i ,s t i,c,m}
[0234] Among them, U i Let s be the set of access points associated with the i-th user equipment. t i,c,m To determine whether the m-th radio resource block of the c-th sub-channel is occupied by the i-th user equipment, p i,c It is the power allocated by the i-th user equipment to the c-th sub-channel.
[0235] In one implementation, the reward model obtained by the user equipment is as follows:
[0236]
[0237] Among them, s t i,c,m To determine whether the m-th radio resource block of the c-th sub-channel is occupied by the i-th user equipment, N b This represents the number of radio resource blocks that all channels can provide within a time slot. For user equipment data within time slot t, The l-th independent sub-part a in the user equipment data l The signal-to-noise ratio, N r This represents the number of radio resource units that all channels can provide within a time slot.
[0238] Alternatively, processor 303, coupled to memory 301, is used to execute programs in memory 301 for:
[0239] Get the current state of the user device;
[0240] The Actor network and Critic network of the user device are obtained, and the Actor network and Critic network are trained according to the model training method described above;
[0241] The current state is input into the Actor network and the Critic network to obtain the output action of the user equipment.
[0242] The uplink resources occupied by the user equipment are determined based on the output action.
[0243] In this application, Figure 9 The diagram only shows some components and does not mean that the terminal device only includes... Figure 9 The components shown.
[0244] The terminal device provided in this embodiment is based on the same inventive concept as the model training method or uplink resource occupation method provided in the embodiments of this application, and has the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.
[0245] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0246] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0247] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0248] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0249] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0250] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0251] This application also provides a computer-readable storage medium corresponding to the model training method or uplink resource occupancy method provided in the foregoing embodiments, wherein a computer program (i.e., a program product) is stored thereon. When the computer program is run by a processor, it executes the model training method provided in any of the foregoing embodiments, or executes the uplink resource occupancy method provided in any of the foregoing embodiments.
[0252] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0253] The computer-readable storage medium provided in the above embodiments of this application and the model training method or uplink resource occupation method provided in the embodiments of this application are based on the same inventive concept and have the same beneficial effects as the methods adopted, run or implemented by the applications stored therein.
[0254] It should be noted that numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known structures and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0255] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0256] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A model training method, characterized in that, include: Modeling of proactive network architecture based on Markov decision-making determines the state model, action model, and reward model of user equipment in the proactive network architecture. The network models constructed from the state model, action model, and reward model of the user device in the active network architecture are trained using a multi-agent reinforcement learning strategy to obtain the trained Actor network and Critic network of the user device. The training of the network model constructed from the state model, action model, and reward model of the user device in the active network architecture using a multi-agent reinforcement learning strategy includes: Construct the Actor network, Critic network, target Actor network, and target Critic network for the user equipment; Initialize the network parameters of the Actor network, Critic network, target Actor network, and target Critic network, as well as the experience replay pool and the maximum number of training iterations. Randomly determine the initial state of each user device; Within each time slot, for each user device, perform an action in the current state, determine the acquired reward value and the next state; The current state, the action, the reward value, and the next state are stored in the experience replay pool; and the current state is updated to the next state. For each user equipment, multiple samples are randomly sampled from the experience replay pool, and the target network estimate for each user equipment is determined based on the samples; Update the Actor network and Critic network of the user equipment based on the target network estimate; After all user equipment's Actor and Critic networks are updated, the target Actor and target Critic networks of the user equipment are soft-updated. Repeat the process of performing the current action in the current state for each user device in each time slot, determining the obtained reward value and the next state, until the maximum number of training iterations is reached; The state model of the user equipment is as follows: Among them, L t For user equipment locations, RB stands for Radio Resource Block, and N stands for... b The number of radio resource blocks that all channels can provide within a time slot; The action model of the user equipment is as follows: A i (t)={a1(t),a2(t),...,a N (t)} a i (t)={p i,c ,U i ,s t i,c,m } Among them, U i Let s be the set of access points associated with the i-th user equipment. t i,c,m To determine whether the m-th radio resource block of the c-th sub-channel is occupied by the i-th user equipment, p i,c It is the power allocated by the i-th user equipment to the c-th sub-channel; The reward model obtained by the user equipment is as follows: Among them, s t i,c,m To determine whether the m-th radio resource block of the c-th sub-channel is occupied by the i-th user equipment, N b This represents the number of radio resource blocks that all channels can provide within a time slot. For user equipment data within time slot t, The l-th independent sub-part a in the user equipment data l The signal-to-noise ratio, N r This represents the number of radio resource units that all channels can provide within a time slot; The Actor network is represented as: m i (S i ;θ i ) Among them, S i Indicates Agent i Observational information, θ i These are the parameters of its Actor network; The Critic network is represented as: q i (S1,S2,...,S N ,A1,A2,...,A N ;ω i ) Where, ω i These are the parameters of its Critic network; The target Actor network is represented as: with me i (S i ;th' i ) Where, θ' i These are the parameters of its Actor network; The target Critic network is represented as: q' i (S1,S2,...,S N ,A1,A2,...,A N ;ω' i ) Where, ω' i These are the parameters of its Critic network.
2. A method for occupying uplink resources, characterized in that, include: Get the current state of the user device; Obtain the Actor network and Critic network of the user device, wherein the Actor network and Critic network are trained according to the model training method described in claim 1; The current state is input into the Actor network and the Critic network to obtain the output action of the user equipment. The uplink resources occupied by the user equipment are determined based on the output action.
3. A model training device, characterized in that, include: The modeling module is used to model the proactive network architecture based on Markov decision-making, and to determine the state model, action model and reward model of user equipment in the proactive network architecture. The training module is used to train the network model constructed by the state model, action model and reward model of the user device in the active network architecture through the multi-agent reinforcement learning strategy, so as to obtain the trained Actor network and Critic network of the user device. The training of the network model constructed from the state model, action model, and reward model of the user device in the active network architecture using a multi-agent reinforcement learning strategy includes: Construct the Actor network, Critic network, target Actor network, and target Critic network for the user equipment; Initialize the network parameters of the Actor network, Critic network, target Actor network, and target Critic network, as well as the experience replay pool and the maximum number of training iterations. Randomly determine the initial state of each user device; Within each time slot, for each user device, perform an action in the current state, determine the acquired reward value and the next state; The current state, the action, the reward value, and the next state are stored in the experience replay pool; and the current state is updated to the next state. For each user equipment, multiple samples are randomly sampled from the experience replay pool, and the target network estimate for each user equipment is determined based on the samples; Update the Actor network and Critic network of the user equipment based on the target network estimate; After all user equipment's Actor and Critic networks are updated, the target Actor and target Critic networks of the user equipment are soft-updated. Repeat the process of performing the current action in the current state for each user device in each time slot, determining the obtained reward value and the next state, until the maximum number of training iterations is reached; The state model of the user equipment is as follows: Among them, L t For user equipment locations, RB stands for Radio Resource Block, and N stands for... b The number of radio resource blocks that all channels can provide within a time slot; The action model of the user equipment is as follows: A i (t)={a1(t),a2(t),...,a N (t)} a i (t)={p i,c ,U i ,s t i,c,m } Among them, U i Let s be the set of access points associated with the i-th user equipment. t i,c,m To determine whether the m-th radio resource block of the c-th sub-channel is occupied by the i-th user equipment, p i,c It is the power allocated by the i-th user equipment to the c-th sub-channel; The reward model obtained by the user equipment is as follows: Among them, s t i,c,m To determine whether the m-th radio resource block of the c-th sub-channel is occupied by the i-th user equipment, N b This represents the number of radio resource blocks that all channels can provide within a time slot. For user equipment data within time slot t, The l-th independent sub-part a in the user equipment data l The signal-to-noise ratio, N r This represents the number of radio resource units that all channels can provide within a time slot; The Actor network is represented as: m i (S i ;θ i ) Among them, S i Indicates Agent i Observational information, θ i These are the parameters of its Actor network; The Critic network is represented as: q i (S1,S2,...,S N ,A1,A2,...,A N ;ω i ) Where, ω i These are the parameters of its Critic network; The target Actor network is represented as: with me i (S i ;th' i ) Where, θ' i These are the parameters of its Actor network; The target Critic network is represented as: q' i (S1,S2,...,S N ,A1,A2,...,A N ;ω' i ) Where, ω' i These are the parameters of its Critic network.
4. An uplink resource occupancy device, characterized in that, include: The status acquisition module is used to acquire the current status of the user equipment; A model acquisition module is used to acquire the Actor network and Critic network of the user device, wherein the Actor network and Critic network are trained according to the model training method described in claim 1; The output acquisition module is used to input the current state into the Actor network and the Critic network to obtain the output action of the user equipment. The resource determination module is used to determine the uplink resources occupied by the user equipment based on the output action.
5. A terminal device, characterized in that, include: Memory and processor; The memory is used to store programs; The processor, coupled to the memory, is used to execute the program for performing the model training method of claim 1, or for performing the uplink resource occupancy method of claim 2.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the model training method of claim 1, or to implement the uplink resource occupancy method of claim 2.