5G network eMBB and URLLC service coexistence scheduling method and system based on federal reinforcement learning

By collaboratively training a global scheduling model on distributed base station nodes, the resource scheduling problem when eMBB and URLLC services coexist is solved, achieving efficient and adaptive resource allocation, improving spectrum efficiency, ensuring low latency and high reliability of URLLC, while protecting data privacy and reducing communication overhead.

CN121968340APending Publication Date: 2026-05-01ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHEJIANG SCI-TECH UNIV
Filing Date
2025-12-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies, when eMBB and URLLC services coexist, suffer from problems such as low spectrum efficiency in resource scheduling, inability to adapt to dynamic changes in services, risk of privacy data leakage, and high communication overhead.

Method used

By employing a federated reinforcement learning approach, a global scheduling model is collaboratively trained on distributed base station nodes. Through Markov decision process modeling and federated reinforcement learning training, efficient, adaptive, and low-overhead dynamic resource scheduling for eMBB and URLLC services is achieved, avoiding raw data aggregation and ensuring data privacy.

Benefits of technology

It achieves high spectrum efficiency, ensures low latency and high reliability for critical services, reduces communication overhead, protects data privacy, and adapts to dynamic changes in services through resource scheduling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121968340A_ABST
    Figure CN121968340A_ABST
Patent Text Reader

Abstract

The invention discloses a 5G network eMBB and URLLC service coexistence scheduling method and system based on federal reinforcement learning. The method comprises the following steps: step 1, modeling a local 5G network eMBB and URLLC service coexistence scheduling problem into a Markov decision process by each base station; 2-1, the central server issues the global model parameters to the base stations participating in training; 2-2, each base station independently trains a model based on local data, and the original data is not out of the local; step 2-3, the base station uploads the updated model parameters to a central server; step 2-4, the central server fuses the parameters of each base station and updates the global model; step 2-5, storing the updated global model parameters in the central server, issuing the updated global model parameters to a new batch of base stations participating in training when the next round of training begins, taking the updated global model parameters as initial parameters of local models of the base stations, and starting the new round of training; and step 3, deploying the trained global model to each base station to realize resource scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication technology and relates to dynamic resource scheduling technology for enhanced mobile broadband (eMBB) and ultra-reliable low-latency communication services (URLLC) in fifth-generation mobile communication networks. Specifically, it is a method and system for coexistence scheduling of eMBB and URLLC services in 5G networks based on federated reinforcement learning. Background Technology

[0002] In core 5G scenarios, eMBB services and URLLC services need to coexist in the same time slot, such as... Figure 1 As shown, this coexistence of business requirements presents a fundamental conflict: eMBB requires long-term, stable broadband resources to ensure throughput, while URLLC demands extremely low transmission latency and extremely high reliability, often acquiring resources instantly through preemption mechanisms. Therefore, how to schedule resources when eMBB and URLLC coexist is a critical technical issue.

[0003] Existing technologies have proposed some solutions to the resource scheduling problem when eMBB and URLLC coexist, but they have some significant shortcomings. For example, static resource allocation strategies cannot adapt to dynamic changes in service flows, resulting in low spectrum efficiency; while centralized intelligent scheduling methods can improve efficiency, the model training process requires uploading data from each base station to a central server, posing risks of data privacy leakage and significant link bandwidth overhead. Therefore, there is an urgent need in this field for an innovative solution that can achieve efficient collaborative scheduling across base stations while protecting privacy and reducing overhead. Summary of the Invention

[0004] To overcome the limitations of existing technologies, this invention provides a method and system for coexistence scheduling of eMBB and URLLC services in 5G networks based on federated reinforcement learning. This invention achieves efficient, adaptive, and low-overhead dynamic resource scheduling of eMBB and URLLC services without aggregating the original data by collaboratively training a global scheduling model on distributed base station nodes.

[0005] The present invention adopts the following technical solution: A method for coexistence scheduling of eMBB and URLLC services in 5G networks based on federated reinforcement learning involves an architecture comprising a central server and multiple distributed base station nodes, with each base station node deploying a local intelligent agent. The specific steps of the method are as follows: Step 1: Distributed problem modeling.

[0006] Each base station models the coexistence scheduling problem of local 5G network eMBB and URLLC services as a Markov Decision Process (MDP). The MDP can be described as a triple consisting of a state space, an action space, and a reward function.

[0007] Step 2: Federated Reinforcement Learning Training. The following is the iterative process: Step 2-1: Global Model Distribution: The central server distributes the global model parameters to the base stations participating in the training.

[0008] Step 2-2: Local Model Training: Each base station independently trains its model based on local data; the original data does not leave the local area. In this step, based on the MDP triples from Step 1, the base station trains its model in state s. t For input, action a t For output, reward R t For feedback, its local model is trained independently.

[0009] Step 2-3: Upload Model Parameters: The base station uploads the updated model parameters to the central server.

[0010] Steps 2-4: Federated Aggregation: The central server merges the parameters of each base station node and updates the global model.

[0011] Steps 2-5: Model Distribution: Updated Global Model Parameters The data is stored on the central server and will be iteratively executed starting from step 2-1 at the beginning of the next round of federated training. It will be distributed to a new batch of base station nodes participating in the training as the initial parameters of their local models, thus starting a new round of collaborative training.

[0012] Step 3: Online distributed scheduling. The trained global model is deployed to each base station to achieve real-time, distributed intelligent resource scheduling.

[0013] Preferably, in step 1: Motion space design: Possible actions a in a mini time slot t t They are collected in the set At={0,1,…,F}, where 0 indicates that no URLLC service packet is being transmitted, and the other data indicates the channel index of the current mini-slot transmitting the URLLC packet.

[0014] State space design: The state of each mini-slot is determined by... express, This represents the state variable of the URLLC service in mini-slot t. This represents the state variable of the eMBB service in mini-slot t. Further, , where Q tQ represents the length of the URLLC queue in the mini-slot t. t ≤Q max , where Q max Indicates the maximum length of the URLLC queue. ,in, Indicates the maximum number of mini slots that the URLLC queue can wait for, l t This represents the number of mini-slots in mini-slot t that represent the current latency of the URLLC service. Note that if... If so, the current mini-slot will immediately send the arriving URLLC data packet. Let S be an F-dimensional vector, representing each variable S in the F spectral channels. t (f), f=1,2,…,F. S t (f) indicates whether the eMBB codeword transmitted on spectrum channel f is interrupted or non-interrupted, S t (f) = -1 indicates an interruption, S t (f)≥0 indicates no interruption. Let p t (w) represents the number of times codeword w is punctured starting from the beginning of the set, and C(w) represents the maximum number of mini-slots allowed to be punctured for codeword w. Then we have... .

[0015] Reward Function Design: The principle of reward function design is to minimize the number of interrupted eMBB codewords while keeping the latency of URLLC packets below a given threshold. For the total set W of codewords in the current time slot t... t The codeword w in the codeword defines the eMBB reward function e. t (w) as follows.

[0016] Define URLLC reward function Here, K0 represents the penalty threshold configured by the system for violating URLLC latency constraints. Therefore, in the mini-slot t, the current state is s. t The current action is 'a', and the next state is 's'. t+1 The current reward function can be expressed as: Among them, R t (s t ,a,s t+1 The following text is abbreviated as R. t .

[0017] This invention also discloses a 5G network eMBB and URLLC service coexistence scheduling system based on federated reinforcement learning, used to execute the above method, which includes the following modules: The distributed problem modeling module is used by each base station to model the coexistence scheduling problem of local 5G network eMBB and URLLC services as a Markov decision process (MDP). The MDP is described as a triple consisting of a state space, an action space, and a reward function. The Federated Reinforcement Learning Training Module includes the following sub-modules: Global Model Distribution Submodule: Used by the central server to distribute global model parameters to the base stations participating in training; The local model training submodule is used for each base station to independently train its model based on local data, ensuring that the original data does not leave the local machine. During the independent training of the local model, the base station uses the triples of the MDP in the distributed problem modeling module to train its model in state s. t For input, action a t For output, reward R t For feedback; Model parameter upload submodule: Used by the base station to upload the updated model parameters to the central server; The federated aggregation submodule is used by the central server to merge parameters from various base stations and update the global model. Model distribution submodule: used for updated global model parameters The data is stored on a central server and distributed to a new batch of base stations participating in the training at the start of the next training round, serving as the initial parameters for their local models and initiating a new round of training. The online distributed scheduling module is used to deploy the trained global model to each base station to achieve real-time, distributed resource scheduling.

[0018] Compared with existing technologies, the resource scheduling method and system of 5G network eMBB and URLLC based on federated reinforcement learning has the following significant technical advantages: (1) High spectrum efficiency: Resources are dynamically allocated on demand through online learning.

[0019] (2) Ensure critical business operations: The model can prioritize meeting the low latency and high reliability requirements of URLLC.

[0020] (3) Privacy protection and security: The original data is stored locally at the base station, which complies with data security regulations.

[0021] (4) Reduce communication overhead: Only model parameters are transmitted instead of raw data, which greatly reduces the burden on the backhaul link. Attached Figure Description

[0022] Figure 1 This is a schematic diagram illustrating the coexistence of eMBB (enhanced mobile broadband) and URLLC (ultra-reliable low-latency communication).

[0023] Figure 2This is a system architecture diagram of the preferred embodiment of the present invention, which relates to the 5G network eMBB and URLLC service coexistence scheduling method based on federated reinforcement learning.

[0024] Figure 3 This is a block diagram of a preferred embodiment of the present invention, which is a 5G network eMBB and URLLC service coexistence scheduling system based on federated reinforcement learning. Detailed Implementation

[0025] To provide a clearer understanding of the technical solution of the present invention, the present invention will now be described in detail with reference to specific implementation examples.

[0026] See Figure 1 This is a schematic diagram of a scenario where eMBB and URLLC coexist.

[0027] like Figure 1 As shown, three eMBB users exist simultaneously in the time-frequency two-dimensional grid, represented by dark green, yellow, and light green respectively, while URLLC services are represented by red. In the figure, the smallest unit in the time domain is a mini-slot, and the smallest unit in the frequency domain is a resource block (RB). It can be seen from the figure that for eMBB User1 and eMBB User2, some mini-slots are perforated by URLLC services, while for eMBB User3, none of the mini-slots are perforated by URLLC services.

[0028] Let the time step of the set be Where T represents the maximum number of mini-slots. The number of channels is denoted as F. At the beginning of each mini-slot, the base station scheduler performs resource allocation for eMBB services, that is, completes the scheduling of each eMBB codeword. Let W... t This represents the set of all eMBB codewords transmitted in a mini-slot t. After eMBB service resource allocation is completed, in each mini-slot, with probability P... u A URLLC packet is generated and placed in the URLLC queue. In each mini-time slot, a URLLC packet is taken from the head of the URLCC queue each time, and the method for allocating resources for coexistence of eMBB and URLLC in this invention determines whether the URLLC packet should be transmitted in the current mini-time slot.

[0029] Throughout the architecture (e.g.) Figure 2 As shown, the model includes a central server and multiple distributed base station nodes. The central server is responsible for maintaining and updating the global scheduling model and coordinating the federated learning process. Each base station node deploys a local agent, which contains a local scheduling model with the same structure as the global model. Each agent makes decisions and is trained based on its local state information.

[0030] This embodiment presents a 5G network eMBB and URLLC service coexistence scheduling method based on federated reinforcement learning. The specific steps are as follows: Step 1: Distributed Problem Modeling Each base station models the coexistence scheduling problem of local 5G network eMBB and URLLC services as a Markov Decision Process (MDP). The MDP is described as a triple consisting of a state space, an action space, and a reward function; where: The motion space design is as follows: possible actions a in mini-time slot t t Collected in set A t In the formula ={0,1,…,F}, 0 indicates that no URLLC service packet is transmitted, and the other values ​​represent the channel index of the current mini-slot transmitting the URLLC packet. The state space design is as follows: the state of each mini-slot is determined by... express, This represents the state variable of the URLLC service in mini-slot t. The state variables represent the eMBB service state in mini-slot t; , where Q t Q represents the length of the URLLC queue in the mini-slot t. t ≤Q max Q max Indicates the maximum length of the URLLC queue; ,in, Indicates the maximum number of mini-slots that the URLLC queue can wait for, l t This represents the number of mini-slots in mini-slot t that indicate the current latency of the URLLC service. If so, the current mini-slot will immediately send the arriving URLLC data packet; Let S be an F-dimensional vector, representing each variable S in the F spectral channels. t (f), f=1,2,…,F; S t (f) indicates whether the eMBB codeword transmitted on spectrum channel f is interrupted, S t (f) = -1 indicates an interruption, S t (f)≥0 indicates non-interruption; let p t Let C(w) represent the number of times codeword w is punctured starting from the beginning of the set, and let C(w) represent the maximum number of mini-slots allowed to be punctured for codeword w. Then we have: The reward function is designed as follows: For the total set W of codewords in the current time slot t... t The codeword w in the codeword defines the eMBB reward function e. t (w) is as follows: Where, p t-1 (w) represents the number of times codeword w has been punched up to the previous time slot; Define URLLC reward function Where K0 represents the penalty threshold configured by the system for violating URLLC latency constraints; therefore, in mini-slot t, the current state is s. t The current action is 'a', and the next state is 's'. t+1 The current reward function is expressed as: .

[0031] Step 2: Federated Reinforcement Learning Training This step employs a federated learning framework based on the proximal policy optimization algorithm, as detailed below: 2-1 Global Model Distribution: The central server distributes the global model parameters to the base stations participating in the training. The local agent is designed as follows: Policy Network (a represents the current action, s represents the current state): Starting with state s... t The input is action a, and the output is action a. t probability distribution Its parameter θ contains the connection weights and bias terms of all layers of the network, which together define the mapping strategy from state to action.

[0032] Value Network : in state s t The input is the state value V(s), and the output is the state value V(s). t Scalar estimation of ). Its parameters It includes the connection weights and biases of all layers in the network, and these parameters are used to learn how to accurately evaluate the long-term expected return of a state.

[0033] Theoretically, the learning objective of a value network is to approximate the ideal state-value function V. π (s), this function is defined as starting from state S and following the policy. π The mathematical expression for the expected cumulative discount return that can be obtained is: Where k represents the k-th future time slot relative to the current time slot t, R t+k This represents the reward function for the k-th future time slot relative to the current time slot t. γ It is the discount factor, a constant between 0 and 1, E π Indicating in strategy π The expected value is taken down. In practice, this ideal value is not calculated directly, but rather through a value network. Learned.

[0034] 2-2. The local training and update process is as follows: 2.1: Data Collection. The agent uses the current policy. Interacting with the environment, in state s t Next, execute action a t Receive reward R t and transition to the next state s t+1 , will the empirical tuple {s t ,a t ,R t ,s t+1 Store it in the local experience replay pool.

[0035] 2.2: Dominance Estimation: A small batch of data is sampled from the pool, and the dominance function is calculated using generalized dominance estimation. Generalized dominance estimation, by incorporating multi-step time-series difference errors, achieves a good balance between bias and variance in dominance estimation. The specific calculation process is as follows: 1) Calculate the timing difference error: For each time step t, first calculate its timing difference error. : First, obtain the value estimate through forward propagation: [The state s is then used for the estimation.] t and s t+1 Input the current value network To obtain the corresponding value estimate and These are value networks based on their current parameters. The calculated numerical value is used to measure the current value of future rewards.

[0036] Then, the timing difference error is calculated: in, γ It is a hyperparameter representing the discount factor, which is a constant between 0 and 1, and is usually set to a number between 0.9 and 0.99.

[0037] 2) Calculate the generalized dominance estimate: Use the calculated series of time-series difference errors The advantage function estimate for time step t is calculated using an exponentially weighted average. : in, λ is the generalized dominance estimation parameter, with a value range of (0,1), used to smoothly control the trade-offs between different time number estimates. T is the maximum time step.

[0038] 2.3: Loss Function Calculation. The loss function of the near-end policy optimization algorithm consists of three parts: Pruning strategy loss: in, , It is a policy network with state s t In the probability distribution of actions after input, the corresponding action 'a' is... t The probability value. This is the probability given by the old policy network when collecting empirical data. The clip function is used to adjust the ratio. t Limited to between. This is a hyperparameter configured in the system, typically set to 0.1. E t This represents the expected value at the current time step t.

[0039] Value function loss: ,in, It is a value network for state s t The estimate. , These are system configuration hyperparameters.

[0040] Entropy reward: Where S is the entropy function: ; in, Is the policy network in state s t Let A be the probability of output for each possible action a, where A is the action space.

[0041] Total loss Among them, c1 and c2 are configurable hyperparameters.

[0042] 2-3. Model parameter update: Minimize using the Adam optimizer This process utilizes gradient descent (see below for minimizing using the Adam optimizer). (Detailed step description), calculating the loss function relative to the parameters The gradients are calculated and these weights and biases are iteratively updated, thereby continuously improving the decision-making ability of the policy network and the evaluation ability of the value network.

[0043] 2-4. The federated aggregation process is as follows: The central server collects the model parameters of each participating base station (the parameters of the i-th base station are represented by the subscript i). This refers to the set of weights and biases of all base station policy networks and value networks, which are then weighted and averaged according to a federated averaging algorithm based on preset weights (such as the amount of local data at each base station) to generate next-generation global model parameters. .

[0044] Through iterative processes described above, a high-performance global scheduling model is finally obtained.

[0045] In steps 2-3, the gradient descent method is specifically as follows: Minimize using the Adam optimizer. The detailed steps are as follows: 1) Initialization. This involves configuring the policy network parameters. and value network parameters Initialize their Adam state variables respectively: Policy network first-order moment vector Second-order moment vector First-order moment vector of value network Second-order moment vector Global time step counter .

[0046] Set Adam hyperparameters: learning rate α, first moment decay rate β 1 (Typically 0.9), second-order moment decay rate β 2 (Typically 0.999), numerical stability constant (Usually 1e-8).

[0047] 2) For each training epoch, execute the following loop: (1) Policy network parameters θ Update: (a) Calculate the gradient: Calculate the total loss Compared to θ gradient ; (b) Update Adam's status: Update time step: ; Update the first-order moment estimate: Update the second-order moment estimate: in, This indicates squaring by elements.

[0048] (c) Deviation correction: (Where t_adm represents the time step) (d) Update parameters: ,in, This indicates taking the square root of each element.

[0049] (2) Value network parameters Update: (a) Calculate the gradient: Calculate the total loss Compared to gradient ;in, The meaning indicates the Find the gradient.

[0050] (b) Update Adam's status: Using the same global time step t adam ; Update the first-order moment estimate: Update the second-order moment estimate: in, Here, squaring is indicated by the element.

[0051] (c) Deviation correction: (d) Update parameters: ,in, This indicates taking the square root of each element.

[0052] 3) Repeat step 2) until the preset number of local training epochs K is completed. In steps 2-4, the federated aggregation process is completed using the federated averaging algorithm, and its specific steps are as follows: 1) The central server collects the local model parameters uploaded by all N base stations participating in this round of training, including: policy network parameters. Value network parameters .

[0053] 2) Weight Calculation: Let n be the number of data samples available for training in the local experience replay pool of base station i. i The weight assigned to base station i is .

[0054] 3) Model aggregation: The central server uses the calculated weights to perform a weighted average of the policy network parameters and the value network parameters to generate a new generation of global model parameters.

[0055] Global policy network parameters updated to Global value network parameters updated to The above operations are performed element-wise.

[0056] 2-5. Model Distribution: Updated global model parameters The data is stored on a central server and will be distributed to a new batch of base station nodes participating in the training when the next round of federated training begins, serving as the initial parameters for their local models and initiating a new round of collaborative training.

[0057] Step 3, Online Distributed Scheduling: Deploy the trained global model to each base station to achieve real-time, distributed intelligent resource scheduling.

[0058] like Figure 3 As shown, this embodiment discloses a 5G network eMBB and URLLC service coexistence scheduling system based on federated reinforcement learning, used to execute the above method, which includes the following modules: The distributed problem modeling module is used by each base station to model the coexistence scheduling problem of local 5G network eMBB and URLLC services as a Markov decision process (MDP). The MDP is described as a triple consisting of a state space, an action space, and a reward function. The Federated Reinforcement Learning Training Module includes the following sub-modules: Global Model Distribution Submodule: Used by the central server to distribute global model parameters to the base stations participating in training; The local model training submodule is used for each base station to independently train its model based on local data, ensuring that the original data does not leave the local machine. During the independent training of the local model, the base station uses the triples of the MDP in the distributed problem modeling module to train its model in state s. t For input, action a t For output, reward R t For feedback; Model parameter upload submodule: Used by the base station to upload the updated model parameters to the central server; The federated aggregation submodule is used by the central server to merge parameters from various base stations and update the global model. Model distribution submodule: used for updated global model parameters The data is stored on a central server and distributed to a new batch of base stations participating in the training at the start of the next training round, serving as the initial parameters for their local models and initiating a new round of training. The online distributed scheduling module is used to deploy the trained global model to each base station to achieve real-time, distributed resource scheduling.

[0059] Other aspects of this embodiment can be found in the above method embodiments.

[0060] The above description is merely a detailed explanation of preferred embodiments and principles of the present invention. For those skilled in the art, there may be changes in specific implementation methods based on the ideas provided by the present invention, and these changes should also be considered within the scope of protection of the present invention.

Claims

1. A scheduling method for coexistence of eMBB and URLLC services in 5G networks based on federated reinforcement learning, characterized by: The specific steps are as follows: Step 1: Distributed Problem Modeling Each base station models the coexistence scheduling problem of local 5G network eMBB and URLLC services as a Markov decision process (MDP). The MDP is described as a triple consisting of a state space, an action space, and a reward function. Step 2: Federated Reinforcement Learning Training Step 2-1, Global Model Distribution: The central server distributes the global model parameters to the base stations participating in the training. Step 2-2, Local Model Training: Each base station independently trains its model based on local data; the original data does not leave the local machine. During the independent training of the local model, based on the triples of the MDP in Step 1, the base station uses state s... t For input, action a t For output, reward R t For feedback; Steps 2-3: Upload Model Parameters: The base station uploads the updated model parameters to the central server; Steps 2-4: Federated Aggregation: The central server merges the parameters of each base station and updates the global model; Steps 2-5: Model Distribution: Updated global model parameters The data is stored on a central server and distributed to a new batch of base stations participating in the training at the start of the next training round, serving as the initial parameters for their local models and initiating a new round of training. Step 3: Online Distributed Scheduling The trained global model is deployed to each base station to achieve real-time, distributed resource scheduling.

2. The 5G network eMBB and URLLC service coexistence scheduling method based on federated reinforcement learning as described in claim 1, characterized in that, In step 1, the action space design is as follows: possible actions a in the mini-time slot t t Collected in set A t In the formula ={0,1,…,F}, 0 indicates that no URLLC service packet is transmitted, and the other values ​​represent the channel index of the current mini-slot transmitting the URLLC packet. The state space design is as follows: the state of each mini-slot is determined by... express, This represents the state variable of the URLLC service in mini-slot t. The state variables represent the eMBB service state in mini-slot t; , where Q t Q represents the length of the URLLC queue in the mini-slot t. t ≤Q max Q max Indicates the maximum length of the URLLC queue; ,in, Indicates the maximum number of mini-slots that the URLLC queue can wait for, l t This represents the number of mini-slots in mini-slot t that indicate the current latency of the URLLC service. If so, the current mini-slot will immediately send the arriving URLLC data packet; Let S be an F-dimensional vector, representing each variable S in the F spectral channels. t (f), f=1,2,…,F; S t (f) indicates whether the eMBB codeword transmitted on spectrum channel f is interrupted, S t (f) = -1 indicates an interruption, S t (f)≥0 indicates non-interruption; let p t Let C(w) represent the number of times codeword w is punctured starting from the beginning of the set, and let C(w) represent the maximum number of mini-slots allowed to be punctured for codeword w. Then we have: The reward function is designed as follows: For the total set W of codewords in the current time slot t... t The codeword w in the codeword defines the eMBB reward function e. t (w) is as follows: Where, p t-1 (w) indicates the number of times codeword w has been punched up to the previous time slot; Define URLLC reward function Where K0 represents the penalty threshold configured by the system for violating URLLC latency constraints; therefore, in mini-slot t, the current state is s. t The current action is 'a', and the next state is 's'. t+1 The current reward function is expressed as: 。 3. The 5G network eMBB and URLLC service coexistence scheduling method based on federated reinforcement learning as described in claim 2, characterized in that, In step 2-1, the local agent in the base station is designed as follows: Policy Network : in state s t The input is action a, and the output is action a. t probability distribution ,parameter θ It includes the connection weights and bias terms of all layers in the network; where a is the current action and s is the current state; Value Network : in state s t The input is the state value V(s), and the output is the state value V(s). t Scalar estimation of ); parameters Includes the connection weights and bias terms of all layers in the network; The learning objective of a value network is to approximate the ideal state-value function V. π (s), this function is defined as starting from state S and following the policy. π The expected cumulative discount return that can be obtained is expressed mathematically as follows: Where k represents the k-th future time slot relative to the current time slot t, R t+k This represents the reward function for the k-th future time slot relative to the current time slot t. γ It is the discount factor, E π Indicating in strategy π Take the mathematical expectation below.

4. The 5G network eMBB and URLLC service coexistence scheduling method based on federated reinforcement learning as described in claim 3, characterized in that, Step 2-2 is as follows: 2-2.1 The agent uses the current policy Interacting with the environment, in state s t Next, execute action a t Receive reward R t and transition to the next state s t+1 , will the empirical tuple {s t ,a t ,R t ,s t+1 Stored in the local experience replay pool; 2-2.

2. Sample small batches of data from the local experience replay pool and calculate the dominance function using generalized dominance estimation. The specific calculation process is as follows: 1) Calculate the timing difference error: For each time step t, first calculate the timing difference error. : First, obtain the value estimate through forward propagation: [The state s is then used for the estimation.] t and s t+1 Input the current value network To obtain the corresponding value estimate and ; Then, the timing difference error is calculated: ; in, γ Indicates the discount factor; 2) Calculate the generalized dominance estimate: using the obtained series of time-series difference errors The advantage function estimate for time step t is calculated using an exponentially weighted average. : in, λ These are the generalized advantage estimation parameters, and T is the maximum time step. 2-2.3 The loss function of the near-end policy optimization algorithm consists of three parts: Pruning strategy loss: in, , It is a policy network with state s t In the probability distribution of actions after input, the corresponding action 'a' is... t The probability value; The ratio is the probability given by the old policy network when collecting empirical data; the clip function is used to adjust the ratio. t Limited to between, It's a hyperparameter, E t This indicates the expected value at the current time step t; Value function loss: in, It is a value network for state s t The estimate, V target = , It's a hyperparameter; Entropy reward: Where S is the entropy function: ; Is the policy network in state s t Let A be the probability of output for each possible action 'a', where A is the action space. Total loss: Here, c1 and c2 are the configured hyperparameters.

5. The 5G network eMBB and URLLC service coexistence scheduling method based on federated reinforcement learning as described in claim 4, characterized in that, In steps 2-3, the Adam optimizer is used to minimize... Calculate the loss function relative to the parameters The gradient is calculated, and the weights and biases are updated iteratively.

6. The 5G network eMBB and URLLC service coexistence scheduling method based on federated reinforcement learning as described in claim 5, characterized in that, In steps 2-3, the Adam optimizer is used to minimize... This is implemented using the gradient descent method, as follows: 1) Policy network parameters and value network parameters Initialize their Adam state variables respectively: Policy network first-order moment vector Second-order moment vector First-order moment vector of value network Second-order moment vector Global time step counter ; Set Adam hyperparameters: learning rate α, first moment decay rate β 1 Second-order moment decay rate β 2 Numerical stability constant ; 2) For each training epoch, execute the following loop: (1) Policy network parameters θ Update: a. Calculate the gradient: Calculate the total loss Compared to θ gradient ; b. Update Adam's status: Update time step: ; Update the first-order moment estimate: ; Update the second-order moment estimate: ; in, This indicates squaring by elements; c. Deviation correction: Where t_adm represents the time step; d. Update parameters: ; (2) Value network parameters Update: a. Calculate the gradient: Calculate the total loss Compared to gradient ;in, The meaning indicates the Find the gradient; b. Update Adam's status: Using the same global time step t adam ; Update the first-order moment estimate: ; Update the second-order moment estimate: ; c. Deviation correction: d. Update parameters: ; 3) Repeat step 2) until the preset number of local training epochs K is completed.

7. The 5G network eMBB and URLLC service coexistence scheduling method based on federated reinforcement learning as described in claim 6, characterized in that, In steps 2-4, the central server collects the model parameters from each participating base station. Then, based on the federated averaging algorithm and according to preset weights, a new generation of global model parameters is generated. .

8. The 5G network eMBB and URLLC service coexistence scheduling method based on federated reinforcement learning as described in claim 7, characterized in that, In steps 2-4, the federated aggregation process is completed using the federated averaging algorithm, as detailed below: 1) The central server collects the local model parameters uploaded by all N base stations participating in this round of training, including: policy network parameters. Value network parameters ; 2) Let n be the number of data samples used for training in the local experience replay pool of base station i. i The weight assigned to base station i is ; 3) The central server uses the calculated weights to perform a weighted average of the policy network parameters and the value network parameters, respectively, to generate a new generation of global model parameters; Global policy network parameters updated to ; Global value network parameters updated to .

9. A 5G network eMBB and URLLC service coexistence scheduling system based on federated reinforcement learning, used to perform the method as described in any one of claims 1-8, characterized in that, Includes the following modules: The distributed problem modeling module is used by each base station to model the coexistence scheduling problem of local 5G network eMBB and URLLC services as a Markov decision process (MDP). The MDP is described as a triple consisting of a state space, an action space, and a reward function. The Federated Reinforcement Learning Training Module includes the following sub-modules: Global Model Distribution Submodule: Used by the central server to distribute global model parameters to the base stations participating in training; Local model training submodule: used for each base station to independently train the model based on local data, without leaving the local area; During the independent training of the local model, based on the triples of the MDP in the distributed problem modeling module, the base station uses state s t For input, action a t For output, reward R t For feedback; Model parameter upload submodule: Used by the base station to upload the updated model parameters to the central server; The federated aggregation submodule is used by the central server to merge parameters from various base stations and update the global model. Model distribution submodule: used for updated global model parameters The data is stored on a central server and distributed to a new batch of base stations participating in the training at the start of the next training round, serving as the initial parameters for their local models and initiating a new round of training. The online distributed scheduling module is used to deploy the trained global model to each base station to achieve real-time, distributed resource scheduling.