Network resource scheduling method, device and storage medium based on near-end policy optimization

By adopting a 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm, the problem of low data transmission efficiency in the converged network of 5G network and time-sensitive network is solved, realizing efficient resource scheduling and data transmission, and improving data reliability and real-time performance in industrial scenarios.

CN121001183BActive Publication Date: 2026-03-06HUNAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511539795.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-03-06
Estimated Expiration
2045-10-27

AI Technical Summary

Technical Problem

Existing 5G networks and time-sensitive networks that are integrated with each other have difficulty efficiently scheduling data transmission for different services, resulting in low real-time performance, reliability, and resource utilization of data transmission.

Method used

A 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm is adopted. By acquiring underlying network information, establishing a 5G wireless channel model, constructing a Markov decision process model, building a policy network and an evaluation network, optimizing resource allocation strategies, and realizing dynamic resource scheduling.

Benefits of technology

When the gating of a time-sensitive network is enabled, priority is given to scheduling ultra-low latency services such as robot control commands. When the gating is disabled, resources are intelligently reused to transmit non-real-time services such as video surveillance, thereby improving the utilization of wireless resources, ensuring reliable data transmission and real-time performance, and balancing the real-time performance, reliability, and resource utilization of data transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121001183B_ABST
    Figure CN121001183B_ABST
Patent Text Reader

Abstract

This invention discloses a 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm. By establishing a 5G wireless channel model and constructing a Markov decision process model, the near-end policy optimization algorithm is used to output action values ​​and optimize the policy, thereby completing 5G-TSN network resource scheduling. The network resource scheduling method provided by this invention implements a dynamic resource allocation strategy based on gating state awareness. It can prioritize the scheduling of ultra-low latency services such as robot control commands when the gating of a time-sensitive network is open, and intelligently reuse resources to transmit non-real-time services such as video surveillance when the gating is closed. This effectively improves wireless resource utilization while ensuring that high-real-time data can be transmitted as quickly as possible, ensuring reliable data transmission in industrial scenarios. It integrates the scheduling mechanisms of 5G wireless networks and time-sensitive networks to efficiently schedule resources, balancing the real-time performance, reliability, and resource utilization of data transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of network resource scheduling methods, specifically relating to a 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm. Background Technology

[0002] With the development of the intelligent manufacturing industry, industrial networks are placing higher demands on the real-time performance, reliability, and flexibility of communication. In existing industrial networks, Time-Sensitive Networking (TSN) can provide high-precision time synchronization and low latency guarantees, but its limited bandwidth and coverage make it difficult to meet the needs of large-scale industrial scenarios. 5G networks, with their high bandwidth, low latency, and wide connectivity, offer a new solution for industrial networks. However, industrial control systems have high requirements for network stability, and the traditional wireless channel characteristics of 5G networks are not conducive to deterministic transmission of industrial control services. Therefore, existing technologies integrate 5G networks with Time-Sensitive Networking, combining the time-sensitive characteristics of Time-Sensitive Networking with the high-performance wireless transmission capabilities of 5G networks to achieve more efficient and reliable communication in fields such as industrial automation, remote control, and intelligent manufacturing.

[0003] In a converged network of 5G and time-sensitive networks, the network needs to carry the transmission of various service data, including time-sensitive services and traditional 5G services. When different service data arrive at the same time, the existing network system has difficulty in efficiently scheduling resources and balancing the real-time performance, reliability and resource utilization of data transmission. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the shortcomings of the existing technology, which makes it difficult to effectively schedule the data transmission of different services in the integrated network of 5G network and time-sensitive network, resulting in low real-time performance, reliability and resource utilization of data transmission. Therefore, the present invention provides a 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm.

[0005] This invention discloses a 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm, comprising the following steps:

[0006] Obtain underlying network information, including DW-TT gating status, queue length of base station users, head-of-queue waiting delay, and channel quality of the 5G system;

[0007] A 5G wireless channel model is established based on the underlying network information, including signal receiving power, number of available resource blocks, and resource block capacity.

[0008] A Markov decision process model is constructed, defining a state space, an action space, and a reward function. The state space includes the underlying topology characteristics of the 5G-TSN network and the characteristics of the deployed service function chain. Actions include resource allocation. The reward function includes the data flow scheduling timeliness value. The state space is established based on the 5G wireless channel model.

[0009] A policy network and an evaluation network are built based on a proximal policy optimization algorithm. The policy network is used to output the mean action value, and the evaluation network is used to output the state value.

[0010] The network is evaluated based on its state output value; the policy network is optimized based on the state value; the policy network is evaluated based on the average action value output by the state output value, and action values ​​are obtained by sampling; 5G-TSN network resource scheduling is implemented based on the action values.

[0011] Furthermore, resource and deadline constraints are considered during data stream scheduling, including:

[0012] ; ; ;

[0013] in, Indicates whether to allocate resources to support data stream scheduling. This represents the amount of resources required for a data stream to be scheduled. Represents the total number of channel resources. Indicates the scheduling time of the time-sensitive stream. For the scheduling time of the video stream, and These represent the allowed cutoff time for the time-sensitive stream and the cutoff time for the video stream, respectively.

[0014] Furthermore, resource block allocation constraints are considered during data flow scheduling, including:

[0015] ;

[0016] in, This indicates the number of resource blocks allocated to the data stream. This indicates the number of resource blocks required for the data stream.

[0017] Furthermore, the signal receiving power is expressed as:

[0018] ;

[0019] in, PL(d) represents the base station's transmitted signal power, and PL(d) represents the path loss.

[0020] ;

[0021] Where d represents the distance between the base station and the terminal, and N represents the distance power loss coefficient. denoted by floor penetration loss factor, and f represents the carrier frequency of the signal.

[0022] Furthermore, the number of available resource blocks is expressed as follows:

[0023] ;

[0024] in, Indicates the number of available resource blocks. Indicates the total bandwidth of the channel. Indicates the guard band bandwidth on one side. This represents the resource block bandwidth of a single resource block.

[0025] The resource block capacity of a single resource block is expressed as:

[0026] ;

[0027] in, Indicates bit error rate, Let r represent the inverse function of the standard normal distribution, and r represent the block length of the resource block. Indicates the signal-to-noise ratio;

[0028] in, .

[0029] Furthermore, the reward function of the Markov decision process model is expressed as:

[0030] ;

[0031] in, Indicates the scheduling time of the time-sensitive stream. For the scheduling time of the video stream, and These represent the allowed cutoff time for the time-sensitive stream and the cutoff time for the video stream, respectively.

[0032] Furthermore, the policy network includes three sequentially connected fully connected layers, which are connected by a ReLU activation function. The last fully connected layer serves as an output layer, used to output the mean action value. A Softmax function is used to calculate the probability distribution of the output action based on the mean output action value. The evaluation network includes three sequentially connected fully connected layers, each of which includes 256 neurons and is used to output the output state value of the current state.

[0033] Furthermore, it also includes: inputting the state into the policy network to obtain the average action value; sampling to obtain action values ​​based on the average action value; executing the action values ​​to obtain reward values; when the number of reward values ​​obtained is greater than a preset threshold, updating the policy network and the evaluation network based on the reward values, and clearing the reward values.

[0034] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the above-described 5G-TSN network resource scheduling method.

[0035] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described 5G-TSN network resource scheduling method.

[0036] Beneficial Effects: This invention discloses a 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm. By establishing a 5G wireless channel model, it constructs the state space, action space, and reward function of a Markov decision process model. Then, through a near-end policy optimization algorithm, it builds a policy network and an evaluation network to achieve action value output and policy optimization, thereby completing 5G-TSN network resource scheduling. The network resource scheduling method provided by this invention implements a dynamic resource allocation strategy based on gating state awareness. It can prioritize the scheduling of ultra-low latency services such as robot control commands when the gating of a time-sensitive network is open, and intelligently reuse resources to transmit non-real-time services such as video surveillance when the gating is closed. This effectively improves wireless resource utilization while ensuring that high-real-time data can be transmitted as quickly as possible, ensuring reliable data transmission in industrial scenarios. It integrates the scheduling mechanisms of 5G wireless networks and time-sensitive networks to efficiently schedule resources, balancing the real-time performance, reliability, and resource utilization of data transmission. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 This is a flowchart illustrating the steps of the method of the present invention. Detailed Implementation

[0039] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the specific embodiments of this application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of this application. Therefore, this application is not limited to the specific embodiments disclosed below.

[0040] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0041] Reference Figure 1 As shown, this invention discloses a 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm, comprising the following steps:

[0042] Step S1: Obtain underlying network information, including DW-TT gating status, queue length of base station users, head-of-queue waiting delay, and channel quality of the 5G system; establish a 5G wireless channel model based on the underlying network information, including signal received power, number of available resource blocks, and resource block capacity;

[0043] Step S2: Construct a Markov decision process model, defining the state space, action space, and reward function; the state space includes the underlying topology characteristics of the 5G-TSN network and the characteristics of the deployed service function chain; actions include resource allocation; the reward function includes the data flow scheduling on-time value; the state space is established based on the 5G wireless channel model.

[0044] Step S3: Construct a policy network and an evaluation network based on the proximal policy optimization algorithm. The policy network is used to output the mean action value, and the evaluation network is used to output the state value.

[0045] Step S4: Evaluate the network's state-based output state value; optimize the policy network based on the state value; obtain action values ​​by sampling based on the average action value output by the policy network based on the state value; and implement 5G-TSN network resource scheduling based on the action values.

[0046] Specifically, in step S1: The underlying network information is acquired, including the DW-TT gating state, the queue length of base station users, the head-of-queue waiting delay, and the channel quality of the 5G system. In this embodiment, the underlying network information is data reflecting the network operating status acquired from the network system. Deep reinforcement learning improves the scheduling performance of the 5G-TSN network system by continuously observing the underlying state and optimizing decision-making actions.

[0047] The underlying network information is represented as follows:

[0048] ;

[0049] in, This indicates the DW-TT gating state, where l represents the queue length of the base station user, and d represents the head-of-queue waiting delay. This indicates the channel quality of the 5G system.

[0050] In this embodiment, the gating state is used to control the scheduling timing of time-sensitive streams. When the gating state is on, scheduling of time-sensitive streams is allowed; when the gating state is off, scheduling of non-time-sensitive streams (such as video streams) is allowed.

[0051] Queue length is used to reflect the current network load. Longer queues may cause data flow timeouts, while shorter queues may lead to resource waste.

[0052] The head-of-line waiting delay includes the head-of-line waiting delay of time-sensitive queues and video queues. During data stream scheduling, data needs to be scheduled within a certain time. The state space records the waiting delay to determine whether the data stream is scheduled on time.

[0053] Channel quality directly affects the carrying capacity of resource blocks. In data flow scheduling, channel quality determines whether resources are sufficient, which is reflected in the number of resource blocks and the data carrying capacity of a single resource block.

[0054] A 5G wireless channel model is established based on the underlying network information, including signal reception power, number of available resource blocks, and resource block capacity.

[0055] In this embodiment, the signal receiving power is expressed as:

[0056] ;

[0057] in, PL(d) represents the base station's transmitted signal power, and PL(d) represents the path loss (unit: dB).

[0058] ;

[0059] Where d represents the distance between the base station and the terminal (in meters), and N represents the distance power loss coefficient. denoted by floor penetration loss factor, and f represents the carrier frequency of the signal.

[0060] In this embodiment, the physical layer channel bandwidth of the 5G network considers the guard band and resource blocks (RBs), and the number of available resource blocks is expressed as follows:

[0061] ;

[0062] in, Indicates the number of available resource blocks. Indicates the total bandwidth of the channel. Indicates the guard band bandwidth on one side. This represents the resource block bandwidth of a single resource block; in this embodiment, the total channel bandwidth of the 5G network is 50MHz, the guard bandwidth is 692.5kHz, and the bandwidth of a single resource block is 180kHz.

[0063] The resource block capacity of a single resource block is expressed as:

[0064] ;

[0065] in, Indicates bit error rate, Let r represent the inverse function of the standard normal distribution, and r represent the block length of the resource block. The signal-to-noise ratio (SNR) is expressed as follows: In this embodiment, the channel noise power is fixed at -90 dBm, and the channel signal-to-noise ratio (SNR) is expressed as a linear value. , Indicates noise power.

[0066] in, .

[0067] In this embodiment, within each time slice (TTI), the capacity of a single resource block and the number of available resource blocks are calculated in real time and used as variables for resource scheduling input into the state space, thereby optimizing resource allocation and latency control for time-sensitive network services.

[0068] Specifically, in step S2, a Markov decision process model is constructed, defining a state space, an action space, and a reward function; the state space includes the underlying topology characteristics of the 5G-TSN network and the characteristics of the deployed service function chain; the actions include resource allocation; and the reward function includes the data flow scheduling timeliness value.

[0069] When scheduling data streams, resource and deadline constraints are considered to ensure that data can complete scheduling tasks within the deadline, including:

[0070] ; ; ;

[0071] in, Indicates whether to allocate resources to support data stream scheduling. This represents the amount of resources required for a data stream to be scheduled. Represents the total number of channel resources. Indicates the scheduling time of the time-sensitive stream. For the scheduling time of the video stream, and These represent the allowed cutoff time for the time-sensitive stream and the cutoff time for the video stream, respectively.

[0072] When scheduling data streams, resource block allocation constraints are considered to fully utilize 5G network channel resources, including:

[0073] ;

[0074] in, This indicates the number of resource blocks allocated to the data stream. This indicates the number of resource blocks required for the data stream.

[0075] In this embodiment, a network resource scheduling optimization strategy is implemented through deep reinforcement learning. Deep reinforcement learning learns the optimal action strategy through the interaction between an agent and the environment. At each moment, the agent selects an action based on the current environmental state and receives a reward signal from the environment as feedback. Through continuous trial and error and feedback, the agent can gradually optimize its strategy to maximize the long-term accumulated reward.

[0076] This embodiment establishes a Markov decision process model for deep reinforcement learning algorithms. Based on the deep reinforcement learning process, the state space, action space, and reward function are designed respectively, thereby supporting the agent's efficient learning and decision-making in the 5G-TSN network environment.

[0077] Specifically, the state space defines the set of all states that an agent may encounter in an environment. Each state is a comprehensive description of the current state of the environment, including various observations, environmental information, and features. In this embodiment, the state space includes the underlying topology features of the 5G-TSN network (including channel quality and queue status), as well as relevant features of the service function chain to be deployed (including data stream type and latency requirements). State information provides the agent with the key inputs needed for decision-making.

[0078] The action space defines the set of all actions an agent can execute at each time step. In this embodiment, the action space includes resource allocation and scheduling priority adjustment operations. The agent influences the dynamic changes of the environment by selecting actions. The action space refers to the set of all possible actions an agent can execute in a given environment. In this embodiment, the decision to schedule the current data stream is determined by outputting action values. Action values ​​are mapped, and the number of resource blocks is allocated to each data stream based on the actual number of resource blocks in the current channel.

[0079] The reward function is a function that evaluates and provides feedback on the behavioral outcomes of an agent in the environment. It defines the reward value obtained by the agent for taking different actions in different states, and is used to guide the agent to choose better behavioral strategies during the learning process. In this embodiment, to achieve the scheduling of the data stream in the shortest possible time with limited resources, a reward function is defined... This is the reward value for the current data stream being scheduled on time. Because video streams have a large amount of data, they are easily selectively scheduled later during the scheduling process, and time-sensitive streams are given priority. Therefore, a penalty is imposed if the data stream fails to complete the scheduling within the deadline.

[0080] Specifically, the reward function of the Markov decision process model is expressed as:

[0081] ;

[0082] in, Indicates the scheduling time of the time-sensitive stream. For the scheduling time of the video stream, and These represent the allowed cutoff time for the time-sensitive stream and the cutoff time for the video stream, respectively.

[0083] Specifically, in step S53: a policy network and an evaluation network are built based on the Proximal Policy Optimization (PPO) algorithm. The policy network is used to output the mean action value, and the evaluation network is used to output the state value.

[0084] In each round, first utilize the existing strategy. Interacting with the environment Each time step, obtain Group data:

[0085] ;

[0086] in, Indicates the first The state of each time slot Indicates the first Actions in a time slot.

[0087] The near-end policy optimization algorithm outputs the mean of actions, ensuring that the action values ​​follow a Gaussian distribution. The standard deviation of the Gaussian distribution decays with the network update frequency. An action is randomly sampled. , This indicates selecting an action from action values ​​that follow a Gaussian distribution. The probability of [the action]. After the action is performed, the environment state is updated to [the desired state]. And generate corresponding rewards. The network parameters will not be updated during this process. When a preset number of... After the data, according to the discount rate Calculate the expected reward and advantage estimate for each time slot in this dataset:

[0088] ;

[0089] ;

[0090] in This indicates the results obtained through network evaluation. The value of a state.

[0091] In this embodiment, the objective function of the near-end policy optimization algorithm is expressed as:

[0092] ;

[0093] in, The entropy of the policy network is represented by... This represents the policy gradient objective function. This represents the objective function for evaluating the network. and This represents a constant coefficient used to adjust the weights of various parts in the network's objective function.

[0094] The policy gradient objective function is expressed as:

[0095] ;

[0096] ;

[0097] The objective function for evaluating the network is expressed as:

[0098] ;

[0099] in, This represents the ratio of the old to the new strategy. The clip function represents the cutoff constant. Set upper and lower bound constraints, The value is limited to [1- ,1+ This range reduces the policy update magnitude and prevents over-updating. Finally, by maximizing... Update network parameters Using the collected Group data on network parameters Continuous updates After that, the parameters Updated to .

[0100] In this embodiment, the near-end policy optimization algorithm is used to solve the 5G-TSN air interface resource scheduling optimization problem. It adopts a network structure composed of a policy network and an evaluation network, and comprehensively considers the characteristics of data streams and channel model features. It models the air interface resource scheduling optimization of video streams and time-sensitive streams as a reinforcement learning process, effectively combining the algorithm with the task.

[0101] Specifically, the policy network comprises three sequentially connected fully connected layers, which are connected by a ReLU activation function. The last fully connected layer serves as the output layer, used to output the mean action value. A Softmax function is used to calculate the probability distribution of the output action based on the mean output action value. The evaluation network comprises three sequentially connected fully connected layers, each of which comprises 256 neurons and is used to output the output state value of the current state.

[0102] Specifically, in step S4: the network outputs state value based on the state; the policy network is optimized based on the state value; the policy network outputs the average action value based on the state value, and obtains the action value by sampling; and 5G-TSN network resource scheduling is implemented based on the action value.

[0103] In this embodiment, the method further includes inputting the state into the policy network to obtain the average action value, sampling to obtain action values ​​based on the average action value, executing the action values ​​to obtain reward values, and when the number of reward values ​​obtained is greater than a preset threshold, updating the policy network and the evaluation network based on the reward values, and clearing the reward values.

[0104] In this embodiment, when scheduling video streams and time-sensitive streams using air interface resources, the algorithm first sets the current state s... k The input is fed into the policy network, and the action value 'a' is sampled from the mean of the output action values. k , will the action value a k Execute in the environment to obtain a reward value r k Enter the next state s k+1The process involves collecting and storing data, then continuing the cycle. During this process, network hyperparameters are not updated. After obtaining a certain batch of K sets of data, the network is continuously iterated and updated. After the update is complete, the previously stored data is cleared. When the current parameter update rounds reach the preset maximum value, and all state requirements are within the preset index range during the iteration cycle, the scheduling optimization of 5G-TSN air interface resources for data flow is completed.

[0105] This embodiment also provides a computer device, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described 5G-TSN network resource scheduling method.

[0106] This embodiment also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described 5G-TSN network resource scheduling method.

[0107] This embodiment provides a 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm. By establishing a 5G wireless channel model, it constructs the state space, action space, and reward function of a Markov decision process model. Then, it uses a near-end policy optimization algorithm to build a policy network and an evaluation network to output action values ​​and optimize policies, thereby completing 5G-TSN network resource scheduling. The network resource scheduling method provided by this invention implements a dynamic resource allocation strategy based on gating state awareness. It can prioritize scheduling ultra-low latency services such as robot control commands when the gating of a time-sensitive network is open, and intelligently reuse resources to transmit non-real-time services such as video surveillance when the gating is closed. This effectively improves wireless resource utilization while ensuring that high-real-time data can be transmitted as quickly as possible, ensuring reliable data transmission in industrial scenarios. It integrates the scheduling mechanisms of 5G wireless networks and time-sensitive networks to efficiently schedule resources, balancing the real-time performance, reliability, and resource utilization of data transmission.

[0108] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0109] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A 5G-TSN network resource scheduling method based on a proximal policy optimization algorithm, characterized in that, The method comprises the following steps: obtaining underlying network information, wherein the underlying network information comprises a DW-TT gating state, a queue length of a base station user, a head-of-line waiting delay, and a channel quality of a 5G system; the DW-TT gating state is used to control a scheduling time of a time-sensitive flow; The gating state is used to control the scheduling time of the time-sensitive flow, and when the gating state is open, the time-sensitive flow is allowed to be scheduled; When the gating state is closed, only the non-time-sensitive flow, including a video flow, is allowed to be scheduled; the head-of-line waiting delay comprises a head-of-line waiting delay of a time-sensitive queue and a video queue; a 5G wireless channel model is established based on the underlying network information, including signal receiving power, a number of available resource blocks, and resource block capacity; a Markov decision process model is constructed, and a state space, an action space, and a reward function are defined; the state space comprises a bottom topology structure feature of the 5G-TSN network and a feature of a deployed service function chain; the action comprises a number of resource blocks for data flow allocation; in each time slice, a single resource block capacity and a number of available resource blocks are calculated in real time and taken as variable inputs of the state space for resource scheduling; the reward function comprises a data flow scheduling punctuality value; when a scheduling time of the time-sensitive flow is less than or equal to an allowed deadline scheduling time of the time-sensitive flow, and a scheduling time of the video flow is less than or equal to a deadline scheduling time of the video flow, the reward function is 1; when the scheduling time of the time-sensitive flow is greater than the allowed deadline scheduling time of the time-sensitive flow, and the scheduling time of the video flow is greater than the deadline scheduling time of the video flow, the reward function is -1; the state space is established based on the 5G wireless channel model; a policy network and an evaluation network are built based on a proximal policy optimization algorithm; the policy network is used to output an action value mean, and the evaluation network is used to output a state value; the evaluation network outputs the state value based on the state; the policy network is optimized based on the state value; the policy network outputs the action value mean based on the state, and an action value is obtained by sampling; the 5G-TSN network resource scheduling is realized based on the action value.

2. The 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm according to claim 1, characterized in that, Constraints of resources and deadlines are considered in data flow scheduling, including: ; ; ; wherein, denotes whether resources are allocated to support the data flow to complete the scheduling, denotes the amount of resources needed by the data flow that can be scheduled, denotes the total number of resources of the channel, denotes the scheduling time of the time-sensitive flow, denotes the scheduling time of the video flow, and denote the allowed deadline scheduling time of the time-sensitive flow and the deadline scheduling time of the video flow, respectively.

3. The 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm according to claim 1, characterized in that, Constraints of resource block quantity allocation are considered in data flow scheduling, including: ; wherein, represents the number of resource blocks allocated for the data stream, represents the number of resource blocks required by the data stream.

4. The 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm according to claim 1, characterized in that, The signal receiving power is represented as: ; wherein, denotes the base station transmit signal power, PL(d) denotes the path loss; ; where d denotes the distance between the base station and the terminal, N denotes the distance power loss coefficient, denotes the floor penetration loss factor, and f denotes the carrier frequency of the signal.

5. The 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm according to claim 1, characterized in that, The number of available resource blocks is represented as: ; wherein, represents the number of available resource blocks, represents the total channel bandwidth, represents the single-side guard band bandwidth, represents the resource block bandwidth of a single resource block; The resource block capacity of a single resource block is represented as: ; wherein denotes the bit error rate, denotes the inverse function of the standard normal distribution, r denotes the block length of the resource blocks, denotes the signal-to-noise ratio; wherein .

6. The 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm according to claim 1, characterized in that, The policy network comprises three fully connected layers connected in sequence, the fully connected layers of the policy network are connected through a ReLU activation function, the last fully connected layer is used as an output layer to output an action value mean, and a Softmax function is used to calculate a probability distribution of an output action based on the output action value mean; the evaluation network comprises three fully connected layers connected in sequence, each fully connected layer of the evaluation network comprises 256 neurons and is used to output an output state value of a current state.

7. The 5G-TSN network resource scheduling method based on a near-end policy optimization algorithm according to claim 1, characterized in that, Further comprising: inputting the state into the policy network to obtain the action value mean, and obtaining the action value based on the action value mean; The action value is executed to obtain a reward value, when the number of obtained reward values is greater than a preset threshold, the policy network and the evaluation network are updated based on the reward value, and the reward value is emptied.

8. A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the 5G-TSN network resource scheduling method in any one of claims 1 to 7. The computer program is executed by the processor to implement the 5G-TSN network resource scheduling method in any one of claims 1 to 7.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, ​

Citation Information

Patent Citations

  • 5G-TSN joint resource scheduling device and method based on DDPG

    CN115811799A

  • Multi-data center task scheduling and data routing method and system under network topology

    CN120811963A