Resource scheduling method in slice packet network

By introducing a mixed granular time slot allocation mechanism and a resource scheduling method of deep reinforcement learning in the sliced ​​packet network, the problem of resource allocation complexity after fine-grained time slot units is solved, and efficient utilization of network resources and satisfaction of business needs is achieved.

CN120166539APending Publication Date: 2025-06-17BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510305373.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In the slice packet network, after the introduction of fine-grained time slot units, the traditional resource allocation model fails, resulting in fragmentation of time slot bandwidth and rigid frame structure, and the coupling complexity between routing computing and resource allocation is significantly increased, making it difficult to meet the needs of diversified business scenarios.

Method used

Using a hybrid granular time slot allocation mechanism and a resource scheduling method based on deep reinforcement learning, the routing path and time slot resource allocation are optimized through the agent decision unit and the policy network to achieve the end-to-end optimal mapping of business requirements and network resources.

Benefits of technology

It effectively solves the problem of high-dimensional discrete action space enumeration, reduces the computational overhead of traditional deep reinforcement learning models, realizes the optimization goal of minimizing network resource costs, and meets the rigid demand for deterministic resources in vertical industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120166539A_ABST
    Figure CN120166539A_ABST
Patent Text Reader

Abstract

The invention relates to a resource scheduling method in a slice packet network. The invention provides a resource scheduling method in a slice packet network, which is characterized in that a mixed granularity time slot allocation mechanism supporting time slot continuity and consistency constraint is constructed aiming at a mixed granularity service bearing requirement in a FlexE hard isolation slice scene, and an asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization is utilized to realize resource scheduling in a slice packet network. A parameterization mapping mechanism of a strategy network and a value network is utilized, the problem of dimension explosion when a traditional reinforcement learning model processes a high-dimensional discrete action space is effectively solved, the problem of resource scheduling in a slice packet network is solved, joint optimal configuration of a routing path and time slot resources is achieved, and the method has the advantages of being simple in structure and convenient to use. The resource utilization rate of the slice packet network is improved, and the network blocking rate and the calculation complexity are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of slice packet network communication, and in particular, to a resource scheduling method in a slice packet network. Background Art

[0002] In recent years, with the integrated innovation of cloud computing, Internet of Things, and artificial intelligence technologies, the Industrial 4.0 revolution has promoted the digital transformation of industries, giving rise to emerging vertical industries such as vehicle-to-everything (V2X), augmented / virtual reality (AR / VR), telemedicine, and smart grid. Diverse application scenarios have put forward multi-dimensional heterogeneous requirements for performance indicators such as communication network latency, transmission capacity, network synchronization, and slice performance, highlighting the contradiction between the traditional network architecture's difficulty in supporting the differentiated requirements of service quality and user experience quality for all scenarios. As the core support of the new generation of information infrastructure, the fifth-generation mobile communication (5G) and the sixth-generation mobile communication (6G) have built virtualized and customized network service paradigms for vertical industries by introducing network slicing technology, and have shown the potential to provide advanced network connection, security, and reliability guarantees for vertical industries.

[0003] Although certain progress has been made in the vertical industry network slice solution based on the slice packet network architecture, in the face of the network requirements of diverse business scenarios and the more flexible resource configuration mechanism in the SPN network, solving the routing decision and resource allocation problems still faces fundamental challenges. Especially when introducing the SPN fine-granularity unit (FGU), the traditional resource allocation model based on coarse-granularity time slot slicing fails, leading to key problems such as time slot bandwidth fragmentation and frame structure rigidity, making the coupling complexity of routing calculation and resource allocation increase exponentially. Therefore, the current research faces two challenges. First, at the theoretical level, existing research mainly focuses on resource optimization in a single scenario and lacks a systematic consideration of the slice full-life cycle management mechanism. Second, at the practical level, the traditional dynamic programming method is difficult to solve the problems caused by the significant increase in the resource allocation decision space brought about by the reduction of time slot bandwidth and the change of the rigid structure of flexible Ethernet frames after introducing FGU slices, resulting in the complication of routing calculation and time slot allocation calculation management. And the heuristic method has a trade-off dilemma between global optimization and convergence efficiency. Especially in a dynamic time-varying network environment, the sudden change of resource demand caused by bursty traffic flows and the unpredictability of network resource status also pose a severe challenge to the real-time response ability of the resource scheduling method, resulting in the difficulty of existing solutions to meet the rigid demand for deterministic resources in industrial control services in multiple scenarios. Summary of the Invention

[0004] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a resource scheduling method in a slice packet network.

[0005] In a first aspect, the present invention provides a resource scheduling method in a slice packet network, comprising:

[0006] Receive a list of services to be transmitted from the sliced ​​packet network, determine the bandwidth requirements of each service in the list of services to be transmitted, and achieve feature matching between service bandwidth and time slot granularity through a hybrid granularity time slot allocation mechanism;

[0007] According to the services in the service list, the bandwidth of the service transmitted by each network node in the sliced ​​packet network is determined, the number of time slots of different levels required for transmission by each network node in the sliced ​​packet network is calculated, and according to the bandwidth requirements of each service, the routing path and time slot resources of each service to be transmitted in the sliced ​​packet network are dynamically scheduled based on the resource scheduling method of deep reinforcement learning. The controller uses the intelligent agent to perform routing calculations and time slot allocations for the different levels of bandwidth services required to be transmitted by each node, accurately schedules network resources, and realizes the end-to-end optimal mapping of service requirements and network resources.

[0008] Furthermore, the method constructs a multi-level time slot system based on FlexE binding technology, and the time slot rate parameters of each granularity level maintain an integer multiple relationship with the physical interface bandwidth of the network node.

[0009] The core process of the method studied in this invention can be divided into three stages, namely, the state perception and decision initialization stage, the strategy evaluation and deep reinforcement learning iteration stage, and the model convergence and solution deployment stage.

[0010] Furthermore, in the state perception and decision initialization phase, the service management unit obtains the network topology state, resource state information and real-time service features through the resource management unit, the routing calculation unit and the service request set, and generates a multi-dimensional feature matrix as the input of the intelligent agent decision unit. The feature matrix covers network configuration parameters, service QoS requirements and historical resource allocation records, and is converted into a state space representation of deep reinforcement learning through the coding layer. At the same time, for the service flow to be observed in the service request set within the current decision window, the intelligent agent decision unit initializes the establishment of a strategy action space mapping relationship for each service.

[0011] In the policy evaluation and learning iteration stage, in the agent decision-making unit, the policy network models the probability distribution mapping from state to action based on a deep neural network, and outputs a subset of feasible actions that satisfy the constraint conditions; the value network evaluates the long-term value expectation of the current state, calculates the deviation between the actual reward and the predicted value, generates an advantage function to quantify the marginal benefit of the action, and then guides the gradient update of the policy network. The policy execution engine combines the resource calculation unit to verify the feasibility of the candidate actions, filters out a subset of valid actions that satisfy the constraint conditions such as time delay and bandwidth, and inputs the result into the reward calculation unit. The reward calculation unit calculates the reward value corresponding to the action set according to the established reward function, and applies a high penalty factor to the actions that cause service blocking or violate the constraint conditions. The intelligent decision-making unit balances exploration and exploitation through a greedy policy, selects the optimal action as the basis for the next state transition, and stores the state-action-reward tuple in the experience replay buffer to complete the update of the experience library data cache.

[0012] In the model convergence and solution deployment stage, when the experience buffer reaches the predetermined capacity threshold, the offline training mode of the deep neural network will be triggered, that is, the network weight parameters are updated using the mini-batch gradient descent algorithm in the experience buffer. When the convergence condition is reached, the target and policy network parameters are solidified as the optimal policy, the decision-making unit generates a global routing path and time slot allocation scheme, and the controller issues it to the underlying network. The resource management unit synchronously adjusts the time slot resource mapping relationship to achieve the optimization goal of minimizing the network resource cost.

[0013] Furthermore, the resource scheduling method based on deep reinforcement learning implemented for calculating the routing path and scheduling time slot resources for the large-bandwidth and small-bandwidth services to be transmitted includes: adopting an asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization.

[0014] For the asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization, by introducing the Actor-Critic framework, it collaboratively integrates the dual mechanisms based on the value function and the policy function. Using a parameterized policy avoids the need to explicitly enumerate all actions, and is more suitable for solving problems in high-dimensional discrete action spaces. The asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization is implemented based on the A3C (Asynchronous Advantage Actor-Critic) algorithm. It interacts with the environment by constructing a global shared neural network and multiple asynchronous agents with the same network structure, and updates the network parameters using an N-step iteration strategy. All network models are of the Actor-Critic structure.

[0015] Furthermore, for the state space, the state space can describe the basic environment in which the agent operates. A reasonable state space design enables it to clearly perceive the current network state and resource usage, enhancing and accelerating the learning process of the agent in the high-dimensional space. In the method, the state representation at a certain moment is represented by a vector array with a length of 2×N V +P×(J + 2×K + 1).

[0016] Furthermore, for the action space, the construction of the action space needs to clearly reflect the combined characteristics of the decision dimensions. For each service request, the parallel agents need to select a physical route from the set of candidate paths that meet the bandwidth and delay constraints, and at the same time select an allocation scheme from the available frequency slot combinations that meet the time slot continuity and consistency. The action space is essentially a two-dimensional discrete space, and its dimension is jointly determined by the number of candidate paths and the number of feasible frequency slot allocation schemes. Therefore, the agent is responsible for selecting a comprehensive solution from the sets of time slot occupancy allocation schemes and the candidate paths that meet the service requests for service routing. For each service request, its action space consists of actions.

[0017] Furthermore, for the reward function, the design of the reward function directly affects the directionality and convergence efficiency of the policy optimization of multiple parallel agents. The present invention introduces a reward function with the successful routing of service requests as the goal. If the service request is successfully routed, a positive reward value of 1 is given, otherwise a negative reward value of -1 is given. This reward function reduces the complexity of the method implementation through a minimalist feedback mechanism, and under the core optimization goal of minimizing the blocking rate, it avoids the problems of policy oscillation or convergence direction deviation caused by multi-dimensional rewards.

[0018] The above technical solutions provided by the embodiments of the present invention have the following advantages compared with the prior art:

[0019] Compared with the prior art, in a dynamic time-varying network environment, the dynamic random arrival characteristics of service requests and the unpredictability of resource status make it difficult for traditional resource optimization strategies based on static network state modeling to meet the service quality requirements of real-time scheduling. Aiming at the hybrid-granularity service bearing requirements based on hard-isolated slices in the FlexE technology-enabled vertical domain scenarios, the present invention proposes a hybrid-granularity time-slot allocation mechanism, which brings time-slot continuity and consistency constraints for minimizing network resource costs and causes the explosion of the discrete state space dimension, resulting in a sharp drop in the exploration efficiency of traditional deep reinforcement learning models relying on value functions and a sharp increase in computational overhead. To solve the SPN routing and time-slot resource dynamic scheduling problems for hybrid-granularity time-slot slicing, the present invention proposes an asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization, innovatively adopting the Actor-Critic framework, and effectively solving the problem of enumerating high-dimensional discrete action spaces through the collaborative optimization mechanism of the value network and the policy network. In addition, the parameterized policy of the policy network reduces the computational overhead of the traditional value function for traversing the explicit action space. The method proposed by the present invention is superior to the existing methods, and can dynamically allocate a routing and time-slot allocation scheme with the minimum network resource cost according to the service requirements in the slice-grouped network, and can maintain a low network blocking rate in different topological structures. At the same time, it can not only meet the service requirements of hard isolation and low latency for real-time control services in the vertical industry, but also improve the network resource utilization rate through resource scheduling, ensuring the QoS (Quality of Service) and SLA (Service-Level Agreement) of different types of services in the slice-grouped network. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic flow chart of a resource scheduling method in a slice-grouped network provided by an embodiment of the present invention.

[0023] Figure 2 It is a model architecture diagram of a resource scheduling method in a slice-grouped network provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0025] It should be noted that in this text, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

[0026] In recent years, with the integration and innovation of cloud computing, Internet of Things, and artificial intelligence technologies driving the digital transformation of industries, emerging vertical industries such as vehicle-to-everything (V2X), augmented / virtual reality (AR / VR), telemedicine, and smart grid have emerged. Diverse application scenarios have put forward multi-dimensional heterogeneous requirements for performance indicators such as the latency, transmission capacity, network synchronization, and slice performance of communication networks, highlighting the contradiction between the traditional network architecture's difficulty in supporting the differentiated requirements of service quality and user experience quality for all scenarios. As the core support of the new generation of information infrastructure, the fifth-generation mobile communication (5G) and the sixth-generation mobile communication (6G) have built a virtualized and customized network service paradigm for vertical industries by introducing network slicing technology, and have shown the potential to provide advanced network connection, security, and reliability guarantees for vertical industries.

[0027] Although the vertical industry network slice solution based on the slice grouping network architecture has made certain progress, in the face of the network requirements of diverse business scenarios and the more flexible resource allocation mechanism in the SPN network, there are still fundamental challenges in solving the routing decision and resource allocation problems. Especially when introducing the SPN fine-granularity unit (FGU), the traditional resource allocation model based on coarse-granularity time slots fails, resulting in key problems such as time slot bandwidth fragmentation and frame structure rigidity, making the coupling complexity of routing calculation and resource allocation increase exponentially. Therefore, the current research faces two challenges. First, at the theoretical level, existing research mainly focuses on resource optimization in a single scenario and lacks a systematic consideration of the slice full-life cycle management mechanism. Second, at the practical level, traditional dynamic programming methods are difficult to solve the problems brought about by the introduction of FGU slices, such as the significant increase in the resource allocation decision space due to the reduction of time slot bandwidth and the change of the rigid structure of flexible Ethernet frames, resulting in the complexity of routing calculation and time slot allocation calculation management. And heuristic methods have a trade-off dilemma between convergence efficiency and global optimization. Especially in a dynamic time-varying network environment, the sudden change in resource demand caused by bursty traffic flows and the unpredictability of network resource status also pose a severe challenge to the real-time response ability of resource scheduling methods, making it difficult for existing solutions to meet the rigid demand for deterministic resources in industrial control services in multiple scenarios. To solve the above problems, the present invention provides a resource scheduling method in a slice grouping network. For the convenience of understanding this embodiment, first, a resource scheduling method in a slice grouping network disclosed in the embodiments of the present invention will be introduced in detail. See Figure 1 As shown, the resource scheduling method in a slice grouping network includes:

[0028] Step S100: The system receives a list of service requests to be transmitted from the slice grouping network, determines the service types and quantities of the services in the list of services to be transmitted, and classifies the services into large-bandwidth services and small-bandwidth services according to the bandwidth size. The bandwidth requirements of each service will be used for subsequent resource allocation calculations.

[0029] Step S200: By introducing a hybrid-granularity time slot allocation mechanism, the system matches the service bandwidth with an appropriate time slot granularity to reduce resource conflicts and scheduling delays and improve resource utilization.

[0030] Based on the existing IEEE802.3 standard, the FlexE slicing technology introduces a FlexE Shim layer between the MAC layer and the PHY layer; this layer implements the core functions of FlexE, mainly responsible for coordinating and distributing the data streams of FlexE Clients to the FlexE PHY, and at the same time supporting rate adaptation between FlexE Clients; in addition, based on the time-division multiplexing technology, FlexE supports the mapping and transmission of any number of sub-interfaces on any FlexE PHY, realizing functions such as port bundling, sub-rate, and channelization, breaking the limitation that the MAC layer only matches the rate of the existing PHY layer interface under the standard Ethernet architecture, and realizing flexible matching of network interfaces; at the same time, it supports dividing one or more bundled Ethernet ports into multiple independent Ethernet elastic hard pipes, realizing end-to-end hard isolation and improving the security of slicing. Specifically, a multi-level hybrid-granularity time slot allocation mechanism is constructed based on the FlexE binding technology. The hybrid-granularity time slot allocation mechanism constructs a multi-level time slot system based on the FlexE binding technology, and the time slot rate parameters of each granularity level maintain an integer multiple relationship with the physical interface bandwidth of the network node..

[0031] In some embodiments, the service bandwidth to be transmitted in the slice packet network is determined by a greedy method, and the number of time slots required for each physical node in the slice packet network to carry each service to be transmitted is calculated. The specific process is as follows:

[0032] For the service request r to be transmitted in any physical node in the slice packet network k The number of hybrid-granularity flexible Ethernet frame time slots occupied Can be calculated by the greedy method.

[0033] Step S300: Based on the resource scheduling method in the slice packet network, the system evaluates the network state through the agent decision-making unit, and optimizes the routing and time slot resource scheduling through the policy network and the value network.

[0034] The core advantage of DRL stems from its unique architecture design. On the one hand, using a deep neural network as a function approximator to save the input of the parameterized network policy enables it to understand and predict complex system states based on a high-dimensional state space, realizing non-linear mapping and feature extraction of the high-dimensional state space; on the other hand, through Markov decision process modeling, through continuous interaction with the target system, it can gradually learn based on the returned rewards and experiences, and adopt an action strategy closest to the ideal target state. The ingenious design of this method enables it to have excellent self-learning ability, can quickly adapt to changes in network conditions, and flexibly handle dynamic optimization problems, showing strong adaptability and flexibility in the process of exploring dynamic environments and resource states.

[0035] In some embodiments, referring to Figure 2 As shown, the core process of the asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization studied in the present invention can be divided into three stages, namely, the state perception and decision initialization stage, the policy evaluation and deep reinforcement learning iteration stage, and the model convergence and solution deployment stage.

[0036] In the state perception and decision initialization stage, the service management unit obtains the network topology state, resource status information, and real-time service characteristics through the resource management unit, routing calculation unit, and service request set, and generates a multi-dimensional feature matrix as the input of the agent decision-making unit. This feature matrix covers network configuration parameters, service QoS requirements, and historical resource allocation records, and is transformed into the state space representation of deep reinforcement learning through the encoding layer. At the same time, for the to-be-observed service flows in the service request set within the current decision window, the agent decision-making unit initializes and establishes a policy action space mapping relationship for each service.

[0037] In the policy evaluation and learning iteration stage, in the agent decision-making unit, the policy network models the probability distribution mapping from state to action based on a deep neural network, and outputs a subset of feasible actions that meet the constraint conditions; the value network evaluates the long-term value expectation of the current state, calculates the deviation between the actual reward and the predicted value, generates an advantage function to quantify the marginal benefit of the action, and then guides the gradient update of the policy network. The policy execution engine combines the resource calculation unit to verify the feasibility of the candidate actions, filters out a subset of effective actions that meet the constraint conditions such as delay and bandwidth, and inputs the result into the reward calculation unit. The reward calculation unit calculates the reward value corresponding to the action set according to the established reward function, and applies a high penalty factor to the actions that cause service blockage or violate the constraint conditions. The intelligent decision-making unit balances exploration and exploitation through the greedy policy, selects the optimal action as the basis for the next state transition, and stores the state-action-reward tuple in the experience replay buffer to complete the update of the experience library data cache.

[0038] In the model convergence and solution deployment stage, when the experience buffer reaches the predetermined capacity threshold, the offline training mode of the deep neural network will be triggered, that is, the network weight parameters are updated using the mini-batch gradient descent algorithm in the experience buffer. When the convergence condition is reached, the target and policy network parameters are solidified as the optimal policy, the decision-making unit generates the global routing path and time slot allocation solution, and the controller sends it to the underlying network. The resource management unit synchronously adjusts the time slot resource mapping relationship, and finally realizes the optimization goal of minimizing the network resource cost.

[0039] The detailed design of the state space, reward function, and training method of the asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization includes:

[0040] The state space can describe the basic environment in which the agent operates. A reasonable state space design enables it to clearly perceive the current network state and resource usage, enhancing and accelerating the agent's learning process in the high-dimensional space. In this method, a vector array of length 2×N V +P×(J + 2×K + 1) represents the state representation at time t

[0041] The state representation includes service request information, resource status information of the network topology, and traffic information of each link. Specifically, the first 2 elements represent service request information, where the source node s k and the destination node d k are represented in a one-hot structure suitable for neural network training. For each link in the P candidate paths, the next 4 elements represent the network resource status information and traffic information, where represents the number of routing hops; represents the one-dimensional matrix vector of the end-to-end delay of the j-th time slot occupancy scheme in the i-th path, which can be expressed as:

[0042] is the one-dimensional matrix vector representing the available time slot resources of the k-th granularity in the i-th path, reflecting the overall distribution of available resources in the network, which can be expressed as:

[0043] where, |p i | represents the number of links in the i-th path. represents the one-dimensional matrix vector of the number of the k-th time slots required by the service in the i-th path, with the same data structure as the same.

[0044] The feature input matrix is composed of the above parameters. By deeply integrating service information with the key features of different candidate paths, the agent can perceive the information of the global network state.

[0045] The construction of the action space needs to clearly reflect the combined characteristics of the decision dimensions. For each service request, the parallel agent needs to select a routing path from the set of candidate paths that meet the bandwidth and delay constraints, and at the same time select an allocation scheme from the available frequency gap combinations that meet the time slot continuity and consistency. The action space is essentially a two-dimensional discrete space, and its dimension is jointly determined by the number of candidate paths and the number of feasible frequency gap allocation schemes. Therefore, the agent is responsible for selecting a comprehensive scheme from the J sets of time slot occupancy allocation schemes and the P candidate paths that meet the service request for service routing. For each service request r k , its action space consists of J×P actions.

[0046] The design of the reward function directly affects the directionality and convergence efficiency of the policy optimization of multiple parallel agents. In the present invention, a reward function is introduced with the goal of whether a service request is successfully routed. If the service request is successfully routed, a positive reward value of 1 is given; otherwise, a negative reward value of -1 is given. This reward function significantly reduces the complexity of method implementation through an extremely simple feedback mechanism. Meanwhile, in a clear scenario with the core optimization goal of minimizing the blocking rate, it avoids the problems of policy oscillation or deviation of the convergence direction that may be caused by multi-dimensional rewards.

[0047] Step S400: Continuously iterate and optimize the resource allocation policy. The system continuously adjusts and improves the policy through the iterative process of deep reinforcement learning to cope with the dynamic changes of network states and service demands. In the offline training phase, the system analyzes historical data and real-time service characteristics and continuously adjusts the network weights and model parameters. Through multiple trainings and evaluations, the system gradually learns the resource scheduling policy that is most suitable for the current network environment and service demands.

[0048] After multiple rounds of training, the system can minimize the network resource cost while ensuring that the service quality requirements are met. Especially in complex network topologies and high-dimensional resource constraints, it can effectively avoid resource conflicts and slot scheduling delays. Finally, the system generates the optimal routing path and slot allocation scheme, which fully utilize the network resources while ensuring the key service requirements such as bandwidth and latency.

[0049] Specifically, when the DNN model is learning and training, the service set is used as the basic training unit. To balance the model learning efficiency and network state diversity, each episode unit contains 1000 dynamic service requests, and their bandwidth requirements follow a uniform distribution within the range of 21 - 260 Mbps. The entire training cycle contains 800 episodes, and a total of about 800,000 service requests are processed, fully covering the complex scenarios of network load fluctuations and resource allocation decisions.

[0050] After the final scheme is generated, the system distributes the optimal routing path and slot allocation scheme to the underlying network through the controller. As the collaborative control center of the network, the controller can dynamically adjust and manage the resource allocation of each node in the network to ensure that each node can execute the assigned tasks as required. This scheme is finally deployed in the network. Through real-time feedback and scheduling optimization, it continuously maintains the efficient utilization of network resources and the stable transmission of services. This process realizes the automation and intelligence of network resource management and improves the adaptability and flexibility of the network in the face of different service demands.

[0051] From the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software and necessary general-purpose hardware. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk, or optical disc of a computer, etc., and includes several instructions to enable an electronic device (which can be a mobile phone, personal computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0052] In the embodiments provided by the present invention, it should be understood that the disclosed structures and methods can be implemented in other ways. For example, the structural embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of structures or units can be in an electrical, mechanical or other form.

[0053] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A resource scheduling method in a slice packet network, characterized in that: include: Receive a list of services to be transmitted from the sliced ​​packet network, determine the bandwidth requirements of each service in the list of services to be transmitted, and achieve feature matching between service bandwidth and time slot granularity through a hybrid granularity time slot allocation mechanism; According to the various services in the service list, the bandwidth of the service transmitted by each network node in the sliced ​​packet network is determined, the number of time slots of different levels required for transmission by each network node in the sliced ​​packet network is calculated, and according to the bandwidth requirements of each service, the routing path and time slot resources of each service to be transmitted in the sliced ​​packet network are dynamically scheduled based on the resource scheduling method of deep reinforcement learning; the controller performs dynamic routing calculation and time slot allocation for different services that need to be transmitted by each node through the intelligent agent, accurately schedules network resources, and realizes end-to-end optimal mapping of service requirements and network resources.

2. The resource scheduling method in a slice packet network according to claim 1, characterized in that: The hybrid granularity time slot allocation mechanism builds a multi-level time slot system based on the FlexE binding technology, and the time slot rate parameters of each granularity level maintain an integer multiple relationship with the physical interface bandwidth of the network node.

3. The resource scheduling method in a slice packet network according to claim 1, characterized in that: The resource scheduling method based on deep reinforcement learning, which is implemented to calculate routing paths and schedule time slot resources for each service in the list of services to be transmitted, includes: adopting an asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization.

4. The resource scheduling method in a slice packet network according to claim 3, characterized in that: The asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization uses a dynamic routing and time slot allocation mechanism to perform state perception and decision initialization by reasonably setting the state space, action space and reward function. Through the strategy evaluation and learning iteration phases of deep reinforcement learning, the system can continuously optimize the time slot allocation scheme, and through the model convergence phase, it finally generates the global optimal routing and time slot allocation scheme, and sends it to the slice grouping network to achieve global optimal resource allocation.

5. The resource scheduling method in a slice packet network according to claim 3, characterized in that: For the asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization, its state space is represented by a length of 2×N V +P×(J+2×K+1) vector array represents the state at time t The state representation includes the service request information, the resource state information of the network topology, and the traffic information of each link. Specifically, the first two elements represent the service request information, where the source node s k and the target node d k It is represented by a one-hot structure suitable for neural network training. For each link in the P candidate paths, the last four elements represent the network resource status information and traffic information, where Represents the number of routing hops; The one-dimensional matrix vector representing the end-to-end delay of the j-th time slot occupancy scheme in the i-th path can be expressed as: is a one-dimensional matrix vector representing the available time slot resources of the kth granularity in the i-th path, reflecting the overall distribution of available resources in the network, which can be expressed as: Among them, |p i | represents the number of links in the i-th path. is a one-dimensional matrix vector representing the number of k-th time slots required by the service on the i-th path. The data structure and same.

6. The resource scheduling method in a slice packet network according to claim 3, characterized in that: For the asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization, the action space is set as a two-dimensional discrete space, and the spatial dimension is determined by the number of candidate paths and the number of feasible frequency slot allocation schemes. Therefore, the agent is responsible for selecting a comprehensive scheme for service routing from the J sets of time slot occupancy allocation schemes and the P candidate paths that meet the service request. For each service request r k , whose action space consists of J×P actions.

7. The resource scheduling method in a slice packet network according to claim 3, characterized in that: For the asynchronous deep reinforcement learning method based on Actor-Critic collaborative optimization, the reward function is introduced with the goal of whether the business request is successfully routed. If the business request is successfully routed, a positive reward value of 1 is given, otherwise a negative reward value of -1 is given. This reward function uses a very simple feedback mechanism and takes minimizing the blocking rate as the core optimization goal, avoiding the problem of strategy oscillation or convergence direction deviation that may be caused by multi-dimensional rewards.

8. The resource scheduling method in a slice packet network according to claim 3, characterized in that: The asynchronous deep reinforcement learning method based on the Actor-Critic collaborative optimization model includes three main processes: state perception and decision initialization, strategy evaluation and learning iteration, and model convergence and solution deployment.

Citation Information

Cited By

  • Dynamic computing power routing computing method and system facing business requirements

    CN122093468A