Network scheduling method and device, computer equipment, storage medium and program product
By determining the optimal edge-to-edge network mode in the SDWAN network and optimizing the routing configuration using a deep reinforcement learning model, the problems of low efficiency and slow response speed in complex task processing of SDWAN networks are solved, and collaborative scheduling of edge nodes and network performance are improved.
Patent Information
- Application Number
- CN202511810213.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-03-06
AI Technical Summary
Existing SDWAN networks suffer from low task processing efficiency and slow system response when faced with complex tasks, and users may experience network performance degradation when selecting network modes and nodes.
By determining the optimal edge-to-edge network mode based on the network status and business requirements of each edge node in the edge cluster, and optimizing the routing configuration using a deep reinforcement learning model, routing configuration information corresponding to the virtual private cloud instance of the target user is generated, thereby realizing the collaborative scheduling of edge nodes and the automated configuration of routing policies.
It improves network performance and the efficiency of automatic routing, avoids performance degradation caused by users selecting the wrong network mode and node, and realizes the collaborative scheduling of computing nodes and the efficient utilization of network resources in edge environments.
Smart Images

Figure CN121619280A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of edge computing technology, and in particular to a network scheduling method, apparatus, computer equipment, storage medium, and program product. Background Technology
[0002] With the development of cloud computing and edge computing, SDWAN (Software-Defined Wide Area Network) has become a widely adopted network architecture in enterprise networks. However, current SDWAN networks often suffer from low task processing efficiency and slow system response when faced with complex tasks. Furthermore, the selection of POP (Point of Presence) nodes directly impacts network performance and user experience.
[0003] In traditional technologies, good network configurations are not typically recommended to users proactively. As a result, users often choose the wrong SDWAN network mode and nodes, leading to a decline in network performance. Summary of the Invention
[0004] Therefore, it is necessary to provide a network scheduling method, apparatus, computer equipment, computer-readable storage medium, and computer program product that can improve network performance in response to the above-mentioned technical problems.
[0005] Firstly, this application provides a network scheduling method, including:
[0006] Based on the network status and service requirements of each edge node in the edge cluster, determine the edge-to-edge network mode that matches the edge cluster.
[0007] Based on the cluster performance description information of each edge node in the edge cluster and the edge-to-edge network pattern that matches the edge cluster, the target cluster is determined from each edge cluster.
[0008] Based on the edge-to-edge network pattern matching the target cluster and the target edge nodes in the target cluster, routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node is generated; the target edge node includes the edge node in the target cluster used to execute tasks;
[0009] The edge forwarding router is distributed to the target edge node, and the routing policy of the target edge node is configured according to the routing configuration information.
[0010] In one embodiment, determining the target cluster from each edge cluster based on the cluster performance description information of each edge node within the edge cluster and the edge-to-edge network pattern matching the edge cluster includes:
[0011] Based on the cluster performance description information of each edge node in the edge cluster, determine the cluster performance quantification value of the edge cluster;
[0012] The network weights of the edge cluster are determined based on the edge-to-edge network pattern that matches the edge cluster.
[0013] The comprehensive score of the edge cluster is determined based on the network weights and the cluster performance quantification values.
[0014] Based on the comprehensive score, the target cluster is determined from each of the edge clusters.
[0015] In one embodiment, the edge-to-edge network mode includes a metropolitan area network mode, a CN2 mode, and a 163 mode; the network weight of the CN2 mode is higher than that of the 163 mode; and the network weight of the 163 mode is higher than that of the metropolitan area network mode.
[0016] In one embodiment, generating routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node based on the edge-to-edge network pattern matching the target cluster and the target edge node in the target cluster includes:
[0017] The prompt routing statement is expanded based on the edge-to-edge network pattern that matches the target cluster and the target edge node in the target cluster.
[0018] Based on the expanded prompt routing statement, generate routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node.
[0019] In one embodiment, before generating routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node based on the edge-edge network pattern matched with the target cluster and the target edge node in the target cluster, the method further includes:
[0020] Deploy the pre-trained deep reinforcement learning model to each edge node in the target cluster;
[0021] The node environment state information of each edge node in the target cluster is input into the pre-trained deep reinforcement learning model to obtain the task scheduling strategy of each edge node.
[0022] According to the task scheduling strategy, the target edge node is determined from each edge node of the target cluster.
[0023] In one embodiment, after determining the target edge node from each edge node of the target cluster according to the task scheduling strategy, the method further includes:
[0024] The experience data of each target edge node after executing the task scheduling strategy is uploaded to the central experience replay pool; the experience data includes at least the executed action, reward value, and new status.
[0025] Based on the reward value of each piece of experience data, target experience data is selected from the central experience replay pool as training samples.
[0026] The pre-trained deep reinforcement learning model is updated and trained using the training samples, and the updated deep reinforcement learning model is deployed to each edge node in the target cluster.
[0027] Secondly, this application also provides a network scheduling device, comprising:
[0028] The mode determination module is used to determine the edge-to-edge network mode that matches the edge cluster based on the network status and service requirements of each edge node in the edge cluster.
[0029] The cluster determination module is used to determine the target cluster from each edge cluster based on the cluster performance description information of each edge node in the edge cluster and the edge-to-edge network pattern that matches the edge cluster.
[0030] The information generation module is used to generate routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node, based on the edge-to-edge network pattern matching the target cluster and the target edge node in the target cluster; the target edge node includes the edge node in the target cluster used to execute tasks.
[0031] The routing configuration module is used to distribute the edge forwarding router to the target edge node and configure the routing policy of the target edge node according to the routing configuration information.
[0032] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.
[0033] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method.
[0034] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the above-described method.
[0035] The aforementioned network scheduling method, apparatus, computer equipment, computer-readable storage medium, and computer program product determine the edge-to-edge network mode matching the edge cluster based on the network status and service requirements of each edge node within the edge cluster; determine the target cluster from each edge cluster based on the cluster performance description information of each edge node within the edge cluster and the edge-to-edge network mode matching the edge cluster; generate routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node based on the edge-to-edge network mode matching the target cluster and the target edge node in the target cluster; the target edge node includes the edge node in the target cluster used to execute tasks; and distribute the edge forwarding router to the target edge node and configure the routing policy of the target edge node according to the routing configuration information. In this way, by comprehensively considering the real-time network status of each node in the edge cluster and specific business requirements, the system intelligently makes decisions and matches the optimal edge-to-edge network mode. Then, by combining cluster performance description information, it accurately selects the target cluster with the best resource status. Subsequently, it automatically generates routing configuration information adapted to the target user's virtual private cloud instance for the selected target edge nodes, and directly deploys the edge forwarding router to the edge nodes with the best execution capabilities and applies routing policies that match the routing configuration information. Through a hierarchical and automated network scheduling process from network mode matching and cluster selection to node scheduling, it realizes the collaborative scheduling of computing nodes in the edge environment, improves network performance, and enhances the efficiency of automatic network routing. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a diagram illustrating the application environment of a network scheduling method in one embodiment.
[0038] Figure 2 This is a flowchart illustrating a network scheduling method in one embodiment;
[0039] Figure 3 This is a logical diagram of a routing instance and a virtual private cloud interaction method in one embodiment;
[0040] Figure 4 This is a flowchart of a deep reinforcement learning model in one embodiment;
[0041] Figure 5 This is a flowchart illustrating a network scheduling method in another embodiment;
[0042] Figure 6 This is a structural block diagram of a network scheduling device in one embodiment;
[0043] Figure 7 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0045] The network scheduling method provided in this application embodiment can be applied to, for example, Figure 1 The application environment shown. The network scheduling system 102 includes an SDN control module, a monitoring and acquisition module, an edge-to-edge network implementation module, a task prediction module, and a deep reinforcement learning network scheduling module.
[0046] Edge-to-Edge Network (EEN) is a key technology concept in the field of edge computing, referring to the secure and efficient intranet interconnection service between different Virtual Private Clouds (VPCs) or edge nodes in an edge cloud environment. This technology solves the complexity of cross-regional interconnection in traditional cloud environments by optimizing network architecture and routing strategies, and is particularly suitable for scenarios that are sensitive to latency and have high data localization requirements.
[0047] SDN (Software-Defined Networking) is an architecture that centrally controls and manages the network through software. It separates the control plane (decision logic) from the data plane (data forwarding) of traditional network devices (such as switches and routers), enabling flexible network configuration, automated management, and rapid iteration.
[0048] In specific implementation, the SDN control module is used to globally perceive network resources and rationally allocate bandwidth and computing resources; the monitoring and acquisition module is used to collect relevant data from each node in the edge cluster; the edge-to-edge network implementation module is used to implement edge-to-edge networks in various modes; the task prediction module is used to select one of the various edge-to-edge network modes based on network status and task requirements; and the deep reinforcement learning network scheduling module is used to implement adaptive router allocation and offloading, optimize the decision-making process through reinforcement learning, and dynamically allocate routes to individual nodes or allow multiple nodes to execute simultaneously.
[0049] In implementing various edge-to-edge network solutions, this system monitors and evaluates network and bandwidth resources to select the optimal edge-to-edge network and designs the execution environment for task prediction modules. During edge-to-edge network instance scheduling, deep learning is used to schedule instances to better nodes, enabling collaborative scheduling and offloading of multiple computing nodes in the edge environment.
[0050] SD-WAN (Software-Defined Wide Area Network) technology plays a crucial role in modern network architecture, especially in scheduling optimization. Combining SDN control modules with SDN technology can significantly improve network resource utilization and the efficiency of scheduling task execution. By leveraging SDN technology's centralized control plane and distributed data plane to achieve flexible control and management of network traffic, multiple edge-to-edge network modes can be implemented. Furthermore, the collaborative work of various nodes improves the efficiency of routing task processing and system response speed. Since multi-node task processing requires significant computing resources, this network scheduling system can address the latency and performance issues that can arise from processing tasks through a single node.
[0051] In practical applications, the network scheduling system 102 determines the edge-to-edge network mode matching the edge cluster based on the network status and service requirements of each edge node within the edge cluster; the network scheduling system 102 identifies the target cluster from each edge cluster based on the cluster performance description information of each edge node within the edge cluster and the edge-to-edge network mode matching the edge cluster; the network scheduling system 102 generates routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node based on the edge-to-edge network mode matching the target cluster and the target edge node in the target cluster; the target edge node includes the edge node in the target cluster used to execute tasks; the network scheduling system 102 distributes the edge forwarding router to the target edge node and configures the routing policy of the target edge node according to the routing configuration information.
[0052] In one exemplary embodiment, such as Figure 2 As shown, a network scheduling method is provided, which can be applied to... Figure 1 The following explanation will be based on the network scheduling system 102 in the example, including:
[0053] Step S202: Determine the edge-to-edge network mode that matches the edge cluster based on the network status and service requirements of each edge node in the edge cluster.
[0054] In this context, an edge cluster can refer to a logical or physical collection of multiple edge nodes deployed at the network edge. These edge nodes can be uniformly managed by a container orchestration system to provide nearby computing, storage, and network services. Optionally, the container orchestration system can be a Kubernetes (K8s) system.
[0055] In practical applications, a Kubernetes system can be used as the basis for deploying controller software such as ONOS, and OpenFlow can be used to realize network topology discovery and communication between the controller and the switch. This can effectively handle network policy formulation, route calculation and traffic scheduling, as well as execute corresponding operations based on actual data forwarding and control plane instructions.
[0056] The network status can include status indicators that reflect the current network communication quality of edge nodes, such as the computing performance, network performance, bandwidth utilization, transmission latency, packet loss rate, network jitter, and link congestion level of edge nodes.
[0057] Among them, business requirements may include the specific requirements of the task to be processed or user instance on the operating environment, such as business type (e.g., real-time task, offline task), latency sensitivity (e.g., low latency requirement), data localization requirements, security level, and cross-regional interconnection requirements (e.g., intra-provincial interconnection or inter-provincial interconnection).
[0058] Edge-to-Edge Network Mode (EEN) refers to the specific network architecture or protocol scheme for enabling intranet communication between different Virtual Private Clouds (VPCs) or edge nodes in an edge cloud environment. These modes represent different routing policies, Quality of Service (QoS) levels, and coverage areas. Then, srv6 can be used to implement three edge-to-edge network modes: metropolitan area network mode, CN2 mode, and 163 mode. Finally, the edge-to-edge network mode that matches the edge cluster can be selected from these three modes.
[0059] The CN2 mode can include a high-quality backbone network connection mode built on the next-generation bearer network architecture. This mode is mainly used for cross-provincial cluster interconnection and is implemented based on Virtual Private Network (VPN) and SRv6 technology.
[0060] The 163 mode can include standard network connection modes built on traditional public computer internet. This mode is achieved by dividing public network segments and binding them with IPv6 addresses, typically associated with a specific VPN leased line.
[0061] The metropolitan area network (MAN) mode can include a local network connection mode built on a local area network (LAN) within a city or region. This mode is mainly used for cluster interconnection within a province, and is implemented by binding and dividing v6 segments and using SRv6 for encapsulation at the Overlay layer.
[0062] In practical applications, the edge-to-edge network acts as a logical mesh. Each VPC instance joining the network deploys a corresponding edge forwarding router, and the interconnection between these routers forms the edge-to-edge network. This enables intranet communication between different clusters within the same cluster, maximizing the utilization of network resources across clusters. Optionally, intra-provincial clusters can utilize a metropolitan area network via VPN. By binding and dividing the V6 segments, V6 addresses are assigned to the VPC routers joined to each edge-to-edge network instance, and SRv6 is used for packet encapsulation and decapsulation at the overlay layer. Inter-provincial clusters are implemented using CN2 mode based on VON, also based on SRv6 mode. In 163 mode, V6 addresses are bound to the divided public network segments for association; these public networks belong to specific VPN leased lines.
[0063] For the convenience of those skilled in the art, Figure 3 An exemplary logical diagram of a method for interaction between a routing instance and a virtual private cloud is provided. It can be seen that... Figure 3 Above, vpc1 and vpc2 can achieve bidirectional routing through edge-to-edge network instances; Figure 3 Below, there are cross-connection links between the four edge network instances, allowing them to communicate with each other.
[0064] For example, the edge-to-edge network mode can be determined based on network conditions and business requirements using preset decision logic. The preset decision logic includes, but is not limited to, the following methods:
[0065] If business needs involve cross-province or cross-regional cluster interconnection and have high requirements for network quality (such as high-priority services), the edge-to-edge network mode can be determined as a high-quality service backbone network mode. This mode achieves high-quality interconnection based on leased lines through VPN and SRv6 technology, such as the CN2 mode.
[0066] If business needs require extensive public network access or binding to specific public network segments, and these are standard priority services, the edge-to-edge network mode can be determined as the general backbone network mode. This mode is suitable for general cross-domain interconnection needs, such as the 163 mode.
[0067] If business needs are mainly concentrated in the same administrative region or province, and there are high requirements for data localization, the edge-to-edge network mode can be determined as the metropolitan area network mode. This mode can effectively reduce detour delays and achieve efficient local networking in the scenario of inter-cluster communication within the province.
[0068] Furthermore, the score of each edge-to-edge network mode can be determined based on the network status of the edge nodes, and then the weight coefficient of each edge-to-edge network mode can be determined based on business requirements. Thus, the matching score of each edge-to-edge network mode can be determined based on the weight coefficient and the score, and the edge-to-edge network mode with the highest matching score can be determined as the edge-to-edge network mode that matches the edge cluster.
[0069] Step S204: Based on the cluster performance description information of each edge node in the edge cluster and the edge-to-edge network pattern that matches the edge cluster, determine the target cluster from each edge cluster.
[0070] The cluster performance description information may include a set of data used to characterize the resource load and processing capabilities of the edge cluster as a whole or its internal nodes, including the computing performance (such as CPU utilization and remaining memory), network performance (such as link latency and packet loss rate) and bandwidth utilization of each edge node in the edge cluster.
[0071] The target cluster is the most suitable cluster selected for deploying the current edge forwarding routers and executing target user services. This target cluster, while meeting the edge-to-edge network requirements of the services, possesses the highest overall cluster performance score; for example, it has ample resources and optimal network quality.
[0072] In practice, the computing performance, network performance, and bandwidth utilization of different edge nodes can be calculated. Different edge-to-edge network modes can be recommended based on different business needs. For different edge nodes, the edge nodes are scored according to the priority of business needs, and the target cluster with the highest score is provided in order of priority. At the same time, network weights can be added. Different edge-to-edge network modes can correspond to different network weights. Then, the target cluster is determined by combining the network weights.
[0073] Step S206: Based on the edge-to-edge network pattern matching the target cluster and the target edge node in the target cluster, generate the routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node.
[0074] The target edge node includes the edge nodes in the target cluster used to execute tasks. Specifically, the target edge node can be the optimal physical or logical node in the target cluster, selected by an intelligent scheduling algorithm, for actually carrying the edge forwarding router instance and executing user computing tasks. This node possesses the optimal resource state under the current environment.
[0075] A Virtual Private Cloud (VPC) instance refers to a logically isolated virtual network environment built on demand for users in an edge cloud environment. Each instance has its own independent IP address range, routing table, and security policy, and is a collection of containers or virtual machines that host user business applications.
[0076] The routing configuration information may include a set of control policies that guide how network traffic is forwarded between different VPC instances, edge nodes, and external networks. Optionally, the routing configuration information may include a set of dynamically generated, machine-executable routing instructions, scripts, or configuration files (such as configurations generated based on prompt statements).
[0077] For example, the `prompt` statement can be used to customize dynamic routes. After determining the edge-to-edge network mode and the target edge node, the `prompt` statement will automatically expand to generate dynamic route statements for the VPC instances of the target user under the edge node, which will be inserted into the database to obtain route configuration information. For instance, when generating routes for multiple VPC instances under the same edge cluster, routes can be generated for VPCs to communicate with each other via edge-to-edge network instances, as well as many-to-many routes between different router instances.
[0078] Step S208: Distribute the edge forwarding router to the target edge node and configure the routing policy of the target edge node according to the routing configuration information.
[0079] Edge forwarding routers can include software-defined network elements or virtualized network function (VNF) instances running on edge computing nodes, and can be logical routers existing in the form of containers, virtual machines, or Pods. Edge forwarding routers can be used to provide traffic ingress and egress services for specific virtual private cloud (VPC) instances, and perform packet encapsulation, decapsulation (such as encapsulation), and forwarding operations.
[0080] In practice, sending the edge forwarding router to the target edge node can be a process of scheduling the edge forwarding router to run on the specified target edge node.
[0081] The routing strategy can include a set of rules for transmission paths in the network, such as data forwarding logic established on the data plane based on routing configuration information.
[0082] For example, edge forwarding routers can be adaptively distributed to high-performance target edge nodes in the target cluster through a deep reinforcement learning network, enabling multiple nodes to work together, reducing the burden on individual nodes, and adapting to changes in the network and tasks.
[0083] In the aforementioned network scheduling method, by comprehensively considering the real-time network status of each node within the edge cluster and specific business requirements, the optimal edge-to-edge network mode is intelligently determined and matched. Then, by combining cluster performance description information, the target cluster with the best resource status is accurately selected. Subsequently, routing configuration information adapted to the target user's virtual private cloud instance is automatically generated for the selected target edge nodes. The edge forwarding router is then directly deployed to the edge nodes with the best execution capabilities and the routing strategy matching the routing configuration information is applied. Through the hierarchical and automated network scheduling process from network mode matching and cluster selection to node scheduling, the collaborative scheduling of computing nodes in the edge environment is realized, improving network performance and the efficiency of automatic network routing deployment.
[0084] In another embodiment, the target cluster is determined from each edge cluster based on the cluster performance description information of each edge node within the edge cluster and the edge-to-edge network pattern matching the edge cluster. This includes: determining the cluster performance quantification value of the edge cluster based on the cluster performance description information of each edge node within the edge cluster; determining the network weight of the edge cluster based on the edge-to-edge network pattern matching the edge cluster; determining the comprehensive score value of the edge cluster based on the network weight and the cluster performance quantification value; and determining the target cluster from each edge cluster based on the comprehensive score value.
[0085] Among them, the cluster performance quantification value includes a standardized numerical indicator that transforms the multi-dimensional cluster performance description information of each edge node in the edge cluster into a standardized numerical indicator.
[0086] The network weight can include a priority coefficient set for different edge-to-edge network modes. This coefficient reflects the differences in transmission quality, coverage, or service level agreements among the different network modes.
[0087] In practical implementation, cluster performance descriptions such as computing performance, network performance, and bandwidth utilization of edge nodes can be standardized (e.g., Min-Max normalization). Each edge node is then scored based on pre-defined business requirements; for example, computing performance is prioritized for compute-intensive services, while bandwidth utilization is prioritized for network-intensive services. The average score, weighted sum, or average score of the top-N best nodes within the edge cluster can then be used as the quantified cluster performance value. Next, network weights are assigned to the edge cluster according to the edge-to-edge network pattern. Finally, the quantified cluster performance values are weighted and summed based on these network weights to calculate the comprehensive score of the edge cluster. The edge cluster with the highest comprehensive score is selected as the target cluster. Therefore, by calculating the computing performance, network performance, and bandwidth utilization of different edge nodes, different edge-to-edge network patterns and edge clusters can be recommended based on different business needs, improving network performance in the edge computing environment.
[0088] The technical solution in this embodiment improves network performance by automatically recommending and providing users with the cluster resources that have the best overall performance in the current environment, thus avoiding the problem of network performance degradation caused by users selecting the wrong network mode and node.
[0089] In another embodiment, the edge-to-edge network mode includes metropolitan area network mode, CN2 mode, and 163 mode; the network weight of CN2 mode is higher than that of 163 mode; the network weight of 163 mode is higher than that of metropolitan area network mode.
[0090] The CN2 mode can include a high-quality backbone network connection mode built on a next-generation bearer network architecture. CN2 mode is mainly used for inter-provincial cluster interconnection, implemented based on Virtual Private Network (VPN) and SRv6 technology. This mode features extremely high transmission stability, low latency, and high QoS guarantees, and therefore receives the highest priority in the network weight evaluation system, making it suitable for critical services that are extremely sensitive to network quality.
[0091] The 163 mode can include standard network connection modes built on traditional public computer internet. This mode is implemented by dividing public network segments and binding them with IPv6 addresses, typically associated with specific VPN leased lines. This mode has wide coverage and abundant bandwidth resources, but it is slightly inferior to the CN2 mode in terms of congestion control and latency jitter. Therefore, the network weight of the 163 mode is lower than that of the CN2 mode.
[0092] The metropolitan area network (MAN) mode can include a local network connection mode built on a local area network (LAN) within a city or region. This mode is mainly used for cluster interconnection within a province, implemented by binding and dividing v6 segments and encapsulating them at the Overlay layer using SRv6. Since this mode mainly serves local traffic and lacks the advantages of wide-area cross-backbone network transmission, or is considered a basic service layer in the global optimization logic of this scheduling strategy, it has the lowest network weight.
[0093] In this embodiment of the application, by implementing three modes of edge-to-edge networks and setting corresponding network weights for each edge-to-edge network mode, the high reliability and network quality of the target cluster can be ensured.
[0094] In another embodiment, based on the edge-to-edge network pattern matching the target cluster and the target edge node in the target cluster, routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node is generated, including: expanding the prompt routing statement based on the edge-to-edge network pattern matching the target cluster and the target edge node in the target cluster; and generating routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node based on the expanded prompt routing statement.
[0095] The prompt routing statement can be a predefined, abstract routing policy template or intent description instruction, containing the basic logical structure of the routing rules (such as source address, destination address, and logical placeholders for the next hop), but without specific parameters yet populated. The prompt routing statement is used to define the intent for network interconnection.
[0096] In practice, extending the prompt routing statement can be achieved by injecting specific contextual parameters into the abstract prompt routing statement, thus instantiating it into concrete instructions. These parameters originate from the selected target edge node (providing physical location / IP information) and the edge-to-edge network mode (providing protocol type / encapsulation rules). Through this process, vague intentions are transformed into precise technical parameters.
[0097] The routing configuration information can be network configuration data that is device-recognizable and executable, generated based on the extended Prompt statement. Specifically, the extended Prompt routing statement can be parsed and converted into a configuration format that edge forwarding routers can read to obtain the routing configuration information.
[0098] The technical solution of this application embodiment utilizes the automatic expansion of the promt statement to generate dynamic routes, which cleverly solves the problem that static configuration cannot adapt to the dynamic changes of edge nodes, realizes automated network orchestration, and improves the efficiency of automatic network routing.
[0099] In another embodiment, before generating routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node based on the edge-to-edge network pattern matched with the target cluster and the target edge node in the target cluster, the method further includes: deploying a pre-trained deep reinforcement learning model to each edge node in the target cluster; inputting the node environment state information of each edge node in the target cluster into the pre-trained deep reinforcement learning model to obtain the task scheduling strategy of each edge node; and determining the target edge node from each edge node of the target cluster according to the task scheduling strategy.
[0100] Deep reinforcement learning models can include decision-making models built on deep neural networks, such as Deep Q-Network (DQN) models. Pre-trained deep reinforcement learning models have undergone initial iterative training on a central server or in the cloud using a large amount of historical data. These models may have undergone lightweight processing (such as model pruning and quantization) to adapt to the limited computing resources of edge nodes. The models can establish a mapping relationship between environmental states and optimal actions (scheduling decisions).
[0101] The node environment status information may include node status (CPU, memory, disk utilization, etc.), network status (bandwidth, latency, packet loss rate, etc.) and task information (task type, size, priority, etc.).
[0102] The task scheduling policy can be the action instructions output by the model based on the node's environmental state information. In the DQN architecture, it corresponds to the selection result of the action space, i.e., determining which specific edge node or combination of nodes (such as cooperative execution) to assign the current task or edge forwarding router to. The target edge node determined by the task scheduling policy is the edge node indicated in the task scheduling policy to execute the task. For example, if the task scheduling policy includes "assign the task to edge node A", then edge node A is the target edge node. After determining the target edge node, dedicated routing configuration information can be generated for that target edge node and distributed to the edge forwarding router.
[0103] For the convenience of those skilled in the art, Figure 4 An exemplary flowchart of a deep reinforcement learning model is provided.
[0104] In practical applications, the monitoring and acquisition module of the network scheduling system can collect information such as node resource utilization, network status, and task queues in real time. Then, a state space, action space, and reward function are defined. The state space includes the current network status (bandwidth, latency, packet loss rate, etc.), node status (CPU, memory, disk utilization, etc.), and task information (type, size, priority, etc.). The action space includes which node or combination of nodes to assign the task to. The reward function is designed based on factors such as task completion time, resource utilization, and network load to encourage DQN to learn efficient scheduling strategies. Positive rewards include: task completion within the deadline: +R1; node load balancing (e.g., reduced CPU utilization variance): +R2; network latency optimization: +R3. Negative penalties include: task timeout: -P; node overload (CPU > 90%): -P2. Reward = α*(R1) + β*(R2) + γ*(R3) - [δ*(P1) + ε*(P2) + ζ*(P3)], α, β, γ, δ, ε, ζ are weight coefficients.
[0105] Then, training data can be collected. Sources of training data can include real-time node resource utilization (CPU, memory, disk, etc.), network status (bandwidth, latency, packet loss rate, etc.), task queues (task type, size, priority, etc.), and historical scheduling records (task allocation nodes, execution results, execution time, etc.). Next, the raw data can be standardized (e.g., Min-Max normalization) to eliminate dimensional differences. A time-series dataset is constructed, aggregating state information within a fixed time window (e.g., 5 seconds) to form a quadruple (s, a, r, s') of state-action-reward-new state. Gaussian noise is then added to simulate network fluctuations, resulting in processed training data. This data augmentation enhances the model's robustness.
[0106] The DQN model structure can include an input layer, hidden layers, and an output layer. The input layer includes a state space dimension (e.g., 10 dimensions: CPU, memory, latency, etc.); the hidden layer consists of a 3-layer fully connected network (256 -> 128 -> 64 nodes) and a ReLU activation function; the output layer includes an action space dimension (e.g., assigning node 1, node 2, and cooperative execution).
[0107] The initialization process of the DQN model may include: creating a main network (Q-network) and a target network with identical structures, and initializing a central experience replay buffer with a capacity of 10,000 samples. Then, the iterative training process of the DQN model may include the following steps: Step 1, selecting actions based on an ε-greedy policy (initial ε = 0.9, decaying by 0.99 per round); Step 2, executing the action, obtaining the reward r and the new state s', and storing (s, a, r, s') in the replay buffer; Step 3, randomly sampling mini-batch data from the replay buffer (batch size = 64); Step 4, calculating the target Q-value. Where γ=0.95 is the discount factor, and θ⁻ is the target network parameter; Step 5, minimize the loss function (mean squared error) through gradient descent: The training termination conditions for the DQN model include: the average reward fluctuation is less than ±2% for 100 consecutive rounds, and / or the task timeout rate drops below 5% and the node overload rate is <3%.
[0108] In the state space design of the DQN model, by fusing multi-dimensional real-time data (such as network latency and node CPU), a globally aware state vector is constructed, which avoids the scheduling deviation problem caused by the lack of global information in traditional SDWAN. For example, when bandwidth utilization > 80% and CPU utilization > 70% are detected, the state vector triggers a high load warning.
[0109] In the action space design of the DQN model, by expanding the action to multi-node collaboration (such as node 1 and node 2), and by splitting tasks through parallel computing, the bottleneck of traditional single-node processing capacity can be avoided. Compared with traditional single-node allocation, the task processing latency is reduced.
[0110] In the design of the reward function for the DQN model, coefficients (α, β, γ...) can be automatically adjusted according to the business type. For example, Meanwhile, an overload accumulation penalty can be introduced for the penalty term of the reward function. Exponentially suppresses repetitive overload.
[0111] Then, the implementation process of the dynamic task scheduling strategy may include: deploying the pre-trained DQN model to the edge nodes, collecting the node environment state information in real time, inputting it into the DQN model, obtaining the task scheduling strategy, executing task allocation according to the task scheduling strategy, and storing new state, action, reward and other experience data in the central experience replay pool.
[0112] Specifically, the model deployment process may include: lightweighting the pre-trained DQN model (model pruning and quantization) and deploying it to edge nodes; designing state preprocessing microservices to standardize input data in real time; and setting up a model update mechanism to synchronize new parameters from the central controller every 24 hours.
[0113] Specifically, the update process of the central experience replay pool may include: after the edge nodes perform scheduling actions, they generate new experiences (s, a, r, s'); after the local cache is full of 50 records, they are encrypted and uploaded to the central experience replay pool; the central pool samples according to priority (high-reward samples have a 3x increased weight). This avoids model oscillations caused by data correlation, and the distributed sample collection accelerates model iteration and training efficiency.
[0114] The technical solution of this application embodiment adaptively distributes edge forwarding routers to high-performance cluster nodes through a deep reinforcement learning model, which can dynamically adapt to network fluctuations and load changes, thereby improving network response speed and resource utilization.
[0115] In other embodiments, after determining the target edge nodes from each edge node of the target cluster according to the task scheduling strategy, the method further includes: uploading the experience data of each target edge node after executing the task scheduling strategy to the central experience replay pool; the experience data includes at least the executed action, reward value, and new state; selecting target experience data from the central experience replay pool according to the reward value of each experience data as training samples; using the training samples to update and train the pre-trained deep reinforcement learning model, and deploying the updated deep reinforcement learning model to each edge node in the target cluster.
[0116] The empirical data may include the feedback quadruple generated by the edge node after executing a complete scheduling action, such as (s, a, r, s') mentioned above. Here, s represents the state before execution, a represents the action executed (i.e., which node was selected), r represents the reward value obtained after executing the action (calculated based on task completion time, resource utilization, etc.), and s' represents the new state after execution.
[0117] The central experience replay pool can include large-capacity storage space deployed on a central controller or in the cloud, used to aggregate distributed experience data from different edge clusters and nodes. Hybrid storage of global data can prevent the model from getting trapped in local optima.
[0118] The target experience data is a subset of samples with higher training value selected from the central experience replay pool. For example, it may include experience data with higher reward values. Experience data with higher reward values will be given a higher sampling probability during training.
[0119] As an example, uploading the experience data after each target edge node executes the task scheduling strategy to the central experience replay pool may include: storing the experience data after the target edge node executes the task scheduling strategy in the local cache of each edge node, and uploading the experience data to the central experience replay pool when the amount of data in the local cache reaches a preset threshold.
[0120] In the implementation, each target edge node can serve as an execution unit for the deep reinforcement learning model, performing task scheduling locally according to the current policy. After each scheduling, the node calculates the reward function, thereby generating an empirical data point. To reduce network overhead, the empirical data is first stored in the node's local cache. When the local cache accumulates to a certain amount, the data can be encrypted and uploaded in batches to the central experience replay pool, achieving efficient collection of distributed samples. Then, data is periodically extracted from the central experience replay pool for model training. To improve training efficiency, a centralized priority sampling strategy can be adopted. Specifically, the sampling weight is set according to the reward value of each piece of experience data. The higher the reward value, the more successful the scheduling is, and the higher the probability that the corresponding experience data will be selected as the target experience data. This allows the model to prioritize learning high-return, high-quality strategies, thereby accelerating the model's convergence speed. Next, the selected target experience data can be used as training samples to update the pre-trained deep reinforcement learning model using gradient descent, resulting in an updated deep reinforcement learning model. The updated deep reinforcement learning model is then deployed to each edge node in the target cluster, allowing each edge node to load the new model parameters and thus enabling the generation of scheduling strategies optimized through real-time experience.
[0121] The technical solution of this application embodiment constructs a dynamic experience replay mechanism and combines distributed collection of state data from edge nodes with centralized priority sampling to solve the problem of low training efficiency of traditional DQN in edge environments, improves the model iteration speed, and ensures the real-time adaptability of the scheduling strategy.
[0122] In another embodiment, such as Figure 5 As shown, a network scheduling method is provided, which can be applied to... Figure 1 Taking the network scheduling system in China as an example, the following steps are included:
[0123] S502 determines the edge-to-edge network mode that matches the edge cluster based on the network status and service requirements of each edge node within the edge cluster.
[0124] S504. Based on the cluster performance description information of each edge node in the edge cluster, determine the cluster performance quantification value of the edge cluster; based on the edge-to-edge network pattern matching the edge cluster, determine the network weight of the edge cluster; based on the network weight and the cluster performance quantification value, determine the comprehensive score value of the edge cluster; based on the comprehensive score value, determine the target cluster from each edge cluster.
[0125] In one embodiment, the edge-to-edge network mode includes metropolitan area network mode, CN2 mode, and 163 mode; the network weight of CN2 mode is higher than that of 163 mode; the network weight of 163 mode is higher than that of metropolitan area network mode.
[0126] S506, deploy the pre-trained deep reinforcement learning model to each edge node in the target cluster; input the node environment state information of each edge node in the target cluster into the pre-trained deep reinforcement learning model to obtain the task scheduling strategy of each edge node; determine the target edge node from each edge node in the target cluster according to the task scheduling strategy.
[0127] S508: Upload the experience data of each target edge node after executing the task scheduling strategy to the central experience replay pool; based on the reward value of each experience data, select target experience data from the central experience replay pool as training samples; use the training samples to update and train the pre-trained deep reinforcement learning model, and deploy the updated deep reinforcement learning model to each edge node in the target cluster.
[0128] The experience data includes at least the actions performed, reward values, and new states.
[0129] S510 expands the prompt routing statement based on the edge-to-edge network pattern matching the target cluster and the target edge node in the target cluster; based on the expanded prompt routing statement, it generates the routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node.
[0130] Among them, the target edge nodes include the edge nodes in the target cluster used to execute tasks.
[0131] S512 distributes the edge forwarding router to the target edge node and configures the routing policy of the target edge node according to the routing configuration information.
[0132] It should be noted that the specific limitations of the above steps can be found in the specific limitations of a network scheduling method described above.
[0133] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0134] Based on the same inventive concept, this application also provides a network scheduling apparatus for implementing the network scheduling method described above. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more network scheduling apparatus embodiments provided below can be found in the limitations of the network scheduling method described above, and will not be repeated here.
[0135] In one exemplary embodiment, such as Figure 6 As shown, a network scheduling device is provided, comprising:
[0136] The mode determination module 610 is used to determine the edge-to-edge network mode that matches the edge cluster based on the network status and service requirements of each edge node in the edge cluster.
[0137] The cluster determination module 620 is used to determine the target cluster from each edge cluster based on the cluster performance description information of each edge node in the edge cluster and the edge-to-edge network pattern that matches the edge cluster.
[0138] The information generation module 630 is used to generate routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node according to the edge-to-edge network mode matched with the target cluster and the target edge node in the target cluster; the target edge node includes the edge node in the target cluster used to perform tasks.
[0139] The routing configuration module 640 is used to distribute the edge forwarding router to the target edge node and configure the routing policy of the target edge node according to the routing configuration information.
[0140] In one embodiment, the cluster determination module 620 is specifically configured to: determine the cluster performance quantification value of the edge cluster based on the cluster performance description information of each edge node in the edge cluster; determine the network weight of the edge cluster based on the edge-to-edge network pattern matching the edge cluster; determine the comprehensive score value of the edge cluster based on the network weight and the cluster performance quantification value; and determine the target cluster from each of the edge clusters based on the comprehensive score value.
[0141] In one embodiment, the edge-to-edge network mode includes a metropolitan area network mode, a CN2 mode, and a 163 mode; the network weight of the CN2 mode is higher than that of the 163 mode; and the network weight of the 163 mode is higher than that of the metropolitan area network mode.
[0142] In one embodiment, the information generation module 630 is specifically used to expand the prompt routing statement according to the edge-to-edge network pattern matching the target cluster and the target edge node in the target cluster; and to generate routing configuration information corresponding to the virtual private cloud instance of the target user under the target edge node according to the expanded prompt routing statement.
[0143] In one embodiment, the network scheduling device further includes a node determination module; the node determination module is specifically used to deploy a pre-trained deep reinforcement learning model to each edge node in the target cluster; input the node environment state information of each edge node in the target cluster into the pre-trained deep reinforcement learning model to obtain the task scheduling strategy of each edge node; and determine the target edge node from each edge node of the target cluster according to the task scheduling strategy.
[0144] In one embodiment, the node determination module is specifically used to upload the experience data of each target edge node after executing the task scheduling strategy to the central experience replay pool; the experience data includes at least the executed action, reward value, and new state; based on the reward value of each experience data, target experience data is selected from the central experience replay pool as training samples; the pre-trained deep reinforcement learning model is updated and trained using the training samples, and the updated deep reinforcement learning model is deployed to each edge node in the target cluster.
[0145] Each module in the aforementioned network scheduling device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0146] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 7 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores node data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a network scheduling method.
[0147] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0148] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0149] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0150] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0151] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0152] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0153] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0154] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A network scheduling method, characterized by, The method comprises: determining an edge-to-edge network mode matched with the edge cluster according to network states and service demands of each edge node in the edge cluster; determining a target cluster from each of the edge clusters according to cluster performance description information of each edge node in the edge cluster and the edge-to-edge network mode matched with the edge cluster; generating routing configuration information corresponding to a virtual private cloud instance of a target user under a target edge node according to the edge-to-edge network mode matched with the target cluster and the target edge node in the target cluster; the target edge node comprises an edge node in the target cluster for executing a task; downloading an edge forwarding router to the target edge node and configuring a routing strategy of the target edge node according to the routing configuration information.
2. The method of claim 1, wherein, The method further comprises: determining a cluster performance quantization value of the edge cluster according to the cluster performance description information of each edge node in the edge cluster; determining a network weight of the edge cluster according to the edge-to-edge network mode matched with the edge cluster; determining a comprehensive score value of the edge cluster according to the network weight and the cluster performance quantization value; determining a target cluster from each of the edge clusters according to the comprehensive score value.
3. The method of claim 2, wherein, The edge-to-edge network mode comprises a metropolitan area network mode, a cn2 mode and a 163 mode; the network weight of the cn2 mode is higher than that of the 163 mode; the network weight of the 163 mode is higher than that of the metropolitan area network mode.
4. The method of claim 1, wherein, The method further comprises: extending a prompt routing statement according to the edge-to-edge network mode matched with the target cluster and the target edge node in the target cluster; generating routing configuration information corresponding to a virtual private cloud instance of a target user under the target edge node according to the extended prompt routing statement.
5. The method of claim 1, wherein, The method further comprises: deploying a pre-trained deep reinforcement learning model to each edge node in the target cluster; inputting node environment state information of each edge node in the target cluster into the pre-trained deep reinforcement learning model to obtain a task scheduling strategy of each edge node; determining the target edge node from each edge node in the target cluster according to the task scheduling strategy.
6. The method of claim 5, wherein, The method further comprises: determining a target edge node from each edge node in the target cluster according to the task scheduling strategy. Uploading experience data of each target edge node after performing the task scheduling strategy to a central experience replay pool; the experience data at least includes an executed action, a reward value and a new state; According to the reward value of each experience data, target experience data is screened from the central experience replay pool as a training sample; The pre-trained deep reinforcement learning model is updated and trained by using the training sample, and the deep reinforcement learning model after the update and training is deployed to each edge node in the target cluster.
7. A network scheduling apparatus characterized by comprising: The apparatus comprises: A mode determination module configured to determine an edge-edge network mode matched with the edge cluster according to network states and service demands of each edge node in the edge cluster; A cluster determination module configured to determine a target cluster from each edge cluster according to cluster performance description information of each edge node in the edge cluster and the edge-edge network mode matched with the edge cluster; An information generation module configured to generate routing configuration information corresponding to a virtual private cloud instance of a target user under a target edge node according to the edge-edge network mode matched with the target cluster and the target edge node in the target cluster; the target edge node includes an edge node in the target cluster for performing a task; A routing configuration module configured to issue an edge forwarding router to the target edge node and configure a routing strategy of the target edge node according to the routing configuration information.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 6.