Multi-dimensional computing power dynamic perception routing decision-making method and system based on SRv6 driving
Through the SRv6-based multi-dimensional computing power dynamic perception routing decision-making method, the shortcomings of the existing SRv6 computing power scheduling in real-time, flexibility and cross-domain collaboration are solved, efficient and intelligent path control and resource allocation are achieved, and resource utilization and business flow scheduling efficiency in cross-cloud and cloud-edge environments are improved.
Patent Information
- Application Number
- CN202510988052.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-03
AI Technical Summary
The existing SRv6 computing power scheduling mechanism has deficiencies in real-time performance, flexibility, multi-dimensional perception, and cross-domain collaboration, making it difficult to achieve efficient and intelligent path control and resource allocation. Especially in cross-cloud and cloud-edge collaborative environments, insufficient perception of network dynamic changes and computing power status leads to delayed path selection, low resource utilization, scheduling decision deviations, and resource island effects.
A multi-dimensional computing power dynamic perception routing decision method based on SRv6 is adopted. By dynamically collecting network and computing node resources, an undirected graph model is constructed. Combined with the improved SPFA algorithm and SRv6 path programmability, the optimal path and target node SID chain are generated to achieve cross-domain resource collaboration and closed-loop scheduling, supporting real-time perception of multi-dimensional resource indicators and path optimization.
It improves the dynamic perception and response capabilities of computing power scheduling, realizes path selection driven by multiple factors, supports collaborative resource scheduling across cloud and cloud-edge environments, enhances the network's adaptability and resource utilization, and ensures efficient forwarding of service requests and intelligent scheduling of business flows.
Smart Images

Figure CN120750832A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computing power scheduling and relates to a multi-dimensional computing power dynamic perception routing decision method and system based on SRv6 drive. Background Art
[0002] With the rapid development of next-generation information technology, cloud computing, edge computing, and artificial intelligence (AI) application scenarios are becoming increasingly diverse. Distributed computing resources are becoming highly heterogeneous and dynamic in terms of spatial distribution and business needs. In this context, achieving efficient computing resource scheduling and dynamic routing across regions, platforms, and networks has become a key issue in improving overall system performance and service quality.
[0003] Traditional computing resource scheduling methods rely primarily on static configurations and centralized container orchestration tools such as Kubernetes and Docker Swarm. While these technologies have achieved promising results in data centers and local cluster environments, they face numerous challenges in wide-area network scenarios such as cross-cloud and cloud-edge collaboration. These challenges include dynamic network changes, path selection insensitivity to computing power status, uneven node load, and insufficient coordination between computing and network resources.
[0004] Especially in wide-area heterogeneous environments, computing resources are not only constrained by the capacity of local computing nodes, but also by multiple factors such as network path accessibility, latency, bandwidth, and link stability. Therefore, resource scheduling methods that rely solely on static path configuration or a single computing power metric can no longer meet the needs of efficient and flexible service scheduling.
[0005] As a new network programming architecture, Segment Routing over IPv6 (SRv6) technology, with its powerful path expression and service chain orchestration capabilities, makes it possible to build a programmable network control system. By defining an extended Segment ID (SID), SRv6 can not only accurately control the forwarding path of data packets, but also carry more semantic information, such as service calls, computing task identification, policy triggering, etc. However, most existing methods are based on static SID configuration and lack the ability to perceive the network status and computing power status in real time, making it difficult to achieve dynamic and efficient resource scheduling strategy generation. In practical applications for multi-dimensional computing power perception and dynamic routing decision-making, there are still the following key deficiencies:
[0006] 1. Limited dynamic perception and response capabilities
[0007] Most current SRv6-based computing scheduling solutions rely on predefined paths and rules, lacking real-time awareness of computing and network status. When node loads, link status, or service requests rapidly change, existing control systems struggle to detect and respond to these dynamic changes. This results in delayed path selection, low computing resource utilization, and difficulty ensuring quality of service.
[0008] 2. Path decision-making lacks a multi-dimensional integration mechanism
[0009] Existing routing algorithms often select paths based on single-dimensional metrics (such as bandwidth, latency, or node CPU utilization), lacking the ability to integrate and jointly optimize multiple key resource metrics (such as GPU utilization, memory usage, service load, and link jitter). This lack of comprehensive consideration often leads to scheduling decision bias, affecting global optimality.
[0010] 3. The selection strategy for computing nodes lacks flexibility and adaptability
[0011] In complex business environments, the computing power requirements requested by users are often non-deterministic and sudden. Existing methods mostly select nodes based on static allocation or rule matching, and cannot flexibly reselect based on the current system status, making it difficult to meet the scheduling requirements of high-concurrency, low-latency businesses.
[0012] 4. Insufficient cross-regional resource coordination capabilities
[0013] Faced with a network environment that is deployed across clouds, cloud-edge, or multiple regions, the existing scheduling mechanism lacks a unified resource view and scheduling interface between different network domains, resulting in an island effect in resource scheduling. Computing resources cannot flow efficiently on a global scale, which in turn limits the scheduling elasticity and service reliability of the overall system.
[0014] 5. Lack of efficient and scalable SID semantic resolution mechanism
[0015] Although SIDs, as key programming elements in SRv6, possess strong expressive power, their scalability and semantic carrying capacity remain insufficient in current systems. Computing gateways are prone to performance bottlenecks when parsing complex SID chains. This is especially true when orchestrating complex service chains. SID parsing accuracy and processing efficiency are difficult to guarantee, impacting the reliability and real-time nature of routing decisions.
[0016] 6. Lack of a closed-loop mechanism between resource perception and strategy generation
[0017] Most existing systems only passively collect computing power status and lack a closed-loop scheduling mechanism that uses real-time perception to adjust paths and provide policy feedback. This disconnects policy generation from actual network status, making continuous optimization and adaptive evolution difficult.
[0018] In summary, the current SRv6-based computing power scheduling and routing mechanisms still have significant shortcomings in terms of dynamic multi-dimensional resource perception, real-time path adjustment, cross-domain collaborative scheduling, and efficient SID resolution. These technical shortcomings limit the wider and more intelligent application of SRv6 in computing power networks.
[0019] Based on this, the present invention proposes a multi-dimensional computing power dynamic perception routing decision method based on SRv6 drive. By constructing an intelligent routing decision mechanism that integrates network status, computing power resource status and multi-dimensional dynamic characteristics, dynamically collecting node and link status information, introducing multi-dimensional perception indicators (such as node computing load, link delay, real-time bandwidth, etc.), and combining the programmable characteristics of SRv6 paths, a joint calculation method is used to generate the optimal path and target node SID chain, realizing the joint optimization of path selection and computing power allocation, thereby realizing efficient forwarding of service requests from the source end to the optimal computing power node. This method not only improves the response efficiency and scheduling quality of service requests, but also enhances the network's ability to adapt to sudden traffic and resource changes, and has broad application prospects. Summary of the Invention
[0020] The present invention aims to solve the key technical bottlenecks of the existing SRv6 computing power scheduling mechanism in terms of real-time performance, flexibility, multi-dimensional perception and cross-domain collaboration, and proposes a multi-dimensional computing power dynamic perception routing decision method based on SRv6 drive to achieve more intelligent, dynamic and adaptive path control and resource allocation in the computing power network environment.
[0021] Specifically, the present invention adopts the following technical solutions:
[0022] A multi-dimensional computing power dynamic perception routing decision method based on SRv6, including:
[0023] S1. Dynamically collect network resource information and computing node resource usage from the cluster environment, and aggregate and persist the acquired data.
[0024] S2. Model the information obtained in S1 as an undirected graph. In this undirected graph, edges represent network status, described by network resources, and nodes represent computing nodes, described by computing power resources. At the same time, the network bandwidth and latency indicators are uniformly converted into edge weights to obtain a dynamic network topology based on converged service sensitivity.
[0025] S3: Based on the resource matching scoring model between business requirements and node real-time load, the optimal computing power node is selected as the target computing power node. Then, based on the dynamic network topology diagram that integrates business sensitivity, the improved SPFA algorithm is used to calculate the optimal network path to the target computing power node.
[0026] S4, based on SRv6-driven dynamic generation and closed-loop execution of service chains, abstracts the target computing nodes and network functions into programmable SIDs, dynamically builds service chains and encapsulates them in the SRH header, implements the execution path, monitors the links, and makes timely adjustments after node failures.
[0027] In the above technical solution, further, in S1, the network resource status including bandwidth and latency information is collected periodically; the collection of computing node resources includes deploying the collector on all nodes in the cluster, and the periodic collection includes key indicators such as CPU utilization, GPU utilization, memory occupancy, disk I / O, network link bandwidth, and latency.
[0028] Furthermore, all collected resource data are aggregated and converted into a structured resource state vector to describe the computing power and network status of the node. The resource state vector is specifically: the computing power of the node:
[0029] S(n i )=[n i (CPU),n i (GPU),n i (Mem),n i (IO)]
[0030] Among them, n i represents the i-th node in the cluster, S(n i ) represents the computing resource set of the i-th node, n i (CPU) represents the CPU resource usage of the i-th node, n i (GPU) represents the GPU resource usage of the i-th node, n i (Mem) represents the memory resource usage of the i-th node, n i (IO) represents the I / O rate of the i-th node; the network status of the node:
[0031] Net(n i ,n j )=[BW(n i ,n j ),RTT(n i ,n j )]
[0032] Among them, n i 、n j Represents the i-th and j-th nodes in the cluster, n i 、n j Satisfy the direct adjacent condition, Net(n i ,n j ) represents n i 、n jThe network status between two nodes, BW(n i ,n j ) represents n i 、n j Bandwidth utilization of network links between nodes, RTT (n i ,n j ) represents n i 、n j The latency of the network link between nodes.
[0033] Furthermore, in the undirected graph G(V,E) in S2, the vertex V represents the computing node, the edge E represents the network link between the nodes, and the attributes of the edge are bandwidth RTT and delay BW. The ratio of the delay from node u to v to the maximum delay is multiplied by an adjustable weight coefficient β1, and the ratio of the delay from node u to v to the maximum bandwidth is multiplied by another adjustable weight coefficient β2. The sum of the two is taken to uniformly convert the two indicators of bandwidth and delay into edge weights, where β1+β2=1, which supports dynamic adjustment based on the service sensitivity to delay / bandwidth.
[0034] Furthermore, the resource matching scoring model is specifically as follows: for each request, the resource status of the request is learned from the request details, and a score is calculated based on the resource status of the request and the real-time load status of the current node. In the resource set to be considered, the ratio of the application status of the request for resource r to the remaining resource amount of resource r on the current node is taken. The sum of the above ratios of all resources in the resource set is the score of the current node.
[0035] Furthermore, all computing nodes are sorted according to their scores, and the computing node with the highest score is the optimal computing power node and serves as the target computing power node.
[0036] Furthermore, the improved SPFA algorithm is specifically as follows: an undirected graph, a source node, and a target computing power node are taken as input, and each edge in the undirected graph has a non-negative edge weight; in the initialization stage, the path distance of all nodes is configured to be infinite, the queue status is not queued, the predecessor node is empty, the queue counter is reset to zero, the source node distance is set to zero and the queue mark is added; in the main loop stage, when the queue is not empty, the node is continuously taken from the head of the queue and its queue mark is cleared, its adjacent nodes are traversed and the current edge weight is dynamically calculated. If the path cost from the current node to the adjacent node is better, the path distance of the adjacent node and the predecessor node are updated. When it is detected that the adjacent node is not queued, it is added to the end of the queue and the status and cumulative queue counter are updated. The termination is triggered when the number of queues exceeds the threshold; when the queue is empty, the algorithm terminates and outputs the minimum path distance of each node and the optimal path reconstructed based on the predecessor node. The dynamic adaptability of the network and the robustness of the algorithm are achieved through dynamic weight calculation and queue count monitoring.
[0037] Furthermore, S4 uses an SDN-based controller to separate the data plane and control plane, implements routing management and path orchestration based on the pyroute2 library, uses a RESTful API for communication between the controller and router nodes, and manages SRv6 paths to achieve real-time dynamic routing planning. Based on S3's optimal network path routing decisions, the controller dynamically adjusts the path in real time to optimize traffic transmission, enabling flexible service programming and path optimization.
[0038] A multi-dimensional computing power dynamic perception routing decision system based on SRv6, which implements the method described in any of the above items, including: a dynamic computing network resource perception module, which is used to dynamically collect network link status and computing node resource usage from a cluster environment, and summarize and persist the acquired data; a computing network resource multi-dimensional decision algorithm module, which is used to route computing power requests to the optimal node through a dynamically generated computing power scheduling strategy based on the information collected by the dynamic computing network resource perception module, and dynamically adjust the path according to the real-time status of the computing network resources to ensure that the optimal route is selected according to changes in network status;
[0039] The SRv6 controller module is used to dynamically adjust paths in real time based on the routing decisions made by the computing network resource multi-dimensional decision-making algorithm module to optimize traffic transmission, enabling flexible service programming and path optimization.
[0040] A computer-readable storage medium stores computer-executable instructions, wherein the instructions are used to implement any of the methods described above when executed.
[0041] The beneficial effects of the present invention are as follows:
[0042] The present invention adopts a core method of decoupling and coordinating the selection of computing power nodes and network path optimization. First, the optimal target computing power node is selected based on the resource matching scoring model of business demand and node real-time load; then, based on the dynamic network topology diagram that integrates business sensitivity (delay / bandwidth preference), the optimized shortest path algorithm (improved SPFA) is used to calculate the optimal network path to the computing power node, realizing the global collaborative decision of "selecting the most appropriate computing point" and "finding the most efficient connection path". At the same time, based on the SRv6-driven service chain dynamic generation and closed-loop execution method, the target computing power node and necessary network functions are abstracted into programmable SIDs, and the SID sequence (service chain) is dynamically constructed according to the two-stage decision results and encapsulated in the SRH header. The programmable controller realizes second-level path distribution and status monitoring, and can have built-in version control and abnormal rollback strategies to form a closed-loop self-healing path control mechanism of "perception-driven decision-making->decision generation SID chain->execution and feedback monitoring", ensuring intelligent and reliable scheduling of cross-domain end-to-end business flows.
[0043] The method based on the present invention can achieve the following core goals:
[0044] 1. Improve the dynamic perception and responsiveness of computing power scheduling
[0045] By introducing SRv6's network programmable mechanism and multi-dimensional resource real-time perception framework, the present invention can continuously track the status of network and computing resources (such as latency, bandwidth, load, GPU utilization, etc.), and quickly generate adaptive path decisions when the business changes, thereby achieving high real-time and high adaptability in computing resource scheduling.
[0046] 2. Build a path selection mechanism driven by multiple factors
[0047] The present invention integrates network indicators and computing resource indicators and proposes a joint decision-making mechanism of "computing power + network". It not only considers the shortest path or the optimal bandwidth, but also conducts a global evaluation based on multi-dimensional factors such as the available computing power and load level of the target node to ensure that the selection of routing paths and computing power nodes is more comprehensive and reasonable.
[0048] 3. Support resource collaborative scheduling across cloud and cloud-edge environments
[0049] Faced with a wide-area computing network with multi-region and multi-type node deployment, the present invention designs a scalable SRv6 path generation and collaborative optimization mechanism, combined with the dynamic SID encoding and decoding and computing power status synchronization capabilities of the computing power gateway, to achieve intelligent distribution of computing power tasks and elastic resource coordination between different network domains.
[0050] 4. Realize SID chain orchestration and service chain dynamic mapping
[0051] Based on the programmability of SRv6, the present invention supports abstracting different service modules into functional SIDs, and uses SID chains to realize dynamic jump and customized orchestration of service flows between multiple service nodes, thereby decoupling the computing power scheduling process from the service path configuration, and enhancing the scalability and orchestration flexibility of the service.
[0052] 5. Build a closed-loop perception-decision-execution mechanism
[0053] The present invention forms a closed-loop control process including data collection, path calculation, strategy generation and path distribution, ensuring that the routing adjustment process is highly coordinated with the system status, and realizing the continuous optimization and intelligent evolution of the computing power network.
[0054] In summary, the present invention introduces the dynamic programming capability driven by SRv6, integrates multi-dimensional computing power and network indicators, realizes the deep coupling and dynamic optimization of computing resources and path decisions, solves the problems of poor real-time performance, single dimension, and weak collaboration capability in the existing technology, and can promote the efficient, stable and intelligent development of future-oriented distributed intelligent computing network systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 Schematic diagram of the overall process of the method of the present invention;
[0056] Figure 2 Schematic diagram of the process of improving the SPFA algorithm in the present invention;
[0057] Figure 3 This is a schematic diagram of the cross-cloud SDN control plane design in the present invention;
[0058] Figure 4 This is a diagram showing the recovery effect of the present invention after a link failure is detected. DETAILED DESCRIPTION
[0059] The following describes the specific implementation of the embodiment of the present invention in detail with reference to the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the embodiment of the present invention and is not used to limit the embodiment of the present invention.
[0060] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features therein may be combined with each other.
[0061] According to a specific embodiment of the present invention, a multi-dimensional computing power dynamic perception routing decision method based on SRv6 driving of the present invention is as follows: Figure 1 As shown, the following steps are included:
[0062] S1. Dynamically collect network resource information and computing node resource usage from the cluster environment, and aggregate and persist the acquired data.
[0063] S2. Model the information obtained in S1 as an undirected graph. In this undirected graph, edges represent network status, described by network resources, and nodes represent computing nodes, described by computing power resources. At the same time, the network bandwidth and latency indicators are uniformly converted into edge weights to obtain a dynamic network topology based on converged service sensitivity.
[0064] S3: Based on the resource matching scoring model between business requirements and node real-time load, the optimal computing power node is selected as the target computing power node. Then, based on the dynamic network topology diagram that integrates business sensitivity, the improved SPFA algorithm is used to calculate the optimal network path to the target computing power node.
[0065] S4, based on SRv6-driven dynamic generation and closed-loop execution of service chains, abstracts the target computing nodes and network functions into programmable SIDs, dynamically builds service chains and encapsulates them in the SRH header to implement the execution path.
[0066] According to a specific embodiment of the present invention, the implementation process of the method may be as follows:
[0067] 1. Overall algorithm design
[0068] This example proposes a multi-dimensional computing power dynamic perception routing decision method based on SRv6. Figure 3 The SDN architecture shown in the figure consists of three modules.
[0069] Module 1, dynamic computing network resource perception module, is mainly responsible for dynamically collecting network link status (i.e. collecting network resources) and computing node resource usage from the cluster environment, and summarizing and persisting the obtained data.
[0070] The iperf3 tool is used to measure and collect network resources, including bandwidth, latency, and other network information. To achieve node resource collection, a custom collector was written in this example and deployed on all compute nodes in the cluster to periodically collect the required information. The collection targets include key indicators such as CPU utilization, GPU utilization, memory usage, disk I / O, network link bandwidth, and latency.
[0071] After collection, data is sent to the collector, which is responsible for integrating information from all nodes and persisting it. The collected computing network data resources will provide the data foundation for the subsequent routing planning module. The data collection period is configurable, with a default collection interval of every 10 seconds, ensuring real-time data while balancing system performance overhead. After the collected data is aggregated, it is first uniformly labeled. After all computing resources are converted to a unified unit, the node computing power description is abstracted into this unified unit indicator. The data is then persisted to provide data for subsequent routing planning.
[0072] Module 2, the multi-dimensional decision-making algorithm module for computing network resources, is designed to route computing power requests to the optimal node through a dynamically generated computing power scheduling strategy, taking into account dimensions including network conditions and computing power resource conditions. This algorithm first models the network information collected in module 1 into an overall network topology. The network topology is an undirected graph. The edges in the undirected graph represent the network status, which is described by the network resource conditions (bandwidth, latency, and other data). The nodes represent computing nodes, which are described by computing power resources. After the modeling is completed, the optimal path from the source node to the target node is calculated based on the improved SPFA algorithm. In practical applications, the algorithm can dynamically adjust the path according to the real-time status of the computing network resources to ensure that the optimal route is selected according to changes in the network status, while taking into account real-time, fault tolerance, and performance optimization.
[0073] Module three, the SRv6 controller module, is a control plane that utilizes an SDN architecture that separates the data and control planes. This controller implements routing management and path orchestration based on the pyroute2 library. The controller communicates directly with router nodes using a RESTful API and manages SRv6 paths, including creation, querying, and deletion, enabling real-time dynamic routing planning. Based on the routing decisions made in module two, the controller dynamically adjusts paths in real time to optimize traffic flow, enabling flexible service programming and path optimization.
[0074] 2. Design of Dynamic Computing Network Resource Perception Module
[0075] (1) Computing network data collection. The core task of computing network data collection is to periodically obtain the computing resource usage and network link status of all nodes in the cluster environment. This module adopts a distributed deployment method. The system has developed a lightweight collector, which is written in Python and deployed to each node in the cluster. The collector performs resource sampling tasks at a fixed period (default every 10 seconds). Since the collector is self-written, the resources that the collector wants to collect can be flexibly configured. In this method, the information collected by the collector includes but is not limited to the following indicators: CPU usage, memory occupancy, disk I / O status, GPU usage, etc. Then, a central collector will be deployed in the cluster at the same time. The node collector pushes the locally collected data to the central collector through the RESTFUL interface. The collector is responsible for integrating the resource information of all nodes and performing unified persistent storage. In terms of network resource collection, the system integrates the iperf3 network test tool to periodically detect the bandwidth and latency of the links between nodes. All network link test tasks are coordinated and initiated by the central scheduler to ensure the stability and representativeness of the sampling results.
[0076] (2) Data description: All collected resource data are aggregated and converted into a structured resource state vector. The state vector is used to describe the computing power and network status of the node, and is the basic input for subsequent computing power scheduling and routing algorithms. The description of the node's computing power is shown in formula (1):
[0077] S(n i )=[n i (CPU),n i (GPU),n i (Mem),n i (IO)] (1)
[0078] Among them, n i represents the i-th node in the cluster, S(n i ) represents the computing resource set of the i-th node, ni (CPU) represents the CPU resource usage of the i-th node, n i (GPU) represents the GPU resource usage of the i-th node, n i (Mem) represents the memory resource usage of the i-th node, n i (IO) represents the IO resource usage (i.e., IO rate) of the i-th node.
[0079] The description of the network status is shown in formula (2):
[0080] Net(n i ,n j )=[BW(n i ,n j ),RTT(n i ,n j )] (2)
[0081] Among them, n i 、n j Represents the i-th and j-th nodes in the cluster, n i 、n j Satisfy the direct adjacent condition, Net(n i ,n j ) represents n i 、n j The network status between two nodes, BW(n i ,n j ) represents n i 、n j Bandwidth utilization of network links between nodes, RTT (n i ,n j ) represents n i 、n j The latency of the network link between nodes.
[0082] 3. Design of multi-dimensional decision-making algorithm module for computing network resources
[0083] (1) Network topology modeling: Combining the node computing power and network data obtained in the dynamic computing network resource perception module, this method constructs a dynamic network topology with these data. The topology is an undirected graph G(V,E), where the vertex V represents the computing node, the edge E represents the network link between the nodes, and the attributes of the edge are bandwidth (RTT) and latency (BW);
[0084] (2) Edge weight calculation. Different types of requests have different resource requirements. For example, in the file transfer scenario, it is a bandwidth-sensitive business, while in the web server scenario, it is a delay-sensitive business. In order to comprehensively consider the bandwidth and delay of the network, the system will uniformly convert the two indicators into edge weights. The calculation formula of edge weight is shown in formula (3):
[0085]
[0086] Where RTT(u,v) and BW(u,v) are the delay and bandwidth from node u to v respectively. max and BW max They represent the maximum delay and available bandwidth among all links, respectively. β1 and β2 are adjustable weight coefficients that satisfy β1+β2=1. Their specific settings can be adjusted as needed, supporting dynamic adjustment based on the service's sensitivity to delay / bandwidth.
[0087] (3) Node selection. This step mainly selects the node that best suits the request. In this example, a scoring method is proposed to select nodes. For each request, the requested resource status (e.g., request for cpu: 500m) can be obtained from the request details. The final score can be obtained based on the matching between the request and the current node. The score calculation formula is shown in formula (4):
[0088]
[0089] Among them, Score represents the final score of the node, R is the resource set that needs to be considered in this algorithm, and request r Indicates the amount of application for a certain resource r, node r Indicates the remaining amount of resource r on the node. The computing node with the highest score, i.e. the optimal computing power node, is selected as the target node.
[0090] (4) Improved SPFA algorithm. This patent uses an improved version of the SPFA (Shortest Path Faster Algorithm) algorithm to calculate the optimal path for the entire graph. Compared with the traditional Dijkstra algorithm, the improved SPFA is more suitable for sparse graphs. When the network topology changes frequently, it has good performance stability and is particularly suitable for wide-area heterogeneous networks. The algorithm dynamically updates the node path cost through a queue-driven mechanism. Its execution process is as follows:
[0091] During the algorithm's initialization phase, all network nodes are configured with their initial states: path distances are initialized to infinity, queue status is marked as unqueued, predecessor nodes are cleared, and the enqueue counter is reset to zero. The source node's path distance is set to zero, it is added to the processing queue, and its queue status is updated to enqueued. The enqueue counter is incremented.
[0092] The algorithm enters the main loop processing stage. When the processing queue is not empty, the following operations are performed in a loop: a node is taken from the head of the queue and its queue status mark is cleared. All adjacent edges of the node are traversed, and for each adjacent node connected by the edge, the current edge weight value is dynamically calculated according to the real-time network parameters. If the path cost of reaching the adjacent node through the current node is less than the optimal path cost recorded for the adjacent node, the optimal path distance data of the adjacent node is updated, and the current node is recorded as its predecessor node. At the same time, the queue status of the adjacent node is detected: if it is not in the queue to be processed, it is added to the tail of the queue according to the preset queue strategy, the queue status mark is updated, and the queue counter accumulation operation is performed. During this process, the number of times each node is queued is continuously monitored. When it is detected that the number of times any node is queued exceeds the preset threshold, it is determined that there is an abnormal loop risk and the algorithm termination mechanism is triggered.
[0093] The algorithm terminates when the processing queue is empty. At this point, a dataset of minimum path costs from the source node to all reachable nodes has been generated, and the complete optimal path can be reconstructed using the predecessor node's records. This process design adapts to network state changes through dynamic weight calculation and enhances algorithm robustness through enqueue count monitoring, effectively addressing the dynamic adaptability and stability issues of optimal path calculation in wide-area heterogeneous network environments.
[0094] The input of the improved SPFA algorithm is as follows:
[0095] 1) A weighted undirected graph G = (V, E), where V is the set of vertices, i.e., the routing nodes in the cluster, and E is the set of edges, i.e., the connections between nodes in the cluster.
[0096] 2) Each edge (u, v) in the graph has a non-negative weight w(u, v), which represents the path cost from node u to node v, i.e., bandwidth, latency, and other data.
[0097] 3) A source node s∈V and a destination node des∈V
[0098] The outputs of the improved SPFA algorithm include:
[0099] 1) The optimal path distance d(s,v) from the source node s to all other nodes v∈V in the graph.
[0100] 2) The predecessor node array path(v) is used to reconstruct the optimal path from s to v.
[0101] The process of the improved SPFA algorithm is as follows Figure 2 shown below.
[0102] According to a specific example of the present invention, the pseudocode of the improved SPFA is provided as follows:
[0103] 1. Initialization
[0104] For all nodes V in the graph:
[0105] · Set the shortest path estimate: d(v) = +∞
[0106] · Set the queue flag: in_queue[v] ← false
[0107] · Set the predecessor node to null: path[v] ← null
[0108] · Set the queue entry counter: cnt[v] ← 0
[0109] · Set the path estimate of the source node s to 0: d[s] ← 0, where d[x] represents the distance from the source node s to the node x
[0110] · Enqueue the source node s and mark it as in the queue: in_queue[s] ← true
[0111] · Initialize the queue Q: Q ← [s]
[0112] 2. Main loop phase
[0113] When the queue Q is not empty, repeat the following operations: <000If v is not in the queue:
[0123] ·Choose to insert v at the end of the queue
[0124] Mark v in the queue: in_queue[v]←true
[0125] Accumulate the number of times v is queued: cnt[v]←cnt[v]+1
[0126] If cnt[v]>|V|, it means there is a negative cycle and you can choose to terminate the algorithm
[0127] 4. SRv6 controller module design
[0128] Based on SRV6 routing control, the path (v) ultimately derived from the improved SPFA algorithm is the final routing path derived by this invention. This path is the fastest routing path from the current node to the most suitable node. The SRv6 controller module proposed in this invention is built on the principles of SDN and aims to achieve intelligent path scheduling and service orchestration in computing-network collaboration scenarios. By decoupling the control plane from the data plane, this controller module supports dynamic configuration and flexible adjustment at the path and service levels, making it suitable for heterogeneous cloud-edge converged network environments.
[0129] The architecture design of the routing control module based on SRV6 is as follows:
[0130] The controller and network devices use a RESTful API as the southbound interface protocol. This allows the controller to issue path control instructions to SRv6-enabled routing devices and retrieve topology status information from them. Each edge device deploys a lightweight REST server program to receive controller requests and perform operations such as path configuration and status reporting. This interface adheres to standard HTTP semantics, making it easy to extend and integrate.
[0131] Secondly, a path management component is designed within the controller. This component provides functions such as path creation, path deletion, and path query. All path information is organized and managed in a dictionary structure. Each path uses the destination address as the key and stores data objects containing fields such as the SID sequence, encapsulation mode, outbound interface, and priority, enabling fast path lookup and status updates.
[0132] The controller encodes the Segment ID list of the path to be configured into SRH (Segment Routing Header), and calls methods such as IPRoute.route() in pyroute2 to complete operations such as encapsulation, update, and deletion of the path, and binds it to the specified IPv6 prefix.
[0133] To support service function path orchestration, the controller designs a SID mapping strategy that maps specific network functions (such as NAT, DPI, and load balancing) to corresponding SID addresses. During the path planning phase, the required service SIDs are automatically inserted into the SID list, ensuring that service requests are processed sequentially along the service chain. This insertion process is achieved through list splicing, without modifying the underlying network structure.
[0134] At the same time, in order to improve the flexibility and reliability of the system, the controller introduces a path version control mechanism. Each path configuration operation records the version number and timestamp. If there are subsequent abnormal situations such as link interruption or path performance degradation, the controller can restore to the historical version through a rollback strategy to ensure business continuity. In addition, to avoid path drift or inconsistent status, a consistency check mechanism is designed to regularly compare and verify the issued paths to promptly detect and correct differences. Figure 4 ,for Figure 3 The embodiment shown shows the recovery effect after a link failure is detected.
[0135] Taking all of the above into account, the SRv6 controller module not only possesses flexible path programming capabilities and powerful policy adaptability, but also enables intelligent orchestration of service function paths through dynamic management of SID sequences. This makes it suitable for a variety of application scenarios, including computing power network resource coordination, service path optimization, and cross-cloud path control. As one of the core control components of this invention, this module significantly improves the control accuracy and intelligence of the computing-network convergence network.
[0136] It can be seen that in terms of computing power perception, the solution of the present invention can freely configure the resources to be obtained from the node through the resource acquisition system, and supports custom interval time to obtain resource information, ensuring that the network topology can be quickly updated and the network can be quickly adapted in a dynamic network topology environment. In terms of resource scheduling strategy, based on the node computing power situation and the network situation, a comprehensive judgment is achieved, and the request can be automatically scheduled to the most appropriate node according to the type of request, thereby improving the matching degree between the request and the resource, and can greatly improve resource utilization as a whole, and the overall efficiency can also be greatly improved. In terms of network control, based on the SDN architecture adopted by SRv6, the data and control planes are separated. Compared with traditional SDN control, SRv6 technology does not need to frequently update the routing table, but directly programs the request header, reducing a lot of network overhead while also achieving flexible control, optimizing resource utilization, and realizing collaborative scheduling of computing network resources in cross-cloud, cloud-edge and other scenarios.
[0137] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0138] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0139] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0140] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0141] The embodiments described above are merely some preferred embodiments of the present invention and are not intended to limit the present invention. Persons skilled in the art may make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, any technical solution obtained by equivalent substitution or equivalent transformation falls within the scope of protection of the present invention.
Claims
1. A multi-dimensional computing power dynamic perception routing decision method based on SRv6, characterized in that: include: S1. Dynamically collect network resource information and computing node resource usage from the cluster environment, and aggregate and persist the acquired data. S2. Model the information obtained in S1 as an undirected graph. In this undirected graph, edges represent network status, described by network resources, and nodes represent computing nodes, described by computing power resources. At the same time, the network bandwidth and latency indicators are uniformly converted into edge weights to obtain a dynamic network topology based on converged service sensitivity. S3: Based on the resource matching scoring model between business requirements and node real-time load, the optimal computing power node is selected as the target computing power node. Then, based on the dynamic network topology diagram that integrates business sensitivity, the improved SPFA algorithm is used to calculate the optimal network path to the target computing power node. S4, based on SRv6-driven dynamic generation and closed-loop execution of service chains, abstracts the target computing nodes and network functions into programmable SIDs, dynamically builds service chains and encapsulates them in the SRH header, implements the execution path, monitors the links, and makes timely adjustments after node failures.
2. The SRv6-driven multi-dimensional computing power dynamic perception routing decision method according to claim 1 is characterized in that: In S1, the network resource status is collected regularly, including bandwidth and latency information; the collection of computing node resources includes deploying collectors on all nodes in the cluster, and the regular collection includes key indicators such as CPU usage, GPU usage, memory occupancy, disk I / O, network link bandwidth, and latency.
3. The multi-dimensional computing power dynamic perception routing decision method based on SRv6 drive according to claim 1 is characterized in that: After aggregation, all collected resource data is uniformly converted into a structured resource state vector to describe the node's computing power and network status. The resource state vector is specifically: the node's computing power: S(n i )=[n i (CPU),n i (GPU),n i (Mem),n i (IO)] Among them, n i represents the i-th node in the cluster, S(n i ) represents the computing resource set of the i-th node, n i (CPU) represents the CPU resource usage of the i-th node, n i (GPU) represents the GPU resource usage of the i-th node, n i (Mem) represents the memory resource usage of the i-th node, n i (IO) represents the I / O rate of the i-th node; The network status of the node: Net(n) i ,a j )=[BW(n i ,a j ),RTT(n i ,a j )] Among them, n i 、n j Represents the i-th and j-th nodes in the cluster, n i 、n j Satisfy the direct adjacent condition, Net(n i ,n j ) represents n i 、n j The network status between two nodes, BW(n i ,n j ) represents n i 、n j Bandwidth utilization of network links between nodes, RTT (n i ,n j ) represents n i 、n j The latency of the network link between nodes.
4. The multi-dimensional computing power dynamic perception routing decision method based on SRv6 drive according to claim 1 is characterized in that: In the undirected graph G(V,E) in S2, vertex V represents a computing node, edge E represents a network link between nodes, and the edge attributes are bandwidth RTT and latency BW. The ratio of the latency from node u to v to the maximum latency is multiplied by an adjustable weight coefficient β1, and the ratio of the latency from node u to v to the maximum bandwidth is multiplied by another adjustable weight coefficient β2. The sum of the two is taken to uniformly convert the bandwidth and latency indicators into edge weights, where β1+β2=1. This supports dynamic adjustment based on the service's sensitivity to latency / bandwidth.
5. The multi-dimensional computing power dynamic perception routing decision method based on SRv6 drive according to claim 1 is characterized in that: The resource matching scoring model is specifically as follows: for each request, the resource status of the request is obtained from the request details, and a score is calculated based on the resource status of the request and the real-time load of the current node. In the resource set to be considered, the ratio of the request's application for resource r to the remaining resource amount of resource r on the current node is taken. The sum of the above ratios of all resources in the resource set is the current node score.
6. The SRv6-driven multi-dimensional computing power dynamic perception routing decision method according to claim 5 is characterized in that: All computing nodes are sorted by score, and the computing node with the highest score is the optimal computing power node and serves as the target computing power node.
7. The SRv6-driven multi-dimensional computing power dynamic perception routing decision method according to claim 5 is characterized in that: The improved SPFA algorithm is specifically as follows: an undirected graph, a source node, and a target computing power node are taken as input, and each edge in the undirected graph has a non-negative edge weight; in the initialization stage, the path distance of all nodes is configured to be infinite, the queue status is not queued, the predecessor node is empty, the queue counter is reset to zero, the source node distance is set to zero and the queue mark is added; in the main loop stage, when the queue is not empty, the node is continuously taken from the head of the queue and its queue mark is cleared, its adjacent nodes are traversed and the current edge weight is dynamically calculated. If the path cost from the current node to the adjacent node is better, the path distance of the adjacent node and the predecessor node are updated. When it is detected that the adjacent node is not queued, it is added to the end of the queue and the status and cumulative queue counter are updated. The termination is triggered when the number of queues exceeds the threshold; when the queue is empty, the algorithm terminates and outputs the minimum path distance of each node and the optimal path reconstructed based on the predecessor node. The dynamic adaptability of the network and the robustness of the algorithm are achieved through dynamic weight calculation and queue count monitoring.
8. The SRv6-driven multi-dimensional computing power dynamic perception routing decision method according to claim 1 is characterized in that: S4 uses an SDN-based controller that separates the data plane from the control plane. It implements routing management and path orchestration based on the pyroute2 library. The controller communicates with router nodes using a RESTful API and manages SRv6 paths to implement real-time dynamic routing planning. Based on S3's optimal network path routing decisions, the controller dynamically adjusts paths in real time to optimize traffic transmission, enabling flexible service programming and path optimization.
9. A multi-dimensional computing power dynamic perception routing decision system based on SRv6, characterized by: Implementing the method according to any one of claims 1 to 8, comprising: a dynamic computing network resource perception module for dynamically collecting network link conditions and computing node resource usage from a cluster environment, and aggregating and persisting the acquired data; The multi-dimensional computing network resource decision-making algorithm module is used to route computing power requests to the optimal node through a dynamically generated computing power scheduling strategy based on the information collected by the dynamic computing network resource perception module. It also dynamically adjusts the path based on the real-time status of the computing network resources to ensure that the optimal route is selected based on changes in network status. The SRv6 controller module is used to dynamically adjust paths in real time based on the routing decisions made by the computing network resource multi-dimensional decision-making algorithm module to optimize traffic transmission, enabling flexible service programming and path optimization.
10. A computer-readable storage medium storing computer-executable instructions, wherein the instructions are used to implement the method according to any one of claims 1 to 8 when executed.
Citation Information
Cited By
Multi-dimensional dynamic sensing and intelligent computing power scheduling method and device for computing power network
CN120994407A
Multi-dimensional dynamic perception and intelligent computing power scheduling method and device for computing power network
CN120994407B
SRv6 computing power network path optimization method and system based on multi-dimensional dynamic scoring
CN121462483A
Hierarchical topological domain weight sensing task scheduling method and system
CN121597348A
Method for supporting heterogeneous signal and heterogeneous network end-to-end fusion scheduling
CN121691149A