Multi-dimensional scheduling method and device based on calculation network speed-up ratio and communication equipment

By introducing the Computing Network Acceleration Ratio (CNAR) and CNAR-ID mechanisms, combined with SRv6 segment routing technology, multi-dimensional scheduling of task flows is achieved, solving the problem that traditional network evaluation methods cannot meet the requirements of complex tasks, and improving the system's business completion rate and resource utilization efficiency.

CN120614397APending Publication Date: 2025-09-09CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510986900.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Traditional network evaluation methods are unable to reflect the multi-dimensional resource requirements of complex tasks, and the existing scheduling mechanism cannot accurately map the real-time load and available path capacity of multi-dimensional task flows, resulting in insufficient scheduling of high-performance nodes, severe congestion on some paths, and increased AI call failure rates.

Method used

By introducing the computing network acceleration ratio (CNAR) as a unified scheduling indicator and combining it with SRv6 segmented routing technology, task feature identification and expected computing network acceleration ratio calculation are performed to generate a computing network task identifier (CNAR-ID). The dynamic scheduling strategy is executed on the SRv6 endpoint node or central computing power node to achieve programmable path selection and dynamic computing power mapping.

Benefits of technology

It significantly improves the business completion rate and resource utilization efficiency of the overall system, and realizes multi-task inference scheduling capabilities with low latency, high throughput and high cost-effectiveness. It is suitable for scenarios such as AI inference and edge-cloud collaborative large model access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120614397A_ABST
    Figure CN120614397A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-dimensional scheduling method and device based on a calculation network speed-up ratio and communication equipment. The method comprises the following steps: receiving a task flow through an SRv6 source node; performing task feature recognition on the task flow, calling a calculation network speed-up ratio calculation engine, and calculating an expected calculation network speed-up ratio of the task flow in the current network and calculation power environment of the system; according to the identified task features and the calculated expected calculation network speed-up ratio, a calculation network task identifier of the task flow is determined, and the calculation network task identifier is used for identifying a scheduling strategy in the task identification tree; the computing network task identifier of the task flow is added to an SRv6 segment route header, the computing network task identifier in the SRv6 segment route header is read through an SRv6 endpoint node, and a scheduling strategy corresponding to the read computing network task identifier is executed; and the service completion rate and the resource utilization efficiency of the whole system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computing-network integration technology, and in particular to a multi-dimensional scheduling method, apparatus, and communication equipment based on computing-network acceleration ratio. Background Art

[0002] Traditional SRv6 uses IPv6-based segment routing (Segment Routing) technology applied to the IPv6 data plane and inserts a Segment Routing Header (SRH) into IPv6 packets. The SRH contains a list of IPv6 addresses (Segment List). The destination address of the packet is updated according to the SID (Segment ID) sequence in the list, and the packet is forwarded segment by segment.

[0003] As services like large-scale AI reasoning, edge computing, and intelligent perception continue to penetrate the terminal side, traditional network evaluation methods based primarily on throughput and fixed latency are unable to reflect the multi-dimensional resource requirements of complex tasks. In particular, in scenarios involving multi-task convergence and heterogeneous resource deployment, there is a significant discrepancy between task scheduling efficiency and actual service performance. Traditional path selection and computing resource allocation models are unable to dynamically reflect the synergistic relationship between task execution quality, reasoning performance, and network transmission. A scheduling method that integrates computing and networking is urgently needed.

[0004] In distributed computing environments like the converged edge, task characteristics are highly diverse. These include diverse AI model execution performance (e.g., token consumption, latency, and load overhead) as well as network resource dynamics (e.g., edge node link changes and SRv6 path switching). Existing scheduling mechanisms, primarily based on static resource awareness or coarse-grained metrics, are unable to accurately map the real-time load and available path capacity of multi-dimensional task flows. This can lead to issues such as insufficient scheduling of high-performance nodes, severe congestion on some paths, and increased AI call failure rates, further impacting the service experience and system efficiency. Summary of the Invention

[0005] Based on this, it is necessary to provide a multi-dimensional scheduling method, device, computing network controller, computer-readable storage medium and computer program product based on computing network acceleration ratio, which can improve the business completion rate and resource utilization efficiency of the overall system in response to the above technical problems.

[0006] In a first aspect, the present application provides a multi-dimensional scheduling method based on computing network acceleration ratio, the method comprising:

[0007] Receive task streams via SRv6 source nodes;

[0008] Identify the task features of the task flow, and call the network speedup ratio calculation engine to calculate the expected network speedup ratio of the task flow under the current network and computing power environment of the system;

[0009] Determine a computing network task identifier for the task flow based on the identified task features and the calculated expected computing network speedup ratio, wherein the computing network task identifier is used to identify a scheduling strategy in a task identification tree;

[0010] The computing network task identifier of the task flow is added to the SRv6 segment routing header, the computing network task identifier in the SRv6 segment routing header is read through the SRv6 endpoint node, and the scheduling policy corresponding to the read computing network task identifier is executed.

[0011] In one embodiment, the method further comprises:

[0012] If the SRv6 endpoint node has the ability to execute the scheduling policy corresponding to the read computing network task identifier, execute the scheduling policy corresponding to the read computing network task identifier through the SRv6 endpoint node;

[0013] In the case that the SRv6 endpoint node does not have the ability to execute the scheduling policy corresponding to the read computing network task identifier, the scheduling policy corresponding to the read computing network task identifier is forwarded to the central computing power node through the SRv6 endpoint node.

[0014] In one embodiment, the method further comprises:

[0015] Receive a path feedback message sent by an SRv6 endpoint node, and update a scheduling policy in the task identification tree according to the path feedback message.

[0016] In one embodiment, the step of identifying task features of the task flow includes:

[0017] The task flow is identified according to service intent, resource type and business level, and is judged according to source IP, destination IP, source port, destination port and transport layer protocol to complete the task feature identification of the task flow.

[0018] In one embodiment, determining the task flow based on the source IP, destination IP, source port, destination port, and transport layer protocol includes:

[0019] Get the prompt structure, model identifier, delay constraint field, and token length in the payload of the historical flow message;

[0020] The task flow is determined in combination with the prompt structure, the model identifier, the delay constraint field, the token length, the source IP, the destination IP, the source port, the destination port, and the transport layer protocol.

[0021] In one embodiment, calculating the expected computing-network speedup ratio of the task flow under the current network and computing power environment of the system includes:

[0022] Calculate the expected computing network speedup ratio based on multiple sub-indicator sampling values ​​under the system's current network and computing power environment and the weight value dynamically assigned to each sub-indicator sampling value;

[0023] Among them, the multiple sub-indicator sampling values ​​include round-trip time sampling value, packet loss rate sampling value, path congestion sampling value, CPU utilization sampling value and computing power node waiting queue length sampling value.

[0024] In a second aspect, the present application further provides a multi-dimensional scheduling device based on computing network acceleration ratio, the device comprising:

[0025] A receiving module, used to receive task flows through SRv6 source nodes;

[0026] A computing network acceleration ratio determination module is used to identify task features of the task flow and call a computing network acceleration ratio calculation engine to calculate the expected computing network acceleration ratio of the task flow under the current network and computing power environment of the system;

[0027] A scheduling strategy determination module is used to determine a computing network task identifier of the task flow based on the identified task characteristics and the calculated expected computing network speedup ratio, wherein the computing network task identifier is used to identify a scheduling strategy in a task identification tree;

[0028] An execution module is used to add the computing network task identifier of the task flow to the SRv6 segment routing header, read the computing network task identifier in the SRv6 segment routing header through the SRv6 endpoint node, and execute the scheduling policy corresponding to the read computing network task identifier.

[0029] In a third aspect, the present application further provides a communication device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method provided in the first aspect when executing the computer program.

[0030] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the method provided in the first aspect when executed by a processor.

[0031] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which implements the method provided in the first aspect when executed by a processor.

[0032] The above-mentioned multi-dimensional scheduling method, apparatus, communication equipment, computer-readable storage medium and computer program product based on computing network acceleration ratio receive task streams through SRv6 source nodes; identify task features of the task streams, and call the computing network acceleration ratio calculation engine to calculate the expected computing network acceleration ratio of the task streams under the current network and computing power environment of the system; determine the computing network task identifier of the task stream based on the identified task features and the calculated expected computing network acceleration ratio, and the computing network task identifier is used to identify the scheduling strategy in the task identification tree; add the computing network task identifier of the task stream to the SRv6 segment routing header, read the computing network task identifier in the SRv6 segment routing header through the SRv6 endpoint node, and execute the scheduling strategy corresponding to the read computing network task identifier; thereby realizing programmable path selection and dynamic computing power mapping for the task stream, providing fine-grained, multi-level, and highly responsive scheduling and optimization capabilities for complex heterogeneous services, and significantly improving the service completion rate and resource utilization efficiency of the overall system. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.

[0034] Figure 1 FIG1 is a flow chart of a multi-dimensional scheduling method based on computing network acceleration ratio in one embodiment;

[0035] Figure 2 A schematic diagram of multi-terminal interaction of a multi-dimensional scheduling method based on computing network acceleration ratio in one embodiment;

[0036] Figure 3 This is a second flow chart of a multi-dimensional scheduling method based on computing network acceleration ratio in one embodiment;

[0037] Figure 4 This is a structural block diagram of a multi-dimensional scheduling device based on network acceleration ratio in one embodiment. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0039] Before introducing the specific embodiments of the present invention, the professional terms involved in the present invention are explained:

[0040] SRv6 (Segment Routing over IPv6): A source-routed segmented path based on IPv6 addresses. Flexible path control and service chain orchestration are achieved by inserting the Segment Routing Header (SRH).

[0041] Segment Routing Header (SRH): A key header structure in SRv6, containing fields such as Segment List, Tag, and Function, used to define and execute path behavior.

[0042] Segment Identifier (SID): An IPv6 address unit used to identify network functions or node locations in SRv6. Nodes parse SIDs sequentially to perform specified forwarding or computing tasks.

[0043] CNAR (Compute Network Acceleration Ratio): Compute network acceleration ratio, which measures the ratio of the collaborative efficiency of computing and transmission resources in a task flow under a specific path, and reflects comprehensive performance such as task response latency, model inference efficiency, and network load.

[0044] CNAR-ID (Computing Network Task Identifier): A unique identifier generated by the controller based on task characteristics. It is used to identify business flow types, AI model requirements, resource strategies, etc., and is used for SRH extension field marking.

[0045] CNAR Task Tree: A task classification structure built by the control plane that maps service characteristics to CNAR-IDs, enabling flow-level identification and policy matching for dynamic decision-making.

[0046] Function + Arg (SRH extension field): An extension field in the SRH used to express advanced network behaviors. "Function" indicates the type of function to be performed, and "Arg" carries computing network parameters such as CNAR-ID and AI model category.

[0047] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0048] In one embodiment, combined Figure 1 and Figure 2 , provides a multi-dimensional scheduling method based on computing network acceleration ratio, the method is performed by, for example but not limited to, a computing network controller, a chip, a communication device or a communication system, and the method includes the following steps S102~S108.

[0049] S102: Receive a task flow through an SRv6 source node.

[0050] When accessing an SRv6 network, an edge gateway or access router can serve as an SRv6 source node, receiving task flows. User terminals or edge intelligent devices can initiate task processing requests to the central cloud or edge service nodes based on service requirements (e.g., video processing, image recognition, real-time inference, etc.). For example, user terminals or edge intelligent devices can initiate AI inference tasks (e.g., image and text generation, image recognition, vector search, etc.).

[0051] S104: Identify the task features of the task flow and call the network acceleration ratio calculation engine to calculate the expected network acceleration ratio of the task flow under the current network and computing power environment of the system.

[0052] The task flow's various flow features can be automatically identified to obtain the task flow's task features. Flow features include, but are not limited to, transmission frequency, packet length, content structure, and task type (e.g., interactive, multimodal, or large-model reasoning).

[0053] CNAR is a quantitative ratio of the efficiency of tasks completed in a collaborative path to the efficiency of local execution in a scenario where computing power and network collaborative scheduling are integrated. CNAR not only considers computational time and network transmission time, but also introduces actual system operation indicators such as AI model processing costs, vector database operation delays, resource queuing, cold start, and error retries. Therefore, CNAR is a multi-dimensional weighted evaluation value of task execution efficiency, which can support edge schedulers in real-time sorting of candidate nodes or scoring of SR policy allocation paths. For example, the general CNAR formula modeling includes:

[0054]

[0055] Among them, Tlocalbase represents the estimated time (benchmark time) for the task to be executed independently locally; Tedgecompute represents the edge computing indicator CPU / memory execution time, etc.; Tnetwork represents the network performance indicator transmission time, including RTT, congestion, packet loss retransmission, etc.; TAI represents the AI ​​model call time and token processing delay, etc.; TDB represents the vector database indicator query / write delay, etc.; Tretry represents the abnormal retry time (including cold start penalty); Tsched represents the controller scheduling or scheduling decision time; Tpenalty comes from the penalty delay estimate of model output errors and invalid responses.

[0056] Based on this, we construct a multi-dimensional, multi-factor computing network acceleration model to derive the expected computing network acceleration ratio. This model not only considers latency but also incorporates the cost-effectiveness of computing resources (computing power consumption vs. results), task processing quality (error rate, task success rate), cost consumption (model call token cost, energy consumption, etc.), AI model load status (queue queuing, inference latency), and data I / O efficiency (such as vector database reading and writing). In other words, the expected computing network acceleration ratio can at least include the CNAR value.

[0057] S106 , determining a network computing task identifier of the task flow based on the identified task features and the calculated expected network computing speedup ratio. The network computing task identifier is used to identify a scheduling strategy in a task identification tree.

[0058] For example, the impact of CNAR on the path selection decision in the scheduling strategy is shown in Table 1 below:

[0059] Table 1

[0060]

[0061] CNAR-ID is a unique identifier used to identify and distinguish the acceleration characteristics and scheduling intent of network tasks in a computing power-network collaborative system. Its essence is a mechanism for clustering, classifying, and assigning unified labels to task flows or access behaviors with similar computing network acceleration characteristics (CNAR). The core function of CNAR-ID is to connect the "task awareness" chain between the forwarding plane and the control plane; finely classify service flow types, support multiple models and path scheduling strategies; implement adaptive task offloading based on dynamic paths, computing power resources, and model collaboration; and support the system's automatic update of the task identification tree (self-learning) based on CNAR-ID.

[0062] CNAR-IDs have the following properties: uniqueness: each task flow type is unique within a time period or policy; decodability: intermediate SRv6 nodes can interpret the task type or resource requirement it represents; and scheduling relevance: a one-to-one correspondence between the CNAR-ID and the policy table, SLA, and model diversion relationship. The control plane binds the CNAR-ID to the Task Identification Tree node for subsequent path delivery and model scheduling.

[0063] For example, the SRv6 control plane uses CNAR-ID to establish a "task identification tree" and maintain the following Table 2:

[0064] Table 2

[0065]

[0066] S108: Add the computing network task identifier of the task flow to the SRv6 segment routing header, read the computing network task identifier in the SRv6 segment routing header through the SRv6 endpoint node, and execute the scheduling policy corresponding to the read computing network task identifier.

[0067] The CNAR-ID is encoded in the SRv6 SRH and passed along with the data packet to the SRv6 endpoint node for decision-making and offloading. The SRv6 endpoint node reads the CNAR-ID field in the SRH and executes the scheduling policy corresponding to the computing network task identifier read.

[0068] The multi-dimensional scheduling method based on the computing network acceleration ratio provided in the embodiment of the present application, by introducing a unified computing network acceleration evaluation index and intelligent identification mechanism, performs full life cycle modeling and identification of task flows, and combines the SRv6 protocol framework to achieve programmable path selection and dynamic computing power mapping for tasks, providing fine-grained, multi-level, and highly responsive scheduling and optimization capabilities for complex heterogeneous services, significantly improving the business completion rate and resource utilization efficiency of the overall system. That is, the multi-dimensional scheduling method based on the computing network acceleration ratio provided in the embodiment of the present application, by introducing the "computing network acceleration ratio" as a unified scheduling metric, combined with the segmented routing mechanism of SRv6 and the multi-source information modeling of the control layer, achieves efficient path selection and resource collaborative optimization of AI tasks between edge nodes and the central cloud.

[0069] In other words, the embodiment of the present application introduces the CNAR-ID identification mechanism into the existing forwarding plane and control plane of SRv6, and builds a "task identification tree" by combining message context fingerprints, AI call metadata and edge computing features to support the dynamic perception and fine classification of controller cross-domain and cross-model access paths. The CNAR indicator is proposed on the control plane. In a specific computing network collaborative environment, the performance acceleration capability of a certain type of task flow relative to the baseline execution path under the current network and computing power joint optimization configuration, as well as the CNAR-ID task tree scheduling algorithm that adapts to multi-model and multi-granularity reasoning task scenarios. It also supports a collaborative scheduling strategy that integrates edge small model reasoning and central large model reasoning, and dynamically guides path selection in combination with the acceleration ratio.

[0070] Optionally, adding the network computing task identifier of the task flow to the SRv6 segment routing header in step S108 includes: adding the network computing task identifier of the task flow to the Function field of the SRv6 segment routing header.

[0071] The multi-dimensional scheduling method based on the computing network acceleration ratio provided in the embodiment of the present application can effectively improve the overall utilization rate of AI and other task flows between computing resources and network resources, and realize low-latency, high-throughput, and cost-effective multi-task reasoning scheduling capabilities. It is widely applicable to scenarios such as AI reasoning, RAG multi-round question and answer, and edge-cloud collaborative large model access. Among them, following the existing SRv6SRH format and basic forwarding method, it adapts to the multi-model access and dynamic scheduling requirements in the computing network collaborative environment, and expands the function / argument field related to the task flow CNAR-ID in the forwarding plane to identify the task type, access characteristics and computing power call request, so that the SRv6 endpoint node performs specific flow classification and resource binding based on the identification information after removing the SRH header. Combined with the programmable path capability of SRv6, the controller uses the mapping relationship between CNAR-ID and access policy to dynamically generate task-oriented TE paths, realizing the fine coordination and integrated forwarding of AI task flows between the edge small model and the central large model node, thereby improving the stability of the reasoning link, shortening the overall response delay, and improving the flexibility of computing power access and the level of intelligent scheduling.

[0072] In one embodiment, Figure 3 As shown, the multi-dimensional scheduling method based on network acceleration ratio also includes the following steps S302~S304.

[0073] S302: When the SRv6 endpoint node has the capability to execute the scheduling policy corresponding to the read network computing task identifier, the scheduling policy corresponding to the read network computing task identifier is executed through the SRv6 endpoint node.

[0074] Among them, it can be determined whether the SRv6 endpoint node has the ability to execute the scheduling policy corresponding to the read computing network task identifier. If it has the ability, the scheduling policy corresponding to the read computing network task identifier is immediately executed to reduce delay.

[0075] S304: When the SRv6 endpoint node does not have the ability to execute the scheduling policy corresponding to the read computing network task identifier, the scheduling policy corresponding to the read computing network task identifier is forwarded to the central computing power node through the SRv6 endpoint node.

[0076] Among them, it can be determined whether the SRv6 endpoint node has the ability to execute the scheduling policy corresponding to the read computing network task identifier. If not, the scheduling policy corresponding to the read computing network task identifier is forwarded to the central computing power node (such as a GPU large model cluster), and the central computing power node executes it in the usual way.

[0077] In the embodiment of the present application, by determining whether the SRv6 endpoint node has the ability to execute the scheduling policy corresponding to the read computing network task identifier, the scheduling policy can be reliably executed, thereby further optimizing the scheduling method.

[0078] In one embodiment, the multi-dimensional scheduling method based on computing network speedup ratio further includes the following steps:

[0079] Receive path feedback messages sent by SRv6 endpoint nodes and update the scheduling policy in the task identification tree based on the path feedback messages.

[0080] In an embodiment of the present application, when the scheduling policy corresponding to the computing network task identifier read is executed by the SRv6 endpoint node, the "path feedback mode" can be enabled, and the SRv6 endpoint node can record and send the path feedback message (i.e., the scheduling policy execution result), so that the computing network controller dynamically updates the scheduling policy in the task identification tree according to the path feedback message, so as to further optimize the scheduling method. In addition, when the scheduling policy corresponding to the computing network task identifier read is executed by the central computing power node, the computing network controller can also receive a data packet with attached "CNAR feedback metadata" (time consumption, success rate, token usage, etc.), so that the computing network controller dynamically updates the scheduling policy in the task identification tree according to the path feedback message, so as to further optimize the scheduling method.

[0081] In one embodiment, the task feature identification of the task flow in step S104 includes the following steps:

[0082] The task flow is identified based on service intent, resource type and business level, and is judged based on source IP, destination IP, source port, destination port and transport layer protocol to complete the task feature identification of the task flow.

[0083] In this embodiment of the present application, the SRv6 source node receives the task flow, introduces the "service intent + resource type + business level" triple as the flow identification basis, and combines it with the basic five-tuple to make a judgment, so as to complete the task characteristics of the task flow accurately, reliably and at low cost.

[0084] In one embodiment, the task feature identification of the task flow in step S104 includes the following steps:

[0085] Obtain the prompt structure, model identifier, delay constraint field, and token length from the payload of the historical flow message; determine the task flow based on the prompt structure, model identifier, delay constraint field, token length, source IP, destination IP, source port, destination port, and transport layer protocol.

[0086] In an embodiment of the present application, an edge prediction engine / lightweight AI model is enabled as a gateway to pre-classify the access task flow, determine whether the session is interactive, multimodal, or large-model reasoning, etc., and extract the prompt structure, model identifier, delay constraint field, and token length in the historical flow message payload as auxiliary judgments to further achieve accurate, reliable, and low-cost identification of the task characteristics of the task flow.

[0087] In one embodiment, calculating the expected computing network speedup ratio of the task flow under the current network and computing power environment of the system in step S104 includes the following steps:

[0088] The expected computing network acceleration ratio is calculated based on multiple sub-indicator sampling values ​​under the system's current network and computing power environment and the weight values ​​dynamically assigned to each sub-indicator sampling value; wherein, the multiple sub-indicator sampling values ​​include round-trip time sampling values, packet loss rate sampling values, path congestion sampling values, CPU utilization sampling values ​​and computing power node waiting queue length sampling values.

[0089] Among them, each sub-indicator sampling value can be multiplied by the corresponding weight value respectively, and the sum is used to determine the network acceleration ratio.

[0090] The embodiment of the present application introduces a dynamic weight adjustment mechanism to automatically adjust the CNAR weight factor combination according to business sensitivity, so as to help achieve dynamic optimization of the scheduling strategy and realize optimized network and computing power fusion scheduling.

[0091] The following exemplifies the application of the technical solutions of the embodiments of the present application to actual scenarios, and lists several application examples of such scenarios to further supplement the technical solutions of the embodiments of the present application:

[0092] Scenario Application Example 1: Edge-Priority Scheduling for Real-Time Voice Interaction Tasks

[0093] In a voice interaction service scenario, a user terminal initiates short-duration voice recognition requests to an edge access device. These short-duration voice recognition requests are referred to as task flows. The system first automatically identifies the task type (i.e., task features) of this task flow based on flow characteristics (such as transmission frequency, packet length, and content structure) and extracts contextual information for CNAR modeling. The controller analyzes and determines that this type of task flow is latency-sensitive and requires moderate computing power, making it suitable for rapid processing by edge nodes.

[0094] The system generates a unique CNAR-ID for the task flow, encapsulates it at the SRv6 source node, and adds it to the Function field in the SRH. Based on the scheduling policy bound to the CNAR-ID, the controller maps the task flow to the "Real-time Interaction" task identification tree branch and dynamically issues an edge-first path orchestration policy.

[0095] When a task flow passes through an SRv6 endpoint, it makes intelligent decisions based on the CNAR-ID information in the SRH. If the SRv6 endpoint has sufficient computing power, it can complete the inference processing and return the results, significantly reducing overall response latency. The system also records performance data for each transaction (such as the number of tokens and model duration) to continuously optimize the scheduling strategy corresponding to the CNAR-ID, achieving self-learning and performance evolution.

[0096] Scenario Application Example 2: Central Node Inference Scheduling for Complex Analytical Requests

[0097] An enterprise's big data platform regularly initiates structured text batch processing requests to the intelligent system for complex analysis tasks such as log parsing and trend modeling. This structured text batch processing request is a task flow.

[0098] After identifying the task flow, the system estimates the CNAR characteristics (i.e., expected computing network acceleration ratio) through parameters such as the number of tokens and context links (i.e., task characteristics), and determines that its reasoning complexity is high and the computing link is long, making it more suitable for processing by a central computing power node.

[0099] Based on this information, the controller generates a unique CNAR-ID and establishes a binding relationship with the "Deep Computation Class" identification tree node. The SRv6 source node, following the control policy, encapsulates the CNAR-ID into the SRH and orchestrates it onto a path with large-scale computing power. During forwarding, the endpoint node is solely responsible for delivery and latency measurement and does not interfere with the execution of inference tasks.

[0100] After receiving a task, the central computing node completes processing and simultaneously reports metrics such as task token consumption and response latency to the controller. The controller uses this feedback to refine the CNAR model's estimation in real time, optimizing task mapping and routing strategies. This process enables coordinated optimization of the control and forwarding planes, supports policy inheritance and generalization, and improves the scheduling efficiency of subsequent tasks of the same type.

[0101] Scenario Application Example 3: Edge-Center Collaborative Offloading Scheduling of Periodic Image Inspection Tasks

[0102] In an industrial monitoring scenario, front-end cameras periodically upload images to an intelligent system for quality inspection. The system identifies this task flow as stable, periodic, and structured. Using a model, it infers that preprocessing tasks can be performed at the edge, but this carries the risk of misjudgment. Therefore, a collaborative reasoning mechanism combining "edge initial judgment and central rejudgment" is required.

[0103] The system generates a CNAR-ID for the task flow and encapsulates it in the SRv6 SRH extension field. The "Fnction" field indicates "cooperative offloading," and the "Arg" field specifies the offloading priority and path mode. The path first includes edge nodes, which are used to quickly complete the initial image judgment.

[0104] If the confidence level of the result is not high, the SRv6 endpoint node will pass the task to the central computing node for further processing according to the strategy.

[0105] The controller continuously optimizes the mapping relationship between CNAR-ID and its recognition tree based on the task completion status, and collects indicators such as misjudgment rate, delay data, and offloading success rate to correct the collaborative offloading strategy.

[0106] This approach reduces the load on central computing nodes while ensuring processing quality, and improves the overall flexibility and accuracy of resource scheduling.

[0107] The technical solutions of the embodiments of this application introduce the "Computation-Network Acceleration Ratio (CNAR)" as a new metric for uniformly evaluating task scheduling efficiency, integrating task computing performance with network transmission costs. This breaks through the traditional scheduling metric based solely on bandwidth or computing power, achieving more comprehensive resource scheduling optimization. A task-feature-based intelligent identification and distribution mechanism (CNAR-ID) is implemented. The CNAR-ID task tree scheduler enables semantic recognition and feature classification of tasks, supporting fine-grained task binding and refined scheduling on the network side, improving the scheduling adaptability and responsiveness of various AI tasks. The controller extension and deep fusion scheduling strategy embeds an enhanced module within the SRv6 network structure, supporting CNAR functions and parameters in the SRH extension field. This enables task-intent-based path matching and control policy injection, significantly improving the dynamic adaptability of network and computing integration. The edge-computing feedback monitoring mechanism, implemented through modules such as the "edge-cloud model coordinator," "edge computing feedback monitor," and "network computing feedback monitor," provides real-time feedback and scheduling corrections, ensuring system stability and adaptability in complex scenarios, surpassing traditional non-feedback strategies. The system's policy engine integrates multiple factors, including CNAR values, task characteristics, and path status, to dynamically formulate optimal scheduling decisions. This engine offers greater generalization and performance improvement than traditional static path or single-dimensional scheduling strategies. This multi-dimensional, multi-factor integration allows for the construction of more accurate computational network acceleration models.

[0108] After applying the technical solutions of the embodiments of this application, the overall utilization rate of AI and other task flows between computing resources and network resources can be effectively improved, and low-latency, high-throughput, and cost-effective multi-task reasoning scheduling capabilities can be achieved. It is widely applicable to scenarios such as AI reasoning, RAG multi-round question and answer, and edge-cloud collaborative large model access.

[0109] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0110] Based on the same inventive concept, embodiments of the present application also provide a multi-dimensional scheduling device based on a computing network speedup ratio for implementing the aforementioned multi-dimensional scheduling method based on a computing network speedup ratio. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations of one or more embodiments of the multi-dimensional scheduling device based on a computing network speedup ratio provided below can be found in the above-mentioned limitations of the multi-dimensional scheduling method based on a computing network speedup ratio, and will not be repeated here.

[0111] In an exemplary embodiment, Figure 4 As shown, a multi-dimensional scheduling device based on computing network acceleration ratio is provided, including:

[0112] The receiving module 410 is configured to receive a task flow through an SRv6 source node.

[0113] The computing network acceleration ratio determination module 420 is used to identify the task characteristics of the task flow and call the computing network acceleration ratio calculation engine to calculate the expected computing network acceleration ratio of the task flow under the current network and computing power environment of the system.

[0114] The computing network acceleration ratio determination module 420 may include a CNAR acceleration ratio analysis and measurement module (CNAREvaluation Module) and a CNAR-ID task tree parsing and scheduling module (CNAR-ID Task TreeScheduler).

[0115] The functional description of the CNAR Acceleration Ratio Analysis and Measurement Module (CNAR Evaluation Module) includes: the CNAR calculation model is responsible for collecting, aggregating and normalizing multi-dimensional indicators in the network and computing power environment, and calculating the CNAR (Compute-Network Acceleration Ratio) as the overall performance acceleration ratio evaluation indicator. CNAR can be composed of sub-indicators such as network latency, bandwidth utilization, computing power utilization, load balancing factor, and energy efficiency ratio. Technical details include: using a windowed indicator sampling method to collect data on the network side (such as RTT, packet loss rate, path congestion) and the computing power side (such as CPU utilization, computing power node waiting queue length); introducing a dynamic weight adjustment mechanism to automatically adjust the CNAR weight factor combination according to business sensitivity; supporting the API interface to output CNAR and historical trends for reference by the controller scheduling.

[0116] The CNAR-ID Task Tree Scheduler module's functional description includes: a CNAR-ID identification tree structure, used to generate and maintain a CNAR-ID-based "task identification tree," clarifying the computing network resource requirements corresponding to each stage of the business process, and fine-tuning task scheduling. It decomposes complex business processes into a directed sequence of task nodes, with each node bound to computing network resource requirements and a CNAR matching policy. Technical details include: each task node has a unique CNAR-ID, carrying labels such as task type, computing intensity, and network priority; the task tree can be used to update SRv6 control plane policies (such as constructing low CNAR cost paths); and the controller matches the current node resource characteristics with the optimal path in the CNAR policy library based on the CNAR-ID when scheduling tasks.

[0117] A scheduling strategy determination module 430 is configured to determine a computing network task identifier for the task flow based on the identified task characteristics and the calculated expected computing network speedup ratio, wherein the computing network task identifier is used to identify a scheduling strategy in a task identification tree;

[0118] The execution module 440 is used to add the computing network task identifier of the task flow to the SRv6 segment routing header, read the computing network task identifier in the SRv6 segment routing header through the SRv6 endpoint node, and execute the scheduling policy corresponding to the read computing network task identifier.

[0119] The execution module 440 may include an Enhanced SRv6 Controller module, an Edge-Cloud Model Coordinator module, a CNAR strategy engine module, and a Net-Compute Feedback Monitor module.

[0120] The functions of the Enhanced SRv6 Controller module include:

[0121] This feature expands existing SRv6 control plane capabilities to support the insertion of new function / argument extensions into the SRH field based on the task CNAR-ID. This feature also enables dynamic SRv6 TE routing based on CNARs, ensuring optimal scheduling for both compute and network collaboration. Technical details include: the controller generates multi-segment TE paths based on the task CNAR-ID and the current global CNAR distribution; the introduction of new fields in the SRH, such as CNAR-Arg and Model-ID, to describe resource preferences and model integration requirements for intermediate nodes; and support for CNAR trend prediction and dynamic path rerouting for path nodes.

[0122] The Edge-Cloud Model Coordinator module's functional description includes: intelligent division of labor, offloading, and collaborative processing between small models (lightweight models at the edge) and large models (complex models at the cloud). It also automatically determines model migration strategies for inference / training tasks based on business characteristics, data volume, and CNAR matching results. Technical details include: support for a model registry and capability annotation: models include metadata such as accuracy, latency tolerance, and required resources; a predictor determines whether to complete tasks at the edge or transfer them to the cloud; and during model handover, it infers the minimum migration cost path based on resource status corresponding to CNAR IDs, avoiding duplicate context loading.

[0123] The CNAR strategy engine and reasoning module's functional description includes: managing the CNAR strategy library, task tree scheduling strategies, and routing path generation rules; providing a rule-based reasoning system, dynamically adjusting strategy templates, and implementing adaptive scheduling behavior. Technical details include: using a rule tree combined with a fuzzy matching algorithm to automatically match the optimal strategy based on CNAR metrics, service type, and model complexity.

[0124] The functional description of the Net-Compute Feedback Monitor module includes: real-time monitoring of the execution effect of Net-Compute collaboration (such as path delay, task execution success rate, and computing power consumption); and providing feedback to the controller and policy modules for path optimization, task reallocation, or model redeployment.

[0125] This embodiment of the present application proposes a novel metric system that integrates the "Computation-Network Acceleration Ratio (CNAR)." Unlike traditional scheduling strategies based on static paths, resource occupancy, or network latency, this system comprehensively perceives the computational consumption, model response time, and vector access characteristics of AI tasks across different paths, enabling a more fine-grained and dynamic multi-dimensional scheduling strategy. By constructing a CNAR-ID task identification system, the system supports automatic classification by task type, path binding, and model distribution, significantly improving the efficiency of integrated computing and network processing. In other words, this embodiment of the present application introduces a joint control mechanism combining a "task identification tree" and SRH extension fields into the SRv6 routing mechanism, enabling the controller to perceive flow-level AI service load characteristics and achieve precise task intent matching and real-time scheduling adjustments through the CNAR function and Arg parameters. Furthermore, the introduction of an edge-cloud collaborative model feedback mechanism and a CNAR learning module equips the system with adaptive optimization capabilities, maintaining efficient and stable service performance in multi-task concurrency and high-concurrency inference scenarios. This system demonstrates significant technological innovation and practical application value.

[0126] Compared with conventional technologies, the technical solutions of the embodiments of the present application have the following beneficial effects:

[0127] 1) The scheduling indicator from the perspective of the computing network (CNAR) introduces the "computation-network acceleration ratio (CNAR)" as a new metric for uniformly evaluating task scheduling efficiency. This integrates task computing performance and network transmission costs, breaking through the traditional scheduling dimension based solely on bandwidth or computing power, and achieving more comprehensive resource scheduling optimization. 2) Intelligent identification and distribution mechanism based on task characteristics (CNAR-ID): The CNAR-ID task tree scheduler implements semantic recognition and feature classification of tasks, supports fine-grained task binding and refined scheduling on the network side, and improves the scheduling adaptability and responsiveness of various types of AI tasks. 3) Controller extension and deep fusion scheduling strategy: An enhanced module is embedded in the SRv6 network structure, supporting the SRH extension field to carry CNAR functions and parameters, realizing path matching and control policy injection based on task intent, significantly improving the dynamic adaptability of network and computing integration. 4) The introduction of an edge-computing feedback monitoring mechanism enables real-time feedback and scheduling corrections through modules such as the "edge-cloud model coordinator," "edge computing feedback monitor," and "network computing feedback monitor," ensuring the system's stability and adaptability in complex scenarios, which is superior to traditional non-feedback strategies. 5) The system policy engine integrates multiple factors such as CNAR values, task characteristics, and path status to dynamically formulate optimal scheduling decisions. Compared with traditional static paths or single-dimensional scheduling strategies, it has stronger generalization capabilities and performance improvement potential. The embodiment of this application refines task types by identifying tree structures and accurately labels CNAR-IDs with SRv6 extended fields to achieve "scheduling by flow and matching on demand." The controller can dynamically match edge or central models based on factors such as token usage, AI model complexity, and path performance feedback, achieving real-time intelligent scheduling for AI application scenarios.

[0128] Each module in the aforementioned multi-dimensional scheduling device based on computing network acceleration ratio can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0129] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0130] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0131] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0132] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.

[0133] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0134] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A multi-dimensional scheduling method based on computing network acceleration ratio, characterized in that: The method comprises: Receive task streams via SRv6 source nodes; Identify the task features of the task flow and call the network speedup ratio calculation engine to calculate the expected network speedup ratio of the task flow under the current network and computing power environment of the system; Determine a computing network task identifier for the task flow based on the identified task features and the calculated expected computing network speedup ratio, wherein the computing network task identifier is used to identify a scheduling strategy in a task identification tree; The computing network task identifier of the task flow is added to the SRv6 segment routing header, the computing network task identifier in the SRv6 segment routing header is read through the SRv6 endpoint node, and the scheduling policy corresponding to the read computing network task identifier is executed.

2. The method according to claim 1, characterized in that The method further comprises: If the SRv6 endpoint node has the ability to execute the scheduling policy corresponding to the read computing network task identifier, execute the scheduling policy corresponding to the read computing network task identifier through the SRv6 endpoint node; In the case that the SRv6 endpoint node does not have the ability to execute the scheduling policy corresponding to the read computing network task identifier, the scheduling policy corresponding to the read computing network task identifier is forwarded to the central computing power node through the SRv6 endpoint node.

3. The method according to claim 1, characterized in that The method further comprises: Receive a path feedback message sent by an SRv6 endpoint node, and update a scheduling policy in the task identification tree according to the path feedback message.

4. The method according to claim 1, wherein The performing task feature identification on the task flow includes: The task flow is identified according to service intent, resource type and business level, and is judged according to source IP, destination IP, source port, destination port and transport layer protocol to complete the task feature identification of the task flow.

5. The method according to claim 4, characterized in that The determining of the task flow according to the source IP, destination IP, source port, destination port and transport layer protocol includes: Get the prompt structure, model identifier, delay constraint field, and token length in the payload of the historical flow message; The task flow is determined in combination with the prompt structure, the model identifier, the delay constraint field, the token length, the source IP, the destination IP, the source port, the destination port, and the transport layer protocol.

6. The method according to claim 1, characterized in that Calculating the expected computing network speedup ratio of the task flow under the current network and computing power environment of the system includes: Calculate the expected computing network speedup ratio based on multiple sub-indicator sampling values ​​under the system's current network and computing power environment and the weight value dynamically assigned to each sub-indicator sampling value; Among them, the multiple sub-indicator sampling values ​​include round-trip time sampling value, packet loss rate sampling value, path congestion sampling value, CPU utilization sampling value and computing power node waiting queue length sampling value.

7. A multi-dimensional scheduling device based on network acceleration ratio, characterized in that: The device comprises: A receiving module, used to receive task flows through SRv6 source nodes; A computing network acceleration ratio determination module is used to identify task features of the task flow and call a computing network acceleration ratio calculation engine to calculate the expected computing network acceleration ratio of the task flow under the current network and computing power environment of the system; A scheduling strategy determination module is used to determine a computing network task identifier of the task flow based on the identified task characteristics and the calculated expected computing network speedup ratio, wherein the computing network task identifier is used to identify a scheduling strategy in a task identification tree; An execution module is used to add the computing network task identifier of the task flow to the SRv6 segment routing header, read the computing network task identifier in the SRv6 segment routing header through the SRv6 endpoint node, and execute the scheduling policy corresponding to the read computing network task identifier.

8. A communication device comprising a memory and a processor, characterized in that: The memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Cross-architecture communication method and system based on SRv6 deterministic network

    CN121664726A