Method, system, medium, program product and terminal for routing and collaborative execution of capability slices of overlay networks based on end-side lightweight models

CN122679441APending Publication Date: 2026-09-01SHANGHAI ZHUIMO TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610835280.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-01

AI Technical Summary

Technical Problem

[0003](1)模型能力与网络调度相互割裂:现有方案通常关注模型如何在单节点或集群内切分,或者关注普通网络路由与服务调度,缺少将模型能力切片直接作为覆盖网络内生能力节点进行统一路由和编排的机制

Benefits of technology

[0022] (1) This application realizes the native organization of model capability slices in the overlay network by abstracting model capability slices into discoverable, filterable and routable capability node resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122679441A_ABST
    Figure CN122679441A_ABST
Patent Text Reader

Abstract

The application provides a method, system, medium, program product and terminal for overlay network capability slice routing and cooperative execution based on end-side lightweight model. The application abstracts model capability slice as discoverable, filterable and routable capability node resources, realizes native organization of model capability slice in overlay network, and completes capability calling sub-step disintegration and slice routing arrangement locally at the end side without completely relying on central control plane, thereby strengthening autonomous processing capability of end-side lightweight model for complex task request. The application comprehensively considers node capability, path state interface constraint, dependency relationship summary and budget state, realizes joint optimization of network state, dependency relationship and model capability slice, and replaces, switches and rearranges local capability slice path when slice node fails, times out, result confidence is insufficient or aggregation slice is abnormal during execution, thereby avoiding overall task failure.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer networks, distributed overlay networks, edge-side intelligent scheduling, and multi-node collaborative execution technology, and in particular to overlay network capability slicing routing and collaborative execution methods, systems, media, program products, and terminals based on edge-side lightweight models. Background Technology

[0002] As model capabilities continue to improve, the approach of a single node carrying the entire model capacity faces significant limitations, including high deployment costs, uneven resource utilization, wide-ranging impact of node failures, and fixed request processing paths. Especially after large model capabilities are broken down into multiple inference slices, the invisibility of tensor dependencies, the disruption of inference pipeline states, the latency of cross-node KVCache synchronization delays, and the recalculation costs caused by the failure of aggregation nodes further amplify latency jitter and failure recovery pressures in the overlay network. Existing technologies typically suffer from the following shortcomings:

[0003] (1) Model capabilities and network scheduling are disconnected: Existing solutions usually focus on how the model is split in a single node or cluster, or focus on ordinary network routing and service scheduling, lacking a mechanism to directly use model capability slices as overlay network endogenous capability nodes for unified routing and orchestration.

[0004] (2) It is difficult to form a unified discovery and calling interface for slice capabilities: When the model is broken down into multiple local capability units, the existing system usually lacks a unified capability profile description method, making it difficult to express the capability type, confidence level, resource boundary, dependency relationship and current availability of each capability slice.

[0005] (3) Lack of linkage between call order and network status: Existing orchestration methods often only call capability units according to task logic, without comprehensively considering path quality, latency, reachability, budget status and relay backoff conditions in the coverage network, resulting in high cross-node call costs, slow convergence and difficulty in recovery after failure.

[0006] (4) Slice dependencies and inference states are not visible: In multi-slice collaborative reasoning scenarios, tensor dependencies, interface constraints and intermediate state boundaries between predecessor slices, successor slices and convergence slices usually lack a structured expression that can directly participate in network routing decisions.

[0007] (5) The execution of capability slices lacks a rerouting and re-arrangement mechanism after failure: When a certain capability node is unreachable, execution times out, the result confidence is insufficient, or the network state degrades, the existing solutions can usually only retry the whole or switch to a fixed one, lacking the ability to switch replacement nodes, reorder or downgrade execution according to the capability slice granularity.

[0008] (6) Lack of local autonomous control over capability slices on the end side: If complex task requests rely entirely on remote centralized scheduling, the end side will find it difficult to make timely adjustments to the capability slice call path based on local network status, slice execution feedback and aggregation status.

[0009] Therefore, a mechanism is needed to unify the description of model capability slices as capability node resources within the overlay network, enabling dependency discovery, task routing, collaborative execution, and reordering after failure. Summary of the Invention

[0010] In view of the shortcomings of the prior art, the present invention provides a method, system, medium, program product and terminal for overlay network capability slicing routing and cooperative execution based on a lightweight end-side model, which is used to solve at least one of the technical problems in the prior art.

[0011] To achieve the above and other related objectives, the first aspect of this application provides a method for overlay network capability slice routing and collaborative execution based on a lightweight end-side model, comprising: abstracting each model capability slice of a distributed deployed lightweight end-side model into capability slice nodes in an overlay network, and generating a capability profile corresponding to each capability slice node; receiving a complex task request, and having the lightweight end-side model decompose the complex task request into a task, generating at least one capability invocation sub-step; collecting the capability profiles and network status information of each capability slice node in the overlay network in real time; based on at least one capability invocation sub-step, the capability profile of each capability slice node, and the network status information, filtering, sorting, and prioritizing the capability slice nodes in the overlay network, determining the capability slice nodes, invocation order, and parallel and / or serial relationships corresponding to each capability invocation sub-step, and orchestrating and generating model capability slice invocation paths for execution of model capability slice routing in the overlay network; and rearranging unexecuted capability invocation sub-steps according to the execution results of the model capability slice routing.

[0012] In some embodiments of the first aspect of this application, based on the capability invocation sub-steps, the capability profiles of the capability slice nodes, and network status information, the capability slice nodes in the coverage network are filtered, sorted, and prioritized to determine the capability slice nodes, invocation order, and parallel and / or serial relationships corresponding to each capability invocation sub-step. The specific process includes: based on the capability profiles and network status information of each capability slice node, combined with the dependencies between each capability invocation sub-step, input / output interface constraints, and convergence slice switching conditions, determining the primary execution capability slice node, backup capability slice node, and convergence capability slice node for each capability invocation sub-step, and obtaining the invocation order between each capability invocation sub-step; determining the serial invocation chain and parallel invocation group based on the dependencies between each capability invocation sub-step, and obtaining the parallel and / or serial relationships between each capability invocation sub-step.

[0013] In some embodiments of the first aspect of this application, the process of reordering unexecuted capability call sub-steps based on the execution result of the execution model capability slice routing includes: if the execution result of the execution model capability slice routing includes at least one of capability slice node failure, capability slice node timeout, insufficient result confidence, insufficient budget, aggregation capability slice node abnormality, excessive cross-node state synchronization cost, or network state degradation, then the unexecuted capability call sub-steps are reordered.

[0014] In some embodiments of the first aspect of this application, the process of rearranging unexecuted capability call sub-steps based on the execution result of the execution model capability slice routing includes: when the confidence level of the result in the execution result of the execution model capability slice routing is lower than the preset confidence level requirement and the budget status allows, then during the rearranging of unexecuted capability call sub-steps, a backup capability slice node switch or a convergence capability slice node switch based on topological distance is performed.

[0015] In some embodiments of the first aspect of this application, the process of reordering unexecuted capability call sub-steps based on the execution result of the execution model capability slice routing includes: when the budget state in the execution result of the execution model capability slice routing is lower than the preset budget requirement or the cross-node state synchronization cost is too high, then during the reordering of unexecuted capability call sub-steps, the capability slice node depth is downgraded, the end-side conservative inference is downgraded, or low-priority sub-steps are skipped.

[0016] In some embodiments of the first aspect of this application, the capability invocation sub-step includes at least one of: sub-step identifier, sub-step type, required capability type, input dependency, execution priority, acceptable delay range, result confidence requirement, or whether parallel execution is allowed.

[0017] To achieve the above and other related objectives, a second aspect of this application provides a capability slice routing and collaborative execution system for overlay networks based on a lightweight end-side model, comprising: a capability slice node abstraction module, used to abstract the capability slices of each model of the distributed deployed lightweight end-side model into capability slice nodes in the overlay network, and generate a capability profile corresponding to each capability slice node; a complex task request decomposition module, used to receive complex task requests, and the lightweight end-side model decomposes the complex task requests to generate at least one capability invocation sub-step; and a profile and information collection module, used to collect information on each capability slice node in the overlay network in real time. The system includes a capability slice node capability profile and network status information; a capability slice routing and execution module, used to filter, sort, and prioritize capability slice nodes in the overlay network based on at least one capability invocation sub-step, the capability profile of each capability slice node, and network status information, to determine the capability slice nodes, invocation order, and parallel and / or serial relationships corresponding to each capability invocation sub-step, and to orchestrate and generate model capability slice invocation paths for execution of model capability slice routing in the overlay network; and a sub-step re-orchestration module, used to re-orchestrate unexecuted capability invocation sub-steps based on the execution results of the model capability slice routing.

[0018] To achieve the above and other related objectives, a third aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the overlay network capability slice routing and cooperative execution method based on an end-side lightweight model.

[0019] To achieve the above and other related objectives, a fourth aspect of this application provides a computer program product comprising computer program code, which, when executed on a computer, enables the computer to implement the overlay network capability slice routing and cooperative execution method based on the end-side lightweight model.

[0020] To achieve the above and other related objectives, a fifth aspect of this application provides an electronic terminal, including a memory, a processor, and a computer program stored in the memory; the processor executes the computer program to implement the overlay network capability slicing routing and cooperative execution method based on the end-side lightweight model.

[0021] As described above, the overlay network capability slicing routing and cooperative execution method, system, media, program product, and terminal based on the end-side lightweight model provided in this application have the following beneficial effects:

[0022] (1) This application realizes the native organization of model capability slices in the overlay network by abstracting model capability slices into discoverable, filterable and routable capability node resources.

[0023] (2) The application can complete the sub-step decomposition of capability invocation and slice routing orchestration locally on the client side, without relying entirely on the central control plane, thus enhancing the autonomous processing capability of the lightweight client-side model for complex task requests.

[0024] (3) This application comprehensively considers node capabilities, path state interface constraints, dependency summary and budget state, etc., to achieve joint optimization of network state, dependency relationship and model capability slice.

[0025] (4) This application reduces redundant slice switching and waiting time by arranging the order of sub-step calls, the location of converged slices and parallel relationships, thereby improving the efficiency of cross-node collaborative execution.

[0026] (5) If there are slice node failures, timeouts, insufficient result confidence or abnormal convergence slices during the execution of this application, the local capability slice paths can be replaced, converged, and rearranged to avoid overall task failure. Attached Figure Description

[0027] Figure 1 The diagram shown is a flowchart illustrating a method for overlay network capability slice routing and collaborative execution based on a lightweight end-side model in one embodiment of this application.

[0028] Figure 2 The diagram shown is a specific embodiment of a method for overlay network capability slicing routing and cooperative execution based on an end-side lightweight model, as described in this application.

[0029] Figure 3 The diagram shown is a structural schematic of a coverage network capability slice routing and cooperative execution system based on an end-side lightweight model according to an embodiment of this application.

[0030] Figure 4 The diagram shown is a specific embodiment of a coverage network capability slice routing and cooperative execution system based on an end-side lightweight model, as described in this application.

[0031] Figure 5 The diagram shown is a structural schematic of an electronic terminal according to an embodiment of this application. Detailed Implementation

[0032] The following specific examples illustrate the implementation of this application. Those skilled in the art can easily understand other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, unless otherwise specified, the following embodiments and features in the embodiments can be combined with each other.

[0033] In the embodiments of this application, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" do not necessarily imply that they are different.

[0034] It should be noted that, in the embodiments of this application, the words "exemplary" or "for example" indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0035] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0036] The method for routing and co-executing overlay network capability slices based on a lightweight end-side model provided in this application discovers, filters, routes, orchestrates, and co-executes distributed model capability slices as capability slice node resources, and adaptively adjusts subsequent capability invocation paths based on network status, capability slice node status, and execution feedback results.

[0037] To facilitate understanding of the embodiments of this application, firstly, in conjunction with Figure 1 and Figure 2 Detailed explanation. Figure 1 This document illustrates a flowchart of a method for overlay network capability slicing routing and cooperative execution based on an edge-side lightweight model, as described in an embodiment of the present invention. The method for overlay network capability slicing routing and cooperative execution based on an edge-side lightweight model in this embodiment mainly includes the following steps:

[0038] Step S11: Abstract the capability slices of each model of the distributed deployed edge lightweight model into capability slice nodes in the overlay network, and generate a capability profile corresponding to each capability slice node.

[0039] It should be noted that the lightweight edge model is deployed on edge devices, including but not limited to client devices, local agent processes, edge node agents, or user-space network components. It enables edge data processing without sending data to the cloud, allowing for data processing in weak network or offline environments, thus demonstrating high practicality. Lightweight edge models can employ lightweight models such as character-level convolutional neural networks, gradient boosting trees, or quantized Transformer models. Deploying lightweight edge models allows for real-time task processing on edge devices, such as image segmentation and object detection.

[0040] In this embodiment, the distributed, lightweight edge models possess various model capabilities. For example, the image-related lightweight models include capabilities such as image preprocessing, object detection, feature extraction, classification and recognition, and defect detection; the speech-related lightweight models include capabilities such as speech denoising, feature extraction, speech recognition, semantic parsing, and instruction conversion; and the time-series analysis lightweight models include capabilities such as data filtering, trend prediction, anomaly detection, and feature fitting. Each model capability is divided into several model capability slices. For example, the image preprocessing model capability is divided into several local image preprocessing model capability slices, and so on. Further details are omitted here.

[0041] Furthermore, in the overlay network formed by the distributed deployment of lightweight edge models, model capability slices are abstracted into capability slice nodes that can be discovered, filtered, and invoked. These capability slice nodes refer to distributed computing entities within the overlay network that undertake local model capability processing and can expose standardized capability input and output interfaces.

[0042] The capability slice nodes include, but are not limited to, local inference capability nodes, local retrieval capability nodes, local verification capability nodes, local generation capability nodes, local filtering capability nodes, or local aggregation capability nodes. Each capability slice node, as a virtual network node in the overlay network, has a unique network identifier and exists independently in the network topology, separate from the original lightweight model carrier on the edge side. This supports single-point addressing, independent invocation, dynamic orchestration, and cross-device collaboration of each model capability slice in the overlay network.

[0043] Specifically, each capability slice node exists independently in the Overlay topology, detached from the original lightweight end-side model carrier. Each capability slice node is configured with a unique Overlay ID. When the lightweight end-side model performs data interaction and route scheduling, it forwards IP packets based on the Overlay ID corresponding to each capability slice node, or encapsulates and transmits data using tunneling protocols such as VXLAN and GRE. Throughout the scheduling process, routing decisions are made based on the capability profile and network location information of each capability slice node.

[0044] In this embodiment, each capability slice node in the overlay network maintains a corresponding capability profile. The capability profile can be jointly generated or incrementally updated in real time from control plane metadata, node registration information, node self-reported capabilities, historical feedback records, capability catalogs, and other equivalent capability description structures. The capability profile can be used to achieve node identification, task matching, intelligent scheduling, and status management in the overlay network. The capability profile includes, but is not limited to: capability slice identifier, capability type, capability summary, processable request category, input / output interface summary, dependency summary, historical success rate, result confidence range, current load, current serviceable concurrency, service scope, accessible data range, budget status, and node health status, as shown in Figure 1, which illustrates some fields in the capability profile. These features enable accurate characterization of the identity attributes, functional boundaries, operational performance, resource status, service scope, and constraints of each capability slice node.

[0045] In this embodiment, the lightweight end-side model collects the running data of each capability slice node based on a sliding time window, and incrementally updates the capability profile of each capability slice node based on the running data; for example, when a capability slice node is detected to have consecutive timeouts or result conflicts, the weight priority of the capability slice node is dynamically reduced based on an exponential decay factor.

[0046] It should be noted that the fields in the capability profile of each capability slice node are not statically configured. Instead, a lightweight model on the edge maintains a sliding time window and incrementally updates the capability profile based on the operational data of the capability slice node collected within this sliding time window. The sliding time window can be set according to time length or task call count to limit the range of recent data participating in the profile update, ensuring that the capability profile reflects the current operational capabilities and service status of the capability slice node. Compared to updating based on full historical data, this method reduces computational overhead.

[0047] Table 1. Field Descriptions in the Capability Profile

[0048] slice_id Capability slice identifier Differentiate slice nodes with different capabilities and participate in directory discovery This can be achieved through capability slice registration identifiers. capability_type Ability Type Determine if it is appropriate to process a certain capability call sub-step This can be achieved through capability enumeration or equivalent tags. capability_summary Capability Summary Describe the scope of tasks that a slice node can handle. This can be achieved through control plane metadata or capability catalog. dependency_profile Dependency Summary Indicates predecessor / successor slice constraints or convergence prerequisites. This can be achieved through dependency graphs or task relationship structures. confidence_range Confidence range of results Assisted result filtering, aggregation verification, and switchover judgment This can be achieved through historical statistics. interface_profile Input / output interface summary Describe the data types and interface boundaries that a slice can receive / output. This can be achieved through a capability catalog or interface description structure. historical success rate Historical success rate Evaluate the reliability of slice nodes This can be achieved by executing feedback logs. current_load Current load Determine whether it is suitable to continue taking orders. This can be achieved through the node's running status. budget_state Budget status Participate in slice node selection, shrinking, and degradation decisions. This can be achieved through session-level or node-level budgeting. service_scope Service scope Determine if the tenant / namespace / service matches. This can be achieved through control plane semantic fields. health_status Node health status Assisted filtering, switching, and rerouting This can be achieved through health checks or heart rate status.

[0049] Step S12: Receive a complex task request, and the lightweight model on the edge side decomposes the complex task request into a task, generating at least one capability invocation sub-step.

[0050] In one embodiment, the complex task request includes, but is not limited to: multi-stage inference requests, retrieval and generation combination requests, classification, verification and summarization requests, composite requests that require the serial processing of multiple capability slices, and model capability invocation requests that require cross-node collaboration.

[0051] In one embodiment, the capability invocation sub-step includes, but is not limited to: sub-step identifier, sub-step type, required capability type, input dependency, execution priority, acceptable latency range, budget constraint, result confidence requirement, whether parallel execution is allowed, whether degraded execution is allowed, and whether switching to a backup capability slice node is allowed.

[0052] In one embodiment, the inputs of the edge-side lightweight model include, but are not limited to: semantic summaries of complex task requests, capability invocation contexts, predecessor-successor dependency summaries, input / output interface constraints, budget states, service scopes, or equivalent slice orchestration features; the outputs of the edge-side lightweight model include, but are not limited to: sets of capability invocation sub-steps, convergent slice switching suggestions, primary / backup slice candidate suggestions, edge-side conservative inference degradation suggestions, or equivalent cooperative execution control parameters.

[0053] Specifically, when the lightweight client-side model outputs capability call sub-steps, it exchanges the capability call sub-step information with the subsequent routing orchestration module through a local state cache, asynchronous call description structure, or equivalent scheduling structure. Specifically, the lightweight client-side model writes the capability call sub-step information into the local state cache or asynchronous call description structure, interacting with the subsequent routing orchestration module through a non-blocking input / output mechanism. This avoids request queuing and throughput reduction caused by synchronous blocking, allowing completion without blocking the original data plane hot path. When there are dependency conflicts, interface constraint mismatches, unmet convergence slice switching conditions, or equivalent collaborative execution anomalies among the capability call sub-steps output by the lightweight client-side model, the system performs dependency verification, interface compatibility verification, or rule correction on the output, and downgrades the corresponding call path to a conservative collaborative execution scheme.

[0054] Step S13: Collect capability profiles and network status information of each capability slice node in real time within the overlay network.

[0055] In one embodiment, the network status information includes, but is not limited to: latency, packet loss rate, reachability, current path status, candidate path quality, bandwidth budget, path degradation status, most recent failure record, most recent recovery record, relay availability, etc.

[0056] Step S14: Based on at least one capability invocation sub-step, the capability profile of each capability slice node, and network status information, filter, sort, and prioritize the capability slice nodes in the overlay network, determine the capability slice nodes, invocation order, and parallel and / or serial relationships corresponding to each capability invocation sub-step, and orchestrate and generate model capability slice invocation paths for execution of model capability slice routing in the overlay network.

[0057] It should be explained that capability profiles and network status information together form the unified input for filtering capability slice nodes. The preset conditions for filtering capability slice nodes in the overlay network include: whether the capability profile meets the capability type required by the capability invocation sub-step; whether the node's current load meets the execution budget; whether the node's network status meets latency, reachability, or budget requirements; whether the node belongs to an allowed tenant, namespace, or service scope; whether the node's historical success rate, result confidence, or health status is higher than preset conditions; whether the node is suitable as a primary execution capability slice node, a backup capability slice node, or a convergence capability slice node; and whether the node meets at least one of the following: predecessor-successor dependency summary, input / output interface constraints, or convergence slice switching conditions.

[0058] In one embodiment of this application, based on the capability invocation sub-step, the capability profile of the capability slice node, and network status information, the capability slice nodes in the coverage network are filtered, sorted, and prioritized to determine the capability slice nodes corresponding to each capability invocation sub-step, the invocation order, and the parallel and / or serial relationships. The specific process includes:

[0059] Based on the capability profiles and network status information of each capability slice node, combined with the dependencies between each capability invocation sub-step, input / output interface constraints, and convergence slice switching conditions, the primary execution capability slice node, backup capability slice node, and convergence capability slice node of each capability invocation sub-step are determined, and the invocation order between each capability invocation sub-step is obtained; based on the dependencies between each capability invocation sub-step, the serial invocation chain and parallel invocation group are determined, and the parallel and / or serial relationships between each capability invocation sub-step are obtained.

[0060] The process of determining serial call chains and parallel call groups based on the dependencies between various capability call sub-steps, and obtaining the parallel and / or serial relationships between various capability call sub-steps, includes: when there are no input dependencies between multiple capability call sub-steps, they are determined to be parallel call groups; when there are predecessor-successor dependencies or output interface coupling between multiple capability call sub-steps, they are determined to be serial call chains based on the dependency summary.

[0061] The dependencies between capability invocation sub-steps refer to relationships where capability invocation sub-step A outputs its result first, and the input of capability invocation sub-step B is the output of capability invocation sub-step A; or capability invocation sub-steps C and D are independent of each other and can be executed in parallel. Input / output interface constraints refer to whether the input format, output format, data protocol, parameter type, etc., of a capability invocation sub-step match a certain capability slice node. The aggregation slice switching condition refers to the capability slice node that acts as the aggregation capability slice node when the execution results of multiple capability slice nodes need to be aggregated to the same capability slice node. Specifically, the primary execution capability slice node is the capability slice node that is primarily responsible for executing the capability invocation sub-step; the backup capability slice node is the capability slice node used to take over execution when the primary execution capability slice is unavailable, its performance degrades, or it does not meet the switching conditions; the aggregation capability slice node is the capability slice node used to receive the execution results of multiple capability invocation sub-steps or multiple capability slice nodes, and to merge, filter, fuse, or forward the results. The order of calling each capability call sub-step refers to determining which capability call sub-steps are executed first and which are executed later; the parallel and / or serial relationship between each capability call sub-step refers to which capability call sub-steps are executed in parallel, or which capability call sub-steps are executed individually or sequentially.

[0062] Specifically, capability profiles, network status information, dependencies between various capability invocation sub-steps, and input / output interface constraints are used for joint decision-making to generate the priority for capability slice node selection. During the selection of capability slice nodes in the overlay network, capability slice nodes can first be constrained and selected based on capability type, scope constraints, or service boundaries. Then, network status information, candidate path quality, budget status, dependency summaries, or input / output interface constraints are combined to sort and prioritize the selected capability slice nodes, thereby obtaining the primary execution capability slice node, backup capability slice node, and convergence capability slice node for each capability invocation sub-step, and simultaneously obtaining the invocation order between each capability invocation sub-step. Then, based on the dependencies between each capability invocation sub-step, serial invocation chains and parallel invocation groups are determined. A serial invocation chain refers to the serial invocation relationship between capability slice nodes, and a parallel invocation group refers to the parallel invocation relationship between capability slice nodes. The dependencies between each capability invocation sub-step form a directed dependency relationship involving predecessor slices, successor slices, and convergence slices, which can constitute a local task graph, invocation order chain, topology dependency graph, or equivalent orchestration structure, etc.

[0063] Furthermore, after determining the capability slice nodes, invocation order, and parallel and / or serial relationships corresponding to each capability invocation sub-step, the model capability slice invocation path is orchestrated and generated. The model capability slice invocation path includes, but is not limited to: the main execution capability slice node corresponding to each capability invocation sub-step, the order of sub-steps, parallel group division, convergence capability slice node, backup capability slice node, degraded execution path, retry order after timeout, sub-step budget constraints, convergence slice switching rules, topological distance constraints, or equivalent slice adjacency constraints, etc.

[0064] Finally, model capability slice routing is executed in the overlay network according to the generated model capability slice invocation path. Executing model capability slice routing specifically includes: serial invocation; parallel invocation; condition-triggered invocation; retrieval before generation; classification before verification; executing local capability slice nodes before executing converged capability slice nodes; executing high-confidence capability slice nodes first, and then deciding whether to add backup capability slice nodes, etc.

[0065] Step S15: Based on the execution results of the capability slice routing of the execution model, rearrange the unexecuted capability call sub-steps.

[0066] In one embodiment of this application, the process of reordering unexecuted capability invocation sub-steps based on the execution result of the execution model capability slice routing includes: if the execution result of the execution model capability slice routing includes at least one of the following: capability slice node failure, capability slice node timeout, insufficient result confidence, insufficient budget, abnormal aggregation capability slice node, excessive cross-node state synchronization cost, or network state degradation, then the unexecuted capability invocation sub-steps are reordered.

[0067] It should be noted that in the overlay network, capability invocation sub-steps are executed according to the capability slice invocation path of the generated model, and the execution results returned by each capability slice node are received, filtered, aggregated, or subsequently scheduled. Execution results include the output results of multiple nodes, result consistency, result confidence, or scope constraints, which can be used to determine the subsequent capability slice routing direction. Specifically, if at least one of the following conditions occurs in the execution results, the subsequent unexecuted capability invocation sub-steps are re-arranged. These conditions include: capability slice node failure, capability slice node timeout, insufficient result confidence, insufficient budget, abnormal aggregation of capability slice nodes, excessive cost of cross-node state synchronization, network state degradation, capability slice node unreachable, and multi-slice result conflicts. The methods for reordering subsequent unexecuted capability call sub-steps include: switching to a backup capability slice node; adjusting the execution order of subsequent capability call sub-steps; reducing parallelism; reducing the depth of subsequent capability slice nodes; abandoning some low-priority sub-steps; switching to single-node conservative execution; rolling back to the basic capability path; requesting a new set of capability slice nodes; performing a topological distance-based convergence capability slice node switch; and performing at least one of the following: edge-side conservative inference degradation or local result-based minimum output.

[0068] For example, when it is found, based on the result confidence level, remaining budget status, cross-node status synchronization cost, estimated call latency, or convergence slice load trend prediction, that the subsequent collaborative execution completion time may exceed the remaining budget, the continued participation of the current slice will lead to a significant increase in tensor synchronization and serialization overhead, or the convergence slice switching cost is lower than the expected retry cost of continuing to maintain the current path, this embodiment can, without waiting for the capability slice node to enter a clear timeout or failure state, actively trigger convergence slice switching, backup slice node switching, parallelism reduction, or end-side conservative inference downgrade to rearrange the unexecuted capability call sub-steps.

[0069] In one embodiment of this application, the process of re-arranging unexecuted capability call sub-steps based on the execution result of the execution model capability slice routing includes: when the confidence level of the result in the execution result of the execution model capability slice routing is lower than the preset confidence level requirement and the budget status allows, then during the re-arranging of unexecuted capability call sub-steps, a backup capability slice node switch or a convergence capability slice node switch based on topological distance is performed.

[0070] The process of switching the aggregation capability slice node based on topology distance includes: when the confidence level of the execution result of the execution model capability slice routing is lower than the preset requirement, and the budget status still allows to continue calling other capability slices, then a new capability slice node is found for the subsequent unexecuted capability call sub-steps. At this time, the corresponding aggregation capability slice node should also be replaced synchronously. The topology distance between the current new capability slice node and other capability slice nodes is calculated, and then the capability slice node with the smallest topology distance is selected as the aggregation capability slice node according to the topology distance result.

[0071] In one embodiment of this application, the process of re-arranging unexecuted capability call sub-steps based on the execution result of the execution model capability slice routing includes: when the budget state in the execution result of the execution model capability slice routing is lower than the preset budget requirement or the cross-node state synchronization cost is too high, the process of re-arranging unexecuted capability call sub-steps involves performing capability slice node depth degradation, end-side conservative inference degradation, or skipping low-priority sub-steps.

[0072] The process of conservative inference degradation on the edge includes: when the budget status in the execution result of the execution of the model capability slice routing is lower than the preset budget requirement, and it is no longer possible to add other models, other capability slice nodes or complex verification processes, in order to avoid interrupting the entire inference process, those complex verification sub-steps that rely on low confidence results will no longer be executed. Instead, an approximate result that can continue to be passed on will be generated using the currently obtained features, context, intermediate representations or local judgments, and transmitted to the subsequent unexecuted capability call sub-steps.

[0073] In one embodiment of this application, based on the execution results of the capability slice routing of the execution model, the node execution effect and network state changes are used as feedback information. This feedback information is then used to correct subsequent capability invocation sub-steps, capability profiles, selection thresholds, or routing orchestration logic. The feedback information includes, but is not limited to: sub-step success rate, number of node failures, number of timeouts, node execution time, consistency of aggregation results, node result confidence, number of backup slice switches, number of degraded executions, budget consumption, and path switching records.

[0074] In one embodiment, the feedback correction includes at least one of the following: updating the capability slice node profile, updating node priority, updating the mapping relationship between capability invocation sub-steps and capability types, updating the routing orchestration threshold, updating the invocation order preference, updating the parallel or serial partitioning strategy, updating the backup slice selection rule, and updating the degradation execution rule. Table 2 shows a description of some fields in the capability invocation sub-steps, execution results, and feedback correction.

[0075] Table 2. Description of fields in Capability Invocation Sub-steps, Execution Results, and Feedback Corrections

[0076] step_type Capability Invocation Sub-Step Type This indicates the steps involved in classification, retrieval, generation and validation, and aggregation. This can be achieved through a task description structure. Step slice mapping Sub-step-slice mapping relationship This indicates which candidate slices can perform a certain capability invocation sub-step. This can be achieved through directory discovery results and capability matching results. dependency_relation Dependency Indicates the order or dependency of capability invocation sub-steps. This can be achieved through task chains or local task graphs. aggregation_slice Aggregation Capability Slice Node This function is responsible for the aggregation, comparison, or final output of multi-slice results. This can be achieved by arranging the results or by using a set of slices. primary_slice Main execution capability slice node Preferred Capability Slice Node This can be determined by filtering and sorting the results. fallback_slice Backup capability slice node Alternate execution when the main slice is unavailable This can be achieved through a candidate slice set. parallel_group Parallel group identifier This indicates that multiple sub-steps or multiple slices can be executed in parallel. Realization can be achieved through grouping structure degrade_path Degraded execution path Simplified Implementation Plan When Budget is Insufficient or Fails This can be achieved through orchestration strategies. result_conflict Result of conflict state Determine whether convergence slice switching or additional verification is needed. This can be achieved through result consistency verification. route_update Routing update action Based on execution feedback, switch nodes, adjust convergence locations, or rearrange sequences. This can be achieved through the local scheduling module.

[0077] In one embodiment, the feedback correction is performed by an edge-side lightweight model, a local orchestration module, or an equivalent correction module, and is completed without blocking the original data plane hot path.

[0078] In one embodiment, the number of node failures, node execution time, convergence result consistency, number of backup slice switching, budget consumption, and number of degradation executions are maintained through a sliding window, time decay factor, or equivalent historical statistics mechanism. This provides feedback correction for subsequent capability invocation sub-steps, capability profiles, selection thresholds, or routing orchestration logic, thereby reducing the impact of outdated slice profiles or outdated execution feedback on subsequent slice selection, convergence switching, and rerouting decisions.

[0079] In one embodiment, when the edge-side lightweight model performs route decision based on the locally cached capability slice profile, network status, aggregation slice switching conditions, or dependency summary, if it finds that the control plane metadata is inconsistent with the actual data plane status, the cache timestamp has expired, the target slice node is offline, the aggregation slice node is unavailable, or the interface compatibility conditions have changed, the system triggers fast rerouting and updates the local cache timestamp, aggregation slice availability status, and subsequent collaborative execution inputs.

[0080] This application supports the distributed deployment, collaborative execution, feedback correction, and conservative inference of model capability slices in the network, and reduces the recalculation overhead after the interruption of the large model inference pipeline through slice dependency discovery, topology-aware rerouting, and end-side conservative inference degradation.

[0081] Combination Figure 2 The specific implementation process of the overlay network capability slice routing and cooperative execution method based on the end-side lightweight model provided in this application includes:

[0082] The system receives complex task requests, parses the requests to generate at least one capability invocation sub-step, obtains each capability slice node and reads its capability profile; constructs a mapping relationship between capability invocation sub-steps and capability slice nodes; generates a slice dependency graph and determines the predecessor slice, successor slice, convergence slice, and serial and / or parallel groups; collects network status information of capability slice nodes and evaluates the primary path, relay path, and backup path, assigning a primary execution capability slice node, backup capability slice node, or convergence capability slice node to each capability invocation sub-step; executes collaborative invocation and aggregates the execution results; and determines whether to continue subsequent slice invocations, switch backup slices, adjust the convergence position, or trigger rerouting based on the execution results.

[0083] To facilitate understanding of the overlay network capability slicing routing and cooperative execution method based on the end-side lightweight model of this application, the following embodiments are provided for illustration.

[0084] Example 1.

[0085] When the lightweight edge model receives a complex task request, it first breaks the request down into a retrieval sub-step and a generation sub-step. Then, based on the capability profile of the capability slice nodes, it selects the first capability slice node to execute the retrieval sub-step, and then inputs the retrieval results into the second capability slice node to execute the generation sub-step. If the first capability slice node times out, it switches to the backup retrieval slice node to continue execution.

[0086] Example 2.

[0087] For the same capability invocation sub-step, the capability profiles and network status information of each capability slice node are collected in real time. If capability slice node A has a high capability confidence but high network latency, and capability slice node B has a slightly lower capability confidence but lower network latency, then a judgment is made based on budget status, task priority, and result confidence requirements to determine whether capability slice node A or capability slice node B should be selected first.

[0088] Example 3.

[0089] When a capability slice node becomes unreachable, fails to execute, or returns a result with insufficient confidence, the current complex request is not abandoned entirely. Instead, subsequent unexecuted capability call sub-steps are paused, and a backup capability slice node is selected, or the execution order of subsequent sub-steps is adjusted. For example, the original "retrieve and generate" call chain can be switched to "continue generation after retrieving backup slices" or changed to a degraded execution path of "directly execute simplified generated slices" when the retrieval of slices fails consecutively.

[0090] Example 4.

[0091] When the current budget is lower than the preset budget requirement, reduce the number of capability slice calls in this round, reduce the size of the parallel execution group, prioritize the retention of high-priority sub-steps, and skip some low-priority verification or aggregation slices, so as to maintain the minimum executable capability of complex task requests under budget constraints.

[0092] It should be emphasized that the overlay network capability slicing routing and cooperative execution method based on the end-side lightweight model provided in this application has the following beneficial effects:

[0093] (1) This application realizes the native organization of model capability slices in the overlay network by abstracting model capability slices into discoverable, filterable and routable capability node resources.

[0094] (2) The application can complete the sub-step decomposition of capability invocation and slice routing orchestration locally on the client side, without relying entirely on the central control plane, thus enhancing the autonomous processing capability of the lightweight client-side model for complex task requests.

[0095] (3) This application comprehensively considers node capabilities, path state interface constraints, dependency summary and budget state, etc., to achieve joint optimization of network state, dependency relationship and model capability slice.

[0096] (4) This application reduces redundant slice switching and waiting time by arranging the order of sub-step calls, the location of converged slices and parallel relationships, thereby improving the efficiency of cross-node collaborative execution.

[0097] (5) If there are slice node failures, timeouts, insufficient result confidence or abnormal convergence slices during the execution of this application, the local capability slice paths can be replaced, converged, and rearranged to avoid overall task failure.

[0098] It is important to emphasize that traditional methods rely on deterministic hardware resources for scheduling and heartbeat messages to detect node status. This application, however, uses the inference capabilities of AI models with dynamic probabilistic characteristics as the scheduling object. It integrates the inference load, cache status, inference confidence, and computing power margin of the model capability slices, along with network link indicators, for joint routing scheduling. This scheduling mode breaks through the capability boundaries of traditional container scheduling and is impossible to achieve with conventional distributed scheduling architectures. It can precisely adapt to the distributed operation requirements of lightweight model capability slices on the edge side.

[0099] Figure 3 This is a schematic block diagram of a coverage network capability slice routing and cooperative execution system based on an end-side lightweight model provided in an embodiment of this application. Figure 3 As shown, the system includes a capability slice node abstraction module 310, a complex task request decomposition module 320, a profile and information collection module 330, a capability slice routing and execution module 340, and a sub-step reordering module 350.

[0100] The capability slice node abstraction module 310 is used to abstract the capability slices of each model of the distributed deployed edge lightweight model into capability slice nodes in the overlay network, and generate a capability profile corresponding to each capability slice node.

[0101] The complex task request decomposition module 320 is used to receive complex task requests and decompose the complex task requests by the end-side lightweight model to generate at least one capability invocation sub-step.

[0102] The profiling and information acquisition module 330 is used to collect capability profiles and network status information of each capability slice node in the overlay network in real time.

[0103] The capability slice routing and execution module 340 is used to filter, sort and prioritize capability slice nodes in the overlay network based on at least one capability call sub-step, the capability profile of each capability slice node and network status information, determine the capability slice nodes, call order and parallel and / or serial relationships corresponding to each capability call sub-step, and arrange and generate model capability slice call paths for execution of model capability slice routing in the overlay network.

[0104] The sub-step reordering module 350 is used to reorder unexecuted capability call sub-steps based on the execution results of the execution model capability slice routing.

[0105] Combination Figure 4 The specific implementation process of the overlay network capability slicing routing and cooperative execution system based on the edge-side lightweight model provided in this application includes:

[0106] The capability slice node abstraction module 310 is used to abstract the capability slices of each model of the distributed deployed edge lightweight model into capability slice nodes in the overlay network, and generate a capability profile corresponding to each capability slice node.

[0107] The complex task request decomposition module 320 includes a complex task request access unit 321 and an edge task request decomposition unit. The complex task request access unit 321 receives complex task requests, and the edge task request decomposition unit generates capability invocation sub-steps based on the edge lightweight model.

[0108] The profiling and information acquisition module 330 includes a network information acquisition unit 331 and a capability profiling acquisition unit 332, used to acquire capability profiling and network status information of each capability slice node. The profiling and information acquisition module 330 is connected to the capability slice node abstraction module 310.

[0109] The capability slice routing and execution module 340 includes a slice call path generation unit 341 and a profile and information acquisition module 330. A slice routing execution unit 342 is also connected. The capability slice routing and execution module 340 is connected to the complex task request decomposition module 320 and the slice call path generation unit 341. The slice call path generation unit 341 evaluates the availability of the primary path, relay path, or backup path based on the capability profile and network status information of each capability slice node. Simultaneously, based on the slice dependency graph, capability profile, and path evaluation results, it determines the primary execution slice, backup slice, convergence slice, and parallel and / or serial call relationships, ultimately orchestrating and generating model capability slice call paths. The slice routing execution unit 342 executes model capability slice routing in the overlay network according to the model capability slice call paths.

[0110] The sub-step re-arrangement module 350 is connected to the capability slice routing and execution module 340. It is used to re-arrange unexecuted capability call sub-steps by performing alternative slice switching, aggregation position adjustment, call order re-arrangement, or degradation and shrinking based on slice execution results, node timeouts, node failures, changes in result confidence, or path degradation.

[0111] It should be understood that the specific process of each module performing the above-mentioned steps has been described in detail in the above method embodiments, and will not be repeated here for the sake of brevity.

[0112] It should also be understood that the module division in the embodiments of this application is illustrative and only represents a logical functional division; in actual implementation, there may be other division methods. Furthermore, the functional modules in the various embodiments of this application can be integrated into a single processor, exist as separate physical entities, or be integrated into a single module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0113] Figure 5 This is a schematic block diagram of the electronic terminal provided in an embodiment of this application. Figure 5 As shown, the electronic terminal 500 includes at least one processor 501, a memory 502, at least one network interface 503, and a user interface 505. The various components in the electronic terminal 500 are coupled together via a bus system 504. It is understood that the bus system 504 is used to implement communication between these components. In addition to a data bus, the bus system 504 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in… Figure 5 The general will label all buses as bus systems.

[0114] The user interface 505 may include a monitor, keyboard, mouse, trackball, clicker, button, touchpad, or touch screen.

[0115] It is understood that memory 502 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.

[0116] In this embodiment of the invention, the memory 502 is used to store various types of data to support the operation of the electronic terminal 500. Examples of this data include: any executable program for operation on the electronic terminal 500, such as the operating system 5021 and application programs 5022; the operating system 5021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 5022 may contain various applications, such as media players, browsers, etc., for implementing various application services. The implementation of the overlay network capability slicing routing and cooperative execution method based on the end-side lightweight model provided in this embodiment of the invention can be included in the application program 5022.

[0117] The methods disclosed in the above embodiments of the present invention can be applied to processor 501, or implemented by processor 501. Processor 501 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 501 or by instructions in the form of software. The processor 501 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 501 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. General-purpose processor 501 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.

[0118] In an exemplary embodiment, the electronic terminal 500 may be used to execute the aforementioned method by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs).

[0119] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to perform the method of any of the embodiments described above.

[0120] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code, which, when run on a computer, causes the computer to perform the method of any of the embodiments described above.

[0121] As used in this specification, the terms "component," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).

[0122] Those skilled in the art will recognize that the various illustrative logical blocks and steps described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.

[0123] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0124] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0125] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0126] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0127] In the above embodiments, the functions of each functional unit can be implemented entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. A computer program product includes one or more computer instructions (programs). When the computer program instructions (programs) are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. Available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0128] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0129] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0130] In summary, the method, system, medium, program product, and terminal for overlay network capability slice routing and collaborative execution based on a lightweight end-side model provided in this application include: abstracting each model capability slice of a distributed lightweight end-side model into capability slice nodes in the overlay network, and generating a capability profile corresponding to each capability slice node; receiving complex task requests, and having the lightweight end-side model decompose the complex task requests to generate at least one capability invocation sub-step; collecting the capability profiles and network status information of each capability slice node in the overlay network in real time; filtering, sorting, and prioritizing the capability slice nodes in the overlay network based on at least one capability invocation sub-step, the capability profiles of each capability slice node, and the network status information, determining the capability slice nodes, invocation order, and parallel and / or serial relationships corresponding to each capability invocation sub-step, and orchestrating and generating model capability slice invocation paths for execution of model capability slice routing in the overlay network; and rearranging unexecuted capability invocation sub-steps according to the execution results of the model capability slice routing.

[0131] This application abstracts model capability slices into discoverable, filterable, and routable capability node resources, enabling the native organization of model capability slices within the overlay network. The application allows for local capability call sub-step decomposition and slice routing orchestration on the client side, eliminating the need for complete reliance on a central control plane and enhancing the autonomous processing capabilities of the lightweight client-side model for complex task requests. This application comprehensively considers node capabilities, path state interface constraints, dependency summaries, and budget states to achieve joint optimization of network state, dependencies, and model capability slices. This application orchestrates the sub-step call order, convergence slice position, and parallel relationships to reduce redundant slice switching and waiting time, improving cross-node collaborative execution efficiency. When slice node failures, timeouts, insufficient result confidence, or convergence slice anomalies occur during execution, this application can replace, switch, and rearrange local capability slice paths to prevent overall task failure. Therefore, this application effectively overcomes various shortcomings of existing technologies and possesses high industrial application value.

[0132] The above embodiments are merely illustrative of the principles and effects of this application and are not intended to limit this application. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in this application should still be covered by the claims of this application.

Claims

1. A method for capability slicing routing and cooperative execution in overlay networks based on a lightweight end-side model, characterized in that, include: The capability slices of each model in the distributed deployment of the lightweight edge model are abstracted into capability slice nodes in the overlay network, and a capability profile corresponding to each capability slice node is generated. Receive complex task requests, and decompose the complex task requests by the lightweight model on the edge to generate at least one capability invocation sub-step; Capability profiles and network status information of each capability slice node are collected in real time within the overlay network; Based on at least one capability invocation sub-step, the capability profile of each capability slice node, and network status information, the capability slice nodes in the overlay network are filtered, sorted, and prioritized to determine the capability slice nodes, invocation order, and parallel and / or serial relationships corresponding to each capability invocation sub-step, and the model capability slice invocation path is orchestrated to be generated for model capability slice routing in the overlay network. Based on the execution results of the capability slice routing of the execution model, the unexecuted capability call sub-steps are rearranged.

2. The method for coverage network capability slicing routing and cooperative execution based on a lightweight end-side model according to claim 1, characterized in that, Based on the capability invocation sub-steps, the capability profiles of the capability slice nodes, and network status information, the capability slice nodes in the coverage network are filtered, sorted, and prioritized to determine the capability slice nodes, invocation order, and parallel and / or serial relationships corresponding to each capability invocation sub-step. The specific process includes: Based on the capability profiles and network status information of each capability slice node, combined with the dependencies between each capability invocation sub-step, input / output interface constraints, and convergence slice switching conditions, the main execution capability slice node, backup capability slice node, and convergence capability slice node of each capability invocation sub-step are determined, and the invocation order between each capability invocation sub-step is obtained. Based on the dependencies between the various capability call sub-steps, the serial call chain and parallel call group are determined, and the parallel and / or serial relationships between the various capability call sub-steps are obtained.

3. The method for coverage network capability slicing routing and cooperative execution based on a lightweight end-side model according to claim 1, characterized in that, The process of re-orchestrating unexecuted capability call sub-steps based on the execution results of the capability slice routing execution model includes: If the execution result of the capability slice routing of the model includes at least one of the following: capability slice node failure, capability slice node timeout, insufficient result confidence, insufficient budget, abnormal aggregation capability slice node, excessive cross-node state synchronization cost, or network state degradation, then the unexecuted capability call sub-steps will be rearranged.

4. The method for coverage network capability slicing routing and cooperative execution based on a lightweight end-side model according to claim 1, characterized in that, The process of re-orchestrating unexecuted capability call sub-steps based on the execution results of the capability slice routing execution model includes: When the confidence level of the execution result of the capability slice routing is lower than the preset confidence level requirement and the budget status allows, the backup capability slice node switch or the aggregation capability slice node switch based on topological distance will be performed during the reordering of the unexecuted capability call sub-steps.

5. The method for coverage network capability slicing routing and cooperative execution based on a lightweight end-side model according to claim 1, characterized in that, The process of re-orchestrating unexecuted capability call sub-steps based on the execution results of the capability slice routing execution model includes: When the budget state in the execution result of the execution model capability slice routing is lower than the preset budget requirement or the cost of cross-node state synchronization is too high, the capability slice node depth degradation, edge-side conservative inference degradation, or low-priority sub-steps are skipped during the re-orchestration of the unexecuted capability call sub-steps.

6. The method for coverage network capability slicing routing and cooperative execution based on a lightweight end-side model according to claim 1, characterized in that, The capability invocation sub-step includes at least one of the following: sub-step identifier, sub-step type, required capability type, input dependency, execution priority, acceptable delay range, result confidence requirement, or whether parallel execution is allowed.

7. A capability slicing routing and cooperative execution system for overlay networks based on a lightweight end-side model, characterized in that, include: The Capability Slice Node Abstraction Module is used to abstract the capability slices of each model in the distributed deployment of the lightweight edge model into capability slice nodes in the overlay network, and generate a capability profile corresponding to each capability slice node. The complex task request decomposition module is used to receive complex task requests and decompose the complex task requests by the edge lightweight model to generate at least one capability invocation sub-step. The profiling and information collection module is used to collect capability profiles and network status information of each capability slice node in the overlay network in real time. The capability slice routing and execution module is used to filter, sort and prioritize capability slice nodes in the overlay network based on at least one capability call sub-step, the capability profile of each capability slice node and network status information, determine the capability slice nodes, call order and parallel and / or serial relationships corresponding to each capability call sub-step, and orchestrate and generate model capability slice call paths for execution of model capability slice routing in the overlay network. The sub-step re-orchestration module is used to re-orchestrate unexecuted capability call sub-steps based on the execution results of the capability slice routing of the execution model.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the overlay network capability slice routing and cooperative execution method based on the end-side lightweight model as described in any one of claims 1 to 6.

9. A computer program product, characterized in that, The computer program product includes computer program code, which, when run on a computer, enables the computer to implement the overlay network capability slice routing and cooperative execution method based on the end-side lightweight model as described in any one of claims 1 to 6.

10. An electronic terminal, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the overlay network capability slice routing and cooperative execution method based on the end-side lightweight model as described in any one of claims 1 to 6.