A semantic cooperative scheduling method and system for wide-area computing power networks to suppress routing oscillations

By employing a hierarchical distributed network architecture and a multi-agent reinforcement learning algorithm, the problems of routing oscillation and privacy protection in wide-area computing networks are solved, enabling collaborative scheduling and performance assurance of cross-domain resources, thereby improving network stability and resource utilization efficiency.

CN122093385APending Publication Date: 2026-05-26NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2026-01-13
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing wide-area computing network collaboration mechanisms suffer from issues such as routing oscillations, privacy protection, and heterogeneous resource silos, making it difficult to achieve efficient and stable resource collaboration in cross-domain scheduling.

Method used

By adopting a hierarchical distributed network architecture, combined with the SLA semantic abstraction mapping model, multi-agent reinforcement learning algorithm and transfer cost awareness mechanism, privacy protection collaboration and end-to-end performance guarantee of cross-domain computing power resources are achieved.

Benefits of technology

It effectively suppressed routing oscillations, achieved global resource optimization and unified description and coordinated scheduling of heterogeneous computing resources under privacy protection conditions, and improved network stability and resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093385A_ABST
    Figure CN122093385A_ABST
Patent Text Reader

Abstract

This invention discloses a semantic collaborative scheduling method and system for wide-area computing power networks (WANs) to suppress routing oscillations. Belonging to the field of intelligent scheduling technology for computing resources, it aims to solve the problems of heterogeneous SLA semantics, conflicts between privacy protection and global optimization, and routing oscillations caused by high-frequency changes in computing power status in existing WANs. The method is based on a hierarchical distributed network architecture, dividing the WAN into several autonomous domains (ADAs). Each ADA deploys an ADA controller containing modules for SLA semantic mapping, agent decision-making, migration cost awareness, and cross-domain coordination. Cross-domain interaction is achieved through boundary coordinators at the domain boundaries, and multiple boundary coordinators converge through inter-domain controllers to form a global control plane. This invention effectively suppresses routing oscillations and improves the service performance and resource utilization of WANs by shielding high-frequency fluctuations in computing power through SLA semantic mapping, balancing privacy and collaboration through a federated multi-agent architecture, and breaking down resource silos through automated semantic intercommunication. It is applicable to cross-ADA collaborative scheduling scenarios in distributed heterogeneous WANs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent scheduling technology for computing resources, and in particular to a semantic collaborative scheduling method for wide-area computing networks that suppresses routing oscillations. Background Technology

[0002] In recent years, with the rapid development of emerging technologies such as cloud computing, edge computing, artificial intelligence, and network slicing, the deployment of computing resources is gradually evolving from centralized cloud centers to distributed, heterogeneous wide-area computing power collaborative systems. Computing Power Network (CPN), as a key form in this evolution, deeply integrates communication networks and computing resources through network programmability and computing power virtualization technologies, realizing an integrated service delivery model of "network as computing power." This system aims to establish a unified collaborative orchestration mechanism among computing resources in different regions and management domains, providing high-performance, low-latency, and highly reliable computing services for various application scenarios.

[0003] In computing network architecture, while a single autonomous system (AS) (such as a specific carrier network domain or an enterprise private cloud) can provide basic computing services, the computing resources within a single AS are objectively limited in terms of capacity limits, computing types, and carrying capacity. When faced with sudden surges in business activity or heterogeneous demands for large-scale, complex computing tasks, the dilemma of exhausting resources within the AS and being unable to schedule tasks often arises. This contradiction between the limited computing resources of a single domain and the unlimited business demands dictates that computing networks must break through geographical limitations and achieve cross-domain collaborative scheduling of computing power—that is, achieving resource complementarity and sharing across a wide area—becoming an inevitable choice to improve the overall service capabilities of the network. However, this process faces multiple challenges, including diverse resource distribution, management autonomy, and heterogeneous technical architecture.

[0004] To address these challenges, academic and industrial communities both domestically and internationally have proposed various resource scheduling and task allocation strategies. Centralized global optimization control schemes are the earliest and most widely used technical approach. For example, Li et al. proposed a centralized scheduling method for satellite-terrestrial integrated networks in 2023, which collects network-wide information through a unified control plane to achieve resource allocation. While such schemes offer strong controllability in small-scale scenarios, they are highly susceptible to information delays, single points of failure, and privacy regulations in wide-area environments, making information sharing across multiple management domains difficult.

[0005] To alleviate the bottleneck of centralized control, distributed collaboration and service-oriented routing mechanisms have become research hotspots. Some existing technologies have proposed the concept of "service identifiers," using abstract identifiers to achieve cross-domain addressing. However, these technologies have significant limitations when dealing with heterogeneous resources. Existing technologies acknowledge that because cross-domain interoperability requires complex negotiation and mapping, the governance of most services, except for a very few basic services that have gained high industry consensus, is restricted to a single operation and management domain, leading to the conclusion that standardization is unnecessary. This effectively indicates that existing mechanisms have proactively abandoned wide-area collaboration of heterogeneous resources to avoid standardization challenges. This self-imposed limitation strategy prevents massive non-standard and personalized computing resources from participating in wide-area scheduling, causing computing networks to degenerate into isolated resource islands in practical applications.

[0006] Meanwhile, while rule-based and heuristic resource scheduling strategies (such as the improved butterfly optimization algorithm proposed by Shanmugam et al.) can respond quickly, they often follow the scheduling thinking of traditional networks, focusing only on network link indicators (such as bandwidth and latency) and lacking comprehensive modeling of the characteristics of computing power itself, making it difficult to adapt to the complex scenarios of deep integration of computing and network in computing power networks.

[0007] Overall, existing wide-area computing power network collaboration mechanisms still have the following prominent problems: (1) Traditional cross-domain scheduling lacks adaptation to the dynamic characteristics of computing power, which can easily cause routing oscillations and make it difficult to ensure scheduling accuracy. Most existing cross-domain scheduling strategies follow traditional thinking and simply map computing power status to network indicators. However, the frequency of changes in computing power load is much higher than that of network topology. If such high-frequency status updates are directly announced, it is very easy to cause severe routing oscillations and table expansion in the underlying routing protocol, which will damage network stability. If the update frequency is reduced, the scheduling decision will be based on outdated information, which will easily assign tasks to nodes with smooth network links but overloaded computing power, resulting in an imbalance in computing network resources and a decline in service performance. (2) There is a structural conflict between privacy protection requirements and the global coordination mechanism required for cross-domain scheduling, making it difficult for existing centralized or distributed scheduling models to balance privacy and scheduling efficiency. Cross-domain coordination requires resource optimization based on a global perspective, but autonomous domains refuse to share complete internal topology and real-time load information due to commercial competition and security / privacy considerations. Existing centralized optimization models are at risk of privacy leakage due to their reliance on global full information (information black box problem), while distributed models, lacking a global perspective, struggle to resolve conflicts of interest and policy coupling between multiple domains, making it impossible to achieve the optimal solution for global resource utilization while protecting privacy. (3) Due to the lack of an automated cross-domain Service Level Agreement (SLA) semantic mapping and alignment mechanism, existing technologies struggle to achieve unified description and collaborative scheduling of heterogeneous computing resources across a wide area. Existing solutions avoid the standardization requirements of most services, restricting these services to a single operation and management domain, effectively abandoning the ability to schedule heterogeneous resources across a wide area. This mechanism is limited to domain governance and cannot resolve the fundamental contradiction of "both cross-domain collaboration and mandatory standardization" in wide-area scenarios. Due to the lack of an algorithm-based automatic SLA semantic mapping mechanism, heterogeneous physical resources from different domains cannot be "translated" into a unified metric standard, resulting in wide-area collaboration being limited to a very small number of basic services and failing to unleash the true value of massive heterogeneous computing power.

[0008] Existing wide-area collaborative mechanisms for computing power networks are still insufficient to effectively suppress routing oscillations caused by high-frequency dynamic changes in computing power, while overcoming multiple barriers related to privacy protection, interaction overhead, and the "islanding" of heterogeneous resources. How to construct a collaborative mechanism with automatic cross-domain SLA semantic mapping capabilities, capable of abstracting and dynamically scheduling resource-side capabilities in partially observable and privacy-constrained environments, is a key technical challenge in the current development of wide-area computing power networks. Therefore, it is necessary to propose a cross-domain privacy-preserving collaborative scheduling method for wide-area computing power networks oriented towards SLA semantic alignment, in order to achieve an optimal balance between autonomy, privacy, and performance guarantees. Summary of the Invention

[0009] Purpose of the invention: This invention addresses the problems existing in the prior art, such as semantic heterogeneity of SLA in wide area computing power networks, conflict between privacy protection and global optimization, and routing oscillation caused by high-frequency changes in computing power status. It provides a semantic collaborative scheduling method and system for wide area computing power networks that suppresses routing oscillation.

[0010] To achieve the above-mentioned objectives, the present invention provides the following technical solution: A semantic collaborative scheduling method for wide-area computing power networks to suppress routing oscillations is disclosed. Based on a hierarchical distributed network architecture, the method utilizes an SLA semantic abstraction mapping model, a multi-agent reinforcement learning collaborative algorithm, and a migration cost-aware reward mechanism to achieve cross-domain computing power resource privacy protection collaboration, end-to-end performance assurance, and routing oscillation suppression. The method includes the following steps: (1) Based on the geographical distribution and management strategy of computing resources, the computing network is divided into several autonomous domains. Autonomous domain controllers and boundary coordinators are deployed in each autonomous domain, and the autonomous domain controllers and boundary coordinators are registered to the global coordination directory. The autonomous domain consists of local computing nodes, communication links and boundary agents. The autonomous domain controller includes an SLA semantic mapping module, an agent decision-making module, a migration cost perception module and a cross-domain coordination module. (2) Each autonomous domain controller obtains the computing power index within the domain through the performance acquisition unit in the SLA semantic mapping module, and converts it into a capability vector through the capability description generation unit to form a domain-level capability description; after constructing a local semantic mapping subgraph in the domain, the semantic mapping unit reports the local semantic mapping subgraph to the boundary coordinator; the boundary coordinator gathers the local semantic mapping subgraphs of each domain, constructs a global SLA mapping map and generates standardized capability labels to realize the semantic alignment and capability mapping of cross-domain performance indicators; (3) The agent decision-making module of each autonomous domain takes the observable state information of its own domain as input and outputs the task placement or migration decision; during the training phase, the agents on each autonomous domain controller perform policy iterative updates based on the multi-agent reinforcement learning algorithm to maximize the composite reward function; the observable state information includes computing power load, task queue length, inter-domain link delay and SLA mapping results, and the task placement or migration decision includes local execution, intra-domain migration and cross-domain migration; (4) The boundary coordinator realizes wide-area collaboration through parameter aggregation and capability label matching mechanism. The inter-domain controller forms a globally optimal task allocation scheme based on the training results and generates a task allocation table. The inter-domain controller is deployed on the global control plane of the computing power network and includes intra-domain agent nodes, boundary coordination control units and monitoring feedback units. (5) The inter-domain controller completes task migration and computing power allocation operations through the agent node within the domain according to the task allocation table, and feeds back the execution results to the agent decision module through the monitoring feedback unit to realize the continuous adaptive update of the strategy.

[0011] Furthermore, the working process of the SLA semantic mapping module in step (2) specifically includes: (21) The performance acquisition unit periodically collects raw performance data from computing power nodes and communication links within the domain through the resource awareness interface M1. The raw performance data includes CPU / GPU utilization, memory usage, available bandwidth, end-to-end latency, reliability level and energy consumption parameters. (22) The capability description generation unit normalizes and discretizes the original performance data and uses a mapping function with hysteresis effect to map the continuously changing original performance data into discrete levels, generating a standardized capability vector containing latency level, bandwidth level and reliability category. (23) The capability description generation unit performs aggregate analysis on the capability vectors of all computing power nodes in the domain, generates a domain-level capability description that can be shared externally through dimensionality reduction and weighted averaging methods, and establishes an index mapping relationship between the domain-level capability description and the original capability vector. (24) The semantic mapping unit uses the BERT pre-trained model to encode the text description of each SLA index and generate the initial feature vector of the node; it uses each feature item as a node and the performance dependency relationship as an edge to construct a local semantic mapping subgraph; it uses the GraphSAGE graph neural network algorithm to aggregate the feature information of neighboring nodes and extract the semantic center point of the domain. (25) The boundary coordinator gathers the local semantic mapping subgraphs and semantic centroids of each autonomous domain, performs cross-domain graph splicing and semantic association learning, and constructs a global SLA mapping graph; calculates the cosine similarity of the embedding vectors of semantic centroids in different domains, and identifies heterogeneous resource descriptions with similarity higher than a preset threshold as semantic alignment relationships; (26) The capability tag management unit generates standardized capability tags corresponding to the domain-level capability description. The capability tags encapsulate latency level, bandwidth level, reliability category and billing constraints. Each autonomous domain publishes the capability tags through the boundary agent.

[0012] Furthermore, the working process of the agent decision-making module in step (3) specifically includes: (31) The state modeling unit constructs the autonomous domain observable state space at time t based on the domain capability labels, neighborhood semantic relationships and real-time operation data generated by the SLA semantic mapping module; the real-time operation data in the domain includes computing node load, task queue length, link bandwidth and latency, and energy consumption level. (32) After receiving the state at time t, the action generation unit inputs the state into the policy network to obtain the action probability distribution of the action space; and selects an action to be executed from the action set based on the preset decision policy. The action set includes local execution, intra-domain migration and cross-domain migration. (33) The policy learning unit periodically samples training samples from the experience replay pool, calculates the gradient using the A3C algorithm, and uses the gradient to update the local policy network parameters to form the local optimal policy of the autonomous region; the experience replay pool stores training samples containing state, action, reward and next time state. (34) The cross-domain coordination module sends the policy update amount, capability label and semantic state summary of the current domain to the boundary coordinator through the cross-domain coordination interface M2; the policy update amount includes gradient information or model weights.

[0013] Furthermore, the working process of the migration cost awareness module specifically includes: After each action is executed, the migration cost calculation unit evaluates the migration cost generated by the task migration. The migration cost includes cross-domain bandwidth usage, migration latency, context switching loss, and task interruption risk. The composite reward generation unit jointly models SLA satisfaction, migration cost, and load balancing into a composite reward function, the expression of which is: R = α R SLA + β (1 - R cost ) + γ R balance

[0014] Where R is the composite reward value, α, β, and γ are weighting coefficients and α + β + γ = 1, R SLA As a reward for SLA satisfaction, R cost R is the normalized migration cost. balanc As a reward for load balancing; The calculated composite reward value is stored in the experience replay pool for policy iteration and updating of the agent.

[0015] Furthermore, the specific process of forming the globally optimal task allocation scheme described in step (4) includes: (41) The boundary coordinator receives the policy update data uploaded by each autonomous system controller, performs asynchronous parameter fusion, and then transmits the aggregation result to the inter-domain controller. (42) The boundary coordination control unit of the inter-domain controller performs policy structure standardization processing on the aggregation results and constructs a global decision space based on the load scale, data credibility, task contribution rate and resource capability weight of each autonomous domain; (43) The inter-domain controller uses the federated averaging algorithm to deeply fuse the standardized policy parameters to generate a preliminary global policy parameter set; (44) The monitoring and feedback unit of the inter-domain controller performs offline evaluation of the preliminary global policy parameter set. The evaluation indicators include the task SLA satisfaction rate, the convergence trend of cross-domain migration delay and the overall network load balancing level. If the policy is unstable, the penalty mechanism is triggered and the corresponding autonomous domain is required to re-collect samples for training. (45) When the global policy training loss tends to stabilize, the transfer overhead shows a clear downward trend, and the global reward continues to increase and meets the convergence condition, the preliminary global policy parameter set is marked as the global optimal policy. (46) Based on the global optimal strategy, the global SLA mapping map and the distribution of computing power resources across the entire network, the inter-domain controller generates the global optimal task allocation scheme and the corresponding task allocation table.

[0016] Furthermore, the process of task migration, computing power allocation, and policy update in step (5) specifically includes: (51) The inter-domain controller synchronizes the task allocation table to the relevant autonomous domain controllers and boundary coordinators by executing the feedback interface M3; (52) The autonomous domain controller performs task context packaging, link scheduling, target domain resource reservation and execution control operations according to the scheduling instructions in the task allocation table; the boundary coordinator is responsible for cross-domain link management and consistency maintenance during the migration process; (53) After the task is completed in the target domain, the autonomous domain controller will report the operation results such as task completion time, resource consumption, migration cost and SLA satisfaction to the inter-domain controller. (54) The monitoring feedback unit of the inter-domain controller periodically performs statistics on the operation results to form a monitoring vector; the monitoring vector includes task completion quality, SLA satisfaction rate, cross-domain link load, computing power utilization and migration efficiency. (55) If the monitoring vector indicates that the autonomous system has experienced performance degradation, the inter-domain controller notifies the corresponding autonomous system controller to re-collect performance indicators and update the capability vector through the execution feedback interface M3, thereby driving the capability label and SLA mapping map to be updated adaptively. (56) Based on the updated semantic information and operational feedback, the inter-domain controller adjusts the task scheduling rules and updates the task allocation table to realize the continuous adaptive optimization of the agent decision-making module strategy.

[0017] Furthermore, the open interfaces involved in the method include: Resource Awareness Interface M1: Located between the underlying computing power resource nodes and the Autonomous System Controller (AS / DC), it is used by the AS / DC's performance acquisition unit to obtain raw status data from physical computing power nodes and network devices. Cross-domain coordination interface M2: Used for information exchange between the boundary coordinator and different autonomous system controllers. The transmitted content includes policy parameters, capability tags, aggregation results and cross-domain load information. Execution Feedback Interface M3: Used for bidirectional communication between the inter-domain controller and the boundary coordinator and autonomous domain controller. On the one hand, it transmits task allocation and migration strategies, and on the other hand, it sends back execution result information.

[0018] Based on the implementation of the above method, the present invention also provides a semantic cooperative scheduling system for wide-area computing power networks that suppresses routing oscillations. The system includes several autonomous domains, a boundary coordinator, and an inter-domain controller. Autonomous Domain: Composed of local computing nodes, communication links, and boundary agents, with an autonomous domain controller deployed internally; the autonomous domain controller includes an SLA semantic mapping module, an agent decision-making module, a migration cost awareness module, and a cross-domain coordination module; SLA Semantic Mapping Module: Used to implement cross-domain SLA indicator abstraction and capability mapping, and generate standardized capability tags; includes a performance acquisition unit, a capability description generation unit, a semantic mapping unit, and a capability tag management unit; The agent decision-making module is used to generate task allocation and transfer decisions based on multi-agent reinforcement learning algorithms; it includes a state modeling unit, an action generation unit, and a policy learning unit. Migration Cost Awareness Module: Used to model the migration cost of cross-domain tasks and generate a composite reward function; includes a migration cost calculation unit and a composite reward generation unit; Cross-domain coordination module: used to interact with the boundary coordinator to achieve policy sharing, parameter aggregation, and global optimization; Boundary Coordinator: Deployed at the edge of the autonomous system, it aggregates the local policy parameters, capability tags and status information of adjacent autonomous system controllers, achieves global approximate convergence through an asynchronous parameter fusion mechanism, and feeds back a unified task allocation policy to the autonomous system controller. Inter-domain controller: Deployed in the global control plane of the computing network, it is used to uniformly execute wide-area task allocation and migration strategies; it includes intra-domain agent nodes, boundary coordination and control units, and monitoring and feedback units; Intra-domain proxy nodes: used to execute task placement and migration commands within an autonomous system; Boundary Coordination and Control Unit: Used to manage the collaborative operation of multiple boundary coordinators, enabling cross-domain task allocation synchronization and policy consistency; Monitoring and feedback unit: used to periodically collect network operation information and feed it back to the intelligent agent decision-making module to achieve adaptive optimization of strategies.

[0019] Furthermore, in the SLA semantic mapping module: The performance acquisition unit collects CPU / GPU utilization, memory usage, available bandwidth, end-to-end latency, reliability level, and energy consumption parameters through the resource-aware interface M1. The capability description generation unit uses a mapping function with hysteresis to map the raw performance data into discrete levels, generating a standardized capability vector. The semantic mapping unit uses the BERT pre-trained model and the GraphSAGE graph neural network algorithm to construct local semantic mapping subgraphs and extract semantic center points; The capability tag management unit generates capability tags with encapsulation latency level, bandwidth level, reliability category, and billing constraints.

[0020] Furthermore, the operation process of the inter-domain controller includes: Receive the parameter aggregation results transmitted by the boundary coordinator and construct the global decision space; The federated averaging algorithm is used to fuse policy parameters and generate a global policy parameter set; The global policy parameter set is evaluated offline to determine the globally optimal policy. A task allocation table is generated based on the globally optimal strategy and distributed to the autonomous system controller and the boundary coordinator. Collect network-wide operation results to drive adaptive updates of capability tags, SLA mapping maps, and task scheduling strategies.

[0021] Beneficial effects: Compared with the prior art, the present invention has the following beneficial effects: 1. To address the routing oscillations and accuracy deficiencies caused by the lack of adaptation to the dynamic characteristics of computing power in traditional cross-domain scheduling, this invention proposes an SLA capability tagging mechanism based on the hysteresis lock-in effect. This mechanism achieves adaptive adaptation to the high-frequency dynamic characteristics of computing power and steady-state routing control. This invention breaks through the traditional approach of directly mapping the original computing power state. Through discretization binning and the hysteresis effect mechanism, it abstracts the high-frequency fluctuating computing power load into relatively stable "capability tags." This mechanism shields non-critical minor fluctuations in computing power at the source, effectively resolving the dilemma between "direct notification leading to routing oscillations" and "reducing frequency leading to information expiration." The upper-layer scheduling strategy makes decisions based on capability tags, decoupling the direct correlation between the underlying computing power state fluctuations and control plane routing updates. While ensuring that the scheduling decision basis does not expire and that SLA constraints are responded to in real time, it significantly suppresses ineffective LSA flooding, achieving a dual guarantee of stability and scheduling accuracy in the wide-area computing power network control plane.

[0022] 2. To address the structural conflict between privacy protection requirements and the global coordination mechanism needed for cross-domain scheduling, this invention constructs a federated multi-agent reinforcement learning architecture based on a boundary coordinator. This architecture achieves optimal global scheduling efficiency while ensuring the privacy and security of each autonomous domain. This invention overcomes the limitations of existing centralized models that rely on comprehensive global information and distributed models that lack a global perspective. Through a collaborative model where "the data remains still while the model moves," each autonomous domain only needs to interact with aggregated policy parameters and abstract capability labels through the boundary coordinator, without sharing sensitive internal topologies and real-time loads. This architecture breaks the "information black box" constraint under partially observable conditions, effectively resolving conflicts of interest and policy coupling among multiple management domains, successfully mitigating the structural conflict between privacy protection and global coordination, and achieving a dialectical unity between the independent decision-making power of each autonomous domain and the efficiency of collaborative optimization of network resources.

[0023] 3. To address the problem of isolated heterogeneous resources across a wide area due to the lack of automated SLA semantic mapping and alignment mechanisms, this invention constructs an automated interoperability mechanism based on SLA mapping graphs and semantic embedding, achieving unified description and collaborative scheduling of heterogeneous computing resources across a wide area. Unlike existing technologies that restrict services to a single domain or support only a small number of standardized services, this invention utilizes graph neural networks and semantic embedding algorithms based on Bidirectional Encoder Representations from Transformers (BERT) to establish a "universal translator" for cross-domain resources. This mechanism can automatically "translate" non-standard and personalized physical resources within different autonomous domains into capability labels with unified metric across the entire network, resolving the fundamental contradiction of "needing cross-domain collaboration but being unable to enforce standardization." This allows massive heterogeneous computing resources to be included in the global scheduling pool without cumbersome manual negotiation, effectively overcoming the resource silo effect of the traditional single-domain governance model and releasing the complementary potential and collaborative value of massive heterogeneous computing power across a wide area. Attached Figure Description

[0024] Figure 1 This is a schematic diagram of the wide-area computing network architecture according to an embodiment of the present invention; Figure 2 Schematic diagram of SLA semantic mapping and resource modeling process for wide-area computing power networks; Figure 3 A schematic diagram of the training process for a cross-domain task scheduling strategy driven by multi-agent reinforcement learning; Figure 4 A schematic diagram of the global policy training process for inter-domain controllers; Figure 5 A schematic diagram of the cross-domain task scheduling and system feedback closed-loop mechanism. Detailed Implementation

[0025] To enable those skilled in the art to better understand the technical solutions provided by the present invention, the embodiments are described in detail below with reference to the accompanying drawings.

[0026] This invention provides a semantic collaborative scheduling method for wide-area computing power networks (WANs) that suppresses routing oscillations. It is a WAN service allocation scheme oriented towards SLA semantic alignment and dynamic optimization. This scheme achieves privacy-preserving collaboration and end-to-end performance assurance for cross-domain computing power resources by designing an SLA semantic abstraction mapping model, a multi-agent reinforcement learning collaborative algorithm, and a migration cost-aware reward mechanism. The core idea of ​​this scheme is to construct independent agent control units within each autonomous region (AGN), perform local decision learning based on locally observable information, and achieve policy fusion and global collaborative optimization among multiple agents through a boundary coordinator. Simultaneously, semantic capability tags are used to shield the high-frequency fluctuations of underlying computing power resources, thereby achieving distributed, self-learning, and effectively oscillating cross-domain dynamic service allocation.

[0027] The aforementioned wide-area computing network comprises several autonomous systems (AS / RS). Each AS / RS consists of local computing nodes, communication links, and boundary agents, used for computing resource coordination and service scheduling in a distributed environment. Each AS / RS has a Domain Controller (DC) to manage its local computing resources, task allocation, and local agent decision-making. AS / RS interact with each other through Boundary Coordinators (BCs) located at the domain boundaries, enabling cross-domain policy coordination and resource collaboration. Multiple Boundary Coordinators converge through Inter-domain Controllers (ICs) to form a unified control plane, enabling global policy distribution and dynamic feedback.

[0028] Each Autonomous System (AS) controller comprises four main functional modules: SLA semantic mapping module, agent decision-making module, migration cost awareness module, and cross-domain coordination module. Their specific configurations are as follows: 1) The SLA semantic mapping module is used to abstract service level agreement indicators and map capabilities between different autonomous systems, and to generate capability labels that shield against high-frequency jitter at the underlying level; 2) The agent decision-making module is used to make task allocation and transfer decisions based on the multi-agent reinforcement learning algorithm and abstract ability labels and state information; 3) The migration cost awareness module is used to model and optimize bandwidth consumption, latency and load changes during cross-domain task migration, and balance scheduling benefits and migration costs. 4) The cross-domain coordination module is used to interact with the boundary coordinator to achieve policy sharing, parameter aggregation, and global optimization.

[0029] Specifically, the SLA semantic mapping module includes a performance acquisition unit, a capability description generation unit, a semantic mapping unit, and a capability tag management unit. The specific configuration is as follows: The performance acquisition unit is responsible for collecting raw metrics such as computing resource utilization, bandwidth usage, latency, and reliability within the domain. The capability description generation unit transforms the above indicators into a unified capability vector representation; The semantic mapping unit constructs an SLA mapping map to achieve semantic conversion and capability matching of performance indicators between different autonomous domains, thereby forming a cross-domain comparable SLA capability table; The capability tag management unit is used to generate capability tags at the semantic level. These tags serve as standardized semantic interfaces, encapsulating abstract capability information such as latency levels, bandwidth tiers, reliability categories, and billing constraints for the domain. By publishing only relatively stable capability tags at the domain boundary proxy, each autonomous region can achieve semantic alignment and resource collaboration at the SLA level while protecting privacy and shielding high-frequency internal state changes.

[0030] Specifically, the agent decision-making module includes a state modeling unit, an action generation unit, and a policy learning unit. Its features include: the state modeling unit defining observable state information for the current domain, including computational load, task queue length, inter-domain link latency, and SLA mapping results; the action generation unit outputting executable actions based on the state information, including local execution, intra-domain migration, and cross-domain migration; and the policy learning unit iteratively updating the state-action-reward relationship based on a deep reinforcement learning algorithm to obtain the optimal task allocation strategy.

[0031] Specifically, the migration cost perception module includes a migration cost calculation unit and a composite reward generation unit. Its features include: the migration cost calculation unit calculating the migration latency, bandwidth consumption, and interruption time incurred when a task migrates from the source domain to the target domain; and the composite reward generation unit jointly modeling SLA satisfaction, migration cost, and load balancing as a reward function to guide the agent in weighing performance against cost.

[0032] Specifically, the cross-domain coordination module works in conjunction with the boundary coordinator to achieve cross-domain information exchange and policy synchronization. The boundary coordinator, located at the edge of the autonomous system (AS), is responsible for aggregating local policy parameters, capability tags, and state information from adjacent AS controllers, and achieving global approximate convergence through an asynchronous parameter fusion mechanism. After obtaining the aggregation results, the boundary coordinator feeds back a unified task allocation policy to the corresponding AS controller, thereby achieving end-to-end collaborative optimization and task consistency control.

[0033] Specifically, the inter-domain controller is deployed on the global control plane of the computing power network to uniformly execute task allocation and migration strategies in a wide-area environment. It coordinates the scheduling behavior of each boundary coordinator to achieve the distribution, execution, and feedback of global strategies. The inter-domain controller internally comprises three parts: intra-domain agent nodes, boundary coordination control units, and a monitoring and feedback unit. Intra-domain agent nodes are responsible for executing task placement and migration commands within each autonomous domain, ensuring consistency between local scheduling and global strategies. The boundary coordination control unit manages the collaborative operations of multiple boundary coordinators, achieving synchronization and strategy consistency in cross-domain task allocation. The monitoring and feedback unit periodically collects information on the entire network's SLA satisfaction rate, migration latency, and task throughput, and transmits the feedback results to the agent decision-making module for continuous adaptive optimization, thus forming a globally closed-loop computing power scheduling and dynamic optimization mechanism.

[0034] The specific open interfaces of each functional entity involved are as follows: 1) Resource Awareness Interface M1: This interface is located between the underlying computing resource nodes and the Autonomous System Controller (AS / RS). Through this interface, the performance acquisition unit in the AS / RS periodically obtains raw state data (such as CPU utilization, memory usage, link bandwidth, and port latency) from the physical computing nodes and network devices, providing basic data support for subsequent SLA semantic mapping and state modeling. 2) Cross-domain coordination interface M2: The interface is characterized by being used for policy and capability information exchange between the boundary coordinator and different autonomous domain controllers. It is mainly responsible for transmitting policy parameters, capability labels, aggregation results and cross-domain load information to realize semantic collaboration and policy convergence among multiple agents. 3) Execution Feedback Interface M3: This interface is characterized by its bidirectional communication between the inter-domain controller, the boundary coordinator, and the autonomous domain controller. On one hand, the inter-domain controller issues task allocation and migration policies through M3, triggering intra-domain agent nodes to execute specific operations. On the other hand, the lower layer transmits execution result information back through the same interface, including indicators such as SLA satisfaction, migration latency, and task throughput, for reward updates and policy corrections in the agent decision-making module. Through the bidirectional information flow of the execution feedback interface M3, a closed loop of policy execution and feedback is achieved.

[0035] This invention enables collaborative optimization of cross-domain computing resources and end-to-end service performance assurance in wide-area computing networks. By introducing an SLA semantic mapping mechanism, a transfer cost-aware model, and a multi-agent reinforcement learning algorithm, this method effectively addresses issues such as heterogeneous SLA definitions, limited information sharing, and high transfer costs in wide-area environments, thereby achieving a balance between privacy protection and global optimization. Overall, the method includes the following steps: Step 1: Based on the geographical distribution and management strategy of computing resources, divide the computing network into several autonomous regions, and deploy controllers and boundary coordinators in each autonomous region and register them to the global coordination directory; Step 2: Each autonomous system controller (AS / RS) acquires its intra-domain computing power metrics through the performance acquisition unit, generates capability vectors, and forms a domain-level capability description. The semantic mapping module constructs a local semantic mapping subgraph within its domain and reports the results to the boundary coordinator. The boundary coordinator aggregates the subgraphs from each domain and constructs a global SLA mapping map, generating standardized capability labels to achieve semantic alignment and capability mapping of cross-domain performance metrics. Step 3: The decision-making modules of each autonomous region take the state information as input and output the task placement or migration decision; during the training phase, the agents on each autonomous region controller perform iterative policy updates through reinforcement learning algorithms to maximize the composite reward function; Step 4: The boundary coordinator achieves wide-area collaboration through parameter aggregation and capability label matching mechanisms. The inter-domain controller forms a globally optimal task allocation scheme based on the training results and generates a task allocation table. Step 5: The inter-domain controller completes task migration and computing power allocation operations according to the task allocation table, and feeds back the execution results to the agent decision-making module to realize continuous adaptive updates of the strategy.

[0036] Example 1: A functional distribution diagram of the overall architecture of the wide-area computing power network according to an embodiment of the present invention is shown below. Figure 1 As shown, this architecture consists of an inter-domain controller, several autonomous system controllers (AS / RS), and computing resource nodes distributed across different geographical locations. Each AS / RS achieves cross-domain interaction and collaborative control through a boundary coordinator deployed at the edge, forming a hierarchical, distributed computing network management system.

[0037] The following details step 1, in which the original computing network is divided into several autonomous domains based on the geographical distribution and management strategy of computing resources. Autonomous domain controllers and boundary coordinators are deployed in each autonomous domain, and registration with the global coordination directory is completed to achieve unified resource identification and scheduling management.

[0038] In this network architecture, multiple service terminals and computing nodes can be connected within the same autonomous system (AS). The end-users' business types can include various heterogeneous computing power requirements such as video rendering, AI inference, vehicle networking, and industrial automation. Each AS has independent control and decision-making capabilities, and can adaptively place and migrate tasks according to business characteristics. The main functional entities involved in this network architecture include: Autonomous System (AS): This describes the domain-based management structure of the entire large-scale wide-area computing network. Based on the geographical distribution of computing resources, network topology, and management strategies, the system divides the network into multiple ASs. Each AS consists of several computing nodes, communication links, and boundary agents, used to execute and migrate computing tasks in a distributed environment. Domain-based management effectively reduces global computational complexity and improves network scalability and manageability.

[0039] Autonomous Domain Controller (ADAS): This is the distributed control core proposed in this invention, possessing four main functions: an SLA semantic mapping module, an agent decision-making module, a migration cost awareness module, and a cross-domain coordination module. Within its domain, it collects raw state data of underlying physical resources in real time through the resource-aware interface M1. Then, using the built-in SLA semantic mapping algorithm, it abstracts the frequently changing raw data into stable capability tags, thereby shielding underlying details and effectively suppressing routing oscillations. Simultaneously, it is responsible for task scheduling and local policy learning within its domain, and exchanges information with the boundary coordinators of neighboring autonomous domains. Furthermore, during initialization, each ADAS registers its own identifier, topology range, and generated capability tag information with the global coordination directory, thereby achieving semantic-based resource publishing rather than direct exposure of raw data.

[0040] Boundary Coordinator (BC): Deployed at the domain boundary, it is a key component for achieving cross-domain privacy collaboration and resilience. Its core feature is its responsibility for aggregating cross-domain information and synchronizing policies. It receives abstract capability tags and policy parameters from autonomous system controllers (AS / RS) through the cross-domain coordination interface M2. This component utilizes an asynchronous parameter fusion mechanism to achieve cross-domain policy collaboration and global approximate convergence, acting as a "stabilizer" to further filter state fluctuations between domains and realize joint optimization and service collaboration of wide-area computing resources. Simultaneously, the Boundary Coordinator maintains a connection with inter-domain controllers on the control plane, receiving policy aggregation results and issuing update instructions to each AS / RS.

[0041] Inter-domain controller: Located in the global control plane of the computing power network, it is used to uniformly manage the task allocation and global policy coordination of multiple autonomous systems. The inter-domain controller comprises three core functional parts: intra-domain agent nodes, boundary coordination control units, and monitoring and feedback units. Intra-domain agent nodes are responsible for coordinating and scheduling task placement and migration operations across all autonomous systems nationwide, ensuring resource matching from a global perspective. Boundary coordination control units manage the collaboration and synchronization between multiple boundary coordinators, maintaining the topology of the global SLA semantic graph. The monitoring and feedback unit periodically collects information on the network-wide SLA satisfaction rate, migration latency, and task throughput through the feedback interface M3, and transmits the feedback results to the decision-making modules of each agent to update the policy. Through the centralized coordination of the inter-domain controller and the distributed decision-making of the autonomous system controllers, this invention achieves globally consistent control and dynamic optimization of the wide-area computing power network.

[0042] The wide-area computing network architecture described in this embodiment achieves collaborative management and dynamic optimization of cross-domain computing resources through a hierarchical distributed control model. While maintaining independent autonomy, each autonomous domain can achieve a globally consistent service allocation strategy through a boundary coordinator. In particular, by introducing capability tags and a boundary coordination mechanism, this architecture effectively suppresses control plane routing oscillations caused by high-frequency changes in computing power status, while ensuring low latency and high reliability, and improves network resource utilization and task execution efficiency. The layered design of this architecture also supports containerized deployment and dynamic expansion, adapting to complex application scenarios such as cloud-edge collaboration and wide-area computing power scheduling, providing fundamental support for subsequent SLA semantic mapping, reinforcement learning decision-making, and system feedback loops.

[0043] Example 2: See Figure 2 The schematic diagram of the wide-area computing power network SLA semantic mapping and resource modeling process in this embodiment is shown in the figure. This embodiment corresponds to step 2, focusing on how each autonomous system controller generates its own domain capability vector and capability label through the SLA semantic mapping module, and constructs an SLA mapping map with the cooperation of the boundary coordinator. The graph neural network algorithm is then used to achieve semantic alignment of cross-domain performance indicators and resource modeling.

[0044] In this embodiment, the overall control system of the computing network is the same as in Embodiment 1, consisting of an inter-domain controller, several autonomous system controllers (AS / RS), and a distributed boundary coordinator. After each AS / RS is deployed and registered to the global coordination directory, its SLA semantic mapping module begins to execute the SLA semantic abstraction and mapping process described in this step. This module includes four units: a performance acquisition unit, a capability description generation unit, a semantic mapping unit, and a capability tag management unit.

[0045] Step 201: The performance acquisition unit within each autonomous system controller periodically collects real-time performance data from the computing nodes and communication links within the domain through the resource-aware interface M1. This data includes CPU / GPU utilization, memory usage, available bandwidth, end-to-end latency, reliability level, and energy consumption parameters. This data constitutes the original computing power state vector of the domain, reflecting the instantaneous operating status of the computing nodes within the domain.

[0046] Step 202: Indicator Standardization and Semantic Quantization. To eliminate differences in measurement dimensions and sampling methods between different autonomous regions and to suppress high-frequency data fluctuations, the capability description generation unit normalizes and discretizes the collected performance indicators into bins. Based on a unified quantization template, this unit uses a mapping function with hysteresis to map continuously fluctuating raw performance data into discrete levels. For example, CPU utilization is divided into four levels: "idle, normal, busy, and overloaded," and the status is updated only when the value breaks through and stabilizes at a new interval threshold. The final standardized capability vector includes components such as latency level, bandwidth level, and reliability category, filtering out minor state fluctuations at the data source.

[0047] Step 203: Domain-level Capability Description Generation. After standardizing performance metrics and semantically quantifying them, the capability description generation unit aggregates and analyzes the capability vectors of all computing power nodes within the domain. This aggregation process comprehensively considers node type, service load category, and computing power feature similarity, generating a domain-level capability description that can be shared externally through methods such as dimensionality reduction and weighted averaging. This description extracts representative performance features at the semantic level and establishes an index mapping relationship between them and the original capability vectors. This preserves the main service characteristics while masking node topology and scheduling strategies, achieving a comparable and non-disclosureable abstract expression.

[0048] Step 204: Local SLA Semantic Mapping and Subgraph Construction. After obtaining the capability vector set and domain-level capability description of their respective domains, the semantic mapping unit within each autonomous system controller (AS / RS) constructs a local semantic mapping subgraph using a semantic embedding algorithm. This unit first encodes the textual descriptions of each SLA indicator using a BERT pre-trained model to generate initial feature vectors for nodes, addressing the issue of literal heterogeneity. Then, it constructs a local subgraph using each feature item as a node and performance dependencies as edges. Finally, it employs the GraphSAGE graph neural network algorithm to aggregate the feature information of neighboring nodes and extract several semantic center points representing the distribution of the main capability features of the domain.

[0049] Step 205: Cross-Domain SLA Mapping Graph Construction and Semantic Alignment. Each autonomous system controller uploads the local semantic subgraph and semantic centroids generated by the semantic mapping unit to the boundary coordinator via the cross-domain coordination interface M2. The boundary coordinator aggregates semantic subgraph information from multiple autonomous systems, performs cross-domain graph stitching and semantic association learning, and constructs a global SLA mapping graph. In the global SLA mapping graph, the boundary coordinator calculates the cosine similarity between the embedding vectors of semantic centroids in different domains, and identifies heterogeneous resource descriptions with similarity higher than a preset threshold as semantic alignment relationships, thereby breaking down the SLA definition barriers between different management domains.

[0050] Step 206: Capability Tag Encapsulation and Publication. After the boundary coordinator completes global graph fusion and alignment, it calls the capability tag management unit of each autonomous system controller to generate formal capability tags. Each capability tag corresponds to a set of semantically aligned domain-level capability descriptions, encapsulating key elements such as latency level, bandwidth level, and reliability category. Ultimately, each autonomous system only needs to publish these semantically aligned and relatively stable capability tags, without exposing internal topology and node details, thereby achieving a balance between privacy protection and cross-domain interoperability while ensuring collaborative feasibility.

[0051] Step 207: After receiving the aggregated results of capability tag sets and SLA mapping maps from multiple boundary coordinators via the execution feedback interface M3, the inter-domain controller constructs a global capability view. This view describes the distribution and performance level of the entire network's computing resources in a unified semantic space, providing standardized input support for subsequent multi-agent reinforcement learning decisions.

[0052] Step 208: Semantic Mapping Feedback and Dynamic Update. When network status or SLA requirements change significantly, the inter-domain controller issues semantic update instructions to each autonomous system controller (AS / RS) via the feedback interface M3. The AS / RS re-executes the performance acquisition and capability update process, generating new capability vectors and mapping subgraphs. The boundary coordinator updates the global semantic structure of the SLA mapping graph accordingly, forming an adaptive semantic closed-loop mechanism to ensure the model dynamically evolves with the system state.

[0053] Through the above steps, this invention establishes a hierarchical and collaborative SLA semantic abstraction and mapping mechanism in a wide-area computing network. Each autonomous domain controller is responsible for local semantic modeling and capability abstraction, the boundary coordinator is responsible for cross-domain semantic fusion and mapping management, and the inter-domain controller achieves global convergence and dynamic optimization. This mechanism not only achieves automated semantic alignment of cross-domain heterogeneous SLAs through BERT and GraphSAGE algorithms, but also shields the underlying resource fluctuations by discretizing capability labels, providing unified, stable, and interpretable semantic support for subsequent reinforcement learning-driven task allocation.

[0054] Example 3: See Figure 3 The schematic diagram of the multi-agent reinforcement learning-driven task allocation policy training process of this embodiment is shown in the figure. This embodiment corresponds to step 3, focusing on the training mechanism of agents within the autonomous system controller, the transfer cost modeling method, and the local policy update process. Following the principle of modular description, this embodiment only focuses on the policy learning closed loop within the autonomous system controller. The collaborative mechanism of cross-domain parameter aggregation and global policy convergence will be described in detail in embodiment 4.

[0055] In the wide-area computing network architecture of this invention, each autonomous system controller (AS / RS) comprises four core functional modules: an SLA semantic mapping module, an agent decision-making module, a transfer cost awareness module, and a cross-domain coordination module. The SLA semantic mapping module provides unified capability labels and semantic index representations for the training process; the agent decision-making module is responsible for constructing reinforcement learning state, action, and policy models; the transfer cost awareness module accurately describes the overhead incurred by cross-domain transfers; and the cross-domain coordination module collaborates with the boundary coordinator during the training phase to achieve parameter sharing and global approximate convergence.

[0056] Before training begins, the system initializes the reinforcement learning environment based on the global capability view generated in Example 2. The Autonomous Domain Controller (ADAS) first obtains the local capability labels and neighborhood semantic relationships generated based on the aforementioned semantic embedding algorithm from the SLA semantic mapping module, and injects them into the state modeling unit of the agent decision-making module for constructing... At any given moment, the observable state space of the autonomous region The state information includes computing node load, queue length, bandwidth and link latency, energy consumption level, as well as semantically abstracted capability tags and neighborhood capability vectors. By introducing this semantic and discretized state representation, the agent can shield itself from high-frequency noise in the underlying metrics, thereby learning robust strategies resistant to oscillations.

[0057] The action generation unit receives the data from the state modeling unit. state of time Then, this state is input into the policy network. To gain actionable space The probability distribution of actions is then analyzed. Subsequently, the action generation unit selects the final action to be executed from the action set based on a preset decision strategy. Action space The action generation unit defines the scheduling behavior into three categories based on the system structure: local execution, intra-domain migration, and cross-domain migration.

[0058] The transfer cost awareness module plays a crucial role in training. Its internal transfer cost calculation unit dynamically evaluates the cost of each action execution, including cross-domain bandwidth consumption, transfer latency, context switching overhead, and task interruption risk. These costs, along with SLA satisfaction and load balancing, constitute the reinforcement learning reward function. This is used to guide policy gradient optimization, enabling the agent to make an adaptive trade-off between performance improvement and cost control. The generated training samples... It is stored in the experience replay pool for subsequent strategy updates.

[0059] The policy learning unit in the agent's decision-making module periodically samples from the experience pool, calculates the gradient using the A3C algorithm, and uses this gradient to adjust the local policy network parameters. Update the strategy gradually to improve its quality, eventually forming a locally optimal strategy for the autonomous region. Once the local policy update is complete, the cross-domain coordination module packages and sends the policy update amount (such as gradient information or model weights), capability labels, and semantic state summary of the local domain to the boundary coordinator through the cross-domain coordination interface M2. This only completes the information reporting action, providing the necessary data input for the federated parameter fusion and global policy convergence performed in Example 4.

[0060] Example 4: The inter-domain controller plays a coordinating, aggregating, evaluating, and re-optimizing role during the global policy training process. This example corresponds to step 4 and follows directly from Example 3, focusing on how, during the training phase, the system constructs a distributed joint training closed loop for cross-domain agents by uniformly fusing the model parameters generated during the policy learning processes of each autonomous domain. See also Figure 4 The global policy training and convergence process includes the following steps: Step S1: In this step, each autonomous system controller (AS / A) uploads its locally trained policy updates (including gradient differences, policy network weights, or behavioral policy probability distributions) to the boundary coordinator via the cross-domain coordination interface M2. After performing initial asynchronous parameter fusion, the boundary coordinator transmits the aggregation result to the inter-domain controller. The boundary coordination control unit within the inter-domain controller standardizes the policy structure and constructs a unified global decision space based on the load scale, data reliability, task contribution rate, and resource capacity weights of each AS / A.

[0061] Step S2: The inter-domain controller uses a federated averaging algorithm to deeply fuse the received parameters. This process aims to address model deviation caused by uneven sample distribution across different autonomous domains. Ultimately, this step generates a preliminary global policy parameter set with cross-domain consistency. This parameter set represents the optimal scheduling logic for the entire network under the current SLA semantic environment.

[0062] Step S3: The inter-domain controller invokes the monitoring feedback unit to perform offline evaluation of the newly generated global policy. Evaluation metrics include changes in task SLA satisfaction rate under simulated conditions, cross-domain migration latency convergence trend, and overall network load balancing level. The evaluation process employs trend modeling or reward distribution analysis to determine the policy's feasibility. If the policy exhibits instability (e.g., migration probability fluctuates drastically within a short period), a penalty mechanism is triggered, requiring specific autonomous regions to re-collect samples for retraining.

[0063] Step S4: When the global policy training loss stabilizes, the transfer overhead decreases significantly, the global reward continues to increase, and the set convergence condition is met, the inter-domain controller marks the policy as the globally optimal policy. Subsequently, the inter-domain controller distributes the policy parameters to each autonomous system controller (ASDC) via the execution feedback interface M3. The agent decision-making modules within each ASDC perform parameter replacement, soft updates, or incremental alignment operations. This process ensures that all agents across the network can make decisions based on a consistent global perspective during the inference phase, providing a unified intelligent foundation for subsequent online operation.

[0064] Example 5: See Figure 5 This embodiment demonstrates the overall operational mechanism of the present invention in the inference stage after the convergence of the reinforcement learning policy, including cross-domain task scheduling, task migration execution, and system-level feedback closed loop. This embodiment corresponds to step 5, focusing on the execution method of the globally optimal policy in a real-world wide-area computing network, and how the system maintains long-term performance stability through monitoring feedback and semantic updates during operation.

[0065] Obtaining the globally optimal strategy in Example 4 Subsequently, the agent decision-making modules within each autonomous system controller enter the inference phase. When a new task arrives at a particular autonomous system, the agent within that system extracts key information from the locally observable state. At this point, the input state is not the original physical indicators, but rather capability labels, cross-domain semantic indexes, and discretized queue lengths generated based on Example 2. The agent inputs these stable features, which mask the subtle fluctuations at the underlying level, into the converged policy network to generate a deterministic scheduling action. This action falls under one of three decision categories: local execution, intra-domain migration, or cross-domain migration. When cross-domain migration is selected, the task enters the cross-domain collaborative decision-making process.

[0066] During the cross-domain collaborative decision-making process, the boundary coordinator performs feasibility matching and performance verification on candidate target domains based on the capability tags and semantic indexes generated in Example 2. The boundary coordinator filters feasible task execution domains and forms a candidate domain set by detecting whether the target domain's capability tags in terms of computing resources, bandwidth, latency, reliability, and energy consumption meet the task SLA constraints. If multiple autonomous regions meet the conditions, the boundary coordinator further generates a candidate domain ranking based on semantic association weights and real-time resource status, and reports the final candidate set to the inter-domain controller.

[0067] The inter-domain controller performs cross-domain scheduling decisions on the global control plane based on the candidate domain set, global capability view, network load status, and task SLA constraints. It uses the global view to generate a global task allocation table, specifying the target execution domain, cross-domain migration path, resource allocation scheme, and required scheduling parameters. The inter-domain controller then synchronizes this table to the relevant autonomous system controllers and corresponding boundary coordinators via the execution feedback interface M3.

[0068] After the task allocation table is issued, each autonomous system controller (AS / RS) executes corresponding task migration operations according to the scheduling instructions, including task context packaging, link scheduling, target domain resource reservation, and execution control. The boundary coordinator is responsible for cross-domain link management and consistency maintenance during the migration process, ensuring that the migration order, resource usage, and execution status of tasks remain synchronized between anonymous domains. After a task is completed in the target domain, the relevant AS / RS reports the task completion time, resource usage, migration cost, and SLA satisfaction to the inter-domain controller.

[0069] After receiving the network-wide operational results, the inter-domain controller enters the system feedback closed-loop phase. Its monitoring feedback unit periodically compiles statistics on task completion quality, SLA satisfaction rate, cross-domain link load, computing power utilization, and migration efficiency, thereby forming a monitoring vector reflecting the overall network operational status. If monitoring data indicates performance degradation in certain autonomous systems (such as link congestion, increased scheduling latency, or increased SLA defaults), the inter-domain controller will invoke the execution feedback interface M3 to notify the controller of that autonomous system to re-collect performance indicators and update the capability vector, thereby driving the adaptive update of the capability label and SLA mapping map.

[0070] After the semantic update is completed, the inter-domain controller will adjust the task scheduling rules in real time based on the new semantic information and operational feedback, so that the generation process of the global task allocation table can reflect the latest system state. This lightweight structured feedback mechanism can restore the scheduling performance of the entire network without changing the reinforcement learning policy parameters, thereby avoiding overall performance degradation caused by short-term fluctuations.

[0071] Through the above-mentioned reasoning execution, candidate domain screening, cross-domain scheduling, migration execution and system feedback closed loop, this invention constructs a continuously evolving cross-domain computing power scheduling system, enabling the entire network to have the ability to adaptively optimize and maintain long-term performance stability under dynamic load conditions.

Claims

1. A semantic cooperative scheduling method for wide-area computing power networks to suppress routing oscillations, characterized in that, The method is based on a hierarchical distributed network architecture and achieves cross-domain computing power resource privacy protection collaboration, end-to-end performance assurance, and routing oscillation suppression through an SLA semantic abstraction mapping model, a multi-agent reinforcement learning collaborative algorithm, and a transfer cost-aware reward mechanism. The method includes the following steps: (1) Based on the geographical distribution and management strategy of computing resources, the computing network is divided into several autonomous domains. Autonomous domain controllers and boundary coordinators are deployed in each autonomous domain, and the autonomous domain controllers and boundary coordinators are registered to the global coordination directory. The autonomous domain consists of local computing nodes, communication links and boundary agents. The autonomous domain controller includes an SLA semantic mapping module, an agent decision-making module, a migration cost perception module and a cross-domain coordination module. (2) Each autonomous domain controller obtains the computing power index within the domain through the performance acquisition unit in the SLA semantic mapping module, and converts it into a capability vector through the capability description generation unit to form a domain-level capability description; after constructing a local semantic mapping subgraph in the domain, the semantic mapping unit reports the local semantic mapping subgraph to the boundary coordinator; the boundary coordinator gathers the local semantic mapping subgraphs of each domain, constructs a global SLA mapping map and generates standardized capability labels to realize the semantic alignment and capability mapping of cross-domain performance indicators; (3) The agent decision module of each autonomous system controller takes the observable state information of its own domain as input and outputs the task placement or migration decision; during the training phase, the agents on each autonomous system controller perform policy iterative updates based on the multi-agent reinforcement learning algorithm to maximize the composite reward function; the observable state information includes computing power load, task queue length, inter-domain link delay and SLA mapping results, and the task placement or migration decision includes local execution, intra-domain migration and cross-domain migration; (4) The boundary coordinator realizes wide-area collaboration through parameter aggregation and capability label matching mechanism. The inter-domain controller forms a globally optimal task allocation scheme based on the training results and generates a task allocation table. The inter-domain controller is deployed on the global control plane of the computing power network and includes intra-domain agent nodes, boundary coordination control units and monitoring feedback units. (5) The inter-domain controller completes task migration and computing power allocation operations through the agent node within the domain according to the task allocation table, and feeds back the execution results to the agent decision module through the monitoring feedback unit to realize the continuous adaptive update of the strategy.

2. The semantic cooperative scheduling method for wide-area computing power networks to suppress routing oscillations according to claim 1, characterized in that, The working process of the SLA semantic mapping module in step (2) specifically includes: (21) The performance acquisition unit periodically collects raw performance data from computing power nodes and communication links within the domain through the resource awareness interface M1. The raw performance data includes CPU / GPU utilization, memory usage, available bandwidth, end-to-end latency, reliability level and energy consumption parameters. (22) The capability description generation unit normalizes and discretizes the original performance data and uses a mapping function with hysteresis effect to map the continuously changing original performance data into discrete levels, generating a standardized capability vector containing latency level, bandwidth level and reliability category. (23) The capability description generation unit performs aggregate analysis on the capability vectors of all computing power nodes in the domain, generates a domain-level capability description that can be shared externally through dimensionality reduction and weighted averaging methods, and establishes an index mapping relationship between the domain-level capability description and the original capability vector. (24) The semantic mapping unit uses the BERT pre-trained model to encode the text description of each SLA index and generate the initial feature vector of the node; it uses each feature item as a node and the performance dependency relationship as an edge to construct a local semantic mapping subgraph; it uses the GraphSAGE graph neural network algorithm to aggregate the feature information of neighboring nodes and extract the semantic center point of the domain. (25) The boundary coordinator gathers the local semantic mapping subgraphs and semantic centroids of each autonomous domain, performs cross-domain graph splicing and semantic association learning, and constructs a global SLA mapping graph; calculates the cosine similarity of the embedding vectors of semantic centroids in different domains, and identifies heterogeneous resource descriptions with similarity higher than a preset threshold as semantic alignment relationships; (26) The capability tag management unit generates standardized capability tags corresponding to the domain-level capability description. The capability tags encapsulate latency level, bandwidth level, reliability category and billing constraints. Each autonomous domain publishes the capability tags through the boundary agent.

3. The semantic cooperative scheduling method for wide-area computing power networks to suppress routing oscillations according to claim 1, characterized in that, The working process of the agent decision-making module in step (3) specifically includes: (31) The state modeling unit constructs the autonomous domain observable state space at time t based on the domain capability labels, neighborhood semantic relationships and real-time operation data generated by the SLA semantic mapping module; the real-time operation data in the domain includes computing node load, task queue length, link bandwidth and latency, and energy consumption level. (32) After receiving the state at time t, the action generation unit inputs the state into the policy network to obtain the action probability distribution of the action space; and selects an action to be executed from the action set based on the preset decision policy. The action set includes local execution, intra-domain migration and cross-domain migration. (33) The policy learning unit periodically samples training samples from the experience replay pool, calculates the gradient using the A3C algorithm, and uses the gradient to update the local policy network parameters to form the local optimal policy of the autonomous region; the experience replay pool stores training samples containing state, action, reward and next time state. (34) The cross-domain coordination module sends the policy update amount, capability label and semantic state summary of the current domain to the boundary coordinator through the cross-domain coordination interface M2; the policy update amount includes gradient information or model weights.

4. The semantic cooperative scheduling method for wide-area computing power networks to suppress routing oscillations according to claim 1, characterized in that, The specific working process of the migration cost awareness module includes: After each action is executed, the migration cost calculation unit evaluates the migration cost generated by the task migration. The migration cost includes cross-domain bandwidth usage, migration latency, context switching loss, and task interruption risk. The composite reward generation unit jointly models SLA satisfaction, migration cost, and load balancing into a composite reward function, the expression of which is: R = a R SLA + b (1 - R cost ) + c R balance Where R is the composite reward value, α, β, and γ are weighting coefficients and α + β + γ = 1, R SLA As a reward for SLA satisfaction, R cost R is the normalized migration cost. balanc As a reward for load balancing; The calculated composite reward value is stored in the experience replay pool for policy iteration and updating of the agent.

5. The semantic cooperative scheduling method for wide-area computing power networks to suppress routing oscillations according to claim 1, characterized in that, The specific process of forming the globally optimal task allocation scheme in step (4) includes: (41) The boundary coordinator receives the policy update data uploaded by each autonomous system controller, performs asynchronous parameter fusion, and then transmits the aggregation result to the inter-domain controller. (42) The boundary coordination control unit of the inter-domain controller performs policy structure standardization processing on the aggregation results and constructs a global decision space based on the load scale, data credibility, task contribution rate and resource capability weight of each autonomous domain; (43) The inter-domain controller uses the federated averaging algorithm to deeply fuse the standardized policy parameters to generate a preliminary global policy parameter set; (44) The monitoring and feedback unit of the inter-domain controller performs offline evaluation of the preliminary global policy parameter set. The evaluation indicators include SLA satisfaction, cross-domain migration delay convergence trend and network load balancing level. If the policy is unstable, the penalty mechanism is triggered and the corresponding autonomous domain is required to re-collect samples for training. (45) When the global policy training loss tends to stabilize, the transfer overhead shows a clear downward trend, and the global reward continues to increase and meets the convergence condition, the preliminary global policy parameter set is marked as the global optimal policy. (46) Based on the global optimal strategy, the global SLA mapping map and the distribution of computing power resources across the entire network, the inter-domain controller generates the global optimal task allocation scheme and the corresponding task allocation table.

6. The semantic cooperative scheduling method for wide-area computing power networks to suppress routing oscillations according to claim 1, characterized in that, The process of task migration, computing power allocation, and policy update in step (5) specifically includes: (51) The inter-domain controller synchronizes the task allocation table to the relevant autonomous domain controllers and boundary coordinators by executing the feedback interface M3; (52) The autonomous domain controller performs task context packaging, link scheduling, target domain resource reservation and execution control operations according to the scheduling instructions in the task allocation table; the boundary coordinator is responsible for cross-domain link management and consistency maintenance during the migration process; (53) After the task is completed in the target domain, the autonomous domain controller will report the operation results such as task completion time, resource consumption, migration cost and SLA satisfaction to the inter-domain controller. (54) The monitoring feedback unit of the inter-domain controller periodically compiles the operation results to form a monitoring vector; the monitoring vector includes task completion quality, SLA satisfaction, cross-domain link load, computing power utilization and migration efficiency. (55) If the monitoring vector indicates that the autonomous system has experienced performance degradation, the inter-domain controller notifies the corresponding autonomous system controller to re-collect performance indicators and update the capability vector through the execution feedback interface M3, thereby driving the capability label and SLA mapping map to be updated adaptively. (56) Based on the updated semantic information and operational feedback, the inter-domain controller adjusts the task scheduling rules and updates the task allocation table to realize the continuous adaptive optimization of the agent decision-making module strategy.

7. The semantic cooperative scheduling method for wide-area computing power networks to suppress routing oscillations according to claim 1, characterized in that, The open interfaces involved in the method include: Resource Awareness Interface M1: Located between the underlying computing power resource nodes and the Autonomous System Controller (AS / DC), it is used by the AS / DC's performance acquisition unit to obtain raw status data from physical computing power nodes and network devices. Cross-domain coordination interface M2: Used for information exchange between the boundary coordinator and different autonomous system controllers. The transmitted content includes policy parameters, capability tags, aggregation results and cross-domain load information. Execution Feedback Interface M3: Used for bidirectional communication between the inter-domain controller and the boundary coordinator and autonomous domain controller. On the one hand, it transmits task allocation and migration strategies, and on the other hand, it sends back execution result information.

8. A semantic cooperative scheduling system for wide-area computing power networks to suppress routing oscillations, characterized in that, The system includes several autonomous domains, boundary coordinators, and inter-domain controllers; Autonomous Domain: Composed of local computing nodes, communication links, and boundary agents, with an autonomous domain controller deployed internally; the autonomous domain controller includes an SLA semantic mapping module, an agent decision-making module, a migration cost awareness module, and a cross-domain coordination module; SLA Semantic Mapping Module: Used to implement cross-domain SLA indicator abstraction and capability mapping, and generate standardized capability labels; It includes a performance acquisition unit, a capability description generation unit, a semantic mapping unit, and a capability tag management unit; Agent Decision Module: Used to generate task allocation and transfer decisions based on multi-agent reinforcement learning algorithms; It includes a state modeling unit, an action generation unit, and a policy learning unit; Migration Cost Awareness Module: Used to model the migration cost of cross-domain tasks and generate a composite reward function; Includes a migration cost calculation unit and a composite reward generation unit; Cross-domain coordination module: used to interact with the boundary coordinator to achieve policy sharing, parameter aggregation, and global optimization; Boundary Coordinator: Deployed at the edge of the autonomous system, it aggregates the local policy parameters, capability tags and status information of adjacent autonomous system controllers, achieves global approximate convergence through an asynchronous parameter fusion mechanism, and feeds back a unified task allocation policy to the autonomous system controller. Inter-domain controller: Deployed in the global control plane of the computing network, it is used to uniformly execute wide-area task allocation and migration strategies; This includes intra-domain agent nodes, boundary coordination and control units, and monitoring and feedback units; Intra-domain proxy nodes: used to execute task placement and migration commands within an autonomous system; Boundary Coordination and Control Unit: Used to manage the collaborative operation of multiple boundary coordinators, enabling cross-domain task allocation synchronization and policy consistency; Monitoring and feedback unit: used to periodically collect network operation information and feed it back to the intelligent agent decision-making module to achieve adaptive optimization of strategies.

9. The wide-area computing power network semantic cooperative scheduling system for suppressing routing oscillations according to claim 8, characterized in that, In the SLA semantic mapping module: The performance acquisition unit collects CPU / GPU utilization, memory usage, available bandwidth, end-to-end latency, reliability level, and energy consumption parameters through the resource-aware interface M1. The capability description generation unit uses a mapping function with hysteresis to map the raw performance data into discrete levels, generating a standardized capability vector. The semantic mapping unit uses the BERT pre-trained model and the GraphSAGE graph neural network algorithm to construct local semantic mapping subgraphs and extract semantic center points; The capability tag management unit generates capability tags with encapsulation latency level, bandwidth level, reliability category, and billing constraints.

10. The wide-area computing power network semantic cooperative scheduling system for suppressing routing oscillations according to claim 8, characterized in that, The operation process of the inter-domain controller includes: Receive the parameter aggregation results transmitted by the boundary coordinator and construct the global decision space; The federated averaging algorithm is used to fuse policy parameters and generate a global policy parameter set; The global policy parameter set is evaluated offline to determine the globally optimal policy. A task allocation table is generated based on the globally optimal strategy and distributed to the autonomous system controller and the boundary coordinator. Collect network-wide operation results to drive adaptive updates of capability tags, SLA mapping maps, and task scheduling strategies.