Cross-domain collaborative multi-task load balancing method and system

By using a joint optimization framework of graph neural networks and non-dominated sorting genetic algorithms, the problems of resource allocation and fault recovery in multi-domain collaborative networks are solved, achieving load balancing with high throughput, high balance and high reliability, which is suitable for scenarios such as 5G/6G and industrial Internet.

CN121644332APending Publication Date: 2026-03-10XIDIAN UNIV +2
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

In multi-domain collaborative networks, existing algorithms struggle to achieve optimal global resource allocation, leading to overload and idle resources in some domains. Furthermore, they lack proactive fault recovery mechanisms and cannot effectively balance throughput maximization, load balancing, and energy consumption optimization. Traditional evolutionary algorithms also have slow convergence speeds and are difficult to respond to environmental changes.

Method used

A joint optimization framework combining graph neural networks and non-dominated sorting genetic algorithms is adopted. Through multi-domain network modeling, communication, computing and fault recovery models are established to perform multi-objective optimization, predict link load and node utilization, and introduce a primary and backup path collaborative fault recovery mechanism to achieve high throughput, balanced and reliable load balancing.

Benefits of technology

It achieves high throughput, high balance, and high reliability cross-domain load balancing in dynamic network environments, improves the convergence speed and solution feasibility of the algorithm, supports real-time operation with hundreds of nodes and thousands of tasks, and is suitable for scenarios such as 5G/6G and industrial internet.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121644332A_ABST
    Figure CN121644332A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-domain collaborative multi-task load balancing method and system which are used for a multi-domain edge core fusion network, and the method comprises the steps: constructing a communication model and a calculation model, respectively depicting the transmission rate and time delay of a task on a main / standby path and the resource distribution between the task and a calculation node, automatically switching to an alternative path through a fault recovery model by utilizing a node / link availability variable when the main path fails; on this basis, a multi-constraint optimization model is established with throughput maximization, load balancing optimization and fault recovery cost minimization as targets, and a joint optimization framework based on a graph neural network and a non-dominated sorting genetic algorithm is proposed; and a non-dominated sorting genetic algorithm is guided to perform population initialization and directed evolution, and a Pareto optimal scheduling strategy is output in combination with a constraint repair mechanism, so that the performance of cross-domain task scheduling, the resource utilization balance and the fault recovery robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of communication technology, in particular to a multi-task load balancing method for cross-domain cooperation.

[0002] The present application also relates to a multi-task load balancing system for cross-domain cooperation. BACKGROUND

[0003] With the rapid development of 5G / 6G networks and the wide application of edge computing technology, the traditional centralized cloud computing architecture is accelerating the evolution of multi-domain cooperative networks with edge-core fusion; under this architecture, computing resources are widely distributed in multiple geographically dispersed edge domains, and each domain is interconnected through the core network to form a cross-domain cooperative computing service system, however, this evolution also brings serious challenges in task scheduling and load balancing.

[0004] Since tasks need to be migrated between different domains, path selection not only involves intra-domain links, but also needs to be optimized in coordination with cross-domain links, while existing algorithms mostly rely on heuristic rules or greedy strategies, making it difficult to achieve optimal configuration of global resources, often leading to overload in some domains while other domains have idle resources; at the same time, actual application scenarios require simultaneous consideration of multiple conflicting goals such as maximum throughput, load balancing, energy optimization and fault recovery, traditional single-objective optimization algorithms cannot effectively trade off, and simple weighted summation is difficult to capture the nonlinear coupling relationship between goals; in addition, in large-scale distributed environments, node failures and link interruptions occur frequently, and existing scheduling algorithms generally lack active prediction and rapid switching mechanisms, often requiring a new scheduling plan after the primary path fails, causing service interruption and a sharp increase in latency. In addition, network topology, task load and resource state are highly dynamic, and traditional evolutionary algorithms such as genetic algorithms or particle swarm optimization have slow convergence speed due to random generation of initial population, making it difficult to respond to environmental changes in a timely manner.

[0005] In recent years, GNN has shown great potential in multi-domain network modeling and feature learning, and can effectively capture the complex dependence between topology structure and nodes, providing a new idea for deep integration of GNN and multi-objective evolutionary algorithms to predict and guide optimization. However, current research still has obvious shortcomings: there is no three-way coupling model for communication-computation-fault recovery in cross-domain scenarios; the integration of GNN and optimization algorithms is relatively shallow, and the predictive ability of GNN for key indicators such as link load and node utilization is not fully utilized; the handling of resource capacity and path uniqueness constraints is relatively rough, making it difficult to ensure the feasibility of the solution; and there is a lack of active fault recovery mechanism for primary and backup path cooperation. SUMMARY

[0006] The application aims to provide a cross-domain collaborative multi-task load balancing method and system, which systematically models cross-domain task scheduling problems, uses GNN to guide population initialization and genetic operation, finely processes constraints and embeds a fast recovery mechanism, thereby achieving high throughput, high balance, high reliability cross-domain load balancing in a dynamic and variable network environment.

[0007] The technical solutions adopted by the application are as follows: A cross-domain collaborative multi-task load balancing method, the method comprising: Multi-domain network and task modeling, dividing the network into multiple domains, each domain including a number of computing nodes, internal links and a task set, all domains being in communication connection with a core network, the core network including a node set and a link set; Modeling, respectively establishing a communication model, a computing model and a fault recovery model; the communication model provides a transmission path during task cross-domain migration, the computing model calculates the load of the target node in task cross-domain migration, and the fault recovery model performs path switching when the transmission path fails; Multi-objective optimization modeling, establishing a multi-objective optimization model according to the cross-domain migration requirements of the tasks, and generating collaborative optimization constraints for cross-domain task scheduling based on the model; Optimization solution, constructing a joint optimization framework, solving the multi-objective optimization model, and outputting the optimal scheduling strategy of the cross-domain tasks.

[0008] Further, the multi-domain network and task modeling specifically includes: dividing the overall network into domains, defined as Each domain has a computing node , an edge set , and an edge cloud cross-domain edge set defined as , the total number of which is equal to the number of domains divided , that is , the node and edge sets in the core network are and , representing the link between two nodes, and the task set in domain is defined as .

[0009] Further, the communication model specifically includes: Let task migrate from node in domain to node in domain , the cross-domain transmission path set is defined as , and in cross-domain task migration, let the routing and forwarding path of the core network be ,in For the task The path hop count, when the main path fails, the task The alternative paths are defined as follows: ,link The capacity is (bit / s), let the variable be... For binary variables, when the task via link During transmission, ,otherwise .Task Upper bound of bottleneck capacity of the main migration path during migration: (1) In the formula, It is a task The main path, given by the above formula, is... The upper bound of the bottleneck capacity of the induced main migration path in the core network is used to characterize the maximum theoretical carrying capacity that the path can provide. Among them, the task The latency consumed by the migration is defined as: (2) In the formula, For link Single-hop transmission delay; link The load is: (3) In the formula, Indicates task The actual allocated and guaranteed transmission rate. The required rate of the task.

[0010] Furthermore, the computational model specifically includes: set up For domain Internal mission Required computing resources (cycles / s), domain internal nodes Total computing power is (cycles / s), variable As a binary variable, when the task In the domain internal nodes When performing calculations, ,otherwise Then the domain Tasks in Migrate to domain Internal node When its target node load is as follows: (4) The total computation latency of the target node is: (5) Where, is the computation data size (bit) of task , is the computation resource (cycles / bit) needed to process 1 bit.

[0011] Further, the fault recovery model specifically includes: Let be the node fault variable, indicates that the domain is available, otherwise ; ; indicates the link fault variable, when , indicates that the link is available, otherwise ; the set of fault links and fault nodes are and ; the set of candidate paths for task migration across domains is , so the state of task migration path is defined as follows: (6) In the formula, when any link in the transmission path is unavailable, , i.e. the above function is used to characterize that the entire migration path is unavailable when any link in the path fails. When the transmission path fails, the variable that selects a candidate path in the set of candidate paths when task migrates is defined as , , which indicates that path is selected, otherwise ; ; When there is a failure in the main path , the candidate path is automatically switched, and the state of the main path is as follows: (7) In summary, the path selection when task migrates satisfies the following formula conditions: (8) (9) Where, indicates the state of the candidate path, For the task The There are several alternative paths. The above constraints are used to ensure that when the primary migration path is available, there is no alternative path competing with it, and when the primary path fails, there is an available alternative migration path for task migration and recovery.

[0012] Furthermore, the multi-objective optimization modeling specifically includes: Define the system's effective throughput as: (10) And define the fault switching cost as: (11) in, Indicates task switching to the first The switching cost required for each alternative migration path; Maximizing throughput is equivalent to " Minimize, and use the maximum link utilization as the upper bound. Upper bound of maximum node utilization Maximum end-to-end delay upper bound As for the remaining objectives, the objective function is as follows: (12) The corresponding multi-objective optimization model is as follows: (13) In the formula, This represents the maximum tolerable latency during the task migration process. and The outgoing link constraints selected for task migration, in middle, Represents a node The sum of outflows, Represents a node The sum of the inflows, when the node When it is the source node, When node When the destination node is When node When it is an intermediate node, ; and These are the constraints for migrating the task to the target computing node. This is a constraint on node computing power, used to limit each task to selecting only one target node. The time consumed during the task migration process must not exceed its maximum tolerable latency. and For task migration path constraints, It is specified that the alternative path must be selected only when the main path is unavailable, It is ensured that only the available alternative path can be selected, and the executability of the recovery path is guaranteed. For the computing node computing capacity constraint, limit the domain The total task demand of the node Does not exceed its computing capacity ; For the forwarding link capacity constraint, it is specified that the sum of the effective rates of tasks carried on the link does not exceed the link capacity , for guaranteeing network side feasibility and avoiding overselling; For multi-objective evaluation linearization.

[0013] Further, the joint optimization framework specifically includes a graph neural network (GNN) prediction module and a non-dominated sorting genetic algorithm II (NSGA-II) optimization module; wherein: The graph neural network prediction module is used to predict the resource state of the core network link and the computing node based on the cross-domain network topology and the monitored link, node and task state. The prediction output includes at least one or more of link utilization or congestion risk, node utilization or overload risk, and can further output the availability discrimination result or failure risk index of the link / node; The NSGA-II optimization module is used to non-dominantly sort the preset multi-objective vector and obtain a Pareto optimal solution set under the condition of satisfying the link capacity, computing capacity, end-to-end delay and failure recovery feasibility constraint; wherein the actual guaranteed rate of the task is represented as a continuous decision variable, and the task is allowed to adopt a speed reduction service strategy to ensure that a feasible solution satisfying the constraint condition can be obtained when the resources are insufficient; Based on the prediction result, the candidate links and candidate nodes are screened, and the unusable links / nodes and the links / nodes with a predicted congestion or overload risk higher than a preset threshold are removed or down-weighted from the candidate set to reduce the search space and improve the proportion of feasible solutions; the threshold is a configurable parameter.

[0014] Further, the GNN prediction module specifically includes a multi-layer graph convolutional network (GCN) and a gated graph neural network (GGNN); wherein: GCN is used to aggregate node and link features of cross-domain networks and generate embedded representations. The input features include at least one or more of the following: link capacity, link latency, and link historical carrying statistics; one or more of the following: node computing power and node historical utilization statistics; and one or more of the following: task demand rate, task data size, and unit data computing demand. GGNN is used to update the embedded representation based on multi-hop neighbor information during task migration events or network state updates, thereby enhancing the ability to represent time-varying loads and fault disturbances. The output layer is set with a multi-task prediction head, which outputs the resource state prediction values ​​or risk indicators of links and nodes respectively. The prediction results are used to provide prior information for NSGA-II, including the basis for initial population construction, genetic operator guidance, and infeasible individual repair.

[0015] Furthermore, the solution process for NSGA-II specifically includes the following steps: (1) Jointly encode the task's computation endpoint decision and the core network routing decision. The encoding shall include at least the following: routing selection variables. Variable allocation for computing nodes Fault switching alternative path selection variables and the actual guaranteed rate variable of the task The initial individuals are heuristically generated based on the congestion / overload risk predicted by GNN, allowing them to prioritize links and nodes with lower risk, and for... Give satisfaction The initial feasible value; (2) The constraint violation is handled by combining the feasibility priority rule with the dynamic penalty function; the penalty function value is adaptively adjusted according to the degree of violation in order to suppress infeasible individuals from entering the elite set.

[0016] (3) Calculate multi-objective vectors for individuals, perform non-dominated sorting to obtain Pareto levels, and use crowding distance to maintain solution set diversity. Combine this with an elite retention strategy to obtain the next generation population.

[0017] (4) In crossover and mutation operations, the GNN output is used to perform targeted perturbation on the gene loci related to the corresponding congested links, overloaded nodes or high-risk paths to guide the search direction.

[0018] (5) When a child individual becomes infeasible, a repair operation is performed based on the constraint violation type; the repair operation includes at least one or more of the following: task relocation, path rerouting, and bandwidth reallocation; when the main path is unavailable, alternative paths that satisfy the availability constraint are selected and updated. This ensures that fault recovery is feasible; (6) Set a convergence criterion. When the Pareto front does not improve significantly in several consecutive iterations or the congestion distance distribution tends to stabilize, terminate the iteration and output the non-dominated solution set. In the output stage, select the final scheduling scheme from the non-dominated solution set according to the business preference and issue it for execution.

[0019] The second technical solution adopted in this invention is a cross-domain collaborative multi-task load balancing system, employing a cross-domain collaborative multi-task load balancing method, the system comprising: The multi-domain network modeling module is used to divide the network into multiple domains and define the computing nodes, links and task sets in each domain, and construct a global topology structure that includes the core network and inter-domain links. The model building module is used to build communication models, computational models, and fault recovery models, which are used to characterize migration paths, node resource load, and fault path switching strategies, respectively. The collaborative optimization modeling module is used to construct a joint optimization model for cross-domain task scheduling based on communication models, computational models, and fault recovery models. The optimization solution module is used to solve multi-objective optimization problems based on graph neural networks and NSGA-II algorithm, and generate Pareto optimal cross-domain task scheduling strategies. The scheduling execution module is used to distribute the optimized scheduling strategy, including task-node allocation and path selection, to network entities.

[0020] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are: This invention introduces a joint optimization framework combining graph neural networks and non-dominated sorting genetic algorithm II (NSGA-II). In multi-domain edge-core fusion network scenarios, it enables collaborative decision-making for multi-task load migration and routing, fully leveraging cross-domain topology and multi-dimensional state information such as links, nodes, and tasks. The graph neural network performs deep modeling and prediction of the global network state, embedding predictions of link utilization, node load, and availability as prior information into the population initialization, genetic operator design, and infeasible solution repair processes of NSGA-II. This ensures that the optimization search focuses on high-quality feasible solution regions from the outset, making it easier to obtain a global Pareto optimal solution set while maintaining constraints. This reduces the risk of the algorithm getting trapped in local optima. In the multi-domain simulation scenarios provided in the examples, it demonstrates faster convergence speed and better throughput, load balancing, and latency metrics compared to traditional heuristic and evolutionary methods.

[0021] In terms of optimizing target design, this invention integrates indicators such as system effective throughput, maximum utilization of links and nodes, end-to-end latency, and fault recovery cost into a multi-objective collaborative optimization framework. It strictly controls key resources and performance indicators such as link capacity, computing capacity, path uniqueness, and service latency through constraints, and distinguishes between the required task rate and the actual guaranteed rate, introducing a configurable rate-reduction service strategy. This ensures model feasibility even under resource constraints or local congestion, while guaranteeing the minimum service quality requirements of the task, thus improving overall resource utilization and scheduling flexibility. Based on a fault recovery model with primary and backup path collaboration and a path selection and switching mechanism based on predicted states, when a link or node failure or availability degradation is detected, the system can automatically select a migration path from a pre-calculated set of alternative paths that meets availability and constraint conditions. This enables rapid task migration and recovery, shortens fault recovery time, and enhances the continuity of cross-domain services and overall network availability.

[0022] Meanwhile, the joint optimization framework proposed in this invention adopts a modular design. The algorithm's time overhead increases exponentially with the number of tasks, computing nodes, and iterations. Under the example settings, it can support the online operation of multi-domain edge-core converged networks with hundreds of nodes and thousands of tasks, demonstrating good scalability and engineering deployment feasibility. Since the core modeling method and optimization process are insensitive to specific service types, access technologies, and network scales, the relevant objective functions and constraints in the framework can be flexibly expanded to various optimization dimensions such as energy consumption, economic cost, security strategies, or service levels according to different application scenarios. It is applicable to various heterogeneous edge computing scenarios such as 5G / 6G mobile communication, industrial internet, and intelligent transportation, thus possessing good versatility and application prospects. Attached Figure Description

[0023] Figure 1 This is a flowchart of a cross-domain collaborative multi-task load balancing method provided in an embodiment of the present invention; Figure 2 This is a structural block diagram of a cross-domain collaborative multi-task load balancing system provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the multi-task load balancing framework provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the cross-domain collaborative multi-task load balancing method architecture provided in an embodiment of the present invention; Figure 5 This is an analysis chart of the optimization progress of NSGA-II provided in the embodiments of the present invention; Figure 6 This is a schematic diagram comparing the performance of different algorithms provided in the embodiments of the present invention.

[0024] In the diagram: 1. Multi-domain network modeling module; 2. Model building module; 3. Collaborative optimization modeling module; 4. Optimization solution module; 5. Scheduling and execution module. Detailed Implementation

[0025] The present invention will now be described in detail with reference to the accompanying drawings.

[0026] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0027] This invention aims to achieve efficient collaborative scheduling and optimal resource allocation for tasks in a multi-domain edge-core converged network. The network architecture is as follows: Figure 3 As shown, a cross-domain collaborative multi-task load balancing method is designed, and its overall architecture is as follows: Figure 4 As shown, the system of this invention adopts a four-stage closed-loop design of "network modeling - model construction - joint optimization - scheduling execution", which realizes the collaborative optimization of communication, computing and fault recovery.

[0028] The technical objectives and advantages of this invention are as follows: Unlike traditional scheduling models that only focus on the computing or communication layers, this invention takes a system perspective and models the transmission rate, node computing power, and link availability involved in cross-domain task migration in a unified manner. It establishes a comprehensive mathematical description that can simultaneously characterize the topological relationships and fault states of multiple domains, providing a quantifiable analytical basis for cross-domain scheduling.

[0029] By introducing GCN and GGNN, this invention can adaptively learn the complex dependencies between nodes and links, and make high-precision predictions of link load distribution, node utilization and resource consumption trends during task migration, providing prior guidance information for subsequent optimization algorithms, thereby improving the accuracy and foresight of scheduling decisions.

[0030] This invention introduces a GNN-driven heuristic initialization and directed mutation mechanism into the evolution process of NSGA-II, dynamically constraining node capacity and link bandwidth to achieve synergistic optimization of maximizing throughput, minimizing load variance, and minimizing recovery cost. This mechanism significantly improves the algorithm's global convergence speed and solution feasibility, effectively avoiding getting trapped in local optima.

[0031] The system of this invention forms a closed-loop control link of "prediction-optimization-execution-feedback" through real-time distribution and execution feedback of optimization results. When node or link anomalies are detected, the algorithm can automatically trigger task migration based on the alternative path set and fault availability variables to ensure task continuity. This mechanism supports real-time operation in multi-domain networks with hundreds of nodes and thousands of tasks, providing a directly deployable technical solution for scenarios such as 5G / 6G, industrial internet, and satellite edge computing.

[0032] The present invention will be described in detail below with reference to embodiments: Example This embodiment provides a cross-domain collaborative multi-task load balancing method, such as... Figure 1 As shown, it specifically includes: Multi-domain Networks and Task Modeling: The multi-domain edge-core fusion network provided in this embodiment is as follows: Figure 3 As shown, the entire network consists of multiple edge domains and a core network. The network is divided into multiple domains, each containing several computing nodes and internal links for processing user-side tasks locally. All domains are connected to the core network via macro base stations. The core network contains a set of nodes and a set of links, responsible for global resource scheduling and cross-domain task coordination, thus forming a multi-domain collaborative network topology that integrates computing and communication. The task sets, task resource requirements, node computing capabilities, and link transmission rates within each domain are defined as follows: Divide the overall network into Each domain is defined as Each domain The computing nodes inside are The edge set is The cross-domain edge set of edge cloud is defined as follows: The total number is equal to the number of domains. ,Right now The sets of nodes and edges in the core network are respectively and , A domain represents a link between two nodes. The task set within is defined as .

[0033] Model building: Establish communication model, computation model and fault recovery model respectively; The communication model is as follows: Set a task From Domain Nodes within Migrate to domain Nodes within Its cross-domain transmission path set is defined as In cross-domain task migration, let the routing and forwarding path of the core network be... ,in For the task The path hop count, when the main path fails, the task The alternative paths are defined as follows: ,link The capacity is (bit / s), let the variable be... For binary variables, when the task via link During transmission, ,otherwise .Task Upper bound of bottleneck capacity of the main migration path during migration: (1) In the formula, It is a task The main path, given by the above formula, is... The upper bound of the bottleneck capacity of the induced main migration path in the core network is used to characterize the maximum theoretical carrying capacity that the path can provide. Among them, the task The latency consumed by the migration is defined as: (2) In the formula, For link Single-hop transmission delay; link The load is: (3) In the formula, Indicates task The actual allocated and guaranteed transmission rate. The required rate of the task.

[0034] The calculation model specifically includes: set up For domain Internal mission Required computing resources (cycles / s), domain internal nodes Total computing power is (cycles / s), variable As a binary variable, when the task In the domain internal nodes When performing calculations, ,otherwise Then the domain Tasks in Migrate to domain internal nodes At that time, the target node load is as follows: (4) The total computation latency of the target node is: (5) in, For the task The size of the calculated data (bits). The computational resources required to process 1 bit (cycles / bit).

[0035] The fault recovery model specifically includes: set up For node fault variables, Representation domain internal nodes Available, otherwise ; Represents a link failure variable, when When, it means Available, otherwise The sets of faulty links and faulty nodes are respectively and ;Task The alternative paths for cross-domain migration are Therefore, the task is defined. The migration path status is as follows: (6) In the formula, when any link in the transmission path is unavailable, In other words, the above function characterizes the unavailability of the entire migration path when any link in the path fails. When the transmission path fails, the task... The variable for selecting a candidate path from the set of candidate paths during migration is defined as follows: , Indicates the selection of a path ,otherwise ; When the main path When a fault occurs, the system automatically switches to an alternative path. The status of the primary path is as follows: (7) In summary, the task The path selection during migration must satisfy the following condition: (8) (9) in, Indicates the status of alternative paths. For the task The There are one alternative path. The above constraints are used to ensure that when the primary migration path is available, there is no alternative path competing with it, and that when the primary path fails, there is an available alternative migration path for task migration and recovery.

[0036] Multi-objective optimization modeling: A multi-objective optimization model is established based on the cross-domain migration requirements of the task, and collaborative optimization constraints for cross-domain task scheduling are generated based on the model, as follows: Centralized computing platforms need to consider both task requirements and load balancing when planning strategies to achieve optimal allocation of network resources; therefore, the effective throughput of the system is defined as: (10) And define the fault switching cost as: (11) in, Indicates task switching to the first The switching cost required for each alternative migration path; Maximizing throughput is equivalent to " Minimize, and use the maximum link utilization as the upper bound. Upper bound of maximum node utilization Maximum end-to-end delay upper bound As for the remaining objectives, the objective function is as follows: (12) Therefore, the specific problem of cross-domain collaborative multi-task load balancing is modeled as follows: (13) In the formula, This represents the maximum tolerable latency during the task migration process. and The outgoing link constraints selected for task migration, in middle, Represents a node The sum of outflows, Represents a node The sum of the inflows, when the node When it is the source node, When node When the destination node is When node When it is an intermediate node, ; and These are the constraints for migrating the task to the target computing node. This is a constraint on node computing power, used to limit each task to selecting only one target node. The time consumed during the task migration process must not exceed its maximum tolerable latency. and For task migration path constraints, The rule stipulates that an alternative path must be selected only when the primary path is unavailable. This ensures that only available alternative paths can be selected, guaranteeing the executability of the recovery path; Constraints on the computing power of computing nodes, restricted domain internal nodes The total task requirements do not exceed its computing power ; To constrain forwarding link capacity, it is stipulated that for any core network link, the sum of the effective rates of the tasks carried on the link shall not exceed the link capacity. This is used to ensure network feasibility and prevent overselling; Used for linearization in multi-objective evaluation.

[0037] Optimization Solution: A joint optimization framework based on graph neural networks and non-dominated sorting genetic algorithm II is constructed to solve the multi-objective optimization problem and output a Pareto optimal scheduling strategy, specifically including: Feature learning and prediction stages in graph neural networks: GCN is used to aggregate node and link features of cross-domain networks and generate embedded representations. The input features include at least one or more of the following: link capacity, link latency, and link historical carrying statistics; one or more of the following: node computing power and node historical utilization statistics; and one or more of the following: task demand rate, task data size, and unit data computing demand. GGNN is used to update the embedded representation based on multi-hop neighbor information during task migration events or network state updates, thereby enhancing the ability to represent time-varying loads and fault disturbances. The output layer is set with a multi-task prediction head, which outputs the resource state prediction values ​​or risk indicators of links and nodes respectively. The prediction results are used to provide prior information for NSGA-II, including the basis for initial population construction, genetic operator guidance, and infeasible individual repair.

[0038] Multi-objective collaborative optimization stage: (1) Jointly encode the task's computation endpoint decision and the core network routing decision. The encoding shall include at least the following: routing selection variables. Variable allocation for computing nodes Fault switching alternative path selection variables and the actual guaranteed rate variable of the task The initial individuals are heuristically generated based on the congestion / overload risk predicted by GNN, allowing them to prioritize links and nodes with lower risk, and for... Give satisfaction The initial feasible value; (2) The constraint violation is handled by combining the feasibility priority rule with the dynamic penalty function; the penalty function value is adaptively adjusted according to the degree of violation in order to suppress infeasible individuals from entering the elite set.

[0039] (3) Calculate multi-objective vectors for individuals, perform non-dominated sorting to obtain Pareto levels, and use crowding distance to maintain solution set diversity. Combine this with an elite retention strategy to obtain the next generation population.

[0040] (4) In crossover and mutation operations, the GNN output is used to perform targeted perturbation on the gene loci related to the corresponding congested links, overloaded nodes or high-risk paths to guide the search direction.

[0041] (5) When a child individual becomes infeasible, a repair operation is performed based on the constraint violation type; the repair operation includes at least one or more of the following: task relocation, path rerouting, and bandwidth reallocation; when the main path is unavailable, alternative paths that satisfy the availability constraint are selected and updated. This ensures that fault recovery is feasible; (6) Set a convergence criterion. When the Pareto front does not improve significantly in several consecutive iterations or the congestion distance distribution tends to stabilize, terminate the iteration and output the non-dominated solution set. In the output stage, select the final scheduling scheme from the non-dominated solution set according to the business preference and issue it for execution.

[0042] This embodiment also provides a cross-domain collaborative multi-task load balancing system, such as Figure 2 As shown, it specifically includes: Multi-domain network modeling module 1 is used to divide the network into multiple domains and define the computing nodes, links and task sets in each domain, and construct a global topology structure including the core network and inter-domain links; Model building module 2 is used to establish communication models, computational models and fault recovery models, and to formalize the constraints of task migration paths, resource allocation and fault switching. Collaborative optimization modeling module 3 is used to construct a joint optimization model for cross-domain task scheduling based on communication models, computational models, and fault recovery models; The optimization and solution module 4 is used to solve multi-objective optimization problems based on graph neural networks and NSGA-II algorithm, and generate Pareto optimal cross-domain task scheduling strategies. The scheduling and execution module 5 is used to distribute the optimized task-node allocation and path selection strategies to network entities, thereby enabling dynamic execution of cross-domain load balancing and fault recovery.

[0043] Simulation results To verify the effectiveness of the method of this invention, a multi-domain fusion network consisting of multiple edge domains and a core network was constructed for simulation experiments. The experimental network contained multiple heterogeneous computing nodes with a wide distribution of node computing power to simulate the real environment, and the link bandwidth was configured in an increasing manner from the intra-domain to the core network. The task set covered different priorities and resource requirements to verify the adaptability of the algorithm in heterogeneous task scenarios. The graph neural network used a multi-layer graph convolutional structure combined with an attention mechanism for feature learning, and the NSGA-II algorithm adopted an adaptive parameter adjustment strategy to balance global search and local optimization capabilities.

[0044] Based on the aforementioned joint optimization method, this embodiment selects three representative performance indicators—system effective throughput, load balancing (characterized by link / node load variance), and average end-to-end latency—and plots their optimization process curves as a function of the NSGA-II iteration number, as shown below. Figure 5 As shown, all three metrics exhibit a generally monotonic or near-monotonic improvement trend during the evolution process: the algorithm rapidly improves the system's effective throughput and significantly reduces load variance and average end-to-end latency in the early stages of evolution; it demonstrates the Pareto trade-off characteristics among multiple objectives in the middle stages; and it gradually converges and tends to stabilize in the later stages. Compared with the initial state, the final scheduling scheme achieves performance improvements of approximately 25%, 67%, and 70% in throughput, load balancing, and average latency, respectively, indicating that the method of this invention has good convergence and stability in terms of throughput improvement, load balancing optimization, and latency control. Simultaneously, the algorithm maintains good population diversity throughout the evolution process, effectively avoiding premature convergence.

[0045] To comprehensively evaluate the algorithm's performance, the method of this invention was compared with typical methods such as random allocation, greedy algorithms, and round-robin allocation. The results are as follows: Figure 6 As shown, the method of this invention significantly outperforms the comparative methods in three key indicators: throughput, load balancing, and average latency. Compared to the random allocation method, the three indicators are improved by 46.5%, 48.8%, and 72.9%, respectively; compared to the greedy algorithm, the three indicators are improved by 23.8%, 4.2%, and 71.7%, respectively; and compared to the round-robin allocation method, the three indicators are improved by 21.9%, 17.5%, and 72.5%, respectively. The comparative results show that the method of this invention, through multi-objective collaborative optimization guided by graph neural networks, effectively overcomes the shortcomings of traditional methods in terms of global optimization capability, load balancing, and latency control, verifying the effectiveness and superiority of the proposed technical solution.

[0046] The cross-domain collaborative multi-task load balancing method provided by this invention can be implemented entirely or partially through software, hardware, firmware, or any other combination. When implemented in software, it can be embodied entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this invention are implemented entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted between different computing devices through a transmission medium. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means such as coaxial cable, optical fiber, digital subscriber line DSL, or wireless means such as infrared, wireless communication, microwave, etc. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. Available media include, but are not limited to, magnetic media such as floppy disks, hard disks, magnetic tapes, optical media such as DVDs and Blu-ray discs, or semiconductor media such as solid-state drives (SSDs) and flash memory.

[0047] This article uses specific embodiments to illustrate the principles and implementation methods of the present invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of the present invention. It should be noted that those skilled in the art can make several improvements and modifications to the present invention without departing from the principles of the present invention, and these improvements and modifications also fall within the protection scope of the claims of the present invention.

Claims

1. A method for multi-task load balancing of cross-domain collaboration, characterized in that, The method comprises: Multi-domain network and task modeling, dividing the network into multiple domains, each domain including a number of computing nodes, internal links and a task set, all domains being in communication connection with a core network, the core network including a node set and a link set; Modeling, respectively establishing a communication model, a computing model and a fault recovery model; the communication model provides a transmission path in the task cross-domain migration process, the computing model calculates the load of the target node in the task cross-domain migration, and the fault recovery model performs path switching when the transmission path fails; Multi-objective optimization modeling, establishing a multi-objective optimization model according to the cross-domain migration requirements of the task, and generating a collaborative optimization constraint of cross-domain task scheduling based on the model; Optimization solving, constructing a joint optimization framework, solving the multi-objective optimization model, and outputting an optimal scheduling strategy of the cross-domain task.

2. The method of claim 1, wherein, The multi-domain network and task modeling specifically includes: dividing the whole network into domains, defined as , each domain has a set of computing nodes , a set of edges , and a set of cross-domain edges defined as , the total number of which is equal to the number of divided domains , that is , the node and edge sets in the core network are respectively and , represents the link between two nodes, and the task set in the domain is defined as .

3. The method of claim 2, wherein, The communication model specifically comprises: Set a task From Domain Nodes within Migrate to domain Nodes within Its cross-domain transmission path set is defined as In cross-domain task migration, let the routing and forwarding path of the core network be... ,in For the task The path hop count, when the main path fails, the task The alternative paths are defined as follows: ,link The capacity is (bit / s), let the variable be... For binary variables, when the task via link During transmission, ,otherwise ,Task Upper bound of bottleneck capacity of the main migration path during migration: (1) wherein is the main path of the task , and formula (1) gives the upper bound of the bottleneck capacity of the main migration path in the core network induced by , which is used to characterize the maximum theoretical carrying capacity that the path can provide. wherein the task The latency consumed by the migration is defined as: (2) In the formula, is the single-hop transmission delay when the link is a single-hop transmission. link The load of the link is: (3) wherein representing a task the actual allocated and guaranteed transfer rate, is the demand rate for the task.

4. The method of claim 2, wherein, The computing model specifically comprises: set up For domain Internal mission Required computing resources (cycles / s), domain internal nodes Total computing power is (cycles / s), variable As a binary variable, when the task In the domain internal nodes When performing calculations, ,otherwise Then the domain Tasks in Migrate to domain internal nodes At that time, the target node load is as follows: (4) The total computing delay of the target node is: (5) wherein is the size of the computation data (bits), is the size of the computation data (bits), is the size of the computation data (bits).

5. The method of claim 2, wherein, The fault recovery model specifically comprises: Let be the node failure variable, denote the domain be the inner node available, otherwise ; be the link failure variable, when denote available, otherwise ; the set of failed links and failed nodes are and respectively; the task is the candidate path set for the inter-domain migration, which is , so the task is defined as follows: the state of the migration path is (6) wherein, when any link in the transmission path is unavailable, i.e. the above function is used to characterize that the entire migration path is unavailable when any link in the path fails; when the transmission path fails, the task The variable that selects a certain alternative path in the set of alternative paths during migration is defined as , represents the selected path , otherwise ; When the primary path fails, the alternative path is automatically switched in, and the status of the primary path is as follows: (7) In summary, the task The path selection at migration satisfies the following equation: (8) (9) wherein, represents an alternative path state, is a task of the first alternative path.

6. The method of claim 1, wherein, The multi-objective optimization modeling specifically comprises: The system effective throughput is defined as: (10) The fault switching cost is defined as: (11) In the formula, represents the task switching to the first switching cost required for the alternative migration path. "maximize throughput" is equivalently translated into "minimize latency" with upper bounds on maximum link utilization , maximum node utilization , maximum end-to-end latency As the remaining objective, the objective function is given by the following equation: (12) The corresponding multi-objective optimization model is specifically as follows: (13) wherein, is the maximum tolerable delay for the task migration process, and is the out-link constraint for task migration selection, in , represents the sum of out-flow traffic of node , represents the sum of in-flow traffic of node , when node is the source node, , when node is the destination node, , when node is the intermediate node, ; and is the constraint condition for the task migrated to the target computing node, is the node computing capability constraint, for limiting each task to select only one target node, constraint the task migration process time consumption not to exceed its tolerable maximum delay, and is the task migration path constraint, provides that only when the main path is unavailable, the alternative path is selected, selects from the available alternative paths, to ensure the executability of the recovery path; is the computing node computing capability constraint, limiting the total task demand of node in the domain not to exceed its computing capability ; is the forwarding link capacity constraint, which provides that for any core network link, the sum of the effective rate of tasks carried on the link does not exceed the link capacity , for ensuring the network side feasibility and avoiding overselling; for multi-objective evaluation linearization.

7. The method of claim 1, wherein, The joint optimization framework specifically comprises a graph neural network prediction module and a non-dominated sorting genetic algorithm II optimization module; wherein: The graph neural network prediction module is used to predict the resource state of the core network link and the computing node based on the cross-domain network topology and the monitored link, node and task state; the prediction output at least includes one or more of link utilization or congestion risk, node utilization or overload risk, and can further output the availability discrimination result or fault risk index of the link / node; The non-dominated sorting genetic algorithm II optimization module is used to non-dominantly sort a preset multi-objective vector and obtain a Pareto optimal solution set under the condition of meeting the link capacity, computing capacity, end-to-end delay and fault recovery feasibility constraints; wherein the actual guaranteed rate of the task is represented as a continuous decision variable, and a speed reduction service strategy is allowed for the task to ensure that a feasible solution meeting the constraint condition can be obtained when resources are insufficient; Based on the prediction result, the candidate links and candidate nodes are screened, and the unusable links / nodes and the links / nodes with a predicted congestion or overload risk higher than a preset threshold are removed or weighted from the candidate set, and the threshold is a configurable parameter.

8. The method of claim 7, wherein, The graph neural network prediction module specifically comprises a multi-layer graph convolution network and a gated graph neural network; wherein: The multi-layer graph convolution network GCN is used to aggregate the node and link features of the cross-domain network and generate embedded representations, and the input features at least include one or more of link capacity, link delay, link historical bearing statistics, one or more of node computing power and node historical utilization statistics, and one or more of task demand rate, task data size and unit data computing demand; A gated graph neural network (GGNN) is used to update embedding representations based on multi-hop neighbor information at task migration events or network state updates to enhance the ability to characterize time-varying loads and fault disturbances; a multi-task prediction head is set in the output layer to output resource state prediction values or risk indicators of links and nodes, respectively, and the prediction results are used to provide prior information for a non-dominated sorting genetic algorithm II (NSGA-II), including initial population construction, genetic operator guidance, and infeasible individual repair.

9. The method of claim 8, wherein, The solving process of the NSGA-II optimization module specifically includes the following steps: (1) Jointly encode the task's computation endpoint decision and the core network routing decision. The encoding shall include at least the following: routing selection variables. Variable allocation for computing nodes Fault switching alternative path selection variables and the actual guaranteed rate variable of the task The initial individuals are heuristically generated based on the congestion / overload risk predicted by GNN, allowing them to prioritize links and nodes with lower risk, and for... Give satisfaction The initial feasible value; (2) The constraint violation is processed by combining the feasibility priority rule with the dynamic penalty function; the penalty function value is adaptively adjusted according to the degree of violation; (3) The multi-objective vector of the individual is calculated, the non-dominated sorting is performed to obtain the Pareto hierarchy, the solution set diversity is maintained by using the crowding distance, and the next generation population is obtained by combining the elite reservation strategy; (4) In the crossover and mutation operations, the GNN output is used to implement directional disturbance on the gene bits related to the corresponding congested link, overloaded node or high-risk path to guide the search direction; (5) When the child individual is infeasible, performing a repair operation based on the constraint violation type; the repair operation at least includes one or more of task relocation, path re-routing, bandwidth reallocation; when the primary path is unavailable, selecting and updating from the alternative path that satisfies the availability constraint , so as to ensure that the fault recovery is feasible; (6) A convergence criterion is set, when the Pareto frontier has no significant improvement or the crowding distance distribution tends to be stable in continuous iterations, the iteration is terminated and the non-dominated solution set is output; in the output stage, the final scheduling scheme is selected from the non-dominated solution set according to the business preference and is executed.

10. A multi-task load balancing system for cross-domain collaboration, employing the multi-task load balancing method for cross-domain collaboration according to any one of claims 1-9. The system comprises: A multi-domain network modeling module is used to divide the network into multiple domains and define the computing nodes, links and task sets in each domain, and to build a global topology structure including the core network and inter-domain links; A model construction module is used to establish a communication model, a computing model and a fault recovery model for representing migration paths, node resource loads and fault path switching strategies, respectively; A collaborative optimization modeling module is used to construct a joint optimization model of cross-domain task scheduling based on the communication model, the computing model and the fault recovery model; An optimization solving module is used to jointly solve the multi-objective optimization problem based on the graph neural network and the NSGA-II algorithm to generate a Pareto optimal cross-domain task scheduling strategy; A scheduling execution module is used to issue the task-node allocation and path selection in the obtained scheduling strategy to network entities.

Citation Information

Cited By

  • An ai-based electric drive support group collaborative power supply and load balancing method

    CN122553116A