Multi-agent system-oriented self-healing graph scheduling system and method
By configuring master-slave nodes and real-time monitoring mechanisms in a multi-agent system, the problem of task interruption caused by node failure is solved, and rapid recovery and business continuity are achieved. It is suitable for high-availability scenarios such as smart finance, industrial Internet, and autonomous driving.
Patent Information
- Application Number
- CN202511277906.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-14
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing multi-agent systems struggle to balance business continuity and response speed when faced with node failures, system upgrades, and time-sensitive scenarios. Traditional solutions cause task interruptions when nodes fail instantaneously, and the retry mechanism cannot meet millisecond-level response requirements, affecting data integrity and business continuity.
It uses a DAG integrity verifier, a hot-swap edge manager, a dual-channel Gossip heartbeat module, and a state snapshot recording unit to ensure task topology acyclicity and data continuity by monitoring node status in real time, configuring master and backup nodes, and quickly restoring node status.
It achieves millisecond-level rapid recovery when a node fails, reduces the average recovery time to 18ms, reduces annual downtime by 99%, ensures business continuity and data integrity, and supports online hot upgrades and seamless switching.
Smart Images

Figure CN120780439A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer science and artificial intelligence, in particular, to a self-healing graph scheduling system and method for multi-agent system (MAS). BACKGROUND
[0002] In the field of multi-agent system (MAS) task scheduling, with the increasing demand for real-time and high availability in scenarios such as intelligent finance, industrial IoT and autonomous driving, the existing technology faces significant challenges. Currently, multi-agent task flow is usually organized in the form of a directed acyclic graph (DAG), but node transient failure (such as GPU memory overflow, container forced termination, etc.) will directly cause the overall DAG to be interrupted, causing the business process to be forced to stop; when online hot upgrade or policy switching, the traditional scheme often needs to pause the global process, seriously affecting the business continuity, and cannot meet the continuous operation needs of key scenarios. In addition, the existing retry mechanism generally takes seconds or even minutes as the time granularity, and in scenarios such as high-frequency trading and industrial control that require millisecond-level response, data loss or decision failure occurs due to insufficient recovery time, making it difficult to adapt to the development needs of real-time business. These technical bottlenecks make it difficult for existing multi-agent task scheduling systems to balance business continuity and response speed when facing node failures, system upgrades and high-time scenarios.
[0003] Through the retrieval of patent documents, it is found that the patent for invention with publication number CN112506627A discloses a directed acyclic graph task scheduling method, system, medium, device and terminal. The method comprises the following steps: S101: obtaining task parameters and system parameters, and user equipment energy consumption constraints and task execution time delay constraints; S102: performing initial processor mapping on all subtasks in the DAG; S103: constructing a subtask priority list; S104: scheduling each subtask in turn according to the priority list, and decomposing the energy consumption and time delay constraints of the entire task to each subtask, and selecting the processor with the maximum execution reliability that meets the decomposed constraints as the processing position of the subtask. The patent focuses on task scheduling under energy consumption and time delay constraints, and lacks a fault self-recovery mechanism.
[0004] In summary, in view of the problems of the existing technology, it is urgent to study a self-healing graph scheduling system and method for multi-agent system (MAS). SUMMARY
[0005] In view of the defects in the prior art, the purpose of the present application is to provide a self-healing graph scheduling system and method for multi-agent system (MAS).
[0006] According to the self-healing graph task scheduling system provided by the present application, the self-healing graph task scheduling system comprises: A DAG integrity verifier for verifying that the task graph topology is still a DAG after a node bypass; A hot-swap edge manager for maintaining a primary node ID and a backup node ID for an edge in the task graph and performing a bypass operation upon a node failure; A dual-channel Gossip heartbeat module for monitoring node latency in real time through a dual channel of UDP and GRPC, with UDP and GRPC serving as backups for each other; A state snapshot recording unit for recording state snapshots of the primary node and the backup node in a periodic manner, with the state snapshots adopting an XOR incremental compression algorithm; A snapshot playback unit for restoring the node state based on the state snapshots.
[0007] The application also provides a self-healing graph task scheduling method, comprising the following steps: Step S1, establishing a directed acyclic graph of a task flow composed of multiple intelligent agent nodes; Step S2, configuring a primary node and a backup node for an edge in the directed acyclic graph to form a hot-swap edge; Step S3, recording state snapshots of the primary node and the backup node in a periodic manner through a state snapshot recording unit; Step S4, detecting the health status of the primary node through a dual-channel Gossip heartbeat module, and determining that the primary node is failed if the heartbeat of the primary node meets a preset failure condition, isolating the primary node and recording a failure event; Step S5, when it is determined that the primary node is failed, redirecting task traffic related to the primary node to the backup node and restoring the running state of the backup node based on the state snapshots, Step S6, if the task traffic redirection and the running state restoration of the backup node are both successful, verifying the acyclic nature of the topology of the switched directed acyclic graph, and if a loop is detected, rolling back to the state before the switching and marking the backup node as unavailable; Step S7, writing the failure event of the primary node, the switching of the backup node, and the topology verification result of the directed acyclic graph into an audit log for monitoring and analysis.
[0008] Preferably, in step S4, the dual-channel Gossip heartbeat module alternately or in parallel sends heartbeat messages using UDP and GRPC.
[0009] Preferably, in step S4, the heartbeat sending interval is not greater than 500 milliseconds.
[0010] Preferably, in step S4, the preset failure condition includes that a single heartbeat delay exceeds a preset threshold, and the preset threshold is not greater than 20 milliseconds.
[0011] Preferably, in step S4, the preset failure condition comprises: no heartbeat response is received for a plurality of times in succession.
[0012] Preferably, in step S3, the current directed acyclic graph topology and node state are recorded.
[0013] Preferably, in step S3, the state snapshot adopts an XOR-delta compression algorithm, and the size of a single node log is not more than 4 kB.
[0014] Preferably, in step S5, if the state recovery of the backup node fails, the backup node is marked as unavailable and an alarm is triggered.
[0015] Preferably, the total fault recovery time T recover satisfies: T recover ≤T detect +T redirect +T replay wherein, T detect is a heartbeat detection time delay, T redirect is a traffic redirection time delay, and T replay is a state playback time delay.
[0016] Compared with the prior art, the present application has the following beneficial effects: 1. The present application effectively avoids the task DAG overall interruption caused by node instantaneous failure by the Hot-Swap Edge hot switching technology, reduces the average recovery time to 18 ms, and reduces the annual downtime by about 99%.
[0017] 2. The present application solves the health detection delay and false negative problem caused by single-channel blockage by the double-channel Gossip heartbeat monitoring technology, improves the fault detection accuracy, and realizes zero abnormal detection rate.
[0018] 3. The present application solves the problem of large state recovery data volume and long playback time by the XOR-delta lightweight snapshot and fast playback technology, can complete node state reconstruction within 10 ms, and ensures zero data loss.
[0019] 4. The present application solves the topology loop and consistency risk after rapid bypass by the DAG integrity verifier and self-healing rollback technology, guarantees 100% continuous business flow, and maintains stable and loop-free topology structure.
[0020] 5. The present application breaks the limitation that system upgrade must be released in downtime by the online hot upgrade mechanism, realizes more than 90% of tasks in the upgrade process "no sense" switching, and realizes zero business interruption. BRIEF DESCRIPTION OF DRAWINGS
[0021] Other features, objects, and advantages of the application will become more apparent from the following detailed description when read in conjunction with the accompanying drawings: Figure 1 Structure diagram of a multi-agent system-oriented self-healing graph scheduling system in an embodiment of the application; Figure 2 Working timing diagram of a Hot-Swap Edge manager in an embodiment of the application; Figure 3 Topological connection diagram of a double-channel Gossip heartbeat module in an embodiment of the application; Figure 4 Workflow diagram of a Snapshot&Replay unit in an embodiment of the application. DETAILED DESCRIPTION
[0022] The application will be described in detail below with specific embodiments. The following embodiments will help those skilled in the art to further understand the application, but do not limit the application in any form. It should be noted that those skilled in the art can make several changes and improvements without departing from the concept of the application. These are within the scope of protection of the application.
[0023] Term explanation: Graph Scheduler: responsible for node arrangement and dependency management of the overall task graph (DAG).
[0024] Hot-Swap Edge: preset primary node and backup node for each dependency edge in the task graph, and can bypass automatically in milliseconds.
[0025] Gossip Heartbeat: point-to-point periodic message passing mechanism between nodes, used for rapid detection of node health status.
[0026] Snapshot&Replay: lightweight persistence of node running state, and reconstruction and replay of unprocessed tasks in case of failure.
[0027] DAG Integrity Validator: checks whether the graph structure remains acyclic and correct after node switching.
[0028] Agent Nodes: independent computing entities that undertake specific business logic or reasoning tasks in the execution graph.
[0029] UDP (User Datagram Protocol): A connectionless, low-overhead network transport protocol suitable for high-frequency heartbeat detection.
[0030] GRPC (Google Remote Procedure Call): A high-performance communication framework based on HTTP / 2, used for heartbeat metadata synchronization.
[0031] XOR-delta: A compression algorithm that only records the difference between adjacent snapshots, significantly reducing log volume and recovery time.
[0032] DAG (Directed Acyclic Graph): A task graph structure without circular dependencies, facilitating parallel processing and state rollback.
[0033] MTTR (Mean Time To Recovery): The average time required for the system to recover from a failure, a key availability indicator.
[0034] Upstream Node: A predecessor node in a directed acyclic graph that provides input data or trigger signals to the current node.
[0035] Scheduler: A control center component responsible for scheduling the execution order of nodes in the task graph, resource allocation, and fault response.
[0036] T_fail (Failure Threshold Time): The upper threshold value of heartbeat delay or packet loss time for determining node failure, usually measured in milliseconds (ms).
[0037] ACK (Acknowledgment): A response packet in communication used to confirm the receipt of a message.
[0038] Cluster: A distributed computing environment for deploying the scheduling system, consisting of multiple computing nodes working together to run the task graph. Figure 3 In the figure, Node 1, Node 2, Node 3, and Node 4 refer to agent nodes, each responsible for a specific business task or logical module, reflecting the master / standby node distribution and switching relationship.
[0039] UDP Heartbeat: A lightweight heartbeat detection message sent based on the UDP protocol.
[0040] GRPC Metadata: A data packet that carries node status and heartbeat supplementary information through the GRPC channel.
[0041] Hot-Swap Edge Manager: A control module that detects node failures and performs switchover operations.
[0042] Replay Queue: Stores messages that were not processed during the switchover, used to resend them in order after a successful switchover.
[0043] Self-Healing: The ability of a system to automatically detect and recover from failures without human intervention.
[0044] Failover: The process of automatically switching to a backup node when the primary node fails.
[0045] Latency: The time taken from detecting a failure to fully restoring service.
[0046] SLA (Service Level Agreement): A quantitative commitment to the availability and quality of service of a system.
[0047] The application discloses a self-healing graph scheduling system and method for a multi-agent system. When any node in the multi-agent task flow graph fails or needs to be upgraded, the system can automatically complete bypassing or hot replacement within a time lower than a preset failure threshold, and ensure that the task topology is continuously loop-free, data is not lost, and service is not interrupted. In the directed acyclic graph of the multi-agent task flow, a master / backup node (Hot-Swap Edge) is configured for each edge, and the health status of the node is monitored in real time through a double-channel Gossip heartbeat. When the master node meets the failure condition, the traffic is automatically redirected to the backup node within milliseconds, and the node state is restored using an XOR-delta snapshot, while the loop-free nature of the topology is verified to ensure task continuity and traceability. Experimental results show that the application reduces the average recovery time to 18 ms, reduces the annual downtime by more than ten times, and is suitable for real-time scenarios such as intelligent finance, industrial internet, and autonomous driving that require parallel agent collaboration and high availability.
[0048] Embodiment 1: Figure 1 The structure diagram of a self-healing graph scheduling system for a multi-agent system in the embodiment of the application.
[0049] As shown in Figure 1 , the embodiment provides a self-healing graph task scheduling system, which comprises: a DAG integrity verifier configured to verify that the task graph topology is still a directed acyclic graph immediately after node bypassing.
[0050] Specifically, when a certain master node in the task graph fails and triggers the backup node to take over, a new task graph structure is generated. The DAG integrity verifier then dynamically checks the structure, uses an improved topological sorting algorithm combined with depth-first search (DFS) to detect loops in the task dependency relationship, and ensures that the task graph still maintains the directed acyclic graph (DAG) property after switching. If the detection finds a loop risk, it automatically rolls back to the state before switching and marks this switching as a failure to prevent the business logic from falling into a dead loop.
[0051] The innovation compared with the prior art is that: The traditional DAG scheduling system uses static compilation period to verify the topological structure, and only performs a one-time check before task arrangement, which cannot cope with dynamic structure changes at runtime. The DAG integrity verifier in the present application supports immediate verification after switching, has the ability of dynamic checking, fast rollback and structure tracing, and effectively solves the problem of potential topological disorder or logic exception after hot switching.
[0052] The hot-swap edge manager is used for maintaining the master node ID and backup node ID for the edges in the task graph, and performing bypass operation when the node fails.
[0053] Figure 2 The working timing diagram of the hot-swap edge manager in the embodiment of the present application.
[0054] As shown in Figure 2 The hot-swap edge manager binds the unique identifiers (ID pairs) of the master node and the backup node for each critical dependency edge in the task graph, and listens to the node state feedback from the double-channel Gossip heartbeat module in real time. When the master node meets the preset failure condition (such as heartbeat delay exceeding threshold or continuous packet loss), the hot-swap edge manager immediately triggers the bypass process to automatically switch the traffic of the original master node to the bound backup node. The switching process does not rely on manual intervention and does not suspend the global scheduling of the task graph, and the traffic redirection operation is completed within milliseconds. At the same time, the hot-swap edge manager also coordinates the call of the Replay unit for state recovery.
[0055] The innovation compared with the prior art is that: The traditional task scheduling framework uses centralized fault management or cold standby strategy. Once the node is down, the recovery process usually needs to be restarted manually or waits for the full graph level rescheduling, resulting in a high average recovery time (MTTR) of several seconds to several minutes. The hot-swap edge manager in the present application can realize asynchronous and non-blocking switching operation based on the edge-level hot standby binding mechanism, and quickly complete the replacement without disturbing other task paths.
[0056] Dual-channel Gossip heartbeat module, used for real-time monitoring of node delay through dual channels of UDP and GRPC, and UDP and GRPC are backup for each other.
[0057] Figure 3 A topological connection diagram of the dual-channel Gossip heartbeat module in the embodiment of the application.
[0058] As shown in Figure 3 , The dual-channel Gossip heartbeat module constructs a point-to-point heartbeat communication link between nodes, the UDP channel is responsible for high-frequency and low-delay heartbeat packet transmission and is used for quickly detecting whether a node is lost; and the GRPC channel is responsible for lower-frequency heartbeat communication carrying metadata information, including node CPU / GPU load, network delay and other indicators. The two channels are redundant, and if any one of the channels is interrupted, the system is automatically degraded to the other channel and the detection frequency is increased. At the same time, the dual-channel Gossip heartbeat module uses a propagation mechanism based on the Gossip protocol, so that each node broadcasts state information to adjacent nodes within a certain period, realizes local calculation of global health perception, and thus avoids a centralized bottleneck.
[0059] Compared with the prior art, the module has the following innovations: The traditional scheduling system often uses a single-channel Ping mechanism (such as ICMP or TCP probe), which is easy to produce false positives or false negatives under network jitter or link congestion. The application significantly improves the accuracy, robustness and response speed of heartbeat detection through the dual-channel fusion mechanism of UDP and GRPC and the distributed Gossip heartbeat propagation strategy, and does not depend on a central node at all, and has good scalability.
[0060] A state snapshot recording unit is configured to record state snapshots of the master node and the backup node in a periodic manner, the state snapshot adopts an XOR-delta algorithm, and only the change data between the last snapshot and the current state is saved, so as to reduce storage and transmission overhead. The snapshot size of each node is not more than 4 kB and is stored in a high-performance log buffer.
[0061] A snapshot replay (Snapshot&Replay) unit is configured to restore the node state based on the state snapshot. In the embodiment, a replay queue is used as a transition cache.
[0062] Figure 4 A workflow diagram of the snapshot replay (Snapshot&Replay) unit in the embodiment of the application.
[0063] As shown in Figure 4As shown, when the master node fails and switches to the backup node, the snapshot replay unit quickly recovers the context state from the latest snapshot and replays the unprocessed input messages to the new node in the order in the replay queue, ensuring that the business processing is seamlessly connected.
[0064] The innovation compared with the prior art is that: The traditional recovery mechanism relies on persistent storage and full read-back, and has high recovery delay and high resource consumption, which is not suitable for high-frequency or edge scenarios. The present application introduces a lightweight and streamable snapshot mechanism, achieving millisecond-level fault recovery and continuous availability.
[0065] The self-healing graph scheduling system displays a graphical monitoring interface module of node state, heartbeat delay and bypass event in real time.
[0066] Embodiment 2: The present embodiment provides a self-healing graph task scheduling method, which is implemented on the self-healing graph task scheduling system in the above embodiment, that is, those skilled in the art can understand the self-healing graph task scheduling method as the running mode of the self-healing graph task scheduling system.
[0067] The method comprises: continuously monitoring the health status of each node by using a dual-channel heartbeat mechanism, triggering a bypass mechanism when detecting that the heartbeat delay of the master node meets a preset failure condition (such as exceeding a threshold); then, redirecting the task flow to the backup node by using the Hot-Swap Edge technology, and performing Snapshot Replay (snapshot replay); after the switching is completed, the DAG integrity verifier checks the topology structure in real time to ensure the correctness of the directed acyclic graph, and if a loop is found, it is automatically rolled back to the state before switching; the event information of the entire process is recorded in the audit log in real time, providing complete and traceable data support for system monitoring.
[0068] Specifically, the self-healing graph task scheduling method comprises the following steps: Step S1, establishing a directed acyclic graph of a task flow composed of a plurality of intelligent agent nodes.
[0069] Step S2, configuring a master node and a backup node for the edges in the directed acyclic graph to form a hot-swappable edge.
[0070] Step S3, recording the state snapshots of the master node and the backup node by the state snapshot recording unit in a periodic manner; the state snapshot adopts an exclusive or incremental compression algorithm, and the single node log size is not more than 4kB.
[0071] The current directed acyclic graph topology and node state are recorded at the same time, which is used for possible rollback.
[0072] Step S4, detecting the health status of the master node by the dual-channel Gossip heartbeat module, if the heartbeat of the master node meets the preset failure condition, determining that the master node fails, isolating the master node and recording the failure event.
[0073] Specifically, the dual-channel Gossip heartbeat module alternately or in parallel sends heartbeat messages by using UDP and GRPC, and the heartbeat sending interval is not greater than 500 milliseconds.
[0074] Further, the preset failure condition includes any one of the following: A, the delay of a single heartbeat exceeds a preset threshold, and the preset threshold is not greater than 20 milliseconds.
[0075] B, no heartbeat response is received for a plurality of times (the number of times is greater than or equal to 3) in succession.
[0076] Step S5, when determining that the master node fails, redirecting the task flow related to the master node to the backup node, and restoring the running state of the backup node based on the state snapshot.
[0077] If the state restoration of the backup node fails, marking the backup node as unavailable and triggering an alarm.
[0078] Step S6, if the task flow redirection and the running state restoration of the backup node are both successful, verifying the loop-freeness of the topology of the directed acyclic graph after switching; if a loop is detected, rolling back to the state before switching and marking the backup node as unavailable.
[0079] In the embodiment, the total fault recovery time delay T recover satisfies: T recover ≤T detect +T redirect +T replay Wherein, T detect is the heartbeat detection time delay, T redirect is the flow redirection time delay, and T replay is the state playback time delay.
[0080] The actual measurement results show that T detect ≈5ms, T redirect ≈7ms, T replay ≈6ms, and the total recovery time delay T recover ≤T detect +T redirect +T replay ≈18ms. By adjusting the preset threshold, the log shard size and other parameters, T recover can be kept within the preset threshold under different hardware and network environments.
[0081] The switching process in this embodiment is not only applicable to the fault recovery scenario (triggering the switching process when detecting node failure), but also applicable to the active upgrade scenario (triggering the switching process when the master node is actively upgraded or policy switching), wherein in the active upgrade scenario, the system realizes the in-sense migration during hot upgrade through the same fault bypass and switching mechanism.
[0082] Step S7, write the fault event of the master node, the switching of the backup node and the topology verification result of the directed acyclic graph into the audit log for monitoring and analysis.
[0083] Embodiment 3: Quantitative trading platform Environment: Kubernetes 20 nodes, NVIDIA A100x10.
[0084] Kubernetes 20 nodes refer to a Kubernetes cluster with 20 worker nodes, which communicate through a high-speed interconnection network (≥10 GbE), and each node runs several containerized agent task modules to form a complete task graph (DAG).
[0085] NVIDIA A100x10 refers to a total of 10 NVIDIA A100 GPU cards configured in the cluster, distributed on different nodes, mainly used to execute delay-sensitive deep learning inference tasks and policy model decision inference tasks.
[0086] In this embodiment, the working process and setting parameters are: 1. Task graph construction: construct an analog quantitative strategy execution DAG, which includes 8 nodes such as price receiving, feature extraction, risk assessment, order execution, etc., and each edge is configured with master / backup nodes.
[0087] 2. Heartbeat mechanism setting: The UDP heartbeat interval is set to 300 ms; τ_fail is set to 20 ms; 3 consecutive heartbeat losses are determined as a fault.
[0088] 3. State snapshot setting: Each node records a snapshot at an interval of 200 ms; The size of a single XOR-delta snapshot is about 3.2 kB; The snapshot is stored in the memory-level Replay buffer.
[0089] 4. Fault simulation method: Simulate the master node instantaneous failure through the Kubernetes kubectl delete pod instruction; The system records the whole process time from heartbeat detection to switching completion.
[0090] 5. Monitoring metrics: Collect the sum of T_detect (the detection latency), T_redirect (the switching latency), and T_replay (the state recovery latency) as T_recover. Evaluate the P95 recovery delay, the node replacement success rate, and the service interruption frequency.
[0091] Experimental results: In 10 rounds of simulation, the average T_recover = 17.8ms (P95); The overall task DAG is not interrupted, and all tasks are automatically bypassed to the backup node; The annualized downtime is less than 30 minutes, which is better than the 99.99% availability SLA requirement.
[0092] Example 4: Industrial IoT edge-cloud collaboration Environment: Jetson AGX edge node + x86 cloud node Wherein: Jetson AGX represents the NVIDIA Jetson AGX Xavier module deployed on the edge device, integrating ARM processors and NVIDIA GPUs, mainly used for performing low-latency data collection, model inference, and event recognition tasks on site. Each Jetson device runs as an agent node in the edge cluster, responsible for primary judgment and rapid response.
[0093] x86 represents a standard x86 architecture server deployed in a cloud data center, equipped with Intel Xeon CPUs and SSD storage, responsible for agent tasks in the task graph that consume more resources and have strong dependence on global state, such as task scheduling, data aggregation, remote updates, etc. In this embodiment, the working process and setting parameters are: 1. Task topology design: The overall task graph includes 6 nodes, of which 4 run on Jetson edge nodes and 2 run on cloud x86 servers; the edge nodes and cloud nodes are bound by Hot-Swap Edge to form a heterogeneous primary-backup path.
[0094] 2. Heartbeat mechanism configuration: The edge node sends a heartbeat every 200ms using a local UDP channel; The cloud node synchronizes the state every 400ms, including node load and cache status; Fault determination: τ_fail = 15ms, or 3 consecutive heartbeat failures.
[0095] 3. Snapshot and sync strategy: Edge node snapshot compressed to ≤ 2.5 kB; Cloud node uses XOR-delta log compression + cloud sync backup; Snapshot frequency: backup every 300 ms, maximum history 3 rounds.
[0096] 4. Experimental operation: Use iptables -A INPUT -s <ip>- j DROP command blocks the edge nodes from the cloud for 30 seconds; Test if the system can automatically cache, switch, and replay tasks.
[0097] 5. Replay logic configuration: Cache task input messages to the Replay Queue; After network recovery, restore node context from the latest snapshot state and replay event queue in original order.
[0098] Experimental results: During the 30-second network interruption: All edge node tasks continue running; Caching and delaying transmission to the cloud messages are 100% successfully replayed; No data loss or control instruction failure occurs; After cloud synchronization is complete, the task graph topology remains acyclic and data continuous, and system business continuity is fully guaranteed.
[0099] Those skilled in the art know that, in addition to implementing the system provided by the present application and each device, module, and unit thereof in a pure computer-readable program code manner, the same functions can be achieved by logically programming the method steps in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers. Therefore, the system provided by the present application and each device, module, and unit thereof can be considered as a hardware component, and the devices, modules, and units included therein for achieving various functions can also be considered as structures within the hardware component. The devices, modules, and units for achieving various functions can also be considered as both software modules for implementing methods and structures within hardware components.
[0100] The specific embodiments of the present application are described above. It should be understood that the present application is not limited to the specific embodiments described above, and various changes or modifications can be made by those skilled in the art within the scope of the claims, without affecting the essential content of the present application. In the case of no conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.< / ip>
Claims
1. A self-healing graph scheduling system for multi-agent systems, characterized by: include: DAG integrity verifier, used to verify that the task graph topology remains a directed acyclic graph immediately after node bypass; A hot-swap edge manager, configured to maintain primary and backup node IDs for edges in the task graph and perform bypass operations in the event of a node failure; A dual-channel Gossip heartbeat module is used to monitor node latency in real time through UDP and GRPC dual channels, with the UDP and GRPC serving as backup for each other. A state snapshot recording unit, configured to periodically record state snapshots of the primary node and the backup node, wherein the state snapshots are compressed using an XOR incremental compression algorithm; The snapshot playback unit is used to restore the node state based on the state snapshot.
2. A self-healing graph scheduling method for a multi-agent system, based on the self-healing graph scheduling system for a multi-agent system according to claim 1, characterized in that: The steps include: Step S1, establishing a directed acyclic graph of a task process consisting of multiple agent nodes; Step S2, configuring a primary node and a backup node for the edge in the directed acyclic graph to form a hot-swappable edge; Step S3, recording the state snapshots of the master node and the backup node in a periodic manner through a state snapshot recording unit; Step S4: Detect the health status of the master node through a dual-channel Gossip heartbeat module. If the heartbeat of the master node meets the preset failure condition, the master node is determined to be failed, the master node is isolated, and the failure event is recorded; Step S5: When it is determined that the master node fails, the task traffic related to the master node is redirected to the backup node, and the operating state of the backup node is restored based on the state snapshot. Step S6: If the task traffic redirection and the operation status recovery of the backup node are both successful, verify the topological acyclicity of the directed acyclic graph after the switch; If no loop is detected, the switch is successful; if a loop is detected, the switch is rolled back to the state before the switch and the backup node is marked as unavailable; Step S7: Write the failure event of the master node, the switching of the backup node, and the topology verification result of the directed acyclic graph into an audit log for monitoring and analysis.
3. The self-healing graph scheduling method for multi-agent systems according to claim 2, characterized in that: In step S4, the dual-channel Gossip heartbeat module uses UDP and GRPC to send heartbeat messages alternately or in parallel.
4. The self-healing graph scheduling method for multi-agent systems according to claim 2, characterized in that: In step S4, the heartbeat sending interval is no more than 500 milliseconds.
5. The self-healing graph scheduling method for multi-agent systems according to claim 2, characterized in that: In step S4, the preset failure condition includes: a single heartbeat delay exceeds a preset threshold, and the preset threshold is no greater than 20 milliseconds.
6. The self-healing graph scheduling method for multi-agent systems according to claim 2, characterized in that: In step S4, the preset failure condition includes: failing to receive a heartbeat response for multiple consecutive times.
7. The self-healing graph scheduling method for multi-agent systems according to claim 2, characterized in that: In step S3, the current directed acyclic graph topology and node status are recorded.
8. The self-healing graph scheduling method for multi-agent systems according to claim 2, characterized in that: In step S3, the state snapshot adopts an XOR incremental compression algorithm, and the log size of a single node does not exceed 4kB.
9. The self-healing graph scheduling method for multi-agent systems according to claim 2, characterized in that: In step S5, if the status recovery of the backup node fails, the backup node is marked as unavailable and an alarm is triggered.
10. The self-healing graph scheduling method for multi-agent systems according to claim 2, characterized in that: Total fault recovery delay T recover satisfy: T recover ≤T detect +T redirect +T replay Among them, T detect is the heartbeat detection delay, T redirect is the traffic redirection delay, T replay The state playback delay.
Citation Information
Patent Citations
Workflow task scheduling method based on directed acyclic graph
CN112346842A
Network quality monitoring method, device, equipment, readable storage medium and program
CN117914752A
Dynamic queue scheduling method and system
CN118331708A
Cloud computing task tracking processing method and system
CN118656200A
Distributed collaborative debugging method and system for industrial robot
CN119292082A
Cited By
Industrial edge intelligent task collaboration method and system based on AI
CN121563145A
An ai-based industrial edge intelligence task collaboration method and system
CN121563145B