Cloud edge collaborative resource management system supporting edge cluster autonomous governance

By designing a cloud-edge collaborative resource management system that supports edge cluster autonomy governance, the problem of edge cluster management dilemma in the case of cloud node failure is solved, and the stable operation and business continuity of edge clusters when cloud node abnormalities are achieved.

CN120066734APending Publication Date: 2025-05-30GUILIN UNIV OF ELECTRONIC TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510234614.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the traditional cloud-edge collaboration framework fails in the cloud node, it leads to edge cluster management difficulties, unable to receive new tasks normally, difficult to monitor the execution progress of assigned tasks, and stagnant metadata updates, affecting system stability and business continuity.

Method used

Design a cloud-edge collaborative resource management system that supports edge cluster autonomous governance, including edge autonomous decision-making module, load collection module, management unit, web backend module, etc. Through the collaborative work of these modules, autonomous task allocation, emergency metadata synchronization and leader election when the cloud is out of touch, ensuring that edge clusters can still operate stably when cloud nodes are abnormal.

Benefits of technology

It realizes the stable operation of edge clusters when cloud nodes are abnormal, ensures business continuity and system stability, and improves the autonomy and high availability of edge clusters through independent decision-making and collaborative management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066734A_ABST
    Figure CN120066734A_ABST
Patent Text Reader

Abstract

The invention discloses a cloud-side collaborative resource management system supporting edge cluster autonomous governance, and relates to the field of cloud-side collaborative edge computing, and the system comprises a cloud-side collaborative resource management framework which comprises an edge autonomous decision module. Tasks are autonomously allocated based on a local candidate list, and an emergency metadata synchronization mode is started; the cloud edge collaborative resource management framework further comprises a load collection module; a management unit; and a Web back-end module. According to the invention, node load data is collected to evaluate the idle degree by using a setting mode that the load collection module, the node role management module, the metadata management module, the Web back-end module and the task management module are matched, the information of the most idle node is sent to the node role management module to update candidates, the most idle node is selected as a leader, the overlarge load is avoided, and the task management efficiency is improved. Even if the cloud node is abnormal, the edge cluster can still operate stably and efficiently, and a solid support is provided for various services.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cloud-edge collaborative edge computing, and in particular to a cloud-edge collaborative resource management system that supports autonomous governance of edge clusters. Background Art

[0002] In the current field of cloud-edge collaborative edge computing, traditional cloud-edge collaborative frameworks such as K8s and its derivatives KubeEdge and K3s are widely used. As a powerful container orchestration system, K8s has demonstrated excellent cluster management capabilities in cloud computing scenarios. It relies on a centralized architecture model and uniformly schedules and controls the entire cluster through cloud nodes. KubeEdge is a framework optimized for edge computing scenarios, dedicated to achieving efficient collaboration between cloud and edge. However, its core management mechanism is still highly dependent on cloud nodes.

[0003] In this architecture system that relies on cloud nodes, once a cloud node fails, it will cause a series of serious problems. For example, when the server room suffers a power outage, causing the cloud node to shut down, or when the network backbone link is interrupted, the communication between the cloud and the edge cluster is blocked, the edge cluster will fall into a management dilemma. Since the traditional framework deploys a large number of key management functions in the cloud, such as task allocation, metadata management, and node status monitoring, after the cloud node fails, the edge cluster loses the instructions and coordination from the cloud and cannot normally receive new tasks. The execution progress of the assigned tasks is difficult to monitor, and metadata updates stagnate, which leads to the stagnation of the overall business. Users cannot effectively manage the edge cluster, which seriously affects the stability of the system and the continuity of business.

[0004] As edge computing application scenarios become increasingly complex and diverse, such as real-time equipment monitoring in the industrial Internet of Things and roadside unit data processing in intelligent transportation systems, the requirements for the autonomy of edge clusters are increasing. These scenarios are often in environments with unstable network conditions and limited cloud support. The traditional cloud-edge collaboration framework exposes vulnerabilities when cloud nodes fail. There is an urgent need for a new general cloud-edge collaborative resource management framework that can support autonomous governance of edge clusters to ensure that even in the event of abnormalities in cloud nodes, edge clusters can still operate stably and efficiently, and continue to provide solid support for various businesses. Summary of the Invention

[0005] The purpose of the present invention is to provide a cloud-edge collaborative resource management system that supports autonomous governance of edge clusters to solve the problems raised in the above background technology.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a cloud-edge collaborative resource management system that supports autonomous governance of edge clusters, comprising:

[0007] A cloud-edge collaborative resource management framework, which includes an edge autonomous decision-making module that takes over control when the cloud loses connectivity, autonomously assigns tasks based on a local candidate list, and initiates an emergency metadata synchronization mode;

[0008] The cloud-edge collaborative resource management framework also includes:

[0009] A load collection module, which is used to collect load data of each node in the cluster at a preset fixed time period;

[0010] A management unit, which is used to manage node roles, metadata, tasks, applications, and data collaboration;

[0011] A web backend module, which is used to build an interactive channel between users and edge clusters;

[0012] The load collection module, management unit and Web backend module communicate with each other through the gRPC framework, and are transmitted through multiplexing and binary frames.

[0013] Preferably, the cloud-edge collaborative resource management framework further includes:

[0014] A decentralized metadata synchronization module, which is used to establish PP metadata synchronization channels between edge nodes and support incremental data merging and conflict resolution in disconnected environments;

[0015] A dual-mode task scheduling module, which relies on cloud scheduling in normal mode and enables edge autonomous scheduling strategies in emergency mode;

[0016] A fault detection and switching module, which is used to monitor the health of the cloud, including network latency and service responsiveness, and to support triggering edge autonomy mode and synchronizing data backhaul after recovery;

[0017] The decentralized metadata synchronization module, the dual-mode task scheduling module, and the fault detection and switching module are connected through data interaction.

[0018] Preferably, the management unit includes:

[0019] A node role management module, which is used to manage the three key node roles in the cluster: leader, candidate, and follower. It also supports receiving the least busy node information sent by the load collection module and updates the cluster candidate list in real time according to established rules, ensuring that the candidates are always in the optimal state and maintaining the dynamic conversion logic of node roles.

[0020] The metadata management module is used to distribute various metadata generated during system operation to each node in the cluster using the Raft protocol as the underlying support for its data distribution. During the distribution process, two guarantee mechanisms, data verification and retransmission, are adopted.

[0021] Preferably, the leader in the node role management module serves as the core control node of the cluster, comprehensively manages the operation status of the edge cluster, and monitors the health status of each node in real time, including whether the node is online and whether the load exceeds the standard, and coordinates the distribution and execution of tasks among the nodes. The candidate in the node role management module plays a key candidate role in the operation of the cluster. The candidate sends an election request signal to the remaining follower nodes in the cluster, organizes and coordinates them to participate in the election voting, collects the voting information fed back by each follower node, and selects the most suitable new leader node according to the established election rules of majority decision.

[0022] Preferably, the follower node in the node role management module follows the commands issued by the leader node, and during the task execution process, the follower node collects its own task execution status information in real time, including the amount of work completed, the execution progress percentage and whether an error is encountered, and regularly uploads the information to the leader node via the communication link.

[0023] Preferably, the management unit further includes:

[0024] A task management module, which is used to receive task requests from the Web backend module, parse and process data;

[0025] An application management collaboration module, which is responsible for the deployment, management, and lifecycle scheduling of edge applications;

[0026] The data collaborative management module is responsible for data interaction and sharing between edge devices and the cloud.

[0027] Preferably, the task management module encapsulates the task into a container by utilizing the Docker API, deploys and executes it on a designated edge node, and during the process, the task management module continuously monitors the running status of each container, regularly collects performance indicators and log information, and feeds the data back to the Web backend module, so that users can obtain the latest running status of the task in real time.

[0028] Preferably, the application management collaboration module supports flexible deployment and migration of applications between edge nodes and the cloud, ensuring high availability and elastic expansion of applications, and collaboration with the application development and testing environment on the cloud. The edge devices responsible for the data collaboration management module can collect data from terminal devices in real time, perform preliminary processing and analysis, and upload the results and related data to the cloud. The cloud stores, analyzes and mines the value of massive data to form a complete data flow path.

[0029] Preferably, the load collection module evaluates node idleness based on the TOPSIS algorithm and screens the best candidate nodes. The load collection module uses a fixed time period of once every 1 minute, and the collected content includes CPU usage, memory occupancy and network bandwidth utilization.

[0030] Preferably, the Web backend module is developed using the Go-Echo framework to provide a RESTful API interface for cluster management for user interaction. The Web backend module supports cluster configuration, monitoring and control operations.

[0031] The technical effects and advantages of the present invention are as follows:

[0032] The present invention utilizes a setting method that cooperates with a load collection module, a node role management module, a metadata management module, a Web backend module and a task management module. The edge cluster nodes are divided into leaders, candidates and followers. The leader manages the cluster, and users can manage data through its Web backend module. The candidate is used to organize followers to elect a leader when the leader fails. The followers receive the leader's command and upload data. The load collection module regularly collects node load data to evaluate the idleness, and sends the information of the idlest node to the node role management module to update the candidate. The leader sends a heartbeat at regular intervals. If the other nodes do not receive it on time, the leader is determined to be faulty. The candidate initiates an election, collects load data again, selects a new leader and starts its Web backend module to ensure that users can manage the cluster normally. The system metadata is sent to each node by the metadata management module to ensure data consistency after a failure. The leader election mechanism allows the edge cluster to adapt to the dynamic environment, selects the idlest node as the leader to avoid excessive load, and ensures efficient operation. This ensures that even if an abnormality occurs in the cloud node, the edge cluster can still operate stably and efficiently, and continue to provide solid support for various businesses. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 This is a block diagram of the cloud-edge collaborative resource management system of the present invention.

[0034] Figure 2 This is a block diagram of the management unit of the present invention.

[0035] Figure 3 This is a timing diagram of the interaction between the load collection module and the node role management module of the present invention.

[0036] Figure 4 This is the leader election timing diagram of the present invention.

[0037] Figure 5 This is the metadata management sequence diagram of the present invention.

[0038] Figure 6 This is the task management timing diagram of the present invention.

[0039] Figure 7 This is a flowchart of the metadata management module of the present invention.

[0040] In the figure: 1. Cloud-edge collaborative resource management framework; 2. Load collection module; 3. Management unit; 301. Node role management module; 302. Metadata management module; 303. Task management module; 304. Application management collaboration module; 305. Data collaborative management module; 4. Web backend module; 5. Edge autonomous decision-making module; 6. Decentralized metadata synchronization module; 7. Dual-mode task scheduling module; 8. Fault detection and switching module. DETAILED DESCRIPTION

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0042] The present invention provides Figure 1-7 A cloud-edge collaborative resource management system that supports autonomous governance of edge clusters is shown, including a cloud-edge collaborative resource management framework 1. The cloud-edge collaborative resource management framework 1 includes an edge autonomous decision-making module 5. The edge autonomous decision-making module 5 is used to take over control when the cloud loses connection, autonomously allocate tasks based on the local candidate list, and start the emergency metadata synchronization mode.

[0043] Furthermore, the cloud-edge collaborative resource management framework 1 also includes a load collection module 2, a management unit 3 and a Web backend module 4. The load collection module 2 evaluates the node idleness based on the TOPSIS algorithm and screens the optimal candidate nodes. The load collection module 2 adopts a fixed time period of once every 1 minute, and the collected content includes CPU usage, memory occupancy and network bandwidth utilization. The load collection module 2 is used to collect the load data of each node in the cluster at a pre-set fixed time period. The management unit 3 is used to manage node roles, metadata, tasks, applications and data collaboration. The Web backend module 4 is developed using the Go-Echo framework and provides a RESTful API interface for cluster management for user interaction. The Web backend module 4 supports cluster configuration, monitoring and control operations. The Web backend module 4 is used to build an interaction channel between the user and the edge cluster. The load collection module 2, the management unit 3 and the Web backend module 4 communicate through the gRPC framework, and are transmitted through multiplexing and binary frames.

[0044] When the TOPSIS node idleness evaluation algorithm in the load collection module 2 is running, it first sends a data collection request to each node in the cluster, and uses a lightweight collection agent to obtain the real-time load information of the node. After the collection is completed, the system will divide the node idleness evaluation into two steps: pre-evaluation stage and optimization evaluation stage.

[0045] Pre-assessment stage

[0046] During the pre-evaluation phase, the system applies preset rules and conditions to filter out nodes that do not meet the election requirements. For example, nodes with a CPU usage exceeding 80% will be excluded. This step helps narrow the range of candidate nodes, reduces the subsequent computing burden, and ensures that the nodes participating in the election have sufficient idle time.

[0047] Excellent evaluation stage

[0048] During the optimization phase, the TOPSIS node idleness evaluation algorithm conducts an in-depth analysis of pre-selected nodes to identify the most idle nodes. The algorithm scores each node based on a series of customizable weights and metrics. These weights and metrics are adjusted according to the characteristics and requirements of the system. For example, in the case of dense deployment of web applications, network latency may become an important consideration; in big data processing scenarios, more emphasis is placed on storage resources and computing power.

[0049] The algorithm scores each pre-selected node according to the set weights and indicators, and generates a node ranking list based on the idleness. The system selects the most idle node from this ranking list, extracts the necessary identification information, and then sends this information to the node role management module 301 for subsequent role allocation and configuration update. The interaction process between the load collection module 2 and the node role management module 301 is as follows: Figure 4 As shown, the specific algorithm pseudo code running process is as follows:

[0050]

[0051]

[0052] Specifically, the cloud-edge collaborative resource management framework 1 also includes a decentralized metadata synchronization module 6, a dual-mode task scheduling module 7 and a fault detection and switching module 8. The decentralized metadata synchronization module 6 is used to establish a P2P metadata synchronization channel between edge nodes and support incremental data merging and conflict resolution in a disconnected environment. The dual-mode task scheduling module 7 is used to rely on cloud scheduling when in normal mode, and enable edge autonomous scheduling strategy in emergency mode. The fault detection and switching module 8 is used to monitor the health status of the cloud, which includes network delay and service response, as well as support for triggering edge autonomous mode and data backhaul synchronization after recovery. The decentralized metadata synchronization module 6, the dual-mode task scheduling module 7 and the fault detection and switching module 8 are connected through data interaction.

[0053] Furthermore, the management unit 3 includes a node role management module 301 and a metadata management module 302. The node role management module 301 is used to manage three types of key node roles in the cluster, including leaders, candidates and followers, and supports receiving the most idle node information sent by the load collection module 2, and updates the cluster candidate list in real time according to established rules to ensure that the candidates are always in the optimal state and maintain the dynamic conversion logic of the node role. The metadata management module 302 is used to use the Raft protocol as the underlying support for its data distribution, and distributes various metadata generated during the system operation to each node in the cluster. In the distribution process, two guarantee mechanisms, data verification and retransmission, are used to ensure data consistency and fault recovery capabilities, ensuring that each metadata All of the metadata can reach each node completely, so as to ensure the data consistency of the entire cluster after any node suddenly fails, and avoid data confusion or loss. The process of the metadata management module 302 is to divide the metadata into fixed-size data packets before sending, and add header information such as sequence number and check code to each data packet, and then send them in sequence according to the sequence number of the data packet. After receiving the data packet, the receiving node will first check the check code. If the check is correct, it will send a confirmation message to the metadata management module 302; if the check is wrong, it will request to resend the data packet. The metadata management module 302 dynamically adjusts the sending strategy according to the feedback of the receiving node to ensure that all data packets can be successfully delivered. When all data packets are received and verified, the receiving node reorganizes them according to the data packet sequence number and restores the complete metadata, thereby ensuring data consistency.

[0054] Specifically, the leader in the node role management module 301 serves as the core control node of the cluster, comprehensively manages the operation status of the edge cluster, and monitors the health status of each node in real time, including whether the node is online and whether the load exceeds the standard, and coordinates the distribution and execution of tasks among the nodes. Users can access the Web backend module 4 address of the leader node, enter the management interface, view the hardware configuration, current load and task list of the node and perform add, delete, modify and query operations on the cluster data, including adjusting the task allocation strategy and cluster operation parameters. The candidate in the node role management module 301 plays a key candidate role in the operation of the cluster. The candidate sends an election request signal to the remaining follower nodes in the cluster, organizes and coordinates them to participate in the election, collects the voting information fed back by the follower nodes, and selects the most suitable new leader node according to the established election rules of majority rule to ensure that the cluster can restore a stable leadership order in the shortest time. When the leader node fails due to hardware failure or network When a node fails to work due to a problem or software error, the candidate node will detect this anomaly and start the leader election process. It sends an election request to other follower nodes, coordinates the election process, collects voting results, and selects a new leader node based on rules such as majority rule to ensure that the cluster can quickly resume stable operation. The follower nodes in the node role management module 301 follow the commands issued by the leader node, and during the task execution process, the follower nodes collect their own task execution status information in real time, including the amount of work completed, the percentage of execution progress and whether errors are encountered, and regularly upload the information to the leader node via the communication link. After receiving the task assignment, the follower will parse and execute the corresponding task, such as running a computing program or processing data storage. During the execution process, the follower node will collect task status information, such as the amount of work completed, the percentage of progress and error conditions, and regularly report this information to the leader node to help the leader monitor task execution and make decisions.

[0055] The node role management module 301 is responsible for the management of the three roles of leader, candidate and follower in the cluster, ensuring that the cluster can run smoothly under various circumstances. On the one hand, this module cooperates with the load collection module 2 to continuously monitor the information of the most idle node and update the cluster candidate list according to established rules, such as sorting and screening by idleness to ensure that the candidate is always the most idle node. On the other hand, it controls the leader election process and monitors the status of the leader node during normal operation; once a leader failure is detected, the leader election process will be quickly started. At this time, the module broadcasts the election notification to all candidate nodes, uses a reliable broadcast mechanism to ensure that the notification is delivered to each candidate node, and collects their load data. Subsequently, the TOPSIS node idleness evaluation algorithm is run to select the most idle node. Finally, a new leader is designated based on the election results, and the identity of the new leader is broadcast to the entire cluster, prompting each node to update its leader identifier, thereby ensuring the smooth operation of the cluster. The leader election process is as follows Figure 4 shown.

[0056] The node role management module 301 of the leader node sends a heartbeat signal to other nodes in the cluster every 5 seconds to ensure the status synchronization and normal communication of all nodes in the network. Each heartbeat data packet contains three key fields: identification information, current timestamp and status check code. These fields work together to ensure the validity and integrity of the heartbeat signal.

[0057] Specifically, the identification information is a 2-byte integer value used as a marker to identify the source of the heartbeat signal. It contains fixed node ID information so that the receiver can confirm the sender of the signal. The current timestamp is an 8-byte integer field that records the exact time when the heartbeat signal is sent. It uses the UNIX timestamp format in seconds to ensure time synchronization between all nodes. Finally, the status check code is a 4-byte integer field calculated by the CRC32 algorithm to verify the integrity of the heartbeat signal during transmission. For example, for a packet containing the identification information 0x1234 and the current timestamp 1704816960 (corresponding to 2 A heartbeat signal packet received at 16:16:00 on Tuesday, January 9, 2024, is combined into a byte sequence. The status check code calculated by the CRC32 algorithm may be 0xCBF43F2D. This check code is appended to the packet and used by the receiving node to verify that the data remains intact during transmission. The format of the heartbeat packet is shown in Table 1. The remaining nodes each run a heartbeat detection thread to monitor. If a node does not receive a heartbeat signal for 15 consecutive seconds, it is determined that the leader has failed and immediately sends a report to the candidate's node role management module 301, triggering the leader election process to ensure business continuity. The heartbeat packet structure is shown in Table 1:

[0058]

[0059] Table 1

[0060] The metadata management module 302 uses the distributed consensus algorithm Raft protocol as the underlying basis for its data distribution to ensure the consistency and high availability of the system metadata in the cluster. The Raft protocol uses mechanisms such as leader election, log replication, and security to enable metadata to reach a consistent state among all nodes in the cluster. Even in the event of failure of some nodes, the normal operation of the system can be maintained. The metadata management process is as follows: Figure 5 shown.

[0061] During the operation of the system, the metadata management module 302 can instantly perceive newly generated metadata and encapsulate these updates into structured data packets. Each data packet not only contains the actual metadata content, but also adds header information such as sequence number, checksum, and other necessary control information. This additional information ensures the correct ordering and integrity verification of the data packets. Specifically, the data packet contains six key fields: sequence number, checksum, data type, metadata, transmission timestamp, and sender ID. Among them, the sequence number field is an integer type value. It assigns a unique serial number to each newly generated metadata data packet to help the receiving node correctly reassemble the original metadata according to the generation order. The checksum field is a string type and is used to verify the integrity of the data packet during transmission. If the receiving node finds that the checksum fails, it will request the sender to retransmit the data packet to prevent data damage or tampering.

[0062] The data type field uses an enumeration to identify the specific data type carried in the packet, such as metadata updates or configuration changes. This allows the receiving node to take appropriate processing actions based on the data type. The core component is the metadata field, which uses a binary format and contains various metadata information required for system operation. In addition, the send timestamp field records the exact time the packet was sent, accurate to the millisecond level, which helps with tracking and debugging, and also facilitates performance analysis and troubleshooting. Finally, the sender ID field uses a string to uniquely identify the node sending the packet, helping to confirm the data source and, when necessary, can be used to track and confirm the data transmission path. The specific format of the metadata packet is shown in Table 4.

[0063] The encapsulated metadata data packet is then sent to other nodes in the cluster in sequence via a pre-configured TCP network channel. To ensure the successful delivery of the data packet, the metadata management module 302 uses the Go language goroutine to start a dedicated feedback listening thread. This thread receives confirmation information from the receiving node through the channel and handles possible retransmission requests. If the listening goroutine detects that the check code returned by the receiving node is incorrect or no confirmation information is received within the specified time, it will notify the sending logic through the channel to immediately retransmit the corresponding data packet. Throughout the process, the system sets the maximum number of retransmissions and the waiting interval before each retransmission to balance efficiency and reliability. In addition, all operations are completed in a non-blocking manner, ensuring the system's high throughput and fault tolerance. This mechanism ensures that each piece of metadata can reach all target nodes completely and accurately. Even if a temporary failure of an individual node occurs, it will not affect the data consistency of the entire cluster. The metadata data packet structure table is shown in Table 2:

[0064]

[0065] Table 2

[0066] The task management module 303 is a core component of the entire distributed system, responsible for efficiently managing and scheduling tasks within the cluster. Developed in the Go language, this module leverages Go's concurrency and efficient network processing capabilities to ensure real-time task scheduling and high throughput. All tasks run as containers, leveraging Docker containerization technology to provide a consistent operating environment and enhance task portability and isolation.

[0067] The task management module 303 maintains a task scheduling queue, where each task element contains several key attributes: task identifier, required resource description, priority, designated deployment node, and other metadata. In particular, the priority of a task is rated according to its resource requirements, and this rating mechanism refers to the QoS classification standard in Kubernetes. Specifically, the QoS level of a task is determined by its request and limit settings for computing resources (such as CPU, memory, etc.). When a task sets both request and limit for each resource type and request = limit, the task is rated as the Guaranteed level, which means the system promises to provide the amount of resources requested by these tasks and will not be preempted or resource-limited due to the requirements of other tasks. Such tasks usually have the highest priority and are suitable for situations with strict performance requirements or intolerant of resource shortages. For tasks with request < limit set, they are classified as the Burstable level. In this case, the task is guaranteed to obtain at least the amount of resources it requests, but can use additional resources that exceed the request but do not exceed the limit when cluster resources permit. Burstable tasks have medium priority and are suitable for most applications. They can flexibly utilize the remaining resources without affecting critical tasks. Finally, if a task does not set request or limit for any resource type, it is rated as the BestEffort level. These tasks have the lowest priority and can only use the remaining resources in the cluster that are not occupied by other high-priority tasks. BestEffort tasks may be restricted or terminated first when resources are紧张, so they are suitable for applications that are insensitive to resource requirements or can tolerate较大延迟. In this way, the task management module 303 can reasonably allocate computing resources according to the resource requirements of different tasks, ensure the efficient execution of critical tasks, and maximize the utilization rate of cluster resources. When a new task enters, the task management module 303 combines the node idle information synchronized from the load collection module 2 and the node performance configuration parameters stored locally and updated regularly, gives priority to high-priority tasks, and allocates them to the designated nodes for execution; for low-priority tasks, they are flexibly diverted to nodes with lighter loads according to the actual situation to balance the overall resource utilization efficiency. The task QoS rating table is shown in Table 3.

[0068] In addition, to simplify operations and improve efficiency, the task management module integrates a service discovery mechanism, enabling each node to automatically identify and connect to other service instances, thus ensuring the elasticity and reliability of the system. To ensure the security and persistence of task data, the module supports persistent volumes, which can maintain the integrity and consistency of data even in the case of container restart or migration.

[0069] During the task execution process, the real-time monitoring system communicates with the task execution agent on the node every 10 seconds to obtain the latest execution status. Once an abnormal situation such as timeout or failure occurs, the system will automatically trigger the retry mechanism and retry to start the task on the specified node. If the specified node is unavailable or the task fails to start multiple times, the module can choose whether to notify the administrator or take other predefined error handling measures based on the configuration. This mechanism effectively improves the success rate of the task and the stability of the system. The task management process is as follows: Figure 6 As shown, the task QoS rating table is shown in Table 3:

[0070]

[0071] Table 3

[0072] Web backend module 4 is the key interface for users to interact with the cluster. It is developed based on the Go language and the Echo framework. This module provides a set of RESTful API interfaces for managing cluster resources. Users enter the access address through the browser, and the server first performs identity authentication to ensure that only authorized users can access the API.

[0073] After the user enters the access address in the browser, the server will ask the user to provide credentials for identity authentication. After the authentication is passed, the user can access the API endpoint and perform related operations. The specific list of provided Web backend APIs is shown in Table 4:

[0074]

[0075]

[0076]

[0077] Table 4

[0078] More specifically, the management unit 3 also includes a task management module 303, an application management collaboration module 304 and a data collaboration management module 305. The task management module 303 is used to receive task requests from the Web backend module 4, parse and process data, the application management collaboration module 304 is responsible for the deployment, management and life cycle scheduling of edge applications, and the data collaboration management module 305 is responsible for data interaction and sharing between edge devices and the cloud. The task management module 303 encapsulates tasks into containers by using the Docker API, deploys and executes them on designated edge nodes, and during the process, the task management module 303 continuously monitors the tasks. It monitors the running status of each container, regularly collects performance indicators and log information, and feeds the data back to the Web backend module 4, so that users can obtain the latest running status of the task in real time. The application management collaboration module 304 supports the flexible deployment and migration of applications between edge nodes and the cloud, ensuring the high availability and elastic expansion of applications, and collaborating with the application development and testing environment on the cloud. The edge devices responsible for the data collaboration management module 305 can collect data from terminal devices in real time, perform preliminary processing and analysis, and upload the results and related data to the cloud. The cloud stores, analyzes and mines the value of massive data to form a complete data flow path.

[0079] Finally, it should be noted that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art can still modify the technical solutions described in the aforementioned embodiments or make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A cloud-edge collaborative resource management system that supports autonomous governance of edge clusters, comprising: A cloud-edge collaborative resource management framework (1), the cloud-edge collaborative resource management framework (1) comprising an edge autonomous decision-making module (5), the edge autonomous decision-making module (5) being used to take over control when the cloud loses connection, autonomously assign tasks based on a local candidate list, and start an emergency metadata synchronization mode; Characterized in that the cloud-edge collaborative resource management framework (1) also includes: A load collection module (2), the load collection module (2) is used to collect load data of each node in the cluster at a preset fixed time period; A management unit (3), the management unit (3) is used to manage node roles, metadata, tasks, applications and data collaboration; A Web backend module (4), wherein the Web backend module (4) is used to build an interactive channel between the user and the edge cluster; The load collection module (2), the management unit (3) and the Web backend module (4) communicate with each other via a gRPC framework, and are transmitted via multiplexing and binary frames.

2. According to claim 1, a cloud-edge collaborative resource management system supporting autonomous governance of edge clusters is characterized in that: The cloud-edge collaborative resource management framework (1) also includes: A decentralized metadata synchronization module (6), the decentralized metadata synchronization module (6) is used to establish a P2P metadata synchronization channel between edge nodes and support incremental data merging and conflict resolution in an off-network environment; A dual-mode task scheduling module (7), wherein the dual-mode task scheduling module (7) is used to rely on cloud scheduling when in normal mode, and to enable edge autonomous scheduling strategy in emergency mode; A fault detection and switching module (8), the fault detection and switching module (8) is used to monitor the health status of the cloud, including network delay and service response, and support triggering edge autonomy mode and synchronization of data return after recovery; The decentralized metadata synchronization module (6), the dual-mode task scheduling module (7) and the fault detection and switching module (8) are connected through data interaction.

3. According to claim 1, a cloud-edge collaborative resource management system supporting autonomous governance of edge clusters is characterized in that: The management unit (3) comprises: A node role management module (301), the node role management module (301) is used to manage three types of key node roles in the cluster, the three types of key node roles include leader, candidate and follower, and supports receiving the most idle node information sent from the load collection module (2), and updates the cluster candidate list in real time according to established rules, ensuring that the candidate is always in the optimal state and maintaining the dynamic conversion logic of the node role; The metadata management module (302) is used to distribute various metadata generated during the system operation to each node in the cluster through the Raft protocol as the underlying support for its data distribution, and in the distribution process, two guarantee mechanisms, data verification and retransmission, are adopted.

4. A cloud-edge collaborative resource management system supporting autonomous governance of edge clusters according to claim 3, characterized in that: The leader in the node role management module (301) serves as the core control node of the cluster, manages the operation status of the edge cluster in all aspects, and monitors the health status of each node in real time, including whether the node is online and whether the load exceeds the standard, and coordinates the distribution and execution of tasks among the nodes. The candidate in the node role management module (301) plays a key candidate role in the operation of the cluster. The candidate sends an election request signal to the remaining follower nodes in the cluster, organizes and coordinates them to participate in the election voting, collects the voting information fed back by the follower nodes, and selects the most suitable new leader node according to the established election rules of majority decision.

5. According to claim 3, a cloud-edge collaborative resource management system supporting autonomous governance of edge clusters is characterized in that: The follower nodes in the node role management module (301) follow the commands issued by the leader node, and during the task execution process, the follower nodes collect their own task execution status information in real time, including the amount of work completed, the execution progress percentage and whether errors are encountered, and regularly upload the information to the leader node via the communication link.

6. A cloud-edge collaborative resource management system supporting autonomous governance of edge clusters according to claim 3, characterized in that: The management unit (3) further comprises: A task management module (303), the task management module (303) is used to receive task requests from the Web backend module (4), parse and process data; An application management collaboration module (304), the application management collaboration module (304) is responsible for the deployment, management and life cycle scheduling of edge applications; A data collaborative management module (305), wherein the data collaborative management module (305) is responsible for data interaction and sharing between the edge device and the cloud.

7. A cloud-edge collaborative resource management system supporting autonomous governance of edge clusters according to claim 6, characterized in that: The task management module (303) encapsulates the task into a container by using the Docker API, deploys and executes it on a designated edge node, and during the process, the task management module (303) continuously monitors the running status of each container, regularly collects performance indicators and log information, and feeds the data back to the Web backend module (4), so that the user can obtain the latest running status of the task in real time.

8. A cloud-edge collaborative resource management system supporting autonomous governance of edge clusters according to claim 6, characterized in that: The application management collaboration module (304) supports the flexible deployment and migration of applications between edge nodes and the cloud, ensuring the high availability and elastic expansion of applications, and collaborating with the application development and testing environment on the cloud. The edge devices responsible for the data collaboration management module (305) can collect data from terminal devices in real time, perform preliminary processing and analysis, and upload the results and related data to the cloud. The cloud stores, analyzes and mines the value of massive data to form a complete data flow path.

9. A cloud-edge collaborative resource management system supporting autonomous governance of edge clusters according to claim 1, characterized in that: The load collection module (2) evaluates the node idleness based on the TOPSIS algorithm and selects the optimal candidate node. The load collection module (2) uses a fixed time period of once every 1 minute, and the collected content includes CPU usage, memory occupancy and network bandwidth utilization.

10. A cloud-edge collaborative resource management system supporting autonomous governance of edge clusters according to claim 1, characterized in that: The Web backend module (4) is developed using the Go-Echo framework and provides a RESTful API interface for cluster management for user interaction. The Web backend module (4) supports cluster configuration, monitoring and control operations.

Citation Information

Cited By

  • Adaptive system and method for end-side cloud dynamic task scheduling of power distribution Internet of Things

    CN120378461A

  • Cross-border e-commerce information risk analysis method in combination with cloud computing

    CN120875888A

  • Communication link dual-mode switching control system based on cloud edge collaboration and broadcast subscription

    CN121193779A

  • Communication link dual-mode switching control system based on cloud edge collaboration and broadcast subscription

    CN121193779B