A distributed peer-to-peer multi-container cluster system

By designing a distributed peer multi-container cluster system, using distributed consistency protocols and dynamic load balancing strategies, the fault tolerance and self-recovery problems of aerospace cloud computing systems in the event of node failures are solved, resource management and task scheduling between nodes are realized, and system destruction resistance and task processing stability are improved.

CN119835284BActive Publication Date: 2025-05-30NO 63921 UNIT OF PLA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510303521.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-30
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The aerospace cloud computing system lacks sufficient fault tolerance and self-recovery capabilities when node failure or disconnection, resulting in task interruption or resource scheduling failure, and cross-node task scheduling and resource sharing face difficulties, affecting the system's damage resistance and task processing stability.

Method used

A distributed peer-to-peer multi-container cluster system is designed, using distributed consistency protocols and dynamic load balancing strategies to realize distributed uncentralized resource management between nodes and unified cloud-side resource management. The system includes the basic resource layer, operating system layer, cloud management layer, data synchronization module and application scheduling module. The data synchronization module maintains data consistency between multiple nodes, and realizes global scheduling of tasks and resources through the application scheduling module.

Benefits of technology

It realizes distributed uncentralized resource management between nodes, improves the system's destructive resistance and task processing stability, supports efficient sharing and dynamic scheduling of multiple nodes and multiple clusters, alleviates the dependence of aerospace cloud applications on the central cloud, and improves maneuverability, flexibility and agile collaboration capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119835284B_ABST
    Figure CN119835284B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of collaborative computing technology, and specifically to a distributed peer-to-peer multi-container cluster system, including: a basic resource layer, which is composed of basic hardware resources of multiple nodes, and the basic hardware resources include computing servers, storage servers, network devices, security devices, and intelligent acceleration hardware; an operating system layer, which is composed of mobile edges corresponding to multiple nodes, and provides resource and service operation support for container application loads; a cloud management layer, which is composed of multiple container cloud management platforms, and is distributed and operated on each node according to a highly available deployment architecture, realizing unified collaborative scheduling of cluster resources and applications; a data synchronization module, which uses a distributed consistency protocol to ensure the consistency of metadata between each container cloud management platform; an application scheduling module, which receives an application deployment request, and distributes the target application and distribution policy to the corresponding target node or node cluster, completing the scheduling of the target application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of collaborative computing, and particularly to a distributed peer multi-container cluster system. Background Art

[0002] In recent years, cloud computing technology has gradually developed from a fixed-end central cloud system to a collaborative fusion system of central cloud and mobile edge cloud. The cloud-edge collaborative system makes full use of the computing power and flexibility of the edge cloud, alleviates the pressure of the central cloud in processing massive data, and improves the response speed of edge-end applications at the same time. In the aerospace field, the cloud-edge collaborative system has become a key technology to meet high real-time and high computing requirements. The edge cloud plays a crucial role in this system. By providing computing power close to the data source, it mainly alleviates the pressure of the central cloud in processing massive data and improves the task response speed at the same time. Aerospace missions, especially those with extremely high requirements for real-time data processing and low latency, have gradually transformed to this kind of cloud-edge collaborative architecture. In this architecture, edge computing nodes can independently perform data analysis and preliminary processing, and only upload tasks that require higher computing power to the central cloud for processing, avoiding the bottleneck of data transmission and improving the real-time performance of tasks and the stability of the system.

[0003] Due to the particularity of aerospace missions, nodes are often located in scenarios with unstable or restricted network environments. In the case of node failures or disconnections, the traditional centralized architecture lacks sufficient fault tolerance mechanisms and self-recovery capabilities, resulting in possible task interruptions or resource scheduling failures in the system. Existing nodes mostly rely on the central cloud for scheduling, which makes it difficult to ensure the overall stability and task continuity of the system when nodes fail, especially under high real-time and high reliability requirements. In addition, under the multi-cluster architecture of aerospace cloud computing, cross-node task scheduling and resource sharing also face difficulties. Nodes cannot fully independently process tasks, leading to excessive dependence of the system on the network and the cloud, and thus affecting the anti-destruction ability of the system. When node damage or system failures occur, the efficiency of task migration and resource rescheduling is low, and the system is difficult to quickly recover after a failure, affecting the overall operation efficiency and the stability of task processing. Therefore, improving the autonomy of nodes, strengthening the flexibility and robustness of task distribution and scheduling are urgent problems to be solved in current aerospace cloud computing, especially in ensuring that the system can continue to operate efficiently under any single-node failure or damage. Summary of the Invention

[0004] The present invention provides a distributed peer multi-container cluster system to solve the problems raised in the above background art.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A distributed peer multi-container cluster system, comprising:

[0007] The basic resource layer is composed of the basic hardware resources of multiple nodes. The basic hardware resources include computing servers, storage servers, network devices, security devices, and intelligent acceleration hardware;

[0008] The operating system layer is composed of mobile edges corresponding to multiple nodes, providing resource and service operation support for container application loads;

[0009] The cloud management layer is composed of multiple container cloud management platforms, which are distributed and run on each node according to a highly available deployment architecture, realizing unified collaborative scheduling of cluster resources and applications;

[0010] The data synchronization module uses a distributed consistency protocol to ensure the consistency of metadata among container cloud management platforms;

[0011] The application scheduling module receives application deployment requests, distributes the target application and distribution policies to the corresponding target nodes or node clusters, and completes the scheduling of the target application.

[0012] Preferably, the data synchronization process of the data synchronization module includes the following steps:

[0013] S1. Data change trigger: When the metadata of a certain cloud management platform changes, the Leader node records the operation log and broadcasts the change to all Follower nodes;

[0014] S2. Data replication: The Follower node will verify and write to the local log and confirm to the Leader node. After the Leader node receives confirmations from the majority of nodes, it submits the change and notifies other nodes to complete the update;

[0015] S3. Conflict resolution: For possible write conflicts, version number management or last write wins strategy can be used to resolve conflicts;

[0016] S4. Fault recovery: If a certain node fails to synchronize in time due to a fault, after recovery, it completes the data change by reading the latest change log.

[0017] Preferably, the Leader node is elected by all nodes. If the current Leader node fails, a new Leader will be elected through re-election and data synchronization will be restored.

[0018] Preferably, the data synchronization module stores metadata in shards:

[0019] It is divided into multiple shards according to resource types or geographical locations, and each shard is managed by a specific node group to avoid a single node bearing too much data;

[0020] The resource types include computing resources, storage resources, and network resources, and the geographical locations include different regions or data centers.

[0021] Preferably, the data synchronization module uses a dynamic load balancing strategy to allocate metadata, and the dynamic load balancing strategy includes:

[0022] Through dynamic shard adjustment combined with the consistent hashing algorithm, redistribute the shards with high load to other nodes;

[0023] By monitoring the node load situation, dynamically adjust the number of virtual nodes or migrate hot data;

[0024] Through request routing optimization technology, direct data requests to the corresponding shards to reduce the overhead of cross-shard communication.

[0025] Preferably, the consistent hashing algorithm includes:

[0026] Abstract the entire hash space as a ring structure, and evenly map each physical node and its multiple virtual nodes to the ring structure through a hash function;

[0027] After the unique identifier of the metadata is calculated by the hash function, it is assigned to the physical node corresponding to the first virtual node on the ring structure that is greater than or equal to its hash value;

[0028] When a node joins or is removed, only a small number of virtual nodes and their corresponding data shards need to be adjusted to ensure the lowest data migration cost.

[0029] Preferably, the application scheduling module includes multiple federated components, and the multiple federated components include a request component API-Server, an execution component Controller-Manager, and a scheduling component Scheduler;

[0030] The scheduling process of the application scheduling module includes the following steps:

[0031] S1. Resource registration and discovery: Each node or node cluster reports and registers its own resource information to the request component API-Server, and the resource information will be stored in the database, and the database details the corresponding relationship between resources and nodes or clusters;

[0032] S2. Resource binding and scheduling: Use a template to create a resource object and a distribution policy, and the execution component Controller-Manager binds the resource object and the distribution policy to generate a binding relationship. The binding relationship includes resource information and corresponding node or node cluster information, and all binding relationships are stored in the database;

[0033] S3. The Scheduler component calculates the optimal resource distribution plan according to the distribution policy and the available status of the corresponding node or node cluster resources, and issues a scheduling command.

[0034] S4. The corresponding node or node cluster receives the scheduling command, independently executes the resource deployment operation, and simultaneously synchronizes the status back to the Controller-Manager component and stores it in the database.

[0035] Preferably, step S3 specifically includes:

[0036] S31. Receive the distribution policy, add the corresponding binding relationship to the work queue, take out a binding relationship from the queue every second for processing, and loop through the work queue until the work queue is processed.

[0037] S32. Obtain the distribution rule of the distribution policy according to the name of the distribution policy in the annotation in the binding relationship.

[0038] S33. Compare the new and old distribution rules, determine whether scheduling is required, and calculate the number of replica copies allocated to the node or node cluster.

[0039] S34. Use the filtering plugin to determine whether the node or node cluster matches the conditions to obtain the matching node or matching cluster; use the scoring plugin to calculate the score of the matching node or matching cluster to obtain the priority of the matching node or matching cluster.

[0040] S35. Obtain the final member cluster to be issued according to the topology distribution constraint rules.

[0041] S36. Adopt the replica scheduling strategy of replication and / or splitting to allocate replicas to the member cluster to be issued.

[0042] S37. Update the cluster field in the binding relationship and set it to the result of the calculated number of replica copies allocated to the member cluster to be issued.

[0043] Preferably, the filtering plugin filters out nodes or node clusters that do not meet the distribution policy through the node filtering interface.

[0044] The scoring plugin calculates the score for each filtered node or node cluster through the node scoring interface, and selects the node or node cluster with the highest score as the scheduling result.

[0045] Preferably, the replica scheduling strategy of replication includes deploying the same number of resource replicas as the template for each member cluster to be issued.

[0046] The split copy scheduling policy includes weight priority and aggregation priority. The weight priority performs splitting according to the weight ratio, and the aggregation priority determines whether to expand or contract copies based on the total number of copies, and then allocates the number of copies according to the maximum available number of copies of the member clusters.

[0047] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0048] The present invention provides a distributed peer-to-peer multi-container cluster system, a new type of mobile edge cloud technology based on a peer-to-peer distribution architecture, which realizes distributed decentralized resource management between nodes, unified resource management between the cloud and edge clusters, forms an agile resource pool with flexible aggregation and dispersion of nodes and efficient cloud-edge collaboration, and supports efficient sharing and dynamic scheduling of multiple nodes and multiple clusters. In particular, the present invention has a certain anti-damage ability against the instability of the edge scenario, realizes anti-damage reconstruction of task applications between nodes, and ensures the normal operation of edge jobs. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 It is a schematic structural diagram of a distributed peer-to-peer multi-container cluster system provided for an embodiment of the present invention;

[0050] Figure 2 It is a framework diagram of an application scheduling module in a distributed peer-to-peer multi-container cluster system provided for an embodiment of the present invention;

[0051] Figure 3 It is a scheduling flowchart of an application scheduling module in a distributed peer-to-peer multi-container cluster system provided for an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0053] Figure 1 It is a schematic structural diagram of a distributed peer-to-peer multi-container cluster system provided for an embodiment of the present invention. As Figure 1 shown, an embodiment of the present invention provides a distributed peer-to-peer multi-container cluster system, including:

[0054] The basic resource layer is composed of the basic hardware resources of multiple nodes. The basic hardware resources include computing servers, storage servers, network devices, security devices, and intelligent acceleration hardware;

[0055] The operating system layer is composed of mobile edges corresponding to multiple nodes, providing resource and service operation support for container application loads;

[0056] The cloud management layer is composed of multiple container cloud management platforms, which are distributed and run on each node according to a highly available deployment architecture, realizing unified collaborative scheduling of cluster resources and applications;

[0057] The data synchronization module uses a distributed consistency protocol to ensure the consistency of metadata among container cloud management platforms;

[0058] The application scheduling module receives application deployment requests, distributes the target application and distribution policies to the corresponding target nodes or node clusters, and completes the scheduling of the target application.

[0059] The present invention realizes the unified scheduling and management capabilities of multi-node and multi-cluster resources by constructing a "peer-to-peer distribution" collaborative architecture based on cloud native, allowing nodes to reuse cloud management capabilities, directly control edge services, improving business agility and resource utilization. When a certain node goes offline, users can roam and log in to the cloud management on other nodes to continue issuing tasks. At the same time, there is no hierarchical relationship between nodes, and any node can be elected as the leader to collaboratively manage all nodes.

[0060] In an embodiment of the present invention, the data synchronization process of the data synchronization module includes the following steps:

[0061] S1. Data change trigger: When the metadata of a certain cloud management platform changes, the Leader node records the operation log and broadcasts the change to all Follower nodes;

[0062] S2. Data replication: The Follower node verifies and writes to the local log and confirms to the Leader node. After the Leader node receives confirmations from the majority of nodes, it submits the change and notifies other nodes to complete the update;

[0063] S3. Conflict resolution: For possible write conflicts, version number management or last write wins strategy can be used to resolve conflicts;

[0064] S4. Fault recovery: If a certain node fails to synchronize in time due to a fault, after recovery, it completes the data change by reading the latest change log.

[0065] In an embodiment of the present invention, the metadata includes:

[0066] Resource metadata: including computing resources, storage resources, network configurations, etc.

[0067] User information: including user's basic information, role, account status, etc.

[0068] User permissions: Include the access and operation permissions of users to resources.

[0069] Furthermore, the Leader node is elected by all nodes. If the current Leader node fails, a new Leader will be elected through re-election and data synchronization will be restored.

[0070] To improve synchronization efficiency and fault tolerance, each node stores multiple copies of data to avoid data loss caused by a single point of failure. If the Leader node fails, the cluster will elect a new Leader through election and restore data synchronization. Meanwhile, the system will detect the survival status of nodes through a health check mechanism. After a failed node recovers, it quickly restores consistency by reading the latest change log or synchronizing the latest copy, ensuring the high reliability and high availability of the distributed system.

[0071] The data synchronization module is a key technical means to maintain data consistency and high availability among multiple nodes in a distributed system through a data synchronization mechanism. To ensure the synchronization of metadata such as user information, resource configuration, and permissions between different nodes, a combined mechanism of distributed consistency protocol and event-driven is required. First, the system achieves strongly consistent data writing through a distributed protocol (such as Raft or Paxos). Each write operation is responsible by the elected Leader node, and the data changes are synchronized to other Follower nodes in the form of logs. The Leader node needs to obtain the confirmation of the majority of nodes to commit the changes, which ensures that all nodes have consistent data copies. When metadata changes, the event queue (such as Kafka) distributes the change events to each node, and each node asynchronously updates the data according to the change log.

[0072] In an implementation manner of the present invention, the data synchronization module stores metadata in a sharded manner:

[0073] It is divided into multiple shards according to resource types or geographical locations, and each shard is managed by a specific node group, avoiding a single node from bearing too much data;

[0074] Resource types include computing resources, storage resources, and network resources, and geographical locations include different regions or data centers.

[0075] Furthermore, the data synchronization module uses a dynamic load balancing strategy to allocate metadata. The dynamic load balancing strategy includes:

[0076] Through dynamic shard adjustment combined with the consistent hashing algorithm, redistribute the shards with high load to other nodes;

[0077] By monitoring the node load situation, dynamically adjust the number of virtual nodes or migrate hot data;

[0078] Through request routing optimization technology, data requests are directed to corresponding shards, reducing the overhead of cross-shard communication.

[0079] Specifically, the consistent hashing algorithm includes:

[0080] The entire hash space is abstracted into a circular structure, and each physical node and its multiple virtual nodes are evenly mapped onto the circular structure through a hash function;

[0081] After the unique identifier of the metadata is calculated by the hash function, it is assigned to the physical node corresponding to the first virtual node on the circular structure that is greater than or equal to its hash value;

[0082] When a node joins or is removed, only a small number of virtual nodes and their corresponding data shards need to be adjusted to ensure the lowest data migration cost.

[0083] Through sharded storage and dynamic load balancing strategies, the present invention not only improves the scalability and data synchronization efficiency of the system, but also effectively avoids the performance bottleneck and data inconsistency problems of a single node, ensures the synchronization performance in a large-scale environment, and also enhances the fault tolerance and response speed of the system, thus adapting to the requirements of complex distributed scenarios.

[0084] Figure 2 It is a framework diagram of an application scheduling module in a distributed peer-to-peer multi-container cluster system provided for an embodiment of the present invention. As Figure 2 shown, in an embodiment of the present invention, the application scheduling module includes multiple federated components, and the multiple federated components include a request component API-Server, an execution component Controller-Manager, and a scheduling component Scheduler; among them, the request component API-Server is responsible for receiving and managing API requests of multi-cluster resource objects; the execution component Controller-Manager is responsible for resource binding, execution of distribution policies, and scheduling decisions; the scheduling component Scheduler is responsible for calculating the best distribution plan of resources according to the distribution policy and cluster status.

[0085] Figure 3 It is a scheduling flowchart of an application scheduling module in a distributed peer-to-peer multi-container cluster system provided for an embodiment of the present invention. As Figure 3 shown, in an embodiment of the present invention, the scheduling process of the application scheduling module includes the following steps:

[0086] S1. Resource registration and discovery: Each node or node cluster reports and registers its own resource information (such as CPU, memory, storage, etc.) to the request component API-Server, and the resource information will be stored in a database, and the database details the corresponding relationship between resources and nodes or clusters;

[0087] S2. Resource Binding and Scheduling: Create resource objects (such as Deployment) and distribution policies (Propagation Policy) using templates. The execution component Controller-Manager binds the resource objects with the distribution policies to generate binding relationships (Resource Binding). The binding relationships include resource information and corresponding node or node cluster information, and all binding relationships are stored in the database.

[0088] S3. The Scheduler component calculates the optimal distribution plan for resources based on the distribution policy and the available status of resources in the corresponding node or node cluster, and issues a scheduling command.

[0089] S4. The corresponding node or node cluster receives the scheduling command, independently executes the resource deployment operation, and simultaneously synchronizes the status back to the execution component Controller-Manager and stores it in the database.

[0090] So far, the distribution and deployment of multi-node or cross-edge applications are completed, and users can access and query applications through the unified control interface.

[0091] In an embodiment of the present invention, step S3 specifically includes:

[0092] S31. Receive the distribution policy, add the corresponding binding relationship to the work queue, take out a binding relationship from the queue every second for processing, and loop through the work queue until the work queue is processed.

[0093] S32. Obtain the distribution rules of the distribution policy according to the name of the distribution policy in the annotation in the binding relationship.

[0094] S33. Compare the new and old distribution rules, determine whether scheduling is required, and calculate the number of replica copies to be allocated to the node or node cluster.

[0095] S34. Determine whether the node or node cluster matches the conditions through the filtering plugin to obtain the matching nodes or matching clusters; calculate the scores of the matching nodes or matching clusters through the scoring plugin to obtain the priorities of the matching nodes or matching clusters.

[0096] S35. Obtain the final member cluster to be issued according to the topology distribution constraint rules.

[0097] S36. Adopt a replica scheduling strategy of replication and / or splitting to allocate replicas to the member cluster to be issued.

[0098] S37. Update the cluster field in the binding relationship and set it to the result of the calculated number of replica copies allocated to the member cluster to be issued.

[0099] Furthermore, the filtering plugin filters out nodes or node clusters that do not meet the distribution policy through the node filtering interface;

[0100] The scoring plugin calculates scores for each filtered node or node cluster through the node scoring interface, and selects nodes or node clusters as the scheduling result according to the score ranking.

[0101] Furthermore, the replicated copy scheduling policy includes deploying the same number of resource copies as the template in each dispatched member cluster;

[0102] The sliced copy scheduling policy includes weight priority and aggregation priority. The weight priority is sliced according to the weight ratio. The aggregation priority determines whether to expand or shrink copies based on the total number of copies, and then allocates the number of copies according to the maximum available number of copies in the member cluster.

[0103] The present invention designs a distributed peer multi-container cluster system, incorporates multi-cloud management, multi-nodes, and edge devices into a unified resource pool, and breaks the limitation of single-cloud centralization. Nodes in this distributed peer multi-container cluster system have the ability to independently process tasks and do not need to rely entirely on the central cloud; at the same time, there is no superior-subordinate relationship between nodes, and any node supports being elected as the leader. Through the data synchronization module and the data synchronization mechanism, data consistency and high availability among multiple nodes in the distributed system are maintained. Through the application scheduling module, global scheduling of tasks and resources is achieved, and it also has a certain degree of anti-destruction ability. The design and implementation of this distributed peer multi-container cluster system can alleviate the computing power pressure and data transmission pressure brought by the current highly dependent on the central cloud for aerospace cloud application data processing, and improve the mobility, flexibility, and agile collaboration ability of aerospace cloud applications.

[0104] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A distributed peer-to-peer multi-container cluster system, characterized in that: include: The basic resource layer is composed of the basic hardware resources of multiple nodes, including computing servers, storage servers, network equipment, security equipment, and intelligent acceleration hardware; The operating system layer consists of mobile edges corresponding to multiple nodes, providing resource and service operation support for container application loads; The cloud management layer consists of multiple container cloud management platforms, which are distributed and run on each node according to the high-availability deployment architecture to achieve unified and coordinated scheduling of cluster resources and applications. The data synchronization module uses a distributed consistency protocol to ensure the consistency of metadata between various container cloud management platforms; The application scheduling module receives application deployment requests, sends the target application and distribution strategy to the corresponding target node or node cluster, and completes the scheduling of the target application.

2. The distributed peer-to-peer multi-container cluster system according to claim 1, characterized in that: The data synchronization process of the data synchronization module includes the following steps: S1. Data change trigger: When the metadata of a cloud management platform changes, the leader node records the operation log and broadcasts the change to all follower nodes; S2, Data replication: The Follower node verifies and writes the local log and confirms with the Leader node. After receiving confirmation from the majority of nodes, the Leader node commits the change and notifies other nodes to complete the update; S3, conflict resolution: For possible write conflicts, version number management or last write priority strategy can be used to resolve the conflicts; S4, Fault recovery: If a node fails to synchronize in time due to a fault, the data change will be completed by reading the latest change log after recovery.

3. The distributed peer-to-peer multi-container cluster system according to claim 2, characterized in that: The Leader node is elected by all nodes. If the current Leader node fails, a new Leader will be selected through re-election and data synchronization will be restored.

4. The distributed peer-to-peer multi-container cluster system according to claim 3, characterized in that: The data synchronization module stores metadata in shards: Divide into multiple shards by resource type or geographic location, and each shard is managed by a specific node group to prevent a single node from carrying too much data; Resource types include computing resources, storage resources, and network resources, and geographical locations include different regions or data centers.

5. The distributed peer-to-peer multi-container cluster system according to claim 4, characterized in that: The data synchronization module uses a dynamic load balancing strategy to distribute metadata. The dynamic load balancing strategy includes: Through dynamic sharding adjustment combined with consistent hashing algorithm, shards with higher load are redistributed to other nodes; Dynamically adjust the number of virtual nodes or migrate hotspot data by monitoring node load conditions; Through request routing optimization technology, data requests are directed to the corresponding shards, reducing the overhead of cross-shard communication.

6. The distributed peer-to-peer multi-container cluster system according to claim 5, characterized in that: The consistent hashing algorithm includes: The entire hash space is abstracted into a ring structure, and each physical node and its multiple virtual nodes are evenly mapped to the ring structure through the hash function; The unique identifier of the metadata is calculated by the hash function and assigned to the physical node corresponding to the first virtual node on the ring structure that is greater than or equal to its hash value; When nodes are added or removed, only a small number of virtual nodes and their corresponding data shards need to be adjusted to ensure the lowest data migration cost.

7. The distributed peer-to-peer multi-container cluster system according to claim 1, characterized in that: The application scheduling module includes multiple federated components, including a request component API-Server, an execution component Controller-Manager, and a scheduling component Scheduler; The scheduling process of the application scheduling module includes the following steps: S1. Resource registration and discovery: Each node or node cluster reports and registers its own resource information to the request component API-Server. The resource information will be stored in the database, which records the correspondence between resources and nodes or clusters in detail. S2, resource binding and scheduling: Use templates to create resource objects and distribution strategies. The execution component Controller-Manager binds resource objects with distribution strategies and generates binding relationships. The binding relationships include resource information and corresponding node or node cluster information. All binding relationships are stored in the database. S3, the scheduling component Scheduler calculates the best resource distribution plan and issues scheduling commands based on the distribution strategy and the available status of the corresponding node or node cluster resources; S4. When the corresponding node or node cluster receives the scheduling command, it will independently execute the resource deployment operation, synchronize the status back to the execution component Controller-Manager, and store it in the database.

8. The distributed peer-to-peer multi-container cluster system according to claim 7, characterized in that: Step S3 specifically includes: S31, receiving the distribution strategy, adding the corresponding binding relationship to the work queue, taking out one binding relationship from the queue for processing every second, and processing the work queue in a loop until the work queue processing is completed; S32. Obtain the distribution rule of the distribution strategy according to the name of the distribution strategy in the annotation in the binding relationship; S33, comparing the new and old distribution rules, determining whether scheduling is required and calculating the number of replicas allocated to the node or node cluster; S34. Determine whether the node or node cluster matches the condition through the filtering plug-in to obtain the matching node or matching cluster; calculate the score of the matching node or matching cluster through the scoring plug-in to obtain the priority of the matching node or matching cluster; S35. According to the topology distribution constraint rule, the final member cluster to be sent is obtained; S36, using a replica scheduling strategy of replication and / or segmentation to distribute replicas to member clusters; S37. Update the cluster field in the binding relationship and set it to the calculated number of copies allocated to the member cluster.

9. The distributed peer-to-peer multi-container cluster system according to claim 8, characterized in that: The filtering plug-in filters out nodes or node clusters that do not meet the distribution strategy through the node filtering interface; The scoring plug-in calculates a score for each filtered node or node cluster through a node scoring interface, and selects a node or node cluster as a scheduling result according to the score.

10. The distributed peer-to-peer multi-container cluster system according to claim 8, characterized in that: The replica scheduling strategy of the replication includes deploying the same number of resource replicas as the template in each distributed member cluster; The replica scheduling strategies for the sharding include weight priority and aggregation priority. The weight priority performs sharding according to the weight ratio, and the aggregation priority determines whether to expand or shrink replicas according to the total number of replicas, and then allocates the number of replicas according to the maximum number of available replicas of the member cluster.

Citation Information

Patent Citations

  • Method for destroy-resistant replacement among multiple cloud platforms

    CN114629782A

  • Cloud edge computing power resource adaptive computing system

    CN116389491A