Platform computing node multi-site cross-regional sharing system

Through a multi-site cross-region sharing system for platform computing nodes that are coordinated across computer rooms and interoperable across networks, the risk of single point failure of the cloud platform is solved, high availability and data consistency are achieved, and the stability and continuity of enterprise business are ensured.

CN120343031APending Publication Date: 2025-07-18GUANGXI POWER GRID CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510339027.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing cloud platform system has a single point of failure risk under the deployment of a single computer room, which leads to the unavailability of the entire platform and business system, lacks data nearby reading capabilities and cross-regional container scheduling strategies, affecting the business continuity and stability of the enterprise.

Method used

The Kubernetes Master control node multi-site management module, SDN network architecture module, global service registration center module and data consistency module are adopted, and the Raft algorithm, Gossip protocol and Kubernetes Federation control plane can achieve cross-computer room coordination, network interoperability, service discovery and data consistency, and combine monitoring and failure recovery modules to ensure high availability and security of the system.

Benefits of technology

It improves the stability, availability and performance of the system, realizes cross-regional data consistency and efficient scheduling of containers, and ensures the sustainable development of enterprise business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343031A_ABST
    Figure CN120343031A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of cross-regional sharing systems, and discloses a platform computing node multi-site cross-regional sharing system, which is characterized by comprising a Kubernetes Master control node multi-site management module, a platform computing node multi-site sharing module, a platform computing node multi-site sharing module, a platform computing node multi-site sharing module, a platform computing node multi-site sharing module, a platform computing node multi-site sharing module, a platform computing node multi-site sharing module, a platform computing node multi-site sharing module and a platform computing node multi-site sharing module, the Kubernetes Master multi-site deployment module is used for realizing the deployment and the operation of a Kubernetes service cluster among a plurality of sites; the SDN network architecture module is used for establishing a software defined network across machine rooms and realizing network intercommunication among a plurality of Kubernetes service clusters; the global service registration center module is used for providing cross-machine-room service discovery and registration functions and ensuring that cloud applications can be connected to available services; and the data consistency module is used for ensuring the final consistency of the multi-site data and preventing the problem of data inconsistency. According to the invention, the stability, availability and performance of the system are improved, and reliable technical support and guarantee are provided for sustainable development of enterprise businesses.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cross-region sharing systems, and particularly to a multi-site cross-region sharing system for platform computing nodes. Background Art

[0002] With the continuous deepening of the construction work of the sub-nodes of the power grid cloud platform, the existing cloud platform has carried a large number of business systems, and the number will continue to increase. Moreover, the data has great value. However, at present, whether it is storage, IaaS, PaaS, middleware, database, or business application systems in the cloud platform are all deployed and run in a single computer room. Once a problem occurs in the single computer room, the entire platform and business application systems will be unavailable, which will seriously affect the normal development of the company's business. Considering the importance of the business systems on the cloud, it is imperative to build a cross-data center multi-active system with continuous and stable guarantee capabilities.

[0003] The existing business systems have the risk of single point of failure in single computer room deployment. Once a problem occurs, the entire platform and business systems will be unavailable, and it is necessary to build a cross-data center multi-active system with continuous and stable guarantee capabilities. The current cloud platform lacks the ability to provide nearby data reading, and it is necessary to study and design the main copy reading priority strategy. The fault domain level setting of the storage cluster is unreasonable. Under the weak network conditions of cross-computer room and cross-region, there are challenges in the distributed deployment and operation of the PaaS platform, and it is necessary to study how to achieve cross-region container automatic computer room level affinity scheduling and orchestration.

[0004] To solve the above problems, a multi-site cross-region sharing system for platform computing nodes is proposed in this application. Summary of the Invention

[0005] (1) Object of the Invention

[0006] To solve the technical problems existing in the background art, the present invention proposes a multi-site cross-region sharing system for platform computing nodes. The present invention improves the stability, availability and performance of the system, and provides reliable technical support and guarantee for the continuous development of enterprise business.

[0007] (2) Technical Solution

[0008] To solve the above problems, the present invention provides a multi-site cross-region sharing system for platform computing nodes, including:

[0009] A multi-site management module for Kubernetes Master control nodes,

[0010] Function: Responsible for the coordination and management between Kubernetes Master clusters across computer rooms.

[0011] Connection: Establish a connection with the Kubernetes Master nodes in each computer room, and implement global scheduling and resource management through a specific algorithm.

[0012] Method steps: Use the Raft algorithm to ensure the consistency of the cluster, and ensure the state synchronization between multiple sites through leader election and distributed log replication.

[0013] Kubernetes Master Multi-site Deployment Module:

[0014] Function: Implement the deployment and operation of a Kubernetes business cluster between multiple sites;

[0015] Connection: Communicate with the Kubernetes Master control nodes in each site to ensure the scheduling and load balancing of business containers among multiple sites.

[0016] Method steps: Use the Kubernetes Federation control plane to extend to multiple sites, and utilize the federated controller to manage the resources of multiple clusters.

[0017] SDN Network Architecture Module:

[0018] Function: Establish a software-defined network across computer rooms to achieve network interconnection between multiple Kubernetes business clusters.

[0019] Connection: Manage the networks of multiple sites through the SDN controller to ensure communication between containers.

[0020] Method steps: Use the Overlay network technology to create a virtual network between different data centers, enabling containers to communicate transparently across physical boundaries.

[0021] Global Service Registry Module:

[0022] Function: Provide service discovery and registration functions across computer rooms to ensure that cloud applications can connect to available services.

[0023] Connection: Communicate with the service instances in each computer room to maintain the global service list.

[0024] Method steps: Use the Gossip protocol to achieve eventual consistency, and ensure the synchronization of service registration information between different data centers by propagating messages.

[0025] Data Consistency Module:

[0026] Function: Ensure the eventual consistency of multi-site data and prevent data inconsistency problems.

[0027] Connection: Interact with the data storage systems of each site to ensure global synchronization of data updates.

[0028] Method steps: Implement distributed consistency based on the Raft algorithm to ensure the same order of data operations among data centers.

[0029] Security authentication and access control module:

[0030] Function: Ensure the security of the system and restrict access to sensitive information and resources.

[0031] Connection: Interact with the authentication and authorization services to ensure that only authorized users can access system resources.

[0032] Method steps: Use a token-based authentication mechanism and combine it with RBAC (Role-Based Access Control) for permission management.

[0033] Monitoring and fault recovery module:

[0034] Function: Monitor the running status of the system, detect and recover faults in a timely manner.

[0035] Connection: Integrate with the monitoring systems of each site to collect running data and logs.

[0036] Method steps: Use Prometheus for metric collection and alerting, and combine it with an automated fault recovery strategy to achieve the reliability and stability of the system.

[0037] Performance optimization and load balancing module:

[0038] Function: Optimize the system performance and achieve reasonable resource allocation and load balancing.

[0039] Connection: Integrate with the load balancer and automated scheduling system to achieve dynamic resource allocation.

[0040] Method steps: Use an automated scheduling algorithm to dynamically adjust the deployment location of containers according to the real-time load situation to achieve load balancing and maximize resource utilization.

[0041] Preferably, the Kubernetes Master control node multi-site management module includes:

[0042] The Raft algorithm unit is used to ensure the consistency among nodes in the cluster and achieve consistent replication of the distributed state machine.

[0043] The cross-site state synchronization unit is responsible for cross-site state synchronization and fault recovery strategies to ensure the consistency and high availability of the global state.

[0044] Preferably, the Kubernetes Master multi-site deployment module includes:

[0045] The Kubernetes Federation control plane, which is used to expand to multiple sites to achieve centralized management and scheduling of resources;

[0046] The deployment toolset unit, including an automated deployment tool and a configuration management tool, which is used to quickly deploy and configure Kubernetes Master nodes.

[0047] Preferably, the SDN network architecture module includes:

[0048] The SDN controller unit, such as OpenSwitch and Netronon, which is used to manage and control the software-defined network across sites.

[0049] The Overlay network configuration unit, which realizes virtual network connections between multiple data centers, enabling containers to communicate transparently across physical boundaries.

[0050] Preferably, the global service registry module includes:

[0051] The Gossip protocol unit, which is used to achieve eventual consistency of the global service registry and ensure the synchronization of service registration information between different data centers;

[0052] The service registration and discovery unit, which is responsible for maintaining the global service list and providing the functions of service registration and discovery.

[0053] Preferably, the data consistency module includes:

[0054] The Raft algorithm implementation unit, which ensures data operation consistency between each data center and avoids data inconsistency problems.

[0055] The data synchronization and replication unit, which is responsible for cross-site data synchronization and replication to ensure eventual data consistency.

[0056] Preferably, the security authentication and access control module includes:

[0057] The authentication unit, which realizes a token-based authentication mechanism to ensure secure access to system resources.

[0058] The authorization management unit, which realizes the RBAC (Role-Based Access Control) permission management policy to restrict access to sensitive information and resources.

[0059] Preferably, the monitoring and fault recovery module includes:

[0060] Monitoring unit, integrated with Prometheus monitoring system, collects operation data and logs, and realizes real-time monitoring and alerting of the system;

[0061] Fault recovery strategy unit, with automated fault recovery strategies, realizes the reliability and stability of the system, and ensures business continuity.

[0062] Preferably, the multi-site management module of the Kubernetes Master control node is respectively connected to the monitoring and fault recovery module, the security authentication and access control module, the global service registry module, the Kubernetes Master multi-site deployment module, the SDN network architecture module, and the data consistency module; the Kubernetes Master multi-site deployment module is respectively connected to the monitoring and fault recovery module, the global service registry module, and the SDN network architecture module.

[0063] The above technical solutions of the present invention have the following beneficial technical effects:

[0064] Enhanced high availability and fault tolerance: By using the Raft algorithm and the cross-site state synchronization unit, the system can achieve highly reliable state synchronization and fault recovery strategies, thereby improving the availability and fault tolerance of the system.

[0065] Maximized resource utilization: Through the Kubernetes Federation control plane and automated scheduling algorithms, the system can achieve centralized management and dynamic allocation of resources, thereby optimizing resource utilization and ensuring load balancing.

[0066] Cross-site service registration and discovery: The global service registry module uses the Gossip protocol to ensure the synchronization of service registration information across data centers, enabling cloud applications to connect to available services, and improving the scalability and flexibility of the system.

[0067] Guaranteed data consistency: The data consistency module uses the Raft algorithm to achieve distributed consistency, ensuring the eventual consistency of multi-site data, and effectively preventing data inconsistency problems.

[0068] Improved security: The security authentication and access control module uses a token-based authentication mechanism and RBAC permission management policies to ensure secure access to system resources, thereby enhancing the overall security of the system.

[0069] Real-time monitoring and fault recovery: The monitoring and fault recovery module integrates the Prometheus monitoring system and implements automated fault recovery strategies, which can timely detect and recover system faults, ensuring the stability and reliability of the system. Description of the Drawings

[0070] Figure 1 This is a schematic diagram of the structure of a multi-site cross-region sharing system for platform computing nodes proposed by the present invention. Specific implementation manners

[0071] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with specific implementation manners and with reference to the accompanying drawings. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In addition, in the following descriptions, the descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0072] As Figure 1 shown, a multi-site cross-region sharing system for platform computing nodes proposed by the present invention includes:

[0073] The multi-site management module of the Kubernetes Master control node is responsible for the coordination and management between Kubernetes Master clusters across computer rooms.

[0074] The multi-site management module of the Kubernetes Master control node establishes connections with the Kubernetes Master nodes in each computer room, and realizes global scheduling and resource management through specific algorithms.

[0075] The Raft algorithm is used to ensure the consistency of the cluster, and the state synchronization between multiple sites is guaranteed through leader election and distributed log replication.

[0076] The multi-site deployment module of the Kubernetes Master is used to realize the deployment and operation of a Kubernetes service cluster between multiple sites;

[0077] The multi-site deployment module of the Kubernetes Master communicates with the Kubernetes Master control nodes of each site to ensure the scheduling and load balancing of business containers in multiple sites.

[0078] The Kubernetes Federation control plane is extended to multiple sites, and the resources of multiple clusters are managed by using a joint controller.

[0079] The multi-site deployment module of the Kubernetes Master is connected to the multi-site management module of the Kubernetes Master control node, and realizes global node management and coordination by establishing connections with the Kubernetes Master nodes deployed in each site;

[0080] SDN network architecture module, used to establish a software-defined network across computer rooms and achieve network interconnection between multiple Kubernetes business clusters.

[0081] The Kubernetes Master multi-site deployment module is connected to the SDN network architecture module to achieve virtual network switching and connection across data centers.

[0082] Manage the networks of multiple sites through the SDN controller to ensure communication between containers.

[0083] Use Overlay network technology to create virtual networks between different data centers, enabling containers to communicate transparently across physical boundaries.

[0084] Global service registry module, used to provide service discovery and registration functions across computer rooms to ensure that cloud applications can connect to available services.

[0085] The global service registry module is connected to the Kubernetes Master control node multi-site management module to receive and update service registration and discovery information from each data center, ensuring the consistency and reliability of the global service registry;

[0086] The global service registry module is connected to the Kubernetes Master multi-site deployment module. By communicating with the Kubernetes Master nodes deployed at each site, it transmits service registration and discovery information to each data center so that containers can connect to available services;

[0087] Communicate with service instances in each computer room to maintain the global service list.

[0088] Use the Gossip protocol to achieve eventual consistency and ensure the synchronization of service registration information between different data centers by spreading messages.

[0089] Data consistency module, used to ensure the eventual consistency of multi-site data and prevent data inconsistency problems.

[0090] Connected to the control node multi-site management module: Through communication with the control node, ensure the consistency of data operation sequences between each data center and avoid data inconsistency problems;

[0091] Interact with the data storage systems of each site to ensure the synchronization of data updates globally.

[0092] Based on the Raft algorithm to achieve distributed consistency and ensure the consistency of data operation sequences between each data center.

[0093] The security authentication and access control module is used to ensure the security of the system and restrict access to sensitive information and resources.

[0094] The security authentication and access control module is connected to the Kubernetes Master control node multi-site management module. Through communication with the control node, it conducts authentication and authorization to restrict access to sensitive information and resources.

[0095] The security authentication and access control module interacts with the authentication and authorization service to ensure that only authorized users can access system resources.

[0096] It uses a token-based authentication mechanism and combines RBAC (Role-Based Access Control) for permission management.

[0097] The monitoring and fault recovery module is used to monitor the running status of the system, detect and recover faults in a timely manner.

[0098] The monitoring and fault recovery module is connected to the Kubernetes Master control node multi-site management module. Through communication with the control node, it obtains the monitoring data and running logs of the system to monitor the system status in real time.

[0099] The monitoring and fault recovery module is connected to the Kubernetes Master multi-site deployment module, communicates with the Kubernetes Master nodes deployed at each site, conducts fault detection and triggers automated fault recovery strategies.

[0100] It is integrated with the monitoring systems at each site to collect running data and logs.

[0101] It uses Prometheus for metric collection and alerts, and combines automated fault recovery strategies to achieve the reliability and stability of the system.

[0102] The performance optimization and load balancing module is used to optimize the system performance, achieve reasonable allocation of resources and load balancing.

[0103] The performance optimization and load balancing module is integrated with the load balancer and automated scheduling system to achieve dynamic allocation of resources.

[0104] It uses an automated scheduling algorithm to dynamically adjust the deployment location of containers according to the real-time load situation to achieve load balancing and maximize resource utilization.

[0105] Specifically, the Kubernetes Master control node multi-site management module includes:

[0106] The Raft algorithm unit is used to ensure the consistency among nodes in the cluster and achieve consistent replication of the distributed state machine.

[0107] Cross-site status synchronization unit, responsible for cross-site status synchronization and fault recovery strategies to ensure the consistency and high availability of the global state.

[0108] Specifically, the Kubernetes Master multi-site deployment module includes:

[0109] Kubernetes Federation control plane, used to extend to multiple sites to achieve centralized management and scheduling of resources;

[0110] Deployment toolset unit, including automated deployment tools and configuration management tools, used to quickly deploy and configure Kubernetes Master nodes.

[0111] In an optional embodiment, the SDN network architecture module includes:

[0112] SDN controller unit, such as OpenSwitch, Netronon, used to manage and control the software-defined network across sites.

[0113] Overlay network configuration unit, which realizes virtual network connections between multiple data centers, enabling containers to communicate transparently across physical boundaries.

[0114] Preferably, the global service registry module includes:

[0115] Gossip protocol unit, used to achieve eventual consistency of the global service registry and ensure the synchronization of service registration information between different data centers;

[0116] Service registration and discovery unit, responsible for maintaining the global service list and providing the functions of service registration and discovery.

[0117] Preferably, the data consistency module includes:

[0118] Raft algorithm implementation unit, ensuring data operation consistency between each data center and avoiding data inconsistency problems.

[0119] Data synchronization and replication unit, responsible for cross-site data synchronization and replication to ensure the eventual consistency of data.

[0120] Preferably, the security authentication and access control module includes:

[0121] Authentication unit, implementing a token-based authentication mechanism to ensure secure access to system resources.

[0122] Authorization management unit, implementing the RBAC (Role-Based Access Control) permission management policy to restrict access to sensitive information and resources.

[0123] Preferably, the monitoring and fault recovery module includes:

[0124] Monitoring unit: Integrating the Prometheus monitoring system to collect operation data and logs, and realizing real-time monitoring and alarming of the system;

[0125] Fault recovery strategy unit, with an automated fault recovery strategy, to realize the reliability and stability of the system and ensure business continuity.

[0126] It should be understood that the above specific embodiments of the present invention are only used for exemplary illustration or explanation of the principle of the present invention, and do not constitute a limitation to the present invention. Therefore, any modifications, equivalent replacements, improvements, etc. made without departing from the spirit and scope of the present invention shall be included within the protection scope of the present invention. In addition, the appended claims of the present invention are intended to cover all changes and modifications that fall within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.

Claims

1. A multi-site cross-region sharing system for platform computing nodes, characterized in that, including: Kubernetes Master control node multi-site management module, used for the coordination and management between Kubernetes Master clusters across data centers; Kubernetes Master multi-site deployment module, used to implement the deployment and operation of a Kubernetes business cluster across multiple sites; SDN network architecture module, used to establish a software-defined network across data centers to achieve network interconnection between multiple Kubernetes business clusters; Global service registry module, used to provide service discovery and registration functions across data centers to ensure that cloud applications can connect to available services; Data consistency module, used to ensure the eventual consistency of multi-site data and prevent data inconsistency problems; Security authentication and access control module, used to ensure the security of the system and restrict access to sensitive information and resources; Monitoring and fault recovery module, used to monitor the system operation status, detect and recover faults in a timely manner; Performance optimization and load balancing module, used to optimize system performance and achieve reasonable resource allocation and load balancing.

2. The multi-site cross-region sharing system for platform computing nodes according to claim 1, wherein The Kubernetes Master control node multi-site management module includes: Raft algorithm unit, used to ensure the consistency between nodes in the cluster and achieve consistent replication of the distributed state machine; Cross-site state synchronization unit, responsible for cross-site state synchronization and fault recovery strategies to ensure the consistency and high availability of the global state.

3. The multi-site cross-region sharing system for platform computing nodes according to claim 2, wherein The Kubernetes Master multi-site deployment module includes: Kubernetes Federation control plane, used to expand to multiple sites to achieve centralized management and scheduling of resources; Deployment toolset unit, including automated deployment tools and configuration management tools, used to quickly deploy and configure Kubernetes Master nodes.

4. A multi-site cross-region sharing system for platform computing nodes according to claim 3, characterized in that, The SDN network architecture module includes: SDN controller unit, used to manage and control the software-defined network across data centers; Overlay network configuration unit, used to implement virtual network connections between multiple data centers to enable containers to communicate transparently across physical boundaries.

5. A multi-site cross-region sharing system for platform computing nodes according to claim 4, characterized in that, The global service registry module includes: Gossip protocol unit, used to achieve the eventual consistency of the global service registry and ensure the synchronization of service registration information between different data centers; Service registration and discovery unit, responsible for maintaining the global service list and providing service registration and discovery functions.

6. The multi-site cross-region sharing system for platform computing nodes according to claim 5, characterized in that, The data consistency module includes: Raft algorithm implementation unit, used to ensure the consistency of data operations between data centers and avoid data inconsistency problems; Data synchronization and replication unit, responsible for cross-site data synchronization and replication to ensure the eventual consistency of data.

7. A multi-site cross-region sharing system for platform computing nodes according to claim 6, characterized in that, The security authentication and access control module includes: Identity authentication unit, used to implement a token-based identity authentication mechanism to ensure secure access to system resources; Authorization management unit, used to implement RBAC permission management policies to restrict access to sensitive information and resources.

8. A multi-site cross-region sharing system for platform computing nodes according to claim 7, characterized in that, The monitoring and fault recovery module includes: Monitoring Unit: Integrate the Prometheus monitoring system to collect operation data and logs, and achieve real-time monitoring and alerts of the system; Fault Recovery Policy Unit, automate the fault recovery policy, achieve the reliability and stability of the system, and ensure business continuity.

9. A multi-site cross-region sharing system for platform computing nodes according to claim 8, characterized in that, The multi-site management module of the Kubernetes Master control node is respectively connected to the monitoring and fault recovery module, the security authentication and access control module, the global service registry module, the Kubernetes Master multi-site deployment module, the SDN network architecture module, and the data consistency module; the Kubernetes Master multi-site deployment module is respectively connected to the monitoring and fault recovery module, the global service registry module, and the SDN network architecture module.