Distributed multi-site operation cloud storage system
Through a distributed multi-site cloud storage system, the problem of cloud platform unavailability caused by the deployment of a single computer room is solved, the distributed storage of data and business continuity is achieved, data availability and read performance are improved, load balancing is optimized, and disaster recovery is supported.
Patent Information
- Application Number
- CN202510376478.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-22
AI Technical Summary
Due to the deployment of a single computer room, the existing cloud platform will become unavailable once there is a problem with the computer room, which will affect the normal development of the business and bring economic losses.
Design a distributed multi-site operation cloud storage system, including data management, storage pool management, replica distribution, network scheduling, fault detection, arbitration and user interface modules, to realize distributed storage of data and cross-computer room network transmission, ensuring data availability and business continuity.
Improves data availability and reliability, enhances business continuity, optimizes read performance, achieves load balancing, and supports disaster recovery, reducing the impact of a single point of failure on the business.
Smart Images

Figure CN120358248A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud storage, and in particular to a distributed multi-site running cloud storage system. Background Art
[0002] In today's cloud computing field, with the continuous deepening of the construction work of the sub-nodes of the power grid cloud platform, the existing cloud platform has carried a large number of business systems, and this number will continue to increase. The data contained in these business systems is of great value, but there is an obvious defect in the current cloud platform. Whether it is storage, IaaS, PaaS, middleware, database, or business application systems, they are all deployed and run in a single computer room;
[0003] This deployment structure of a single computer room has a significant risk, that is, once a problem occurs in this computer room, the entire cloud platform and the business application systems it carries will be in an unavailable state, which will seriously affect the normal development of the company's business and may cause serious economic losses and damage to the reputation;
[0004] Therefore, a distributed multi-site running cloud storage system is proposed. Summary of the Invention
[0005] Based on the technical problems existing in the background art, the present invention proposes a distributed multi-site running cloud storage system.
[0006] A distributed multi-site running cloud storage system proposed by the present invention includes:
[0007] A data management module, which is responsible for data storage, retrieval, and replica management to ensure data integrity and availability;
[0008] A storage pool management module, which is used to manage the resources of the storage pool;
[0009] A replica distribution module, which is used to automatically calculate the placement rules of data replicas to achieve distributed storage of data;
[0010] A network scheduling module, which is responsible for data transmission and scheduling in the cross-computer room network;
[0011] A fault detection module, which is used to monitor the running state of the storage system, detect faults and abnormal situations, and issue alarms and handle them in a timely manner;
[0012] An arbitration module, which is responsible for decision-making and preventing brain split in the event of a fault scenario to ensure data consistency and system stability;
[0013] A control plane service module, which is used to provide control plane services for the distributed storage system;
[0014] The user interface module is used to provide an interface for users to interact with the distributed storage system.
[0015] Preferably, the data management module is communicatively connected to the storage pool management module and the replica distribution module respectively to implement data storage and replica management. The storage pool management module is communicatively connected to the data management module and the replica distribution module respectively to implement the management and optimization of the storage pool. The replica distribution module is communicatively connected to the data management module and the storage pool management module respectively to implement the automatic distribution of data replicas. The network scheduling module is communicatively connected to the data management module and the storage pool management module respectively to implement the scheduling and transmission of data in the network. The fault detection module is communicatively connected to the data management module, the storage pool management module, and the replica distribution module respectively to provide fault detection and handling functions. The arbitration module is communicatively connected to the fault detection module, the data management module, and the storage pool management module respectively to implement the decision-making process for fault scenarios. The control plane service module is communicatively connected to the data management module, the storage pool management module, and the replica distribution module respectively to implement the management and services of the control plane. The user interface module is communicatively connected to the data management module, the storage pool management module, and the replica distribution module respectively to implement the interaction and monitoring between users and the system.
[0016] Preferably, the data management module includes a data storage unit, a data retrieval unit, and a replica management unit. The data storage unit is responsible for storing data in the distributed storage system. The data retrieval unit is used to provide a data retrieval function, allowing users to search for and obtain data through the interface. The replica management unit is responsible for managing data replicas to ensure the availability and reliability of the data.
[0017] Preferably, the storage pool management module includes a storage space allocation unit and a data replica distribution optimization unit. The storage space allocation unit is responsible for allocating storage space for data to ensure that there are sufficient storage resources available. The data replica distribution optimization unit optimizes the distribution of data according to the replica distribution algorithm to improve data access performance and availability.
[0018] Preferably, the replica distribution module includes a data block allocation unit and a data replica management unit. The data block allocation unit assigns a unique hash value to each data block according to the decentralized scalable hashing algorithm to determine its storage location. The data replica management unit is responsible for managing data replicas to ensure that data is backed up at multiple locations to improve data availability and reliability.
[0019] Preferably, the network scheduling module includes a data transmission unit and a data scheduling optimization unit. The data transmission unit is responsible for transmitting data in the cross-data center network. The data scheduling optimization unit optimizes the scheduling of data according to the network conditions and data access requirements to improve data access performance.
[0020] Preferably, the fault detection module includes a status monitoring unit and a fault warning unit. The status monitoring unit monitors the running status of the storage system in real time, including the health status of each node and network connections. The fault warning unit is used to issue an alarm in a timely manner when a fault or abnormal situation is detected, so as to handle it in a timely manner.
[0021] Preferably, the arbitration module includes a decision-making judgment unit and a data consistency maintenance unit. The decision-making judgment unit is responsible for making decision judgments when a fault or conflict occurs. The data consistency maintenance unit ensures the data consistency between multiple data centers, guaranteeing the security and reliability of the data.
[0022] Preferably, the control plane service module includes a configuration management unit and a permission control unit. The configuration management unit is responsible for managing the configuration information of the system, including node configuration and storage configuration. The permission control unit is responsible for managing user permissions to ensure that only authorized users can access and operate the data.
[0023] Preferably, the user interface module includes a user login authentication unit and a data operation interface unit. The user login authentication unit provides user login and identity authentication functions to ensure that only authorized users can use the system interface. The data operation interface unit provides data storage, retrieval, and monitoring operation interfaces for the convenience of users.
[0024] The present invention has the following beneficial effects:
[0025] 1. Improve data availability and reliability: Through the design of a multi-active system across data centers, distributed storage of data is achieved. Even when an anomaly occurs at a single site, replicas in other data centers can ensure the integrity and availability of the data, thereby improving the data availability and reliability of the entire cloud platform.
[0026] 2. Enhance business continuity: Adopting a cross-data center distribution strategy with multiple levels of fault domains ensures that when a local fault occurs, the business system can quickly switch to other healthy data centers, thus guaranteeing business continuity and reducing the business interruption time caused by single-point failures.
[0027] 3. Optimize read performance: The strategy of giving priority to reading the primary replica can achieve the near-source reading of data, significantly improving the data reading performance, reducing the data access latency, and enhancing the user experience.
[0028] 4. Achieve load balancing: The storage controlled replica distribution algorithm based on centerless scalable hashing can automatically calculate the replica placement rules for various loads, achieving load balancing, reducing the pressure on a single data center, and improving the overall system performance.
[0029] 5. Support for disaster recovery: The design of multi-active data center-level storage in the same city and disaster recovery enables quick switching to other data centers in the same city in the event of a regional disaster, rapid restoration of business operations, and reduction of the impact of disasters on business. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 It is a flowchart of a distributed multi-site operating cloud storage system proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] Referring to Figure 1 , the present invention proposes a distributed multi-site operating cloud storage system, including:
[0032] A data management module, which is responsible for data storage, retrieval, and replica management to ensure data integrity and availability;
[0033] A storage pool management module, which is used to manage the resources of the storage pool;
[0034] A replica distribution module, which automatically calculates the placement rules of data replicas to achieve distributed storage of data;
[0035] A network scheduling module, which is responsible for data transmission and scheduling in the cross-data center network;
[0036] A fault detection module, which monitors the operating status of the storage system, detects faults and abnormal conditions, and issues alarms and handles them in a timely manner;
[0037] An arbitration module, which is responsible for decision-making and preventing split-brain in the event of a fault scenario to ensure data consistency and system stability;
[0038] A control plane service module, which provides control plane services for the distributed storage system;
[0039] A user interface module, which provides an interface for users to interact with the distributed storage system.
[0040] In a specific embodiment, the data management module is communicatively connected to the storage pool management module and the replica distribution module respectively to implement data storage and replica management. The storage pool management module is communicatively connected to the data management module and the replica distribution module respectively to implement the management and optimization of the storage pool. The replica distribution module is communicatively connected to the data management module and the storage pool management module respectively to implement the automatic distribution of data replicas. The network scheduling module is communicatively connected to the data management module and the storage pool management module respectively to implement the scheduling and transmission of data in the network. The fault detection module is communicatively connected to the data management module, the storage pool management module, and the replica distribution module respectively to provide fault detection and handling functions. The arbitration module is communicatively connected to the fault detection module, the data management module, and the storage pool management module respectively to implement the decision-making process for fault scenarios. The control plane service module is communicatively connected to the data management module, the storage pool management module, and the replica distribution module respectively to implement the management and services of the control plane. The user interface module is communicatively connected to the data management module, the storage pool management module, and the replica distribution module respectively to implement the interaction and monitoring between the user and the system.
[0041] In a specific embodiment, the data management module includes a data storage unit, a data retrieval unit, and a replica management unit. The data storage unit is responsible for storing data in a distributed storage system. The data retrieval unit is used to provide a data retrieval function, allowing users to search for and obtain data through an interface. The replica management unit is responsible for managing data replicas to ensure the availability and reliability of data.
[0042] It should be noted that: data storage and replica management algorithms:
[0043] Formula: Data = {D1, D2,..., Dn}, where Di represents a data block and n represents the number of data blocks;
[0044] Symbol description: D represents data, {} represents a set, = represents equivalence, and {D1, D2,..., Dn} represents a data block set.
[0045] In a specific embodiment, the storage pool management module includes a storage space allocation unit and a data replica distribution optimization unit. The storage space allocation unit is responsible for allocating storage space for data to ensure that there are sufficient storage resources available. The data replica distribution optimization unit optimizes the distribution of data according to the replica distribution algorithm to improve data access performance and availability.
[0046] It should be noted that: storage pool management algorithms:
[0047] Formula: StoragePool = {SP1, SP2,..., SPm}, where SPi represents a storage pool and m represents the number of storage pools;
[0048] Symbol Explanation: SP represents the storage pool, {} represents a set, = represents equivalence, and {SP1, SP2,..., SPm} represents a set of storage pools.
[0049] In a specific embodiment, the replica distribution module includes a data block allocation unit and a data replica management unit. The data block allocation unit assigns a unique hash value to each data block according to the decentralized scalable hashing algorithm to determine its storage location. The data replica management unit is responsible for managing data replicas to ensure that data is backed up at multiple locations, thereby improving data availability and reliability.
[0050] It should be noted that: Replica distribution algorithm:
[0051] Formula: Replicas={R1, R2,..., Rp}, where Ri represents a replica and p represents the number of replicas;
[0052] Symbol Explanation: R represents a data replica.
[0053] In a specific embodiment, the network scheduling module includes a data transmission unit and a data scheduling optimization unit. The data transmission unit is responsible for transmitting data in the cross-data center network. The data scheduling optimization unit optimizes the scheduling of data according to the network conditions and data access requirements to improve data access performance.
[0054] In a specific embodiment, the fault detection module includes a status monitoring unit and a fault warning unit. The status monitoring unit monitors the running status of the storage system in real time, including the health status of each node and network connections. The fault warning unit is used to issue an alarm in a timely manner when a fault or abnormal situation is detected for timely handling.
[0055] In a specific embodiment, the arbitration module includes a decision-making judgment unit and a data consistency maintenance unit. The decision-making judgment unit is responsible for making decision judgments in case of faults or conflicts. The data consistency maintenance unit ensures data consistency among multiple data centers to guarantee data security and reliability.
[0056] In a specific embodiment, the control plane service module includes a configuration management unit and a permission control unit. The configuration management unit is responsible for managing the configuration information of the system, including node configuration and storage configuration. The permission control unit is responsible for managing user permissions to ensure that only authorized users can access and operate data.
[0057] In a specific embodiment, the user interface module includes a user login authentication unit and a data operation interface unit. The user login authentication unit provides user login and identity authentication functions to ensure that only authorized users can use the system interface. The data operation interface unit provides data storage, retrieval, and monitoring operation interfaces for the convenience of users.
[0058] Working principle and usage process of the present invention: By constructing a cloud storage system operating in a distributed multi-site manner and based on a cross-data center distribution strategy with multiple levels of failure domains, distributed storage of data and load balancing are achieved. The data replica distribution algorithm in the system can automatically calculate the placement rules of data replicas to ensure the uniform distribution of data among multiple data centers, avoiding the risk of single point of failure. At the same time, the system adopts a strategy of giving priority to reading the primary replica to achieve the nearby reading of data and improve the performance of data access. When a failure occurs, the system can automatically switch to other healthy data centers to ensure the continuity of business. Through the design of an arbitration site, the system can prevent the occurrence of split-brain and ensure data consistency and system stability. Overall, through distributed storage and multi-active design, the system effectively solves the problems of data availability and business continuity existing in the single data center deployment of existing cloud platforms.
[0059] The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, making equivalent substitutions or changes should be covered within the protection scope of the present invention.
Claims
1. A distributed multi-site running cloud storage system, characterized in that, It includes: A data management module, which is responsible for data storage, retrieval, and replica management to ensure data integrity and availability; A storage pool management module, which is used to manage the resources of the storage pool; A replica distribution module, which automatically calculates the placement rules for data replicas to achieve distributed storage of data; A network scheduling module, which is responsible for data transmission and scheduling in the cross-data center network; A fault detection module, which monitors the operating status of the storage system, detects faults and anomalies, and issues alarms and performs processing in a timely manner; An arbitration module, which is responsible for decision-making and preventing split-brain in the event of a fault scenario to ensure data consistency and system stability; A control plane service module, which provides control plane services for the distributed storage system; A user interface module, which provides an interface for users to interact with the distributed storage system.
2. A distributed multi-site operating cloud storage system according to claim 1, wherein: The data management module is communicatively connected to the storage pool management module and the replica distribution module respectively to implement data storage and replica management. The storage pool management module is communicatively connected to the data management module and the replica distribution module respectively to implement the management and optimization of the storage pool. The replica distribution module is communicatively connected to the data management module and the storage pool management module respectively to achieve automatic distribution of data replicas. The network scheduling module is communicatively connected to the data management module and the storage pool management module respectively to implement data scheduling and transmission in the network. The fault detection module is communicatively connected to the data management module, the storage pool management module, and the replica distribution module respectively to provide fault detection and processing functions. The arbitration module is communicatively connected to the fault detection module, the data management module, and the storage pool management module respectively to implement decision-making processing in the fault scenario. The control plane service module is communicatively connected to the data management module, the storage pool management module, and the replica distribution module respectively to implement control plane management and services. The user interface module is communicatively connected to the data management module, the storage pool management module, and the replica distribution module respectively to implement user-system interaction and monitoring.
3. A distributed multi-site operating cloud storage system according to claim 2, wherein: The data management module includes a data storage unit, a data retrieval unit, and a replica management unit. The data storage unit is responsible for storing data in the distributed storage system. The data retrieval unit provides a data retrieval function, allowing users to search for and obtain data through an interface. The replica management unit is responsible for managing data replicas to ensure data availability and reliability.
4. A distributed multi-site operating cloud storage system according to claim 3, wherein: The storage pool management module includes a storage space allocation unit and a data replica distribution optimization unit. The storage space allocation unit is responsible for allocating storage space for data to ensure that there are sufficient storage resources available. The data replica distribution optimization unit optimizes the distribution of data according to the replica distribution algorithm to improve data access performance and availability.
5. A distributed multi-site running cloud storage system according to claim 4, characterized in that: The replica distribution module includes a data block allocation unit and a data replica management unit. The data block allocation unit assigns a unique hash value to each data block according to the decentralized scalable hashing algorithm to determine its storage location. The data replica management unit is responsible for managing data replicas to ensure that data is backed up at multiple locations to improve data availability and reliability.
6. A distributed multi-site operating cloud storage system according to claim 5, characterized in that: The network scheduling module includes a data transmission unit and a data scheduling optimization unit. The data transmission unit is responsible for transmitting data in the cross-rack network. The data scheduling optimization unit optimizes the scheduling of data according to the network conditions and data access requirements to improve the data access performance.
7. A distributed multi-site running cloud storage system according to claim 6, characterized in that: The fault detection module includes a status monitoring unit and a fault warning unit. The status monitoring unit monitors the running status of the storage system in real time, including the health status of each node and the network connection. The fault warning unit is used to issue an alarm in a timely manner when a fault or abnormal situation is detected for timely handling.
8. A distributed multi-site operating cloud storage system according to claim 7, characterized in that: The arbitration module includes a decision-making judgment unit and a data consistency maintenance unit. The decision-making judgment unit is responsible for making decision judgments in case of faults or conflicts. The data consistency maintenance unit ensures the consistency of data among multiple data centers to guarantee the security and reliability of data.
9. A distributed multi-site running cloud storage system according to claim 8, characterized in that: The control plane service module includes a configuration management unit and a permission control unit. The configuration management unit is responsible for managing the configuration information of the system, including node configuration and storage configuration. The permission control unit is responsible for managing user permissions to ensure that only authorized users can access and operate data.
10. A distributed multi-site operating cloud storage system according to claim 9, characterized in that: The user interface module includes a user login authentication unit and a data operation interface unit. The user login authentication unit provides user login and identity authentication functions to ensure that only authorized users can use the system interface. The data operation interface unit provides data storage, retrieval, and monitoring operation interfaces for the convenience of users.