Data storage method and device of server cluster and storage medium
By combining remote direct memory access and power-protected memory in the server cluster, the performance limitation caused by high latency is solved, achieving efficient data storage and fast persistence, and improving the efficiency and performance of data storage.
Patent Information
- Application Number
- CN202510990152.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-11-07
AI Technical Summary
Existing data storage methods for server clusters suffer from high latency, resulting in performance limitations and failing to meet the low-latency, high-throughput file operation requirements of scenarios such as high-performance computing, real-time data analysis, and cloud storage services.
The data storage method using a server cluster includes a primary storage node and N backup storage nodes in the storage partition. Data is written to the pre-allocated memory of the backup storage node through a one-sided write operation via remote direct memory access. After confirming the write, it is determined that the data has been written to the pre-allocated memory of the primary storage node. The high-speed characteristics of power-saving memory are used to reduce persistence operations.
By reducing frequent persistent operations to hard drives or solid-state drives, data storage efficiency is improved, data storage performance is optimized, and the performance requirements of high-performance computing and real-time data processing are met.
Smart Images

Figure CN120909506A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer storage, and particularly relates to a data storage method and device of a server cluster and a storage medium. BACKGROUND
[0002] Small file creation operations usually involve multiple data updates and data storage processes, wherein each data storage process may involve multiple data persistence operations and data persistence to underlying storage media. In conventional storage architecture, metadata and data persistence are mostly dependent on hard disk or solid state disk storage pools. For large-scale small file high-frequency creation scenarios, the input / output operation delay is high, and the accumulation of such delay limits the overall performance of the file system, which cannot meet the growing demand for low-latency, high-throughput file operations such as high-performance computing, real-time data analysis, cloud storage services, etc.
[0003] Therefore, the data storage method of the server cluster in the related art has the problem of limited performance due to high latency. SUMMARY
[0004] The present application provides a data storage method and device of a server cluster and a storage medium to at least solve the problem of limited performance due to high latency in the related art.
[0005] The present application provides a data storage method of a server cluster, the server cluster comprising a storage partition, storage nodes in the storage partition comprising one master storage node and N backup storage nodes, each storage node in the storage partition being pre-allocated with pre-allocated memory, and N being a positive integer greater than or equal to 1. The method comprises: in the case that a current storage node is the master storage node, receiving a data write request from a client by the current storage node, wherein the data write request is used to request to write specified data; performing a one-sided write operation of remote direct memory access by the current storage node to write the specified data into the pre-allocated memory of a backup storage node in the N backup storage nodes; and in the case that it is determined that the specified data has been written into the pre-allocated memory of the backup storage node in the N backup storage nodes, determining by the current storage node that the specified data has been written into the pre-allocated memory of the current storage node.
[0006] The application further provides a data storage device of a server cluster, the server cluster comprising a storage partition, storage nodes in the storage partition comprising one master storage node and N backup storage nodes, the storage nodes in the storage partition each being pre-allocated with pre-allocated memory, N being a positive integer greater than or equal to 1; the device comprising: a receiving unit configured to receive, by a current storage node, a data write request from a client in a case where the current storage node is the master storage node, wherein the data write request is used to request writing of specified data; a first executing unit configured to perform, by the current storage node, a one-side write operation of a remote direct memory access to write the specified data into the pre-allocated memory of a backup storage node of the N backup storage nodes; and a first determining unit configured to determine, by the current storage node, that the specified data has been written into the pre-allocated memory of the current storage node in a case where it is determined that the specified data has been written into the pre-allocated memory of the backup storage node of the N backup storage nodes.
[0007] The application further provides an electronic device comprising: a memory configured to store a computer program; and a processor configured to implement the steps of any of the data storage methods of the server cluster when executing the computer program.
[0008] The application further provides a computer-readable storage medium, the computer-readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement the steps of any of the data storage methods of the server cluster.
[0009] The application further provides a computer program product comprising a computer program, the computer program being executed by a processor to implement the steps of any of the data storage methods of the server cluster.
[0010] According to the application, since the server cluster comprises a storage partition, the storage nodes in the storage partition comprise one master storage node and N standby storage nodes, the storage nodes in the storage partition are pre-allocated with pre-allocated memories, N is a positive integer greater than or equal to 1; in the case that the current storage node is the master storage node, the current storage node receives a data write request from a client, wherein the data write request is used to request to write specified data; the current storage node performs a one-sided write operation of remote direct memory access to write the specified data into the pre-allocated memory of the standby storage node in the N standby storage nodes, and flexibly uses the pre-allocated memory; in the case that it is determined that the specified data has been written into the pre-allocated memory of the standby storage node in the N standby storage nodes, the current storage node determines that the specified data has been written into the pre-allocated memory of the current storage node, so that the efficient cache utilization mechanism can be realized by using the data fast persistence technology, and the data does not need to be frequently persisted into a hard disk or a solid state disk by performing a persistence operation, thereby the problem that the data storage method of the server cluster in the related art is limited in performance due to high delay can be solved, and the technical effects of improving the efficiency of data storage and optimizing the data storage performance are achieved. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0012] Figure 1 An application scenario diagram of a data storage method of a server cluster according to an embodiment of the present application.
[0013] Figure 2 A flow diagram of an optional data storage method of a server cluster according to an embodiment of the present application.
[0014] Figure 3 A structure diagram of an optional data storage method of a server cluster according to an embodiment of the present application.
[0015] Figure 4 A flow diagram of another optional data storage method of a server cluster according to an embodiment of the present application.
[0016] Figure 5 A flow diagram of still another optional data storage method of a server cluster according to an embodiment of the present application.
[0017] Figure 6 A structure block diagram of an optional data storage device of a server cluster according to an embodiment of the present application. DETAILED DESCRIPTION
[0018] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0019] It should be noted that, in the description of the present application, the terms "comprising", "containing" or any other variants thereof are intended to cover non-exclusive containing, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. The terms "first", "second" and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0020] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0021] According to an aspect of the embodiments of the present application, a data storage method of a server cluster is provided. Optionally, in the present embodiment, the above-mentioned data storage method of the server cluster can be applied to a hardware environment composed of a terminal device 102 and a server 104 as shown in the figure. Figure 1 As shown in the figure, the server 104 is connected with the terminal device 102 through a network, and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal, and a database can be set on the server or independently of the server, which is used to provide data storage services for the server 104. Figure 1 As shown in the figure, the server 104 is connected with the terminal device 102 through a network, and can be used to provide services (such as application services, etc.) for the terminal or the client installed on the terminal, and a database can be set on the server or independently of the server, which is used to provide data storage services for the server 104.
[0022] The above-mentioned network can include but is not limited to at least one of the following: a wired network, a wireless network. The above-mentioned wired network can include but is not limited to at least one of the following: a wide area network, a metropolitan area network, a local area network, and the above-mentioned wireless network can include but is not limited to at least one of the following: WIFI (Wireless Fidelity), Bluetooth.
[0023] The data storage method of the server cluster of the embodiments of the present application can be executed by the server 104, or can be executed by the terminal device 102, or can be executed by the server 104 and the terminal device 102 together. Wherein, the terminal device 102 executing the data storage method of the server cluster of the embodiments of the present application can also be executed by the client installed thereon.
[0024] Taking the data storage method of the server cluster executed by the server 104 in this embodiment as an example, Figure 2 is a flowchart of an optional data interaction method according to an embodiment of the present application, as shown in the figure, the flow of the method includes the following steps: Figure 2
[0025] Step S202, in the case where the current storage node is the primary storage node, receiving, by the current storage node, a data write request from a client, wherein the data write request is used to request to write specified data;
[0026] Step S204, performing, by the current storage node, a one-sided write operation of remote direct memory access to write the specified data into the pre-allocated memory of the backup storage node in the N backup storage nodes;
[0027] Step S206, in the case where it is determined that the specified data has been written into the pre-allocated memory of the backup storage node in the N backup storage nodes, determining, by the current storage node, that the specified data has been written into the pre-allocated memory of the current storage node.
[0028] The method in this embodiment can be applied to the field of computer storage technology and can be applied to the scene of data storage using power retention memory. Here, the power retention memory refers to a special type of memory, which has the characteristic of non-volatility, that is, it can still keep the previously stored data from being lost after the power is disconnected.
[0029] In the context of the big data era, one of the biggest challenges faced by file storage systems, especially those applied to high-performance computing, real-time data analysis and cloud storage services, is to handle the rapid creation and writing of a large number of small files. These small files may only have a few tens of kilobytes (KB) or less, but the number is extremely large, such as millions, tens of millions or even more. Conventional storage architectures mainly rely on hard disk drives (HDD) and solid state drives (SDD) as the underlying data storage medium, and in the context of handling such large-scale small file frequent read and write, these data storage solutions begin to show their low efficiency.
[0030] Taking a distributed file system as an example, the creation process of a small file can include: a client initiates a file creation request; after the file system receives the request, it allocates necessary metadata space for the new file in the metadata storage area, which involves creating a unique identifier for the file, permission information, etc.; the file system updates the directory table and adds an entry pointing to the new file, so that users can find the file through the path name; when the file is created and ready to receive data, the client writes data into the file, and each data write triggers a series of persistence operations to ensure that data is not lost due to system crashes or sudden power failures, which includes writing metadata and actual data blocks related to the file to storage media.
[0031] In this process, the metadata and data persistence operations become the performance bottleneck of the entire small file creation process, especially when using HDD and SSD as storage media. Although HDD has large capacity, its rotation and seek time results in high I / O delay, generally between a few milliseconds to tens of milliseconds. While SSD greatly reduces seek time and improves read / write speed, the delay of a single operation is still difficult to ignore, about 2 milliseconds. When the file system needs to frequently perform metadata updates and data block persistence operations, these seemingly small delays will accumulate under large-scale concurrent operations, ultimately seriously affecting the overall file system performance.
[0032] Here, the metadata persistence operation is a data management process that ensures that the metadata of the file system is not lost in the event of a sudden power failure or system failure, thereby maintaining the consistency and integrity of the file system. When the file system performs file creation, deletion, renaming, attribute modification, or data update, the corresponding metadata will also change. In order to prevent these changes from being lost in the event of a system failure, the file system needs to write the changed metadata to a persistent storage medium, and this process is the metadata persistence.
[0033] Specifically, the persistence of the metadata log (Metadata Log, MDLog for short) to the SSD storage pool takes about 2 milliseconds, which is used in the file system to record metadata updates and ensure consistency. Similarly, the persistence of data blocks to SSD also takes about 2 milliseconds. When processing a single small file, such a delay may not be important, but if it is to process thousands or even more small files, each operation needs to wait at least 4 milliseconds (two 2-millisecond delays), plus other system overhead, the actual waiting time may be longer.
[0034] In the field of high-performance computing, real-time data analysis, etc., the performance of the file system directly affects the running efficiency of the application program and the user experience. These systems often need to process a large amount of data input and output, which includes a large number of small file read and write. For example, a scientific simulation may need to generate or read hundreds of millions of data files, and the size of each file may only be a few. In this scenario, every millisecond of delay can significantly extend the completion time of the task. This not only reduces the throughput of the file system, that is, the number of file operations processed per second, but also increases the user's waiting time, because each write request is delayed until all persistent operations are completed.
[0035] Persistent Memory (PMem) is a new type of non-volatile storage technology that combines the fast read and write performance of Random Access Memory (RAM) and the persistence of a disk, and can save data in memory after power failure. The read and write speed of persistent memory can reach nanoseconds, which is much faster than the millisecond-level delay of HDD and SSD.
[0036] However, the current file system and storage architecture do not effectively use this medium as a cache layer to optimize the data persistence process. When designing a regular file system, the main consideration is still based on the storage environment of HDD and SSD, so the persistence strategy and data management mechanism are not optimized for the characteristics of persistent memory. For example, the file system may frequently flush data and metadata to disk without considering the caching capabilities of persistent memory, which results in the high-speed read and write capabilities of persistent memory not being fully utilized, and the system is still limited by the I / O performance of SSD and HDD.
[0037] Therefore, the data storage method of the server cluster in the related art has the problem of limited performance due to high latency.
[0038] To at least partially solve the above technical problems, the embodiment provides a data storage method of a server cluster, wherein the server cluster comprises a storage partition, storage nodes in the storage partition comprise one master storage node and N backup storage nodes, the storage nodes in the storage partition are each previously allocated with pre-allocated memory, N is a positive integer greater than or equal to 1; the method comprises: in the case that a current storage node is the master storage node, receiving, by the current storage node, a data write request from a client, wherein the data write request is used to request to write specified data; performing, by the current storage node, a one-sided write operation of a remote direct memory access, to write the specified data into the pre-allocated memory of a backup storage node of the N backup storage nodes, and flexibly use the pre-allocated memory; in the case that it is determined that the specified data has been written into the pre-allocated memory of the backup storage node of the N backup storage nodes, determining, by the current storage node, that the specified data has been written into the pre-allocated memory of the current storage node, so that the efficient cache utilization mechanism can be realized by using the data fast persistence technology, and the data does not need to be frequently persisted into a hard disk or a solid state disk by performing a persistence operation, therefore, the problem that the data storage method of the server cluster in the related art has limited performance due to high latency can be solved, and the technical effects of improving the efficiency of data storage and optimizing the data storage performance are achieved.
[0039] In the embodiment, the server cluster can comprise a plurality of storage partitions (Pt for short), the server cluster is a network composed of a plurality of independent servers, which can provide services together through a shared resource and a load sharing mechanism, the storage partition can be a basic unit of data storage and management in the server cluster, one server in the server cluster can correspond to one storage node, each storage partition can contain one master storage node and at least one backup storage node, to realize redundant storage and fast access of data, and improve the reliability of data and the overall performance of the storage system, wherein the master storage node can be the main read-write node of the storage partition, responsible for processing read-write requests of the client, and writing data into the pre-allocated memory, and the backup storage node can be used to store a copy of the master node data, to provide data redundancy and recovery capability when the master node fails.
[0040] For example, each storage partition can contain three nodes (one master storage node and two backup storage nodes), forming a three-copy data storage mode, which can realize that the remaining two nodes can continue to provide services when any node fails or is offline, and maintain the availability and consistency of data.
[0041] Optionally, each storage node in a storage partition can be pre-allocated with pre-allocated memory, the size of which can be pre-set, and the memory can be power-on memory, which is used to temporarily store data from the client, so as to take advantage of the high-speed characteristics of the power-on memory while avoiding the performance loss caused by frequent disk flushing operations.
[0042] Optionally, one storage node can belong to multiple partitions, which means that the power-on memory resources on one storage node can be shared by multiple partitions, improving resource utilization, reducing hardware costs, and enabling uniform distribution of load among different partitions to avoid individual partitions from becoming bottlenecks due to insufficient resources.
[0043] Optionally, one storage node can be responsible for different content in different storage partitions, for example, one storage node can be a primary storage node in one storage partition and a backup storage node in another storage partition.
[0044] Optionally, a data write request can be received by the current storage node from the client when the current storage node is a primary storage node, wherein the data write request is used to request to write specified data.
[0045] Optionally, when the client needs to store or update data, it can create a data write request and send it to the storage partition in the server cluster, which can include the data to be written and metadata information such as file path, data size, operation type, etc. The request can be received by the primary storage node in the storage partition and the specified data can be written to its own pre-allocated memory, which can be power-on memory.
[0046] Optionally, the allocation of data write requests can be based on a hash algorithm-based partition allocation, where the hash algorithm is a method of converting input data into a fixed-length output (i.e. hash value). The purpose of the hash algorithm is to produce different hash values for different inputs while ensuring that the same input produces the same hash value. The hash algorithm can uniformly distribute the input in the entire value range of the output to avoid hash collisions. In a distributed storage system, when a write request occurs, a hash algorithm can be used to hash a certain key attribute (such as file name, object key, etc.) in the request to produce a hash value, which is then mapped to a specific storage partition to determine which node should handle the request, thereby achieving load balancing among servers, i.e. ensuring that the write request load of different nodes is as uniform as possible.
[0047] Optionally, a one-sided write operation of remote direct memory access can be performed by the current storage node to write the specified data to the pre-allocated memory of the backup storage node in the N backup storage nodes.
[0048] Here, Remote Direct Memory Access (RDMA) allows a node to directly write data into the memory of another node without the processor of the target node participating in data processing. In the conventional network communication mode, when data is sent from one node to another node, the processor of the target node must receive the data and write it into the memory or disk, which brings additional processing delay and resource consumption. However, the RDMA technology skips the participation of the processor of the target node by directly accessing the memory, and the data is directly transmitted from the memory of the sending node to the memory of the receiving node, thereby greatly reducing the network transmission delay and improving the efficiency of data replication and storage.
[0049] Optionally, when the primary storage node receives a data write request from the client, it not only stores the data in the local pre-allocated memory (which can be the power-on memory), but also immediately writes the same data into the pre-allocated memory of the N backup storage nodes in the storage partition to which it belongs through the one-sided write operation of RDMA, thereby realizing multi-copy storage of data and improving the reliability of data.
[0050] Optionally, after receiving the first write request (such as 100KB size), the data can be written into the 0-100KB interval of the pre-allocated memory, and the subsequent write requests are sequentially written into the subsequent space until the pre-allocated memory is filled or a certain condition is reached.
[0051] Optionally, in the case where it is determined that the specified data has been written into the pre-allocated memory of the backup storage node in the N backup storage nodes, the current storage node can determine that the specified data has been written into the pre-allocated memory of the current storage node.
[0052] In this embodiment, after the data is synchronized to the backup storage node, the primary storage node needs to confirm whether the data has been successfully written into the pre-allocated memory of the backup node. It can wait for feedback from the N backup storage nodes in the storage partition to confirm that the data has arrived safely. The feedback can be automatically sent by the backup storage node after the data is written, or reported periodically through a certain heartbeat mechanism.
[0053] By migrating the MDLog and data persistence from the conventional non-volatile storage medium to the power-on memory cache, the time consumed by a single persistence can be reduced from milliseconds to within 0.2ms, greatly shortening the waiting time for multiple persistences in the small file creation process, improving the system throughput, and meeting the performance requirements of file operations in high-performance computing, real-time data processing, and other scenarios.
[0054] According to the embodiments provided in the present application, the server cluster includes a storage partition, the storage nodes in the storage partition include one master storage node and N backup storage nodes, the storage nodes in the storage partition are each pre-allocated with pre-allocated memory, N is a positive integer greater than or equal to 1; the method includes: in the case that the current storage node is the master storage node, receiving, by the current storage node, a data write request from a client, wherein the data write request is used to request to write specified data; performing, by the current storage node, a one-sided write operation of a remote direct memory access, to write the specified data into the pre-allocated memory of a backup storage node of the N backup storage nodes; in the case that it is determined that the specified data has been written into the pre-allocated memory of the backup storage node of the N backup storage nodes, determining, by the current storage node, that the specified data has been written into the pre-allocated memory of the current storage node, which can solve the problem that the data storage method of the server cluster in the related art is limited in performance due to high latency, and achieve the technical effects of improving the efficiency of data storage and optimizing the performance of data storage.
[0055] In one example embodiment, the server cluster further includes an interface layer component, the storage nodes in the storage partition run a data storage process, and the data storage process running on the storage nodes in the storage partition and the interface layer component use shared memory for lock-free queue communication; receiving, by the current storage node, a data write request from a client includes: reading, by the data storage process running on the current storage node, a request data packet encapsulated by the data write request from a specified queue of the shared memory in a non-blocking manner, wherein the request data packet is obtained by encapsulating the received data write request by the interface layer component, and the request data packet is written into the specified queue by the interface layer component in an atomic manner.
[0056] Optionally, in the above architecture design, a fast persistence layer (FPL) can be deployed on each server of the cluster as an independent process, i.e., a data storage process, so that the management of the power preservation memory and the data persistence operation can be dispersed to each server node, and distributed processing and storage of data can be achieved.
[0057] Optionally, the data storage process (which can be an FPL process) can be used to manage the power preservation memory resource on the node.
[0058] To enable upper-layer applications to seamlessly utilize the fast persistence capabilities of the data storage process, the server cluster can also include an interface layer component. For example, corresponding to the FPL process, the Fast Persistence Layer Client (FPLC) can serve as the interface layer component. It can act as a transparent fast persistence interface, eliminating the need for the application to interact directly with the power-saving memory. Instead, it operates indirectly through the FPLC, simplifying the complexity of the data retrieval process.
[0059] Optionally, the interface layer component can act as a bridge between the Distributed File System (DFS) or network file system and the underlying storage nodes. This decouples the direct interaction between upper-layer applications and storage nodes, translating application-layer data write requests into a format that storage nodes can understand and process. In this way, the interface layer component can hide the underlying storage details, providing a more consistent and standardized interface to upper-layer applications. It transforms complex persistence operations into simple API calls, making them transparent to upper-layer applications, allowing applications to focus on business logic without worrying about the details of data storage.
[0060] For example, such as Figure 3 As shown, the interface layer component provides an interface for communication with the data storage process (which can be the FPL process), allowing unstructured integration of file system services and storing data in the storage pool. The underlying layer can include system platform, network hardware, power-protected memory, and remote direct memory access network, etc. For example, the DFS component writes data to the FPL, and also supports the Metadata Server (MDS) to write log records to power-protected memory through the interface layer component when processing metadata logs. It can work in conjunction with file system services (e.g., network file system, distributed file system, file system interface, etc.).
[0061] Optionally, the data storage processes running on the storage nodes within the storage partition can communicate with the interface layer components via lock-free queues using shared memory. Shared memory is an inter-process communication mechanism that allows multiple processes or threads to share the same memory space, thereby enabling fast data exchange and transfer. In a server cluster, shared memory can be used as a data buffer pool between the interface layer components and the data storage processes on the storage nodes to reduce data transmission latency and improve processing speed.
[0062] Here, the conventional queue often uses a mutex lock to ensure data consistency and integrity in a multi-threaded environment, but this will bring additional processor overhead and potential lock competition problems, thereby affecting performance. A lock-free queue can ensure safe data exchange in high concurrency by avoiding the use of traditional lock mechanisms and using atomic operations and other implementation mechanisms, while avoiding the performance bottleneck caused by traditional locks.
[0063] Optionally, the data storage process can read the request data packet from the specified queue of the shared memory in a non-blocking manner, which means that the process will not stop or wait because the queue is empty, but will return immediately and then try to read again at a later time to improve the concurrency and response speed of data processing, avoid process idle waiting, and improve resource utilization.
[0064] Optionally, the request data packet can be obtained by encapsulating the received data write request by the interface layer component, which can include object identifiers, offsets, data lengths, data contents, and other parameters; the interface layer component can use atomic methods when encapsulating data write requests and writing to the specified queue. Atomic writing means that the write operation is either completely executed or completely not executed, and it will not be interrupted by other processes or threads during this period, which not only ensures data consistency, but also reduces system overhead caused by lock mechanisms and improves write efficiency.
[0065] Through the embodiment, by using a shared memory lock-free queue communication mechanism between the data storage process and the interface layer component, the data exchange process between the data storage process and the interface layer component can be optimized, and the performance of data storage can be improved.
[0066] In one example embodiment, the pre-allocated memory of the storage nodes in the storage partition has a specified capacity, and the physical address of the pre-allocated memory of the master storage node has a corresponding relationship with the physical address of the pre-allocated memory of the standby storage node in the N standby storage nodes.
[0067] The current storage node performs a one-sided write operation of remote direct memory access to write the specified data into the pre-allocated memory of the standby storage node in the N standby storage nodes, including:
[0068] The current storage node writes the specified data into the first physical address of the pre-allocated memory of the current storage node;
[0069] The current storage node uses the address translation service of remote direct memory access to write the specified data from the first physical address into the second physical address corresponding to the pre-allocated memory of the standby storage node in the N standby storage nodes and the first physical address, to write the specified data into the pre-allocated memory of the standby storage node in the N standby storage nodes.
[0070] In the embodiment, the pre-allocated memory of the storage node in the storage partition can have a specified capacity, and the part of the memory can be non-volatile, for example, power-protected memory, which can keep data from being lost after power failure, and provide higher data persistence speed and reliability than traditional HDD or SSD.
[0071] The physical address of the pre-allocated memory of the master storage node can have a corresponding relationship with the physical address of the pre-allocated memory of the standby storage node in the N standby storage nodes.
[0072] Optionally, the physical address of the pre-allocated memory of the master storage node can have a one-to-one corresponding relationship with the physical address of the pre-allocated memory of the same partition in the N standby storage nodes, so that the data stored in a specific physical address in the master storage node memory can also find its copy at the corresponding physical address of each standby storage node, thereby simplifying the data replication and management process. When the master storage node updates the data, only the specific physical address of the data needs to be concerned, and the RDMA technology can be used to directly write the updated data to the corresponding address of the standby storage node.
[0073] Optionally, the master storage node can use the address conversion service of RDMA to directly write the data stored in the first physical address to the same physical address, i.e., the second physical address, in the pre-allocated memory of the N standby storage nodes. Since the RDMA write operation does not depend on the processor of the target node, the data transmission delay is extremely low, which is very suitable for high-concurrency storage environment.
[0074] Through the embodiment, by combining the strategy of pre-allocating corresponding memory for the master storage node and the standby storage node and the one-sided write operation of RDMA, a fast and reliable data replication mechanism can be provided, the performance of data replication is optimized, the system delay is reduced, and the data redundancy and overall reliability of the system are enhanced.
[0075] In one example embodiment, the above method further comprises:
[0076] In the case that the data written in the pre-allocated memory of the current storage node is full or reaches a preset condition, performing a disk flushing operation by the current storage node to persist the data written in the pre-allocated memory of the current storage node to the first storage medium;
[0077] Controlling the standby storage node in the N standby storage nodes to perform a disk flushing operation by the current storage node to persist the data written in the pre-allocated memory of the standby storage node in the N standby storage nodes to the second storage medium, wherein the first storage medium and the second storage medium are the same storage medium or different storage media.
[0078] Pre-allocated memory (e.g., battery-backed memory) can protect data from loss in the event of power failure, while providing high-speed data read and write capabilities. However, the memory resource is limited, and when it is filled with data written to it or reaches a certain preset condition, a flush operation can be performed to persist the data to other storage media.
[0079] Optionally, when the amount of data stored in the pre-allocated memory reaches its upper capacity limit, a flush operation can be automatically triggered to write the data in the memory to the storage media to free up the pre-allocated memory to receive new write requests.
[0080] Optionally, in addition to the capacity being filled, the flush operation can also be triggered by a preset strategy, such as the time for which the data resides in the memory reaching a certain time threshold.
[0081] Optionally, after triggering the flush operation, the current storage node (e.g., the primary storage node) can write the data in its pre-allocated memory to the first storage media. The first storage media can be the same type as the pre-allocated memory (e.g., battery-backed memory) or another type of storage media (e.g., HDD).
[0082] Optionally, the primary storage node not only needs to persist its own data, but can also be used to control N backup storage nodes to perform flush operations to persist the data in their respective pre-allocated memories to the second storage media. The type of the second storage media can be the same as or different from the first storage media.
[0083] It can be understood that this batch flush method can significantly reduce the flush frequency, reduce system overhead, and improve storage performance compared to flushing after a single write request.
[0084] Optionally, the flush operation can be synchronous (the primary storage node waits for all backup storage nodes to complete the flush before continuing) or asynchronous (the primary storage node initiates the flush operation and continues to process other requests).
[0085] Through this embodiment, by triggering the flush operation when the data written to the pre-allocated memory of the current storage node is filled or reaches a preset condition, the cache resource can be effectively utilized, and the redundant storage and consistency of the data can be ensured.
[0086] In one example embodiment, after performing the single-sided write operation of the remote direct memory access by the current storage node, the above method further includes:
[0087] performing a single-sided read operation of the remote direct memory access on the last byte of the specified data written to the pre-allocated memory of the backup storage node in the N backup storage nodes by the current storage node;
[0088] In case the single-sided read operation is successfully executed, it is determined that the specified data has been written into the pre-allocated memory of the backup storage node among the N backup storage nodes.
[0089] Although the RDMA single-sided write operation can significantly accelerate data transmission, its callback mechanism is not sufficient to ensure that the data has been completely written into the power-on memory of the target node, because the callback of the completion of the RDMA single-sided write operation only indicates that the data has been sent from the network interface of the sending node, but cannot ensure that the data has been stably written into the memory of the receiving node, especially on a non-volatile storage medium such as power-on memory.
[0090] Optionally, after performing the single-sided write operation of remote direct memory access through the current storage node, i.e., before the master storage node performs the callback of the single-sided write operation, a single-sided read operation of remote direct memory access can be immediately performed on the last byte of the specified data written into the pre-allocated memory of the backup storage node among the N backup storage nodes, thereby indirectly verifying whether the entire piece of data has been successfully written. If the read operation can correctly return the last byte of the data, it can be inferred that the entire data block has been completely written into the target memory region, i.e., it can be confirmed that the data has been completely written into the pre-allocated memory of the backup storage node without data corruption or loss.
[0091] Optionally, in case the single-sided read operation is successfully executed, it is determined that the specified data has been written into the pre-allocated memory of the backup storage node among the N backup storage nodes.
[0092] Optionally, after completing the local persistence, the master node can send a write completion notification to the backup node in an asynchronous manner, and the notification can include address information of the write request data in the power-on memory of the backup node.
[0093] Optionally, after confirming the completion of the local persistence operation, the master node can also send a response message to the interface layer component (e.g., FPLC) through shared memory, notifying that the write operation has been completed, thereby ensuring that the client can timely learn the status of the write operation.
[0094] For example, as Figure 4As shown, the interface layer component can be called by the upper layer application to write interface, send a write request through shared memory communication, the master storage node (i.e. the master node) can pre-apply power protection memory for the standby storage node (i.e. the standby node), write data into the pre-apply power protection memory of the standby node through the one-way write operation of RDMA, and the master storage node confirms that the data is written successfully after collecting the RDMA callback of the two standby storage nodes (indicating that the data has been written into the memory of the standby storage node), and completes the three-copy persistence. Subsequently, the master storage node can asynchronously send a write completion notification to the standby storage node, carrying the address information of the write request in the standby storage node power protection memory, so as to facilitate the standby storage node to manage the write request data according to the object identifier, and prepare for subsequent data flushing or data recovery during master-standby switching.
[0095] Through the embodiment, the integrity of data writing in the pre-allocated memory of the N standby storage nodes is verified through the one-way read operation of remote direct memory access performed by the current storage node, which can provide an efficient and low-latency data verification mechanism, and improve the overall performance and reliability of the system.
[0096] In one example embodiment, a data storage process and a monitoring component are running on the storage nodes in the storage partition, the data storage process running on the storage nodes in the storage partition is used to process the received write request, and the monitoring component running on the storage nodes in the storage partition is used to monitor the state of the data storage process running on the storage nodes in the storage partition; the above method further comprises:
[0097] In the case that the monitoring component running on the current storage node detects that the data storage process running on the current storage node is abnormal, the node state of the current storage node recorded in the first state mapping table is updated to the failure state, wherein the first state mapping table is used to record the node state of the storage nodes in the storage partition;
[0098] Send fault notification information to the standby storage node in the N standby storage nodes, wherein the fault notification information is used to notify that the current storage node has failed.
[0099] In the embodiment, a monitoring component can be running on each storage node (i.e. on each server) in the server cluster for monitoring the storage node.
[0100] Optionally, the data storage process running on each storage node can be used to process all write requests from the front-end client, including but not limited to storing data to the local power protection memory, synchronizing data with other storage nodes, flushing data to the persistent storage device, etc.; correspondingly, the monitoring component running on the storage node can be used to monitor the state of the corresponding data storage process, including the running status, performance indicators, error logs, etc.
[0101] Optionally, when the monitoring component detects that the data storage process running on the current storage node is abnormal, it can update the state of the current storage node saved in the first state mapping table from a normal running state to a failure state, where the first state mapping table can be a global data structure used to record state information of storage nodes in all storage partitions in the storage system.
[0102] Optionally, the monitoring component can also send failure notification information to the standby storage node related to the failed current storage node, that is, it can send the failure notification information to the standby storage node in the N standby storage nodes in the storage partition to which the failed current storage node belongs.
[0103] Through the embodiment, the monitoring component detects the state of the storage node, triggers state update and failure notification in the case of state abnormality, which can effectively isolate and deal with failures and ensure data continuity and consistency.
[0104] In one example embodiment, the above method further comprises:
[0105] In the case where the current storage node is one of the N standby storage nodes and the primary storage node fails, the current storage node determines whether the promotion condition is met according to the second state mapping table, where the second state mapping table is used to record the node state of the N standby storage nodes and the standby storage node in the N standby storage nodes in chronological order of completing data state recovery, and the promotion condition includes that the current storage node is the first standby storage node recorded in the second state mapping table that has completed data state recovery and is in a valid state.
[0106] In the case where the current storage node meets the promotion condition, the current storage node performs a promotion operation to promote the current storage node to a new primary storage node.
[0107] The current storage node sends an information acquisition request to other standby storage nodes except the current storage node in the N standby storage nodes, where the information acquisition request is used to request the latest operation record and incomplete operation of the other standby storage nodes.
[0108] In the case where the current storage node receives information acquisition response information returned by the other standby storage nodes corresponding to the information acquisition request, the current storage node performs a data synchronization operation on the other standby storage nodes.
[0109] In a distributed storage system, in order to improve the reliability of data and the availability of the system, a primary-standby architecture can also be used, in which one node acts as a primary node to handle read and write operations, and other nodes act as standby nodes to take over services when the primary node fails. When the primary node fails, the system needs to perform primary-standby switchover, which includes data synchronization and state update.
[0110] During the master-slave switchover, the data on the master node needs to be synchronized to the standby node to ensure that the data of all nodes is consistent. For a large-scale distributed storage system, this process can involve thousands or even more files and data blocks. Conventional data synchronization mechanisms are inefficient when synchronizing a large amount of data, requiring a large amount of time and network bandwidth, which can result in a long business interruption time. In addition, if the data replication is not complete or the synchronization operation is interrupted during data writing, it can cause data inconsistency between different nodes. For example, if the standby node receives a new write request during data synchronization, but does not receive the updated data from the master node in time, it can cause the data on some nodes to be inconsistent with other nodes in the system. This not only affects the integrity and reliability of the data, but also increases the management complexity and maintenance cost of the system.
[0111] As can be seen, the complexity and risk of the master-slave switchover and data synchronization process directly affect the business continuity of the system. During the switchover, the system can need to suspend external services until the data synchronization is complete, which cannot meet the needs of business scenarios that require high real-time performance and cannot tolerate long interruptions.
[0112] In this embodiment, a master-slave switchover process can be provided, allowing the standby storage node to upgrade to the master node according to certain conditions and perform necessary data synchronization to restore the service state of the system.
[0113] Optionally, the storage node can also include a second state mapping table, which is an auxiliary structure for tracking the state of the standby storage node. Unlike the first state mapping table, it can record the node state in chronological order according to the completion of data state recovery by the standby storage node.
[0114] Optionally, the current standby storage node needs to meet the master upgrade condition to become the new master storage node. The master upgrade condition can include that the current storage node is the first standby storage node in the second state mapping table that has completed data state recovery and is in a valid state. It can be the first standby node in the second state mapping table that is marked as having completed data state recovery, i.e., it can be the first storage node that has updated the missing or inconsistent data state to the latest storage node through the data synchronization or recovery process after the master storage node fails, and is currently in a valid running state, for example, the partition includes nodes 1 (master), 2, and 3, the second state mapping table is nodes 2 and 3, and when node 1 fails, node 2 upgrades to the master node.
[0115] Optionally, the standby storage node can bypass the master storage node and directly communicate with the interface layer component to obtain data in the case of master storage node failure, or can recover data from a pre-set data storage node, which is not limited in this embodiment.
[0116] Optionally, once the backup storage node determines that it meets the above master upgrade condition, it can immediately perform the master upgrade operation, including but not limited to updating its identity in each state mapping table, adjusting configuration parameters, and changing system roles, and officially becoming the new master storage node.
[0117] Optionally, the new master storage node can send an information acquisition request to the remaining backup storage nodes after completing the identity transition to collect the latest operation records and incomplete operation conditions of each node.
[0118] Optionally, based on the collected information, the new master storage node can perform a data synchronization operation, which can include but is not limited to identifying data differences by comparing the latest data records of the backup storage nodes, formulating a data synchronization strategy, and synchronizing data from other backup storage nodes to the new master storage node according to the strategy to ensure that the data states of all nodes are synchronized to the latest and achieve data consistency. After completing data synchronization, the new master storage node can also send a state update notification to all backup storage nodes to inform the latter of the latest data state and operation record, so that the backup nodes can update their own data view to ensure the overall coordination and consistency of the system.
[0119] For example, as Figure 5 the monitoring component can monitor the master node failure, update the first state mapping table and push the state mapping table to the two backup nodes in the storage partition to which the failed master node belongs, wherein according to the partition state in the second state mapping table, one backup node can determine to upgrade the master and become the new master node, the new master node sends a data acquisition request to the backup nodes in the partition, the backup nodes parse the write requests in the pre-application power preservation memory, and reply with a response containing partition information (such as the latest operation record, incomplete operation condition, and write request details between them). The new master node synchronizes the data of the nodes in the second state mapping table according to the response data, quickly restores the master-slave data consistency, enables the system to quickly recover the foreground business processing capability, and the monitoring component can also update the recorded partition master node. Since the backup storage nodes that have completed all write request data state recovery have very small data differences (such as the old master node failure before the partition master node only handles 2 concurrent write requests, and the backup storage nodes that have completed all write request data state recovery at most differ by 2 write requests), the synchronization process takes a short time, and the impact of the failure on the business is minimized.
[0120] Through this embodiment, by the backup storage node determining the master upgrade condition and performing the master upgrade operation and data synchronization according to the second state mapping table, the service state of the system can be quickly restored, ensuring the coordination and consistency of the data and improving the reliability of the data storage.
[0121] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software on a general hardware platform, and of course can also be realized by hardware, but in many cases the former is a better implementation.
[0122] According to another aspect of the embodiments of the present application, a device for implementing the data storage method of the above-mentioned server cluster is further provided. Figure 6 is a structural block diagram of an optional data storage device of a server cluster according to an embodiment of the present application, as shown in the figure, the device can include: Figure 6
[0123] The receiving unit 602 is configured to receive, by the current storage node, a data write request from the client in a case where the current storage node is a primary storage node, wherein the data write request is used to request to write specified data.
[0124] The first execution unit 602 is configured to perform, by the current storage node, a one-sided write operation of a remote direct memory access to write the specified data into the pre-allocated memory of the backup storage node in the N backup storage nodes.
[0125] The first determination unit 606 is configured to determine, by the current storage node, that the specified data has been written into the pre-allocated memory of the current storage node in a case where it is determined that the specified data has been written into the pre-allocated memory of the backup storage node in the N backup storage nodes.
[0126] It should be noted that the receiving unit 602 in this embodiment can be used to execute the above step S202, the first execution unit 604 in this embodiment can be used to execute the above step S204, and the first determination unit 606 in this embodiment can be used to execute the above step S206.
[0127] The server cluster includes a storage partition, the storage nodes in the storage partition include one master storage node and N backup storage nodes, the storage nodes in the storage partition are each pre-allocated with pre-allocated memory, N is a positive integer greater than or equal to 1; the method includes: in the case that the current storage node is the master storage node, receiving, by the current storage node, a data write request from a client, wherein the data write request is used to request to write specified data; performing, by the current storage node, a one-sided write operation of a remote direct memory access to write the specified data into the pre-allocated memory of the backup storage node of the N backup storage nodes; in the case that it is determined that the specified data has been written into the pre-allocated memory of the backup storage node of the N backup storage nodes, determining, by the current storage node, that the specified data has been written into the pre-allocated memory of the current storage node, which can solve the problem that the data storage method of the server cluster in the related art is limited in performance due to high latency, and achieve the technical effects of improving the efficiency of data storage and optimizing the performance of data storage.
[0128] The features of the embodiments of the data storage device of the server cluster can be referred to the related descriptions of the embodiments of the data storage method of the server cluster, which will not be repeated here.
[0129] In one example embodiment, the server cluster further includes an interface layer component, the data storage processes run on the storage nodes in the storage partition, and the data storage processes running on the storage nodes in the storage partition and the interface layer component use shared memory for non-lock queue communication; the receiving unit includes: a reading module configured to read, in a non-blocking manner, a data write request encapsulated into a request data packet by the data storage process running on the current storage node from a specified queue of the shared memory, wherein the request data packet is obtained by encapsulating the received data write request by the interface layer component, and the request data packet is written into the specified queue by the interface layer component in an atomic manner.
[0130] In one example embodiment, the pre-allocated memory of the storage nodes in the storage partition each has a specified capacity, the physical address of the pre-allocated memory of the master storage node has a corresponding relationship with the physical address of the pre-allocated memory of the backup storage node of the N backup storage nodes; the first execution unit includes: a first writing module configured to write, by the current storage node, the specified data into a first physical address of the pre-allocated memory of the current storage node; a second writing module configured to write, by the current storage node, the specified data from the first physical address into a second physical address of the pre-allocated memory of the backup storage node of the N backup storage nodes corresponding to the first physical address using an address translation service of a remote direct memory access, so as to write the specified data into the pre-allocated memory of the backup storage node of the N backup storage nodes.
[0131] In an example embodiment, the apparatus further includes a second execution unit configured to perform, by the current storage node, a flush operation to persist the data written in the pre-allocated memory of the current storage node to the first storage medium in a case that the data written in the pre-allocated memory of the current storage node is full or reaches a preset condition; and a third execution unit configured to control, by the current storage node, a backup storage node of the N backup storage nodes to perform the flush operation to persist the data written in the pre-allocated memory of the backup storage node of the N backup storage nodes to the second storage medium, wherein the first storage medium and the second storage medium are the same storage medium or different storage media.
[0132] In an example embodiment, the apparatus further includes a fourth execution unit configured to perform, by the current storage node, a single-side read operation of the remote direct memory access after performing the single-side write operation of the remote direct memory access by the current storage node, wherein the single-side read operation is performed on a last byte of the specified data written in the pre-allocated memory of the backup storage node of the N backup storage nodes; and a second determination unit configured to determine that the specified data is written in the pre-allocated memory of the backup storage node of the N backup storage nodes in a case that the single-side read operation is performed successfully.
[0133] In an example embodiment, the storage nodes in the storage partition run a data storage process and a monitoring component, the data storage process running on the storage nodes in the storage partition is configured to process a received write request, and the monitoring component running on the storage nodes in the storage partition is configured to monitor a state of the data storage process running on the storage nodes in the storage partition; the apparatus further includes an updating unit configured to update, in a case that the monitoring component running on the current storage node monitors that the data storage process running on the current storage node is abnormal, a node state of the current storage node recorded in a first state mapping table to an invalid state, wherein the first state mapping table is configured to record the node state of the storage nodes in the storage partition; and a first sending unit configured to send, to a backup storage node of the N backup storage nodes, failure notification information, wherein the failure notification information is configured to notify that the current storage node fails.
[0134] In an example embodiment, the apparatus further includes: a third determining unit configured to, in a case where the current storage node is one of the N backup storage nodes and the primary storage node fails, determine whether an upgrade condition is met by the current storage node according to the second state mapping table, wherein the first state mapping table is configured to record the N backup storage nodes and the node states of the backup storage nodes in the N backup storage nodes in chronological order of completing data state recovery, and the upgrade condition includes that the current storage node is the first backup storage node recorded in the second state mapping table that has completed data state recovery and is in a valid state; a fifth executing unit configured to, in a case where the current storage node meets the upgrade condition, execute an upgrade operation by the current storage node to upgrade the current storage node to a new primary storage node; a second sending unit configured to send, by the current storage node, an information acquisition request to the other backup storage nodes except the current storage node among the N backup storage nodes, wherein the information acquisition request is configured to request to acquire the latest operation record and the unfinished operation condition of the other backup storage nodes; and a sixth executing unit configured to, in a case where information acquisition response information corresponding to the information acquisition request is returned by the other backup storage nodes, execute a data synchronization operation by the current storage node on the other backup storage nodes.
[0135] Embodiments of the present application further provide an electronic device, comprising a memory and a processor, the memory storing a computer program, and the processor being configured to execute the computer program to perform the steps in any of the above-mentioned data storage method embodiments of the server cluster.
[0136] Embodiments of the present application further provide a computer readable storage medium, the computer readable storage medium storing a computer program, wherein the computer program is configured to perform the steps in any of the above-mentioned data storage method embodiments of the server cluster when executed.
[0137] In an example embodiment, the computer readable storage medium can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.
[0138] Embodiments of the present application further provide a computer program product, the computer program product comprising a computer program, the computer program being executed by a processor to implement the steps in any of the above-mentioned data storage method embodiments of the server cluster.
[0139] The embodiment of the present application further provides another computer program product comprising a nonvolatile computer readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-mentioned data storage method embodiments of the server cluster.
[0140] Those skilled in the art can further understand that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in the above description in a general manner. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0141] The above describes in detail the data storage method, device and storage medium of the server cluster provided by the present application. The principles and implementation modes of the present application are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that, for those skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A data storage method of a server cluster, characterized by, The server cluster comprises a storage partition, storage nodes in the storage partition comprise one master storage node and N backup storage nodes, the storage nodes in the storage partition are each pre-allocated with pre-allocated memory, N is a positive integer greater than or equal to 1; the method comprises: In the case that the current storage node is the master storage node, receiving, by the current storage node, a data write request from a client, wherein the data write request is used to request to write specified data; Performing, by the current storage node, a one-sided write operation of remote direct memory access to write the specified data into the pre-allocated memory of a backup storage node in the N backup storage nodes; In the case that it is determined that the specified data has been written into the pre-allocated memory of a backup storage node in the N backup storage nodes, determining, by the current storage node, that the specified data has been written into the pre-allocated memory of the current storage node.
2. The method of claim 1, wherein, The server cluster further comprises an interface layer component, a data storage process runs on the storage nodes in the storage partition, and the data storage process running on the storage nodes in the storage partition and the interface layer component communicate through a shared memory in a lock-free queue manner; The receiving, by the current storage node, of the data write request from the client comprises: Reading, by the data storage process running on the current storage node, in a non-blocking manner, a request data packet encapsulated by the data write request from a specified queue of the shared memory, wherein the request data packet is obtained by encapsulating the received data write request by the interface layer component, and the request data packet is written into the specified queue by the interface layer component in an atomic manner.
3. The method of claim 1, wherein, The pre-allocated memory of the storage nodes in the storage partition each has a specified capacity, and the physical address of the pre-allocated memory of the master storage node has a corresponding relationship with the physical address of the pre-allocated memory of a backup storage node in the N backup storage nodes; The performing, by the current storage node, of the one-sided write operation of remote direct memory access to write the specified data into the pre-allocated memory of a backup storage node in the N backup storage nodes comprises: Writing, by the current storage node, the specified data into a first physical address of the pre-allocated memory of the current storage node; Using, by the current storage node, an address translation service of remote direct memory access to write the specified data from the first physical address into a second physical address corresponding to the pre-allocated memory of a backup storage node in the N backup storage nodes and the first physical address, to write the specified data into the pre-allocated memory of a backup storage node in the N backup storage nodes.
4. The method of claim 3, wherein, The method further comprises: In the case that the data written in the pre-allocated memory of the current storage node is full or reaches a preset condition, performing, by the current storage node, a disk flushing operation to persist the data written in the pre-allocated memory of the current storage node into a first storage medium. The current storage node controls a backup storage node in the N backup storage nodes to perform a disk flushing operation to persist data written in the pre-allocated memory of the backup storage node in the N backup storage nodes to a second storage medium, wherein the first storage medium and the second storage medium are the same storage medium or different storage media.
5. The method of claim 1, wherein, After the single-side write operation of the remote direct memory access is performed by the current storage node, the method further comprises: performing, by the current storage node, a single-side read operation of the remote direct memory access on the last byte of the specified data written in the pre-allocated memory of the backup storage node in the N backup storage nodes; in the case that the single-side read operation is successfully performed, determining that the specified data has been written in the pre-allocated memory of the backup storage node in the N backup storage nodes.
6. The method of claim 1, wherein, The storage nodes in the storage partition run a data storage process and a monitoring component, the data storage process running on the storage nodes in the storage partition is used to process received write requests, and the monitoring component running on the storage nodes in the storage partition is used to monitor the state of the data storage process running on the storage nodes in the storage partition. The method further comprises: in the case that the monitoring component running on the current storage node monitors that the data storage process running on the current storage node is abnormal, updating the node state of the current storage node recorded in the first state mapping table to an invalid state, wherein the first state mapping table is used to record the node state of the storage nodes in the storage partition; sending fault notification information to the backup storage nodes in the N backup storage nodes, wherein the fault notification information is used to notify that the current storage node has failed.
7. The method according to any one of claims 1 to 6, characterized in that, The method further comprises: in the case that the current storage node is one of the N backup storage nodes and the primary storage node has failed, determining whether the promotion condition is met by the current storage node according to a second state mapping table, wherein the second state mapping table is used to record the N backup storage nodes and the node state of the backup storage nodes in the N backup storage nodes in chronological order of completing data state recovery, and the promotion condition includes that the current storage node is the first backup storage node recorded in the second state mapping table that has completed data state recovery and is in a valid state; in the case that the current storage node meets the promotion condition, performing a promotion operation by the current storage node to promote the current storage node to a new primary storage node; sending, by the current storage node, an information acquisition request to the other backup storage nodes in the N backup storage nodes except the current storage node, wherein the information acquisition request is used to request to acquire the latest operation record and uncompleted operation condition of the other backup storage nodes; in the case that the information acquisition response information corresponding to the information acquisition request returned by the other backup storage nodes is received, performing, by the current storage node, a data synchronization operation on the other backup storage nodes.
8. A data storage apparatus of a server cluster, characterized by, The server cluster comprises a storage partition, storage nodes in the storage partition comprise one master storage node and N backup storage nodes, the storage nodes in the storage partition are each pre-allocated with pre-allocated memory, N is a positive integer greater than or equal to 1; the device comprises: A receiving unit configured to, in a case where a current storage node is the master storage node, receive, by the current storage node, a data write request from a client, wherein the data write request is used to request writing of specified data; A first execution unit configured to perform, by the current storage node, a one-sided write operation of remote direct memory access to write the specified data into pre-allocated memory of a backup storage node of the N backup storage nodes; A first determination unit configured to, in a case where it is determined that the specified data has been written into the pre-allocated memory of the backup storage node of the N backup storage nodes, determine, by the current storage node, that the specified data has been written into the pre-allocated memory of the current storage node.
9. An electronic device, comprising: comprise: a memory configured to store a computer program; a processor configured to implement the steps of the data storage method of the server cluster according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, wherein the computer program is executed by the processor to implement the steps of the data storage method of the server cluster according to any one of claims 1 to 7.
Citation Information
Patent Citations
Distributed block storage data processing method and device, equipment and storage medium
CN114443364A
Data processing method and device of distributed storage system, equipment and medium
CN117255101A
Distributed data storage system and data storage method
CN117632034A
Data access device, method and system, data processing unit and network card
CN117667761A
High-concurrency read-write optimization system for distributed file system, and medium and device
WO2025001603A1