Data replication method for distributed file system, and distributed file system
By generating and replaying operation logs at the file semantic layer, the high complexity of distributed file systems in synchronous and asynchronous replication modes is solved, thus improving data replication performance.
Patent Information
- Application Number
- PCT/CN2025/073101
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-19
- Filing Date
- 2025-01-17
- Publication Date
- 2025-12-04
AI Technical Summary
Existing distributed file systems require different block layer schemes for synchronous and asynchronous replication modes, resulting in high development and maintenance complexity and low data replication performance.
By generating operation logs at the file semantic layer and replaying them, synchronous and asynchronous replication capabilities are achieved, reducing dependence on the block layer and adopting a single data replication scheme.
It reduces the complexity of system development and maintenance, improves data replication performance, and simplifies the implementation process of synchronous and asynchronous replication.
Smart Images

Figure CN2025073101_04122025_PF_FP_ABST
Abstract
Description
Data replication methods and distributed file systems
[0001] This application claims priority to Chinese Patent Application No. 202410703584.2, filed on May 31, 2024, entitled "Method for Processing File Operation Logs and a File System", and Chinese Patent Application No. 202410800089.3, filed on June 19, 2024, entitled "Data Replication Method and Distributed File System of a Distributed File System", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of storage technology, and in particular to a data replication method for a distributed file system and a distributed file system. Background Technology
[0003] A distributed file system consists of multiple storage nodes, which store files in units of physical blocks. Considering the high reliability and stability of the business, the storage nodes of a distributed file system can be networked in a master-slave configuration. That is, the distributed file system includes a master storage cluster and a slave storage cluster. Based on this network configuration, data replication between the master and slave can be based on physical blocks and can employ synchronous or asynchronous replication modes.
[0004] However, the aforementioned synchronous or asynchronous replication modes are based on physical blocks. When a storage node executes an operation request for a file in a distributed file system, it may decompose it into multiple operations on physical blocks. The backup storage cluster synchronizes the physical block operations of the primary storage cluster. In synchronous replication mode, all physical block operations are copied from the primary storage cluster to the backup storage cluster in real time, without needing to consider historical data versions. In asynchronous replication mode, it is necessary to ensure that the operations received by the backup storage cluster can be restored to a consistent data version at a certain point in history (i.e., the complete set of physical block operations corresponding to the file operation request at that time). Therefore, the physical block-based replication method requires different schemes to be designed for synchronous and asynchronous replication modes, which increases the complexity of distributed file system development and maintenance. In addition, the replication capability of a distributed file system depends on the system's replication capability of physical blocks, resulting in low data replication performance between the primary and backup. Summary of the Invention
[0005] This application provides a data replication method and a distributed file system, which can reduce the complexity of developing and maintaining a distributed file system and improve data replication performance. The technical solution is as follows:
[0006] Firstly, a data replication method in a distributed file system is provided, executed by a first storage cluster in the distributed file system, which also includes a second storage cluster. The method includes:
[0007] The first storage cluster receives a first operation request for a file, executes the first operation request based on the semantics of the first operation request, and generates an operation log corresponding to the semantics of the first operation request. The operation log is used to indicate at least one of the execution result of the file data and the execution result of the file metadata. The first storage cluster sends the operation log to the second storage cluster, which then replays the first operation request based on the operation log.
[0008] The semantics of the operation request indicate whether it is a request to manipulate the file's metadata or its data. The file's data describes its actual content, which can be text, audio, video, etc. The file's metadata describes its characteristics, including at least one of the following: filename, file size, inode, storage address of data blocks, file creation time, owner, and permissions. The storage address of data blocks can be represented by offset and length parameters. This application does not limit the content of the file's data and metadata. Requests to manipulate the file's metadata include requests to create and delete files, while requests to manipulate the file's data include requests to write data to the file.
[0009] In the above method, generating operation logs based on the semantics of file operation requests and replaying operation requests based on these operation logs achieves data replication at the file semantic layer. Compared with data replication at the block layer, it does not rely on the replication capabilities of the block layer and can use a single data replication scheme to achieve synchronous and asynchronous replication capabilities. Therefore, it is not necessary to design different block layer schemes to achieve synchronous and asynchronous replication, thereby reducing the complexity of system development and maintenance and improving data replication performance.
[0010] In one possible implementation, based on the semantics of the first operation request, the first operation request is executed to obtain an execution result corresponding to the semantics of the first operation request. Based on the execution result and the execution order of the first operation request, an operation log of the first operation request is generated, including:
[0011] If the first operation request is a request to operate on the file's metadata, then the first operation request is executed, and the execution result of the file's metadata is obtained. Based on the execution result of the file's metadata, a metadata operation log is generated. Based on the metadata operation log and the execution order of the first operation request, an operation log for the first operation request is generated. If the first operation request is a request to operate on the file's data, then the first operation request is executed, and the execution result of the file's data and the file's metadata are obtained. Based on the execution result of the file's data, a data operation log is generated. Based on the execution result of the file's metadata, a metadata operation log is generated. Based on the metadata operation log, the data operation log, and the execution order of the first operation request, an operation log for the first operation request is generated.
[0012] In one possible implementation, the first storage cluster includes a first network attached storage (NAS) server cluster and a first metadata storage node;
[0013] If the first operation request is a request to operate on the file's metadata, then the first operation request is executed, the execution result of the file's metadata is obtained, a metadata operation log is generated based on the execution result of the file's metadata, and an operation log for the first operation request is generated based on the metadata operation log and the execution order of the first operation request, including:
[0014] If the first operation request is an operation request on the metadata of a file, the first NAS server cluster sends the metadata targeted by the first operation request to the first metadata storage node; the first metadata storage node assigns a logical sequence number to the first operation request, executes the first operation request on the metadata, obtains the execution result of the metadata of the file, and generates a metadata operation log based on the execution result of the metadata of the file. The logical sequence number is used to indicate the execution order of the first operation request; the first NAS server generates an operation log for the first operation request based on the metadata operation log and the execution order of the first operation request.
[0015] In one possible implementation, the first storage cluster includes a first NAS server cluster, a first metadata storage node, and a first data storage cluster;
[0016] If the first operation request is a data operation request on a file, then the first operation request is executed, obtaining the execution results for the file data and the file metadata. Based on the execution results for the file data, a data operation log is generated; based on the execution results for the file metadata, a metadata operation log is generated; and based on the execution order of the metadata operation log, the data operation log, and the first operation request, an operation log for the first operation request is generated, including:
[0017] If the first operation request is a data operation request on a file, the first NAS server cluster sends the data carried by the first operation request to the first data storage cluster; the first data storage cluster executes the first operation request on the data carried by the first operation request, obtains the execution result on the file data, and generates a data operation log based on the execution result on the file data; the first NAS server cluster sends the metadata targeted by the first operation request to the first metadata storage node; the first metadata storage node assigns a logical sequence number to the first operation request, executes the first operation request on the metadata, obtains the execution result on the file metadata, and generates a metadata operation log based on the execution result on the file metadata, with the logical sequence number used to indicate the execution order of the first operation request; the first NAS server cluster generates an operation log for the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request.
[0018] In one possible implementation, the first NAS server cluster includes at least one first NAS server and at least one second NAS server, wherein the first NAS server is used to receive and forward operation requests, and the second NAS server is used to generate and forward operation logs of the operation requests.
[0019] The first NAS server cluster generates an operation log for the first operation request based on the metadata operation log and the logical sequence number of the first operation request, including any one of the following:
[0020] The first NAS server receives the metadata operation log and the logical sequence number of the first operation request sent by the first metadata storage node, and sends the metadata operation log and the logical sequence number of the first operation request to the second NAS server. The second NAS server generates the operation log of the first operation request based on the metadata operation log and the logical sequence number of the first operation request. The second NAS server retrieves the metadata operation log from the first metadata storage node according to the logical sequence number of the first operation request, and generates the operation log of the first operation request based on the metadata operation log and the logical sequence number of the first operation request.
[0021] In one possible implementation, the first NAS server cluster includes multiple first NAS servers and at least one second NAS server, wherein the first NAS servers are used to receive and forward operation requests, and the second NAS servers are used to generate and forward operation logs of the operation requests.
[0022] The first NAS server cluster generates an operation log for the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request, including any one of the following:
[0023] The first NAS server receives the metadata operation log and the logical sequence number of the first operation request sent by the first metadata storage node. It then sends the metadata operation log, the logical sequence number of the first operation request, and the data carried in the first operation request to the second NAS server. The second NAS server generates the operation log of the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request. Based on the logical sequence number of the first operation request, the second NAS server retrieves the metadata operation log from the first metadata storage node. Based on the metadata operation log, it retrieves the data carried in the first operation request from the first data storage cluster. Finally, based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request, the second NAS server generates the operation log of the first operation request.
[0024] In one possible implementation, each first NAS server is associated with a second NAS server, each first NAS server is used to receive and forward a set of operation requests, and the generation and forwarding of operation logs for the set of operation requests received and forwarded by each first NAS server are performed by the second NAS server associated with the first NAS server.
[0025] In one possible implementation, the first NAS server and the second NAS server are the same NAS server.
[0026] Secondly, a data replication method in a distributed file system is provided, executed by a second storage cluster in the distributed file system, which also includes a first storage cluster. The method includes:
[0027] The system receives operation logs sent by the first storage cluster. The operation logs indicate the execution results and execution order of the first operation request corresponding to the semantics of the first operation request. The first operation request is an operation request for a file. The execution results include at least one of the execution results for the file data and the execution results for the file's metadata. Based on the operation logs, the system replays the first operation request.
[0028] In one possible implementation, the first operation request is replayed based on the operation log, including:
[0029] Based on the execution order of the first operation request, the operation log is stored; based on the operation log, the first operation request is replayed.
[0030] In one possible implementation, the operation log is stored based on the execution order of the first operation requests, including:
[0031] If the operation log includes a metadata operation log, the metadata operation log is stored based on the execution order of the first operation request, and the metadata operation log indicates the execution result of the file's metadata; if the operation log includes both a metadata operation log and a data operation log, the metadata operation log and the data operation log are stored based on the execution order of the first operation request, and the data operation log indicates the execution result of the file's data.
[0032] In one possible implementation, the second storage cluster includes a second NAS server cluster and a second data storage cluster. The second NAS server cluster is used to receive operation logs sent by the first storage cluster. The second data storage cluster includes multiple data storage nodes, and each data storage node includes at least one storage partition.
[0033] If the operation log includes a metadata operation log, then the metadata operation log is stored based on the execution order of the first operation request. The metadata operation log indicates the execution result of the file's metadata, including:
[0034] If the operation log includes metadata operation logs, the second NAS server cluster determines the first storage partition in the second data storage cluster based on the logical sequence number of the first operation request in the operation log, and stores the metadata operation log in the first storage partition. The logical sequence number is used to indicate the execution order of the first operation request.
[0035] In one possible implementation, the second storage cluster includes a second NAS server cluster and a second data storage cluster. The second NAS server cluster is used to receive operation logs sent by the first storage cluster. The second data storage cluster includes multiple data storage nodes, and each data storage node includes at least one storage partition.
[0036] If the operation log includes a metadata operation log and a data operation log, then the metadata operation log and the data operation log are stored based on the execution order of the first operation request. The data operation log indicates the execution result of the data on the file, including:
[0037] If the operation log includes metadata operation logs and data operation logs, the second NAS server cluster determines the second storage partition in the second data storage cluster based on the logical sequence number of the first operation request in the operation log, and stores the metadata operation log in the second storage partition; based on the metadata operation log, the second NAS server cluster divides the data carried by the first operation request in the data operation log into at least one data block, determines at least one third storage partition in the second storage cluster, and stores one data block and the logical sequence number of the first operation request in one third storage partition.
[0038] In one possible implementation, the second NAS server cluster includes multiple third NAS servers, each of which is used to execute a stored procedure for the operation logs of a set of operation requests, and the stored procedures for the operation logs of different sets of operation requests can be executed in parallel.
[0039] In one possible implementation, the first operation request is replayed based on the operation log, including:
[0040] According to the execution order of the first operation requests, the first operation requests are added to the waiting replay queue. If the first operation request is a file metadata operation request, when the first operation request meets the first concurrent replay condition, the first operation request is added to the metadata concurrent replay queue according to the execution order of the first operation requests, metadata replay is performed on the first operation request, and the first operation request is added from the metadata replay queue to the replay queue. The first concurrent replay condition indicates that the first operation request and the operation requests in the metadata concurrent replay queue are not mutually exclusive, and multiple operation requests in the metadata concurrent replay queue can be replayed concurrently. If the first operation request is a file data operation request... When the first operation request meets the second concurrent replay condition, the first operation request is added to the data concurrent replay queue according to the execution order of the first operation request, and data replay is performed on the first operation request. When the first operation request meets the first concurrent replay condition, the first operation request is added from the data concurrent replay queue to the metadata concurrent replay queue according to the execution order of the first operation request, and metadata replay is performed on the first operation request. The first operation request is added from the metadata replay queue to the replay queue. The second concurrent replay condition indicates that the first operation request and the operation requests in the data concurrent replay queue are not mutually exclusive, and multiple operation requests in the data concurrent replay queue can be replayed concurrently.
[0041] The condition that operation requests in the concurrent replay queue of metadata are not mutually exclusive means that the first operation request and the operation requests in the concurrent replay queue of metadata can be executed in parallel. For example, the inode information corresponding to the first operation request is different from the inode information of the operation requests in the concurrent replay queue of metadata; that is, the first operation request and the operation requests in the concurrent replay queue of metadata cannot be operation requests targeting the same metadata.
[0042] In one possible implementation, the second storage cluster also includes a second metadata storage node;
[0043] Perform metadata replay on the first operation request, including:
[0044] The second NAS server cluster retrieves metadata operation logs from the second data storage cluster and sends the metadata operation logs to the second metadata storage node; the second metadata storage node performs metadata replay of the first operation request based on the metadata operation logs.
[0045] In one possible implementation, the second storage cluster further includes a second metadata storage node, each storage partition includes a data storage partition and a metadata storage partition, and the data carried by the first operation request and the logical sequence number of the first operation request are stored in the metadata storage partition in the third storage partition;
[0046] Data replay for the first operation request includes:
[0047] The second NAS server cluster retrieves metadata operation logs from the second data storage cluster and sends the metadata operation logs to the second metadata storage node. Based on the metadata operation logs, the second metadata storage node determines multiple third storage partitions in the second storage cluster and writes the data carried by the first operation request in the third storage partition from the metadata storage partition of the third storage partition to the data storage partition of the third storage partition.
[0048] In one possible implementation, the second NAS server cluster includes multiple fourth NAS servers, each of which is used to perform a playback process for a set of operation requests, and the playback processes for different sets of operation requests can be executed in parallel.
[0049] In one possible implementation, the third NAS server and the fourth NAS server are the same NAS server.
[0050] Thirdly, a distributed file method is provided for use in a distributed file system, which includes a first storage cluster and a second storage cluster. The method includes:
[0051] The first storage cluster receives a first operation request for a file, executes the first operation request based on its semantics, and obtains an execution result corresponding to the semantics of the first operation request. The execution result includes at least one of the execution result for the file's data and the execution result for the file's metadata. The first storage cluster generates an operation log for the first operation request based on the execution result and the execution order of the first operation request. The first storage cluster sends the operation log to the second storage cluster. The second storage cluster receives the operation log sent by the first storage cluster and replays the first operation request based on the operation log.
[0052] In one possible implementation, the first storage cluster executes the first operation request based on the semantics of the first operation request, obtains an execution result corresponding to the semantics of the first operation request, and generates an operation log for the first operation request based on the execution result and the execution order of the first operation request, including:
[0053] If the first operation request is a request to operate on the file's metadata, then the first storage cluster executes the first operation request, obtains the execution result of the file's metadata, generates a metadata operation log based on the execution result of the file's metadata, and generates an operation log for the first operation request based on the metadata operation log and the execution order of the first operation request. If the first operation request is a request to operate on the file's data, then the first storage cluster executes the first operation request, obtains the execution result of the file's data and the execution result of the file's metadata, generates a data operation log based on the execution result of the file's metadata, generates a metadata operation log based on the execution result of the file's metadata, and generates an operation log for the first operation request based on the metadata operation log, the data operation log, and the execution order of the first operation request.
[0054] In one possible implementation, the first storage cluster includes a first NAS server cluster and a first metadata storage node;
[0055] If the first operation request is a request to operate on the file's metadata, then the first storage cluster executes the first operation request, obtains the execution result of the file's metadata, generates a metadata operation log based on the execution result of the file's metadata, and generates an operation log for the first operation request based on the metadata operation log and the execution order of the first operation request, including:
[0056] If the first operation request is an operation request on the metadata of a file, the first NAS server cluster sends the metadata targeted by the first operation request to the first metadata storage node; the first metadata storage node assigns a logical sequence number to the first operation request, executes the first operation request on the metadata, obtains the execution result of the metadata of the file, and generates a metadata operation log based on the execution result of the metadata of the file. The logical sequence number is used to indicate the execution order of the first operation request; the first NAS server generates an operation log for the first operation request based on the metadata operation log and the execution order of the first operation request.
[0057] In one possible implementation, the first storage cluster includes a first NAS server cluster, a first metadata storage node, and a first data storage cluster;
[0058] If the first operation request is a data operation request on a file, then the first storage cluster executes the first operation request, obtains the execution result of the data operation on the file and the execution result of the metadata operation on the file, generates a data operation log based on the execution result of the metadata operation on the file, generates a metadata operation log based on the execution result of the metadata operation log, the data operation log, and the execution order of the first operation request, and generates an operation log for the first operation request, including:
[0059] If the first operation request is a data operation request on a file, the first NAS server cluster sends the data carried by the first operation request to the first data storage cluster; the first data storage cluster executes the first operation request on the data carried by the first operation request, obtains the execution result on the file data, and generates a data operation log based on the execution result on the file data; the first NAS server cluster sends the metadata targeted by the first operation request to the first metadata storage node; the first metadata storage node assigns a logical sequence number to the first operation request, executes the first operation request on the metadata, obtains the execution result on the file metadata, and generates a metadata operation log based on the execution result on the file metadata, with the logical sequence number used to indicate the execution order of the first operation request; the first NAS server cluster generates an operation log for the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request.
[0060] In one possible implementation, the first NAS server cluster includes at least one first NAS server and at least one second NAS server, wherein the first NAS server is used to receive and forward operation requests, and the second NAS server is used to generate and forward operation logs of the operation requests.
[0061] The first NAS server cluster generates an operation log for the first operation request based on the metadata operation log and the logical sequence number of the first operation request, including any one of the following:
[0062] The first NAS server receives the metadata operation log and the logical sequence number of the first operation request sent by the first metadata storage node, and sends the metadata operation log and the logical sequence number of the first operation request to the second NAS server. The second NAS server generates the operation log of the first operation request based on the metadata operation log and the logical sequence number of the first operation request. The second NAS server retrieves the metadata operation log from the first metadata storage node according to the logical sequence number of the first operation request, and generates the operation log of the first operation request based on the metadata operation log and the logical sequence number of the first operation request.
[0063] In one possible implementation, the first NAS server cluster includes multiple first NAS servers and at least one second NAS server, wherein the first NAS servers are used to receive and forward operation requests, and the second NAS servers are used to generate and forward operation logs of the operation requests.
[0064] The first NAS server cluster generates an operation log for the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request, including any one of the following:
[0065] The first NAS server receives the metadata operation log and the logical sequence number of the first operation request sent by the first metadata storage node. It then sends the metadata operation log, the logical sequence number of the first operation request, and the data carried in the first operation request to the second NAS server. The second NAS server generates the operation log of the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request. Based on the logical sequence number of the first operation request, the second NAS server retrieves the metadata operation log from the first metadata storage node. Based on the metadata operation log, it retrieves the data carried in the first operation request from the first data storage cluster. Finally, based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request, the second NAS server generates the operation log of the first operation request.
[0066] In one possible implementation, each first NAS server is associated with a second NAS server, each first NAS server is used to receive and forward a set of operation requests, and the generation and forwarding of operation logs for the set of operation requests received and forwarded by each first NAS server are performed by the second NAS server associated with the first NAS server.
[0067] In one possible implementation, the first NAS server and the second NAS server are the same NAS server.
[0068] In one possible implementation, the second storage cluster replays the first operation request based on the operation log, including:
[0069] The second storage cluster stores the operation logs based on the execution order of the first operation request; and replays the first operation request based on the operation logs.
[0070] In one possible implementation, the second storage cluster stores the operation logs based on the execution order of the first operation requests, including:
[0071] If the operation log includes a metadata operation log, the second storage cluster stores the metadata operation log based on the execution order of the first operation request, and the metadata operation log indicates the execution result of the file's metadata; if the operation log includes both metadata operation logs and data operation logs, the metadata operation logs and data operation logs are stored based on the execution order of the first operation request, and the data operation log indicates the execution result of the file's data.
[0072] In one possible implementation, the second storage cluster includes a second NAS server cluster and a second data storage cluster. The second NAS server cluster is used to receive operation logs sent by the first storage cluster. The second data storage cluster includes multiple data storage nodes, and each data storage node includes at least one storage partition.
[0073] If the operation log includes a metadata operation log, the second storage cluster stores the metadata operation log based on the execution order of the first operation requests. The metadata operation log indicates the execution result of the file's metadata, including:
[0074] If the operation log includes metadata operation logs, the second NAS server cluster determines the first storage partition in the second data storage cluster based on the logical sequence number of the first operation request in the operation log, and stores the metadata operation log in the first storage partition. The logical sequence number is used to indicate the execution order of the first operation request.
[0075] In one possible implementation, the second storage cluster includes a second NAS server cluster and a second data storage cluster. The second NAS server cluster is used to receive operation logs sent by the first storage cluster. The second data storage cluster includes multiple data storage nodes, and each data storage node includes at least one storage partition.
[0076] If the operation log includes both metadata operation logs and data operation logs, the second storage cluster stores the metadata operation logs and data operation logs based on the execution order of the first operation requests. The data operation logs indicate the execution results of the data on the file, including:
[0077] If the operation log includes metadata operation logs and data operation logs, the second NAS server cluster determines the second storage partition in the second data storage cluster based on the logical sequence number of the first operation request in the operation log, and stores the metadata operation log in the second storage partition; based on the metadata operation log, the second NAS server cluster divides the data carried by the first operation request in the data operation log into at least one data block, determines at least one third storage partition in the second storage cluster, and stores one data block and the logical sequence number of the first operation request in one third storage partition.
[0078] In one possible implementation, the second NAS server cluster includes multiple third NAS servers, each of which is used to execute a stored procedure for the operation logs of a set of operation requests, and the stored procedures for the operation logs of different sets of operation requests can be executed in parallel.
[0079] In one possible implementation, the second storage cluster replays the first operation request based on the operation log, including:
[0080] According to the execution order of the first operation requests, the first operation requests are added to the waiting replay queue. If the first operation request is a file metadata operation request, when the first operation request meets the first concurrent replay condition, the first operation request is added to the metadata concurrent replay queue according to the execution order of the first operation requests, metadata replay is performed on the first operation request, and the first operation request is added from the metadata replay queue to the replay queue. The first concurrent replay condition indicates that the first operation request and the operation requests in the metadata concurrent replay queue are not mutually exclusive, and multiple operation requests in the metadata concurrent replay queue can be replayed concurrently. If the first operation request is a file data operation request... When the first operation request meets the second concurrent replay condition, the first operation request is added to the data concurrent replay queue according to the execution order of the first operation request, and data replay is performed on the first operation request. When the first operation request meets the first concurrent replay condition, the first operation request is added from the data concurrent replay queue to the metadata concurrent replay queue according to the execution order of the first operation request, and metadata replay is performed on the first operation request. The first operation request is added from the metadata replay queue to the replay queue. The second concurrent replay condition indicates that the first operation request and the operation requests in the data concurrent replay queue are not mutually exclusive, and multiple operation requests in the data concurrent replay queue can be replayed concurrently.
[0081] In one possible implementation, the second storage cluster also includes a second metadata storage node;
[0082] The second storage cluster performs metadata replay on the first operation request, including:
[0083] The second NAS server cluster retrieves metadata operation logs from the second data storage cluster and sends the metadata operation logs to the second metadata storage node; the second metadata storage node performs metadata replay of the first operation request based on the metadata operation logs.
[0084] In one possible implementation, the second storage cluster further includes a second metadata storage node, each storage partition includes a data storage partition and a metadata storage partition, and the data carried by the first operation request and the logical sequence number of the first operation request are stored in the metadata storage partition in the third storage partition;
[0085] The second storage cluster replays the data for the first operation request, including:
[0086] The second NAS server cluster retrieves metadata operation logs from the second data storage cluster and sends the metadata operation logs to the second metadata storage node. Based on the metadata operation logs, the second metadata storage node determines multiple third storage partitions in the second data storage cluster and writes the data carried by the first operation request in the third storage partition from the metadata storage partition of the third storage partition to the data storage partition of the third storage partition.
[0087] In one possible implementation, the second NAS server cluster includes multiple fourth NAS servers, each of which is used to perform a playback process for a set of operation requests, and the playback processes for different sets of operation requests can be executed in parallel.
[0088] In one possible implementation, the third NAS server and the fourth NAS server are the same NAS server.
[0089] Fourthly, a distributed file system is provided, which includes:
[0090] The first storage cluster is used for:
[0091] Receive a first operation request for a file, execute the first operation request based on the semantics of the first operation request, and obtain an execution result corresponding to the semantics of the first operation request. The execution result includes at least one of the execution result of the file's data and the execution result of the file's metadata.
[0092] Based on the execution results and the execution order of the first operation request, an operation log for the first operation request is generated;
[0093] Send operation logs to the second storage cluster;
[0094] The second storage cluster is used for:
[0095] Receive the operation log sent by the first storage cluster, and replay the first operation request based on the operation log.
[0096] In one possible implementation, the first storage cluster is used for:
[0097] If the first operation request is a request to operate on the file's metadata, then the first storage cluster executes the first operation request, obtains the execution result of the file's metadata, generates a metadata operation log based on the execution result of the file's metadata, and generates an operation log for the first operation request based on the metadata operation log and the execution order of the first operation request. If the first operation request is a request to operate on the file's data, then the first storage cluster executes the first operation request, obtains the execution result of the file's data and the execution result of the file's metadata, generates a data operation log based on the execution result of the file's metadata, generates a metadata operation log based on the execution result of the file's metadata, and generates an operation log for the first operation request based on the metadata operation log, the data operation log, and the execution order of the first operation request.
[0098] In one possible implementation, the first storage cluster includes a first NAS server cluster and a first metadata storage node;
[0099] The first NAS server cluster is used to: if the first operation request is an operation request on the metadata of a file, send the metadata targeted by the first operation request to the first metadata storage node;
[0100] The first metadata storage node is used to: assign a logical sequence number to the first operation request, execute the first operation request on the metadata, obtain the execution result of the metadata of the file, and generate a metadata operation log based on the execution result of the metadata of the file. The logical sequence number is used to indicate the execution order of the first operation request. The first NAS server generates the operation log of the first operation request based on the metadata operation log and the execution order of the first operation request.
[0101] In one possible implementation, the first storage cluster includes a first NAS server cluster, a first metadata storage node, and a first data storage cluster;
[0102] The first NAS server cluster is used for:
[0103] If the first operation request is a data operation request on a file, the data carried by the first operation request is sent to the first data storage cluster; the first data storage cluster executes the first operation request on the data carried by the first operation request, obtains the execution result on the file data, and generates a data operation log based on the execution result on the file data;
[0104] Send the metadata targeted by the first operation request to the first metadata storage node;
[0105] The first metadata storage node is used to: assign a logical sequence number to the first operation request, execute the first operation request on the metadata, obtain the execution result of the metadata of the file, and generate a metadata operation log based on the execution result of the metadata of the file. The logical sequence number is used to indicate the execution order of the first operation request.
[0106] The first NAS server cluster is also used to generate the operation log of the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request.
[0107] In one possible implementation, the first NAS server cluster includes any of the following:
[0108] The first NAS server is used to: receive the metadata operation log and the logical sequence number of the first operation request sent by the first metadata storage node, send the metadata operation log and the logical sequence number of the first operation request to the second NAS server, and generate the operation log of the first operation request based on the metadata operation log and the logical sequence number of the first operation request.
[0109] The second NAS server is used to: retrieve metadata operation logs from the first metadata storage node according to the logical sequence number of the first operation request, and generate operation logs for the first operation request based on the metadata operation logs and the logical sequence number of the first operation request.
[0110] In one possible implementation, the first NAS server cluster includes any of the following:
[0111] The first NAS server is used to: receive the metadata operation log and the logical sequence number of the first operation request sent by the first metadata storage node, send the metadata operation log, the logical sequence number of the first operation request and the data carried by the first operation request to the second NAS server, and the second NAS server generates the operation log of the first operation request based on the data operation log, the metadata operation log and the logical sequence number of the first operation request.
[0112] The second NAS server is used to: obtain metadata operation logs from the first metadata storage node based on the logical sequence number of the first operation request; obtain the data carried by the first operation request from the first data storage cluster based on the metadata operation logs; and generate operation logs for the first operation request based on the data operation logs, the metadata operation logs, and the logical sequence number of the first operation request.
[0113] In one possible implementation, each first NAS server is associated with a second NAS server, each first NAS server is used to receive and forward a set of operation requests, and the generation and forwarding of operation logs for the set of operation requests received and forwarded by each first NAS server are performed by the second NAS server associated with the first NAS server.
[0114] In one possible implementation, the first NAS server and the second NAS server are the same NAS server.
[0115] In one possible implementation, the second storage cluster is used for:
[0116] The second storage cluster stores the operation logs based on the execution order of the first operation request;
[0117] Based on the operation log, the first operation request is replayed.
[0118] In one possible implementation, the second storage cluster is used for:
[0119] If the operation log includes a metadata operation log, the metadata operation log is stored based on the execution order of the first operation request. The metadata operation log indicates the execution result of the file's metadata.
[0120] If the operation log includes a metadata operation log and a data operation log, then the metadata operation log and the data operation log are stored based on the execution order of the first operation request, and the data operation log indicates the execution result of the data on the file.
[0121] In one possible implementation, the second storage cluster includes a second NAS server cluster and a second data storage cluster. The second NAS server cluster is used to receive operation logs sent by the first storage cluster. The second data storage cluster includes multiple data storage nodes, and each data storage node includes at least one storage partition.
[0122] The second NAS server cluster is used to: if the operation log includes metadata operation logs, determine the first storage partition in the second data storage cluster based on the logical sequence number of the first operation request in the operation log, and store the metadata operation log in the first storage partition. The logical sequence number is used to indicate the execution order of the first operation request.
[0123] In one possible implementation, the second storage cluster includes a second NAS server cluster and a second data storage cluster. The second NAS server cluster is used to receive operation logs sent by the first storage cluster. The second data storage cluster includes multiple data storage nodes, and each data storage node includes at least one storage partition.
[0124] The second NAS server cluster is used for:
[0125] If the operation log includes metadata operation log and data operation log, then based on the logical sequence number of the first operation request in the operation log, the second storage partition in the second data storage cluster is determined, and the metadata operation log is stored in the second storage partition.
[0126] Based on the metadata operation log, the data carried by the first operation request in the data operation log is divided into at least one data block, at least one third storage partition in the second storage cluster is determined, and a data block and the logical sequence number of the first operation request are stored in a third storage partition.
[0127] In one possible implementation, the second NAS server cluster includes multiple third NAS servers, each of which is used to execute a stored procedure for the operation logs of a set of operation requests, and the stored procedures for the operation logs of different sets of operation requests can be executed in parallel.
[0128] In one possible implementation, the second storage cluster is used for:
[0129] According to the execution order of the first operation requests, the first operation requests are added to the waiting replay queue. If the first operation request is a file metadata operation request, when the first operation request meets the first concurrent replay condition, the first operation request is added to the metadata concurrent replay queue according to the execution order of the first operation requests, metadata replay is performed on the first operation request, and the first operation request is added from the metadata replay queue to the replay queue. The first concurrent replay condition indicates that the first operation request and the operation requests in the metadata concurrent replay queue are not mutually exclusive, and multiple operation requests in the metadata concurrent replay queue can be replayed concurrently. If the first operation request is a file data operation request... When the first operation request meets the second concurrent replay condition, the first operation request is added to the data concurrent replay queue according to the execution order of the first operation request, and data replay is performed on the first operation request. When the first operation request meets the first concurrent replay condition, the first operation request is added from the data concurrent replay queue to the metadata concurrent replay queue according to the execution order of the first operation request, and metadata replay is performed on the first operation request. The first operation request is added from the metadata replay queue to the replay queue. The second concurrent replay condition indicates that the first operation request and the operation requests in the data concurrent replay queue are not mutually exclusive, and multiple operation requests in the data concurrent replay queue can be replayed concurrently.
[0130] In one possible implementation, the second storage cluster also includes a second metadata storage node;
[0131] The second NAS server cluster is used to: retrieve metadata operation logs from the second data storage cluster and send the metadata operation logs to the second metadata storage node;
[0132] The second metadata storage node is used to: replay the metadata of the first operation request based on the metadata operation log.
[0133] In one possible implementation, the second storage cluster further includes a second metadata storage node, each storage partition includes a data storage partition and a metadata storage partition, and the data carried by the first operation request and the logical sequence number of the first operation request are stored in the metadata storage partition in the third storage partition;
[0134] The second NAS server cluster is used to: retrieve metadata operation logs from the second data storage cluster and send the metadata operation logs to the second metadata storage node;
[0135] The second metadata storage node is used to: determine multiple third storage partitions in the second storage cluster based on the metadata operation log, and write the data carried by the first operation request in the third storage partition from the metadata storage partition of the third storage partition to the data storage partition of the third storage partition.
[0136] In one possible implementation, the second NAS server cluster includes multiple fourth NAS servers, each of which is used to perform a playback process for a set of operation requests, and the playback processes for different sets of operation requests can be executed in parallel.
[0137] In one possible implementation, the third NAS server and the fourth NAS server are the same NAS server.
[0138] Fifthly, a data replication method apparatus for a distributed file system is provided, applied to a first storage cluster of the distributed file system. The apparatus includes at least one functional module for performing the data replication method for the distributed file system provided by the first aspect or any possible implementation thereof.
[0139] In a sixth aspect, a data replication method apparatus for a distributed file system is provided, applied to a second storage cluster of the distributed file system. The apparatus includes at least one functional module for performing the data replication method for the distributed file system provided in the second aspect or any possible implementation thereof.
[0140] In a seventh aspect, a storage node is provided, the storage node including a memory and a processor;
[0141] The processor is used to execute instructions stored in the memory to cause the storage node to perform a data replication method of a distributed file system as provided by any of the possible implementations of the first or second aspect described above.
[0142] Eighthly, a storage cluster is provided, the storage cluster including at least one storage node, each of the storage nodes including a memory and a processor;
[0143] The processor of the at least one storage node is used to execute instructions stored in the memory of the at least one storage node, so that the storage cluster performs the data replication method of the distributed file system provided by any of the possible implementations of the first or second aspect described above.
[0144] In a ninth aspect, a computer program product containing instructions is provided, which, when executed by a storage node, causes the storage node to perform a data replication method of a distributed file system as provided in the first aspect or any possible implementation thereof, or, when executed by a storage cluster, causes the storage cluster to perform a data replication method of a distributed file system as provided in the second aspect or any possible implementation thereof.
[0145] In a tenth aspect, a computer-readable storage medium is provided, comprising computer program instructions that, when executed by a storage node, enable the storage node to perform a data replication method of a distributed file system as provided in the first aspect or any possible implementation thereof, or, when executed by a storage cluster, enable the storage cluster to perform a data replication method of a distributed file system as provided in the second aspect or any possible implementation thereof.
[0146] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0147] Figure 1 is a schematic diagram of a distributed file system provided in an embodiment of this application;
[0148] Figure 2 is a schematic diagram of a distributed file system provided in an embodiment of this application;
[0149] Figure 3 is a flowchart of a data replication method for a distributed file system provided in an embodiment of this application;
[0150] Figure 4 is a flowchart of copying an operation log according to an embodiment of this application;
[0151] Figure 5 is a schematic diagram of a concurrent playback window provided in an embodiment of this application;
[0152] Figure 6 is a flowchart of a data replication method for a distributed file system provided in an embodiment of this application;
[0153] Figure 7 is a schematic diagram of a playback process provided in an embodiment of this application;
[0154] Figure 8 is a schematic diagram of the structure of a data replication device in a distributed file system provided in an embodiment of this application;
[0155] Figure 9 is a schematic diagram of the structure of a data replication device in a distributed file system provided in an embodiment of this application;
[0156] Figure 10 is a schematic diagram of the structure of a storage node provided in an embodiment of this application;
[0157] Figure 11 is a schematic diagram of a storage cluster provided in an embodiment of this application;
[0158] Figure 12 is a schematic diagram of a possible implementation of a storage cluster provided in an embodiment of this application. Detailed Implementation
[0159] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0160] First, let's introduce some of the technical terms used in this application.
[0161] File System: The file system provides users with a multi-level structure called a directory tree for managing files. In this tree structure, each file or directory contains metadata information. Metadata operations typically involve two main steps: path decoding and metadata processing. When a user accesses a file by pathname, the metadata is accessed as follows: First, the file system performs path decoding, locates the target file, and checks if the user has the correct permissions; second, the file system performs metadata operations, atomically updating the metadata of the accessed file.
[0162] Distributed file systems (NAS) distribute data across multiple storage nodes. The concurrency capabilities of these nodes improve system access performance, while the overall system capacity is increased through the capacity of the nodes. A typical NAS consists of multiple NAS servers, multiple data storage nodes, and at least one metadata storage node. The file writing process involves first dividing the data into multiple data blocks, writing these blocks to the data storage nodes, and then updating the metadata on the metadata storage node. The file reading process involves first reading the metadata information from the metadata storage node to obtain the file's index node (inode) information and the storage location of the data blocks, and then reading the data from the data storage node.
[0163] The implementation environment of the embodiments of this application is described below.
[0164] Figure 1 is a schematic diagram of a distributed file system provided in an embodiment of this application. As shown in Figure 1, the distributed file system includes a first storage cluster 101 and a second storage cluster 102, and data is replicated between the first storage cluster and the second storage cluster based on a replication link. It should be noted that Figure 1 only shows two storage clusters in the distributed file system, and the first storage cluster 101 is used as the primary storage cluster and the second storage cluster 102 as the backup storage cluster for example. However, the number of storage clusters in the distributed file system is not limited, and the number of storage clusters can be greater than 2. When the number of storage clusters is greater than 2, the number of primary and backup storage clusters is not limited. For example, the primary and backup storage clusters can be one-to-many or many-to-many, and this embodiment of the application does not limit this.
[0165] The first storage cluster 101 includes a first NAS server cluster 1011, a first metadata storage node 1012, and a first data storage cluster 1013.
[0166] The first NAS server cluster 1011 includes multiple NAS servers. These NAS servers can be Windows Server Message Block (SMB) servers, or Linux or Unix Network File System (NFS) servers; this embodiment does not limit the specific type. The first function of the first NAS server cluster is to receive file operation requests sent by clients and, based on the semantics of the operation requests, generate at least one of a metadata request and a data request. The metadata request is used to interact with the first metadata storage node 1012 to generate a metadata operation log; the data request is used to interact with the first data storage cluster 1013 to generate a data operation log. The semantics of the operation request indicate whether the operation request is a request for file metadata or a request for file data. Requests for file metadata include operations such as creating and deleting files, while requests for file data include operations such as writing data. The second function of the first NAS server cluster is to generate operation logs for the operation requests and send these logs to the second storage cluster via a replication link, thereby triggering the second storage cluster to synchronize the operation requests, ensuring data consistency between the primary and backup ends. As shown in Figure 1, the first NAS server cluster 1011 includes multiple first NAS servers and one second NAS server. The first NAS servers are used to perform a first function of the first NAS server cluster 1011, and the second NAS server is used to perform a second function of the first NAS server cluster 1011. By assigning different functions of the first NAS server cluster 1011 to different NAS servers, the processing of different operation requests and the sending of operation logs can be performed synchronously, thereby improving the performance of the first NAS server cluster. In some embodiments, the first function is implemented by a processing module, and the second function is implemented by a replication module. That is, the processing module is deployed on the first NAS server, and the second replication module is deployed on the second NAS server. This application embodiment does not limit this.
[0167] In some embodiments, as shown in FIG2, FIG2 is a schematic diagram of a distributed file system provided in an embodiment of the present application. The first NAS server cluster includes multiple first NAS servers and multiple second NAS servers. One first NAS server and one second NAS server are bound together to form multiple NAS server groups. Each NAS server group is responsible for a set of operation requests, so that the first NAS server cluster can concurrently process multiple sets of operation requests, thereby improving the throughput of the first NAS server for operation requests and the data replication efficiency.
[0168] In some embodiments, the first NAS server and the second NAS server are the same NAS server, that is, the above two functions of the first storage cluster are provided on a single NAS server, and different NAS servers in the first NAS server cluster concurrently process multiple sets of operation requests.
[0169] The first metadata storage node 1012 can be a metadata server (MDS). This first metadata storage node 1012 is used to process metadata requests, store metadata, and metadata operation logs. It assigns a globally unique logical sequence number (LSN) to each operation request, which indicates the execution order of the operation requests. The first metadata storage node 1012 includes a metadata area and a metadata operation log area. The metadata area stores metadata, and the metadata operation log area stores metadata operation logs. In some embodiments, the first metadata storage node 1012 uses a key-value log (KVLOG) format to store the metadata operation logs in the metadata operation log area, thereby improving the readability and maintainability of the metadata operation logs. It should be noted that Figure 1 uses one first metadata storage node as an example. The number of first metadata storage nodes in the first storage cluster can be greater, and this embodiment does not limit the number of first metadata storage nodes in the first storage cluster.
[0170] The first data storage cluster 1013 includes multiple data storage nodes, each of which includes at least one storage partition (PT). Figure 1 illustrates this with an example of the first data storage cluster having M storage partitions, where M is an integer greater than 1. This embodiment does not limit the number of storage partitions in the first data storage cluster. Each storage partition includes a log storage partition and a data storage partition. The log storage partition stores data operation logs, and the data storage partition stores data. In some embodiments, the data operation log is a write-ahead log (WAL). When processing a data request, the data storage node does not directly write the data carried by the data request to the data storage partition. Instead, it writes the data to a file called the WAL in the log storage partition, and then writes the data in the WAL from the log storage partition to the data storage partition.
[0171] In this system, storage partitions store data at the physical block level. Therefore, before writing the data carried by an operation request to a storage partition, the data needs to be divided into at least one data block to fill one physical block of the storage partition. This data block is then stored in different storage partitions. For example, the first data storage cluster includes three storage partitions (ACs), and each physical block has a capacity of 4 megabytes (MB). First, upon receiving operation request Q1, which carries 10MB of data, the data is divided into three data blocks: 4MB, 4MB, and 2MB. These three data blocks are then stored in storage partition AC. Next, upon receiving operation request Q2, which carries 5MB of data, the data is divided into two data blocks: 2MB and 3MB. The 2MB data block is written to storage partition C to fill one physical block in partition C. The 3MB data block is written to storage partition A, and so on, following the order of storage partitions ACA. Once one physical block in one storage partition is filled, the next physical block in the next storage partition is written.
[0172] The second storage cluster 102 includes a second NAS server cluster 1021, a second metadata storage node 1022, and a second data storage cluster 1023.
[0173] The second NAS server cluster 1021 includes multiple NAS servers. These NAS servers can be Windows Server Message Block (SMB) servers, or Linux or Unix Network File System (NFS) servers; this embodiment does not limit the specific type. The first function of the second NAS server cluster is to receive operation logs sent by the first storage cluster. The second function is to determine the storage location of the operation logs in the second data storage cluster 1023 and store the operation logs in the corresponding storage location. The third function is to retrieve the operation logs of the operation requests from the second storage cluster 1023 according to the execution order of the operation requests and send the operation logs to the second metadata storage node 1022 for playback. As shown in Figure 1, the second NAS server cluster 1021 includes one third NAS server and multiple fourth NAS servers. The third NAS server performs the first function of the second NAS server cluster 1021, and the fourth NAS servers perform the second and third functions of the second NAS server cluster 1021. In some embodiments, as shown in Figure 2, the second NAS server cluster includes multiple third NAS servers and multiple fourth NAS servers. One third NAS server and one fourth NAS server are bound together to form multiple NAS server groups. Each NAS server group is responsible for a set of operation requests. Each NAS server group is associated with a set of storage partitions in the second storage cluster (that is, the operation logs of the set of operation requests handled by the NAS server group are stored in the set of storage partitions associated with the NAS server group). Thus, multiple NAS server groups can concurrently process the operation log storage and replay process of multiple sets of operation requests, thereby improving the data replication performance of the distributed file system.
[0174] The second data storage cluster 1023 includes multiple data storage nodes, each of which includes at least one storage partition. Figure 1 illustrates this using a first data storage cluster with M storage partitions as an example, where M is an integer greater than 1. This embodiment does not limit the number of storage partitions in the second data storage cluster. Each storage partition includes a log storage partition and a data storage partition. The log storage partition stores operation logs to be replayed, and the data storage partition stores data. In some embodiments, the operation logs to be replayed are stored in the log storage partition in redo log (RDL) format. The process of writing the data carried by the operation request into the storage partition of the second data storage cluster 1023 is the same as that of the first data storage cluster and will not be described again.
[0175] The second metadata storage node 1022 is used to control the playback of operation requests and store metadata. The first metadata storage node 1012 includes a metadata area for storing metadata. It should be noted that Figure 1 is an example using one second metadata storage node, and this embodiment does not limit the number of first metadata storage nodes in the first storage cluster.
[0176] The replication link between the first storage cluster 101 and the second storage cluster 102 can be an Ethernet link, a Fibre Channel, or an InfiniBand, etc., and this application embodiment does not limit this.
[0177] This application provides a data replication method for a distributed file system. In this method, a first storage cluster executes a first operation request based on the semantics of that first operation request for a file and generates an operation log corresponding to the semantics of the first operation request. This operation log indicates at least one of the execution results for the file's data and the execution results for the file's metadata. The first storage cluster sends the operation log to a second storage cluster. The second storage cluster then replays the first operation request based on the operation log. In this method, generating the operation log based on the semantics of the file operation request and replaying the operation request based on the operation log achieves data replication at the file semantic layer. Compared to implementing data replication at the block layer, this method does not rely on the replication capabilities of the block layer and can use a single data replication scheme to achieve both synchronous and asynchronous replication capabilities. Therefore, it eliminates the need for different block layer schemes to implement synchronous and asynchronous replication, thereby reducing the complexity of developing and maintaining the distributed file system and improving data replication performance.
[0178] The semantics of the operation request indicate whether it is a request to manipulate the file's metadata or the file's data. Requests to manipulate the file's metadata include file creation and deletion requests, while requests to manipulate the file's data include data writing requests. The first storage cluster determines the semantics of the first operation request using different methods depending on the actual situation. For example, the first storage cluster can determine the semantics of the first operation request based on the request's Uniform Resource Locator (URL), request method, request header information, request parameters, and request body content. This embodiment of the application does not limit this approach.
[0179] The following section will further elaborate on the above method by taking the first operation request as an operation request on the file's metadata and the first operation request on the file's data as examples.
[0180] First, based on the distributed file system shown in Figure 1, the process of the above method is described using the first operation request as an example of an operation request on the file's metadata. Figure 3 is a flowchart of a data replication method for a distributed file system provided in an embodiment of this application. This method is executed by the distributed file system, which includes a first storage cluster and a second storage cluster. The first storage cluster includes a first NAS server cluster, a first metadata storage node, and a first data storage cluster. The second storage cluster includes a second NAS server cluster, a second metadata storage node, and a second data storage cluster. The method includes the following steps 301 to 311.
[0181] Step 301: The first storage cluster receives the first operation request for the file sent by the client.
[0182] In this embodiment, the first NAS server cluster receives a first operation request for a file sent by the client. In some embodiments, any NAS server in the first NAS server cluster can receive the first operation request. In other embodiments, only a portion of the NAS servers in the first NAS server cluster can receive the first operation request; this embodiment does not limit the scope of the application.
[0183] Step 302: If the first operation request is an operation request on the file's metadata, then the first storage cluster executes the first operation request and obtains the execution result on the file's metadata.
[0184] In this process, the first NAS server cluster determines that the first operation request is a request to manipulate the file's metadata, generates a metadata request, and sends this request to the first metadata storage node in the first storage cluster. The first metadata storage node executes the metadata request and obtains the execution result of the file's metadata. Taking a file creation operation request as an example, the execution result of the file's metadata includes information such as the file creation time, the file owner, and permissions. In some embodiments, the first metadata storage node stores the execution result of the file's metadata in its metadata area.
[0185] Step 303: The first storage cluster assigns a logical sequence number to the first operation request. The logical sequence number of the first operation request is used to indicate the execution order of the first operation request.
[0186] In this process, the first metadata storage node assigns a globally unique logical sequence number to the first operation request. When the second storage cluster subsequently replays the first operation request, it executes the operation requests in the order of their logical sequence numbers. This ensures that the second storage cluster can replay the operation requests in the correct order, thereby guaranteeing the correctness of data replication and improving the data reliability of the distributed file system.
[0187] Step 304: The first storage cluster generates a metadata operation log based on the execution results of the file's metadata.
[0188] In some embodiments, the first metadata storage node generates a metadata operation log based on the KVLOG format, and stores the metadata operation log in the metadata operation log area of the first metadata storage node.
[0189] Step 305: The first storage cluster generates the operation log of the first operation request based on the metadata operation log and the execution order of the first operation request.
[0190] The first NAS server cluster includes at least one first NAS server and at least one second NAS server. The first NAS server is used to receive and forward operation requests, and the second NAS server is used to generate and forward operation logs for the operation requests. By assigning different functions of the first NAS server cluster 1011 to different NAS servers, the processing of different operation requests and the sending of operation logs can be executed synchronously, thereby improving the performance of the first NAS server cluster.
[0191] The process of the first storage cluster generating the operation log of the first operation request includes: the first NAS server receiving the metadata operation log and the logical sequence number of the first operation request sent by the first metadata storage node, sending the metadata operation log and the logical sequence number of the first operation request to the second NAS server, and the second NAS server generating the operation log of the first operation request based on the metadata operation log and the logical sequence number of the first operation request.
[0192] In some embodiments, each first NAS server is associated with a second NAS server. Each first NAS server receives and forwards a set of operation requests. The generation and forwarding of operation logs for the set of operation requests received and forwarded by each first NAS server are performed by the second NAS server associated with the first NAS server. In the above embodiments, by grouping operation requests and having the generation and forwarding of operation logs for each set of operation requests performed by a group of NAS servers, the generation and forwarding of operation logs for different sets of operation requests can be performed in parallel among different groups of NAS servers, thereby improving the efficiency of data replication and the concurrency performance of the distributed file system.
[0193] In some embodiments, the first NAS server and the second NAS server are the same NAS server. This application does not limit this aspect.
[0194] Step 306: The first storage cluster sends the operation log to the second storage cluster.
[0195] The second NAS server sends the operation log to the second storage cluster.
[0196] Step 307: The second storage cluster receives the operation log sent by the first storage cluster.
[0197] In this embodiment, the second NAS server cluster within the second storage cluster receives the operation log. In some embodiments, any NAS server in the second NAS server cluster can receive the first operation request sent by the first storage cluster. In other embodiments, a third NAS server in the second NAS server cluster can receive the second operation request sent by the first storage cluster; that is, some NAS servers in the second NAS server cluster can receive the second operation request. This application does not limit this aspect.
[0198] Step 308: The second storage cluster stores the operation log based on the execution order of the first operation request.
[0199] The second storage cluster stores metadata operation logs based on the execution order of the first operation request. These metadata operation logs indicate the execution results of the metadata of the file.
[0200] The second storage cluster includes a second NAS server cluster and a second data storage cluster. The second NAS server cluster is used to receive operation logs sent by the first storage cluster. The second data storage cluster includes multiple data storage nodes, and each data storage node includes at least one storage partition.
[0201] In some embodiments, each storage partition is associated with an operation log for storing a set of operation requests, that is, each storage partition is associated with a set of logical sequence numbers. The process of the second storage cluster storing the operation log includes: the second NAS server cluster determining the first storage partition in the second data storage cluster based on the logical sequence number of the first operation request in the operation log, and storing the metadata operation log in the first storage partition. Each NAS server in the second NAS server cluster is bound to at least one set of logical sequence numbers, and each NAS server is also bound to a set of storage partitions. The second NAS server cluster determines the NAS server associated with the logical sequence number based on the logical sequence number of the first operation request, and then determines the storage partition bound to that NAS server. In some embodiments, after the second storage cluster receives the operation log sent by the first storage cluster, the first storage cluster randomly allocates the storage partition corresponding to the logical sequence number; this application embodiment does not limit this.
[0202] Step 309: The second storage cluster sends a replication completion message for the operation log to the first storage cluster.
[0203] In steps 308 and 309 above, the second storage cluster returns a replication completion message to the first storage node immediately after storing the operation log, without waiting for the operation log to be synchronized. Under the synchronous replication method, the waiting time of the first storage cluster can be reduced.
[0204] Step 310: After receiving the replication completion message, the first storage cluster returns an operation completion message for the first operation request to the client.
[0205] It should be noted that steps 305 to 310 above are illustrated using synchronous replication as an example. In some embodiments, asynchronous replication is used. That is, after the first storage cluster sends the operation log to the second storage cluster, it immediately returns an operation completion message for the first operation request to the client. In asynchronous replication, the process of the first storage cluster generating the operation log for the first operation request includes: the second NAS server retrieves the metadata operation log from the first metadata storage node according to the logical sequence number of the first operation request, and generates the operation log for the first operation request based on the metadata operation log and the logical sequence number of the first operation request. The second NAS server then sends this operation log to the second storage cluster. Specifically, the second NAS server retrieves the metadata operation log from the first metadata storage node using a round-robin method according to the logical sequence number.
[0206] Step 311: The second storage cluster replays the first operation request based on the operation log.
[0207] The process by which the second storage cluster replays the first operation request based on the operation log includes: adding the first operation request to a waiting replay queue according to its execution order; when the first operation request meets a first concurrent replay condition, adding the first operation request to a metadata concurrent replay queue according to its execution order, performing metadata replay on the first operation request, and adding the first operation request from the metadata replay queue to the replay queue. The first concurrent replay condition indicates that the first operation request and the operation requests in the metadata concurrent replay queue are not mutually exclusive, and multiple operation requests in the metadata concurrent replay queue can be replayed concurrently. "Not mutually exclusive" means that the first operation request and the operation requests in the metadata concurrent replay queue can be executed in parallel. For example, the inode information corresponding to the first operation request is different from the inode information of the operation requests in the metadata concurrent replay queue; that is, the first operation request and the operation requests in the metadata concurrent replay queue cannot be operation requests targeting the same metadata.
[0208] The second storage cluster performs metadata replay on the first operation request, which includes: the second NAS server cluster obtains metadata operation logs from the second data storage cluster and sends the metadata operation logs to the second metadata storage node; the second metadata storage node performs metadata replay on the first operation request based on the metadata operation logs.
[0209] In the above method, the first storage cluster executes the first operation request based on its semantics and generates an operation log corresponding to the semantics of the first operation request; the first storage cluster sends the operation log to the second storage cluster; and the second storage cluster replays the first operation request based on the operation log. In this method, generating the operation log based on the semantics of the operation request and replaying the operation request based on the operation log achieves data replication at the file semantic layer. Compared to implementing data replication at the block layer, this method does not rely on the replication capabilities of the block layer and can use a single data replication scheme to achieve both synchronous and asynchronous replication capabilities. Therefore, it eliminates the need for different block layer schemes to implement synchronous and asynchronous replication, thereby reducing the complexity of developing and maintaining the distributed file system and improving data replication performance. Furthermore, using a single scheme to implement synchronous and asynchronous replication reduces system complexity.
[0210] The following example illustrates the process shown in steps 301 to 311 above, using the first operation request as a file creation request.
[0211] Example 1: Based on the distributed file system shown in Figure 1, the process of the first storage cluster executing a create operation request includes: After receiving the create request, the master NAS Server (i.e., the first NAS server cluster) first processes it using the processing module (deployed on the first NAS server in the first NAS server cluster), and then sends the relevant metadata to the master MDS (i.e., the first metadata storage node); the master MDS allocates a globally unique LSN for this create operation request to ensure that the operation can be replayed in the correct order on the backup, then writes the metadata to the metadata area, generates a metadata operation log for the result, writes it to the KVLOG, and after execution, returns the LSN and KVLOG to the NAS Server node; after receiving the LSN and KVLOG corresponding to the create operation request, the NAS Server: if it is synchronous replication, NAS... The server forwards the LSN and KVLOG to the replication module (deployed on the second NAS server in the first NAS server cluster). The replication module assembles the complete log (i.e., the operation log of the first operation request) and sends it to the backup (i.e., the second storage cluster) for processing via the replication link. Once the backup has finished processing, the create request ends. If it is asynchronous replication, the create operation request ends here. The replication module itself reads the KVLOG from the primary MDS node (e.g., it can obtain it by polling in LSN order) and assembles it into a complete log before sending it to the backup for processing via the replication link.
[0212] Example 2: Based on the distributed file system shown in Figure 2, the process of the first storage cluster executing a create operation includes: After receiving the create operation request, the master NAS Server (i.e., the first NAS server cluster) first selects a shard (group) within the current NAS Server node. The processing module (deployed on the first NAS server in the first NAS server cluster) processes the create operation request and sends the relevant metadata and shard information to the master MDS (i.e., the first metadata storage node). The master MDS allocates a globally unique LSN for this create operation to ensure that the operation can be replayed in the correct order on the backup. Then, it writes the metadata to the metadata area, generates a metadata operation log, writes it to the KVLOG according to the shard, and after execution, returns the LSN and KVLOG to the NAS Server node. After receiving the LSN and KVLOG, if it is synchronous replication, the NAS Server... The server forwards the LSN and KVLOG to the replication module (deployed on the second NAS server in the first NAS server cluster). The replication module assembles the complete log (i.e., the operation log of the first operation request) and sends it to the backup (i.e., the second storage cluster) for processing via the replication link. Once the backup has finished processing, the create request ends. If it is asynchronous replication, the create operation request ends. The replication modules of each shard in all NAS servers read their own LSN and KVLOG from the MDS in shard order, assemble them into a complete log, and send it to the backup for processing.
[0213] As can be seen from the above examples, compared with Example 1, this embodiment maps LSN to multiple Shards according to certain rules. In MDS, metadata operation logs are stored at the Shard granularity. In NAS Server, each replication module reads metadata operation logs from MDS at the Shard granularity and replicates them concurrently to the backup for processing, which significantly improves the efficiency and throughput of log replication.
[0214] Example 3: Building upon Example 2, in the second storage cluster, LSNs are mapped to N logical shards according to certain rules, and these N logical shards are then mapped to M PTs. This ensures that each shard knows which PT its logs are stored in for subsequent playback processing (M and N are both positive integers). In the second storage cluster, the second NAS server cluster includes a replication module (deployed on the third NAS server within the second NAS server cluster), responsible for receiving replication messages sent by the first storage cluster and deserializing them into metadata operation logs and data. Additionally, the second NAS server cluster includes a log module (deployed on the fourth NAS server within the second NAS server cluster), responsible for calculating the assigned PT for metadata operation logs based on their LSNs and offsets relative to the first PT, thus achieving log distribution balance. The second data storage cluster includes a log module and a playback module, and divides the storage area into an RDL area (i.e., a metadata storage partition) and a data area (i.e., a data storage partition). The RDL area stores logs, while the data area stores the actual data of user files (i.e., the data carried by operation requests).
[0215] Referring to Figure 4, which is a flowchart of an operation log replication process provided in an embodiment of this application, taking a create operation request as an example, the replication process of the second storage cluster includes: the replication module of a node in the NAS Server cluster receives the replication message of the create operation request sent by the master and deserializes it into the metadata operation log of the create operation request; the log module obtains the identifier (ID) of the belonging shard by taking the LSN modulo the total number of shards of the metadata operation log of this create operation request, and then obtains the ID of the PT to which the metadata operation log belongs by taking the shard ID modulo the total number of PTs, and then sends the metadata operation log to the Space node (i.e., the data storage node) where the PT ID is located; after receiving the create metadata log, the Space node appends it to the redo log in the RDL area, and the backup replication process is completed.
[0216] Example 4: Based on Example 3, the process of replaying operation logs in the second storage cluster is as follows: The replay module of the NAS Server loads the metadata operation logs of its own shard into the replay module according to the LSN order, and forwards them to the MDS node to add them to the waiting replay queue (in Example 3, we have mapped the LSN to N shards according to certain rules, and each shard is bound to a replay module of a NAS Server, so that the replay module knows which LSNs it should be responsible for replaying); the MDS, as the control center of replay, after receiving the metadata logs and LSNs forwarded by the NAS Server, maintains two sliding windows in its internal replay module: the metadata concurrent replay window (equivalent to the metadata concurrent replay queue) and the data concurrent replay window (equivalent to the data concurrent replay queue), as shown in Figure 5. Figure 5 is a schematic diagram of a concurrent replay window provided in the embodiment of this application, which strictly arranges the replays in ascending order of LSN. In this process, LSNs first enter the data concurrent playback window (LSNs within the concurrent playback window must be non-exclusive; for example, different data operation requests targeting different memory address ranges cannot overlap, and metadata operation requests are skipped in the data concurrent playback window). Only after data playback is complete can LSNs enter the metadata concurrent playback window (LSNs within the concurrent playback window must also be mutually exclusive; for example, inodes cannot be identical). The size of the sliding window controls the degree of concurrent playback, and the non-exclusive mechanism for LSNs entering the window also ensures the correctness of the playback.
[0217] Below, based on the distributed file system shown in Figure 1, the process of the above method will be described using the first operation request as an example of an operation request on the metadata of a file. Figure 6 is a flowchart of a data replication method for a distributed file system provided in an embodiment of this application. The method is executed by the distributed file system, which includes a first storage cluster and a second storage cluster. The first storage cluster includes a first NAS server cluster, a first metadata storage node, and a first data storage cluster. The second storage cluster includes a second NAS server cluster, a second metadata storage node, and a second data storage cluster. The method includes the following steps 601 to 609.
[0218] Step 601: The first storage cluster receives the first operation request sent by the client.
[0219] In this embodiment, the first NAS server cluster receives a first operation request sent by the client. In some embodiments, any NAS server in the first NAS server cluster can receive the first operation request. In other embodiments, only the first NAS server in the first NAS server cluster can receive the first operation request, that is, only some NAS servers in the first NAS server cluster can receive the first operation request. This application does not limit this aspect.
[0220] Step 602: If the first operation request is a data operation request for the file, the first storage cluster executes the first operation request and obtains the execution result of the data operation on the file and the execution result of the metadata operation on the file.
[0221] The process by which the first storage cluster obtains the execution results of the file data and the file metadata includes: the first NAS server cluster sending the data carried by the first operation request to the first data storage cluster; the first data storage cluster executing the first operation request on the data carried by the first operation request to obtain the execution results of the file data, and generating a data operation log based on the execution results of the file data; the first NAS server cluster sending the metadata targeted by the first operation request to the first metadata storage node; the first metadata storage node assigning a logical sequence number to the first operation request, executing the first operation request on the metadata, and obtaining the execution results of the file metadata, wherein the logical sequence number is used to indicate the execution order of the first operation request.
[0222] Step 603: The first storage cluster generates an operation log for the first operation request based on the execution results of the file data, the execution results of the file metadata, and the execution order of the first operation request.
[0223] The process of the first storage cluster generating the operation log of the first operation request includes: the first NAS server cluster generating a metadata operation log based on the execution result of the file's metadata; and the first NAS server cluster generating the operation log of the first operation request based on the metadata operation log and the execution order of the first operation request.
[0224] In some embodiments, the first metadata storage node generates a metadata operation log based on the KVLOG format, and stores the metadata operation log in the metadata operation log area of the first metadata storage node.
[0225] The first NAS server cluster includes at least one first NAS server and at least one second NAS server. The first NAS server is used to receive and forward operation requests, and the second NAS server is used to generate and forward operation logs for the operation requests. The process of the first storage cluster generating the operation log for the first operation request includes: the first NAS server receiving the metadata operation log and the logical sequence number of the first operation request sent by the first metadata storage node; sending the metadata operation log, the logical sequence number of the first operation request, and the data carried in the first operation request to the second NAS server; and the second NAS server generating the operation log for the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request.
[0226] In some embodiments, each first NAS server is associated with a second NAS server. Each first NAS server receives and forwards a set of operation requests. The generation and forwarding of operation logs for the set of operation requests received and forwarded by each first NAS server are performed by the second NAS server associated with the first NAS server. In the above embodiments, by grouping operation requests and having the generation and forwarding of operation logs for each set of operation requests performed by a group of NAS servers, the generation and forwarding of operation logs for different sets of operation requests can be performed in parallel among different groups of NAS servers, thereby improving the efficiency of data replication and the concurrency performance of the distributed file system.
[0227] In some embodiments, the first NAS server and the second NAS server are the same NAS server. This application does not limit this aspect.
[0228] Step 604: The first storage cluster sends the operation log to the second storage cluster.
[0229] The second NAS server sends the operation log to the second storage cluster.
[0230] Step 605: The second storage cluster receives the operation log sent by the first storage cluster.
[0231] In this embodiment, the second NAS server cluster within the second storage cluster receives the operation log. In some embodiments, any NAS server in the second NAS server cluster can receive the first operation request sent by the first storage cluster. In other embodiments, a third NAS server in the second NAS server cluster can receive the second operation request sent by the first storage cluster; that is, some NAS servers in the second NAS server cluster can receive the second operation request. This application does not limit this aspect.
[0232] Step 606: The second storage cluster stores the operation log based on the execution order of the first operation request.
[0233] The second storage cluster stores the metadata operation log and data operation log in the operation log based on the execution order of the first operation request.
[0234] The second storage cluster includes a second NAS server cluster and a second data storage cluster. The second NAS server cluster receives operation logs sent by the first storage cluster. The second data storage cluster includes multiple data storage nodes, each of which includes at least one storage partition. In some embodiments, each storage partition is associated with a metadata operation log for storing a set of operation requests; that is, each storage partition is associated with a set of logical sequence numbers. The process of the second storage cluster storing the operation logs includes: the second NAS server cluster determining a second storage partition in the second data storage cluster based on the logical sequence number of a first operation request in the operation log, and storing the metadata operation log in the second storage partition; the second NAS server cluster dividing the data carried by the first operation request in the data operation log into at least one data block based on the metadata operation log, determining at least one third storage partition in the second storage cluster, and storing one data block and the logical sequence number of the first operation request in one third storage partition.
[0235] In this configuration, each NAS server in the second NAS server cluster is bound to at least one set of logical sequence numbers, and each NAS server is further bound to a set of storage partitions. The second NAS server cluster determines the NAS server associated with the logical sequence number of the first operation request, and then determines the storage partition bound to that NAS server. In some embodiments, after the second storage cluster receives the operation log sent by the first storage cluster, the first storage cluster randomly allocates the storage partition corresponding to the logical sequence number; this application embodiment does not limit this approach.
[0236] The first storage cluster, based on the principle of filling a physical block in a storage partition, needs to divide the data into at least one data block and store the at least one data block in different storage partitions. The first storage cluster stores the inode, offset, length and other information of the data in the metadata operation log. The second storage cluster, based on the metadata operation log, obtains the inode, offset, length and other information, and divides the data into at least one data block in the same way as the first storage node, thereby determining the storage partition where each data block is located.
[0237] Step 607: The second storage cluster sends a replication completion message for the operation log to the first storage cluster.
[0238] In steps 606 and 607 above, the second storage cluster returns a replication completion message to the first storage node immediately after storing the operation log, without waiting for the operation log to be synchronized. Under the synchronous replication method, the waiting time of the first storage cluster can be reduced.
[0239] Step 608: After receiving the replication complete message, the first storage cluster returns an operation complete message to the client for the first operation request.
[0240] It should be noted that steps 604 to 608 above are illustrated using synchronous replication as an example. In some embodiments, asynchronous replication is used. That is, after the first storage cluster sends the operation log to the second storage cluster, it immediately returns an operation completion message for the first operation request to the client. In asynchronous replication, the process of the first storage cluster generating the operation log for the first operation request includes: the second NAS server, based on the logical sequence number of the first operation request, obtains the metadata operation log from the first metadata storage node; based on the metadata operation log, obtains the data carried by the first operation request from the first data storage cluster; and based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request, generates the operation log for the first operation request. Specifically, the second NAS server obtains the metadata operation log from the first metadata storage node in a round-robin manner according to the logical sequence number.
[0241] Step 609: The second storage cluster replays the first operation request based on the operation log.
[0242] The process by which the second storage cluster replays the first operation request based on the operation log includes: when the first operation request meets the second concurrent replay condition, adding the first operation request to the data concurrent replay queue according to the execution order of the first operation request, performing data replay on the first operation request; when the first operation request meets the first concurrent replay condition, adding the first operation request from the data concurrent replay queue to the metadata concurrent replay queue according to the execution order of the first operation request, performing metadata replay on the first operation request, and adding the first operation request from the metadata replay queue to the replayed queue. The second concurrent replay condition indicates that the first operation request and the operation requests in the data concurrent replay queue are not mutually exclusive, and multiple operation requests in the data concurrent replay queue can be replayed concurrently. "Not mutually exclusive with the operation requests in the data concurrent replay queue" means that the first operation request and the operation requests in the data concurrent replay queue can be executed in parallel. For example, the data address corresponding to the first operation request is different from the data address corresponding to the operation request in the metadata concurrent replay queue. That is, the data address ranges corresponding to the first operation request and the operation requests in the data concurrent replay queue cannot overlap.
[0243] The second storage cluster performs metadata replay on the first operation request, which includes: the second NAS server cluster obtains metadata operation logs from the second data storage cluster and sends the metadata operation logs to the second metadata storage node; the second metadata storage node performs metadata replay on the first operation request based on the metadata operation logs.
[0244] The second storage cluster also includes a second metadata storage node. Each storage partition includes a data storage partition and a metadata storage partition. The data carried by the first operation request and the logical sequence number of the first operation request are stored in the metadata storage partition of the third storage partition. The second storage cluster performs data replay on the first operation request, including: the second NAS server cluster obtaining the metadata operation log from the second data storage cluster and sending the metadata operation log to the second metadata storage node; the second metadata storage node determining multiple third storage partitions in the second storage cluster based on the metadata operation log, and writing the data carried by the first operation request from the metadata storage partition of the third storage partition to the data storage partition of the third storage partition.
[0245] In the above method, the first storage cluster executes the first operation request based on its semantics and generates an operation log corresponding to the semantics of the first operation request; the first storage cluster sends the operation log to the second storage cluster; and the second storage cluster replays the first operation request based on the operation log. In this method, generating the operation log based on the semantics of the operation request and replaying the operation request based on the operation log achieves data replication at the file semantic layer. Compared to implementing data replication at the block layer, this method does not rely on the replication capabilities of the block layer and can use a single data replication scheme to achieve both synchronous and asynchronous replication capabilities. Therefore, it eliminates the need for different block layer schemes to implement synchronous and asynchronous replication, thereby reducing the complexity of developing and maintaining the distributed file system and improving data replication performance. Furthermore, using a single scheme to implement synchronous and asynchronous replication reduces system complexity.
[0246] The following example illustrates the process shown in steps 601 to 609 above, taking the first operation request as a data write request.
[0247] Example 5: After the master NAS Server (i.e., the first NAS server cluster) receives a write operation request, the processing module (deployed on the first NAS server in the first NAS server cluster) distributes the data according to the PT distribution rules, and then forwards the distributed data blocks to the corresponding PT's Space node (i.e., the data storage node in the first storage cluster). After receiving the data blocks, the master Space node appends the write write AL to the corresponding PT. This data request is processed at the Space node (the background process moves the data from the WAL to the data area to reduce operation latency). After the master processing module finishes writing the data blocks, it sends the metadata related to the write operation request to the master MDS node (i.e., the first metadata storage node in the first NAS server cluster). The master MDS node allocates a globally unique LSN for this write operation request and persists the KVLOG. After execution, it returns the LSN and KVLOG to the NAS Server node. After receiving the LSN and KVLOG corresponding to the write operation request, the NAS Server: If it is synchronous replication, NAS... The server forwards the LSN, KVLOG, and data to the replication module. The replication module (deployed on the first NAS server in the first NAS server cluster) assembles the complete log and sends it to the backup server for processing via the replication link. After processing, this Write request ends. If it is asynchronous replication, this Write request ends, and the replication module itself reads the KVLOG from the primary MDS node in LSN order (e.g., it can obtain it by polling in LSN order). Then, based on the KVLOG, it reads the WAL from the corresponding PT according to the same PT shuffling rules as in step 1, assembles it into a complete log (i.e., the operation log of the first operation request), and sends it to the backup server for processing via the replication link. Here, the complete log is the operation log, and the LSN indicates the execution order of the operation requests.
[0248] Example 6: Building upon Example 5, the replication module is bound to the processing module in the NAS Server of the master NAS cluster. All LSNs are divided into N logical Shards (groups), and each Shard is mapped to the replication module of a NAS Server node (each NAS Server will not have the same Shard), so that multiple nodes can replicate concurrently to improve replication performance. Based on the distributed file system shown in Figure 2, the process of executing a write operation request in the first storage cluster includes: After receiving the write request, the master NAS Server first selects a shard within the current NAS Server node, then shuffles the data according to the PT shuffling rules, and then forwards the shuffled data blocks to the corresponding PT's Space node; After receiving the data blocks, each node in the master Space appends the write AL to the corresponding PT, and the data request is processed at the Space node (the background process will move the data from the WAL to the data area to reduce operation latency); After the master processing module finishes writing the data blocks, it sends the operation-related metadata and corresponding shard information to the master MDS node; The master MDS node allocates a globally unique LSN for this write operation request and persists the KVLOG to the shard. After execution, it returns the LSN and KVLOG to the NAS Server node; After receiving the LSN and KVLOG corresponding to the write operation request, the NAS Server: If it is synchronous replication, NAS... The server processing module directly assembles the LSN, KVLOG, and data into a complete log and sends it to the backup server for processing via the replication link. Once processing is complete, the write operation request ends. In asynchronous replication, the write operation request ends here. Each shard's replication module in all NAS Servers reads its own LSN and KVLOG from the MDS in shard order. Then, based on the KVLOG, using the same sharding rules as the processing module, it reads the WAL log from the corresponding PT, assembles the LSN, metadata operation log, and data into a complete log, and sends it to the backup server for processing via the replication link. As can be seen from the above example, compared to Example 5, Example 6 significantly improves log replication efficiency and throughput by mapping the LSN to N shards according to certain rules, storing the metadata operation log at the shard level in the MDS, and having each replication module in the NAS Server read the metadata operation log from the MDS at the shard level and concurrently replicate it to the backup server for processing.
[0249] Example 7: Building upon Example 6, in the second storage cluster, LSNs are mapped to N logical shards according to certain rules, and these N logical shards are then mapped to M PTs. This ensures that each shard knows which PT its logs are stored in for subsequent playback processing (M and N are both positive integers). In the second storage cluster, the second NAS server cluster includes a replication module responsible for receiving replication messages sent by the first storage cluster and deserializing them into metadata operation logs and data. It also includes a log module responsible for calculating the assigned PT for metadata operation logs based on their LSNs and data offsets to achieve log distribution balance. The second data storage cluster includes a log module and a playback module, and divides the storage area into an RDL area (i.e., a metadata storage partition) and a data area (i.e., a data storage partition). The RDL area is used to store logs, and the data area is used to store the actual data of user files (i.e., the data carried by operation requests).
[0250] Referring to Figure 4, which is a flowchart of an operation log replication process provided in this application embodiment, taking a write operation request as an example, the process of the second storage cluster replicating the operation log includes: the replication module of a node in the NAS Server cluster receives the replication message of the write operation request sent by the master and deserializes it into the metadata log and data of the write operation request; the log module obtains the belonging shard ID by taking the modulo of the LSN of the metadata log of this write operation request with the total number of shards, and then obtains the ID of the PT to which the metadata log (excluding data) belongs by taking the modulo of the shard ID with the total number of PTs; at the same time, the data of this write operation request is broken into multiple data blocks according to the inode and offset, and each data block also has its own belonging PT ID. Finally, a write operation request is divided into a metadata operation log (excluding data, only recording metadata information such as LSN, inode, and offset) and multiple data blocks (including LSN, data, etc.). Finally, the log module sends these to the Space node where their respective belonging PT is located; after receiving the log, each Space node appends it to the redo log in the RDL area. In the log, the backup replication process is now complete. In Example 7, any node in the NAS Server cluster can receive any replication message sent by the master. The log module maps the log LSN to N shards, and then maps the N shards to M PTs, thus balancing the metadata logs across different storage areas. This leverages multiple nodes to improve replication performance and facilitates subsequent replay processing.
[0251] Meanwhile, for write operation requests, in addition to storing their metadata information in the associated PT according to the mapping rules, the data is also broken down into multiple data blocks according to the PT's shuffling rules. The data blocks and LSNs are stored in the RDL area of the corresponding PT. When replaying this write operation request, it is only necessary to first read the inode, offset, length and other information from the PT where the metadata information is located, and then calculate the PT where each data block is located according to the PT shuffling rules based on this information. The replay data request can then be sent to the PT where each data block is located. Finally, the local log read in the PT is moved to the data area to complete the data replay operation. During replay, there is no need to read and write data blocks across nodes, which greatly reduces the network traffic between nodes and improves the replay performance.
[0252] Example 8: Based on Example 7, the process of replaying operation logs in the second storage cluster is as follows: The replay module of the NAS Server loads the metadata logs of its own shard into the replay module according to the LSN order, and forwards them to the MDS node to add them to the waiting replay queue (in Example 3, we have mapped the LSN to N shards according to certain rules, and each shard is bound to a replay module of the NAS Server, so that the replay module knows which LSNs it should be responsible for replaying); The MDS, as the control center of replay, after receiving the metadata logs and LSNs forwarded by the NAS Server, maintains two sliding windows in its internal replay module: the metadata concurrent replay window (equivalent to the metadata concurrent replay queue) and the data concurrent replay window (equivalent to the data concurrent replay queue), as shown in Figure 5. Figure 5 is a schematic diagram of a concurrent replay window provided in the embodiment of this application, which strictly arranges the replay in ascending order of LSN. In this process, LSNs first enter the data concurrent playback window (LSNs within the concurrent playback window must be non-exclusive; for example, different data operation requests targeting different memory address ranges cannot overlap, and requests for metadata operations are skipped in the data playback window). Only after data playback is complete can LSNs enter the metadata concurrent playback window (LSNs within the concurrent playback window must also be mutually exclusive; for example, inodes cannot be identical). The size of the sliding window controls the degree of concurrent playback, and the non-exclusive mechanism for LSNs entering the window ensures the correctness of the playback.
[0253] Referring to Figure 7, which is a schematic diagram of a playback process provided in an embodiment of this application, the concurrent data playback window process and the concurrent metadata playback window process are as follows. First, after entering the concurrent data playback window (non-write operation requests are directly marked as completed data playback), the MDS broadcasts a list of LSNs for all write operation requests within the playback window to all NAS Server nodes. Upon receiving the list of playable write operation request LSNs, each NAS Server is responsible for calculating the PT (Primary Time Limit) of the data block within this list of playable write operation request LSNs according to the PT shuffling rules, and then sends a data playback request to the corresponding PT. Upon receiving the data playback request, the Space node of the corresponding PT reads the data block from its local RDL area and moves it to the data area, completing the data playback. Then, the NAS Server reports the information that the corresponding LSN has completed data playback to the MDS. The MDS marks the LSNs that have completed data playback as completed data playback and waits to enter the concurrent metadata playback window.
[0254] In Example 8, each NAS Server's replay module is bound to one or more shards. The LSN that this NAS Server needs to load for replay is obtained through the mapping rules between LSN and shard. Then, the PT where the metadata log of this LSN is located is calculated through the LSN. After reading the metadata log from the corresponding PT, it is forwarded to the MDS node. Then, through the data concurrent replay window and metadata concurrent replay window mechanism, the replay concurrency is significantly improved and the replay speed is accelerated while ensuring the correctness of the replay.
[0255] The log-based file semantic layer replication method and technology proposed in this application can also be applied to master-slave replication scenarios in distributed clusters, such as distributed databases and distributed key-value systems.
[0256] Figure 8 is a schematic diagram of a data replication device in a distributed file system provided in an embodiment of this application. The device includes a receiving module 801, an execution module 802, and a sending module 803.
[0257] The receiving module 801 is configured to: receive a first operation request for a file, execute the first operation request based on the semantics of the first operation request, and obtain an execution result corresponding to the semantics of the first operation request, wherein the execution result includes at least one of the execution result of the file data and the execution result of the file metadata;
[0258] The generation module 802 is used to: generate an operation log for the first operation request based on the execution result and the execution order of the first operation request;
[0259] The sending module 803 is used to send an operation log to the second storage cluster, which is used by the second storage cluster to replay the first operation request.
[0260] In one possible implementation, the generation module 802 is used for:
[0261] If the first operation request is an operation request on the file's metadata, then the first operation request is executed, the execution result of the file's metadata is obtained, a metadata operation log is generated based on the execution result of the file's metadata, and an operation log for the first operation request is generated based on the metadata operation log and the execution order of the first operation request.
[0262] If the first operation request is a data operation request for a file, then the first operation request is executed, and the execution results for the file data and the file metadata are obtained. Based on the execution results for the file data, a data operation log is generated. Based on the execution results for the file metadata, a metadata operation log is generated. Based on the execution order of the metadata operation log, the data operation log, and the first operation request, an operation log for the first operation request is generated.
[0263] Figure 9 is a schematic diagram of a data replication device in a distributed file system provided in an embodiment of this application. The device includes a receiving module 901 and a playback module 902.
[0264] The receiving module 901 is used to receive the operation log sent by the first storage cluster. The operation log indicates the execution result and the execution order of the first operation request corresponding to the semantics of the first operation request. The first operation request is an operation request for a file. The execution result includes at least one of the execution result of the file data and the execution result of the file metadata.
[0265] The replay module 902 is used to replay the first operation request based on the operation log.
[0266] In one possible implementation, the playback module 902 includes:
[0267] A storage unit is used to store operation logs based on the execution order of the first operation request;
[0268] The replay unit is used to replay the first operation request based on the operation log.
[0269] In one possible implementation, the storage unit is used for:
[0270] If the operation log includes a metadata operation log, the metadata operation log is stored based on the execution order of the first operation request. The metadata operation log indicates the execution result of the file's metadata.
[0271] If the operation log includes a metadata operation log and a data operation log, then the metadata operation log and the data operation log are stored based on the execution order of the first operation request, and the data operation log indicates the execution result of the data on the file.
[0272] In one possible implementation, the playback unit is used for:
[0273] According to the execution order of the first operation requests, the first operation requests are added to the waiting replay queue. If the first operation request is a file metadata operation request, when the first operation request meets the first concurrent replay condition, the first operation request is added to the metadata concurrent replay queue according to the execution order of the first operation requests, metadata replay is performed on the first operation request, and the first operation request is added from the metadata replay queue to the replay queue. The first concurrent replay condition indicates that the first operation request and the operation requests in the metadata concurrent replay queue are not mutually exclusive, and multiple operation requests in the metadata concurrent replay queue can be replayed concurrently. If the first operation request is a file data operation request... When the first operation request meets the second concurrent replay condition, the first operation request is added to the data concurrent replay queue according to the execution order of the first operation request, and data replay is performed on the first operation request. When the first operation request meets the first concurrent replay condition, the first operation request is added from the data concurrent replay queue to the metadata concurrent replay queue according to the execution order of the first operation request, and metadata replay is performed on the first operation request. The first operation request is added from the metadata replay queue to the replay queue. The second concurrent replay condition indicates that the first operation request and the operation requests in the data concurrent replay queue are not mutually exclusive, and multiple operation requests in the data concurrent replay queue can be replayed concurrently.
[0274] The receiving module 801, execution module 802, sending module 803, receiving module 901, and playback module 902 can all be implemented in software or in hardware. For example, the implementation of the receiving module 801 will be described below. Similarly, the implementation of the execution module 802, sending module 803, receiving module 901, and playback module 902 can refer to the implementation of the receiving module 801.
[0275] As an example of a software functional unit, the receiving module 801 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the receiving module 801 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0276] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0277] As an example of a hardware functional unit, the receiving module 801 may include at least one computing device, such as a server. Alternatively, the receiving module 801 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0278] The multiple computing devices included in the receiving module 801 can be distributed in the same region or in different regions. Similarly, the multiple computing devices included in the receiving module 801 can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the multiple computing devices included in the receiving module 801 can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0279] It should be noted that, in other embodiments, the receiving module 801 can be used to execute any step in the data replication method in the distributed file system, the execution module 802 can be used to execute any step in the data replication method in the distributed file system, the sending module 803 can be used to execute any step in the data replication method in the distributed file system, and the receiving module 901 and the playback module 902 can be used to execute any step in the data replication method in the distributed file system. The steps implemented by the receiving module 801, the execution module 802, the sending module 803, the receiving module 901, and the playback module 902 can be specified as needed. The receiving module 801 and the execution module 802 respectively implement all the functions of the data replication device in the distributed file system shown in FIG8; or, the sending module 803, the receiving module 901, and the playback module 902 respectively implement all the functions of the data replication device in the distributed file system shown in FIG9.
[0280] This application also provides a storage node 1000. Figure 10 is a schematic diagram of the structure of a storage node provided in this application embodiment. As shown in Figure 10, the storage node 1000 includes: a bus 1001, a processor 1002, a memory 1003, and a communication interface 1004. The processor 1002, the memory 1003, and the communication interface 1004 communicate with each other through the bus 1001. The storage node 1000 can be a storage node or a terminal device. It should be understood that this application does not limit the number of processors and memories in the storage node 1000.
[0281] Bus 1001 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 10, but this does not imply that there is only one bus or one type of bus. Bus 1001 can include pathways for transmitting information between various components of storage node 1000 (e.g., memory 1003, processor 1002, communication interface 1004).
[0282] The processor 1002 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0283] The memory 1003 may include volatile memory, such as random access memory (RAM). The memory 1003 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0284] The memory 1003 stores executable program code, which the processor 1002 executes to implement the functions of the aforementioned receiving module 801, generating module 802, and sending module 803, or to implement the functions of the aforementioned receiving module 901 and playback module 902, thereby implementing the steps in the data replication method of the distributed file system. That is, the memory 1003 stores instructions for executing the data replication method of the distributed file system. Figure 10 only exemplarily illustrates the program code stored in the memory 1003 that implements the functions of the aforementioned receiving module 901 and recycling module 902.
[0285] The communication interface 1004 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the storage node 1000 and other devices or communication networks.
[0286] This application also provides a storage cluster. Figure 11 is a schematic diagram of a storage cluster provided in this application embodiment. As shown in Figure 11, the storage cluster includes at least one storage node 1000. The memory 1003 of one or more storage nodes 1000 in the storage cluster may store the same instructions for executing data replication methods in a distributed file system.
[0287] In some possible implementations, the memory 1003 of one or more storage nodes 1000 in the storage cluster may also store partial instructions for executing data replication methods in a distributed file system. In other words, a combination of one or more storage nodes 1000 can jointly execute instructions for executing data replication methods in a distributed file system.
[0288] It should be noted that the memory 1003 in different storage nodes 1000 in the storage cluster can store different instructions, which are used to execute part of the functions of the data replication device in the distributed file system. That is, the instructions stored in the memory 1003 in different storage nodes 1000 can realize the functions of one or more of the aforementioned receiving module 801, generating module 802, sending module 803, receiving module 901, and playback module 902.
[0289] It should be understood that the function of storage node 1000 shown in Figure 11 can also be performed by multiple storage nodes 1000.
[0290] In some possible implementations, one or more storage nodes in a storage cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc. Figure 12 illustrates one possible implementation. Figure 12 is a schematic diagram of a possible implementation of a storage cluster provided in an embodiment of this application. As shown in Figure 12, storage node 1000A and storage node 1000B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each storage node. In this type of possible implementation, the memory 1003 in storage node 1000A stores instructions for executing the function of the receiving module 901. Figure 12 uses the example of the memory 1003 in storage node 1000A storing instructions for executing the function of the receiving module 901. Simultaneously, the memory 1003 in storage node 1000B stores instructions for executing the function of the playback module 902. Figure 12 uses the example of the memory 1003 in storage node 1000B storing instructions for executing the function of the playback module 902.
[0291] This application embodiment also provides another storage cluster. The connection relationship between the storage nodes in this storage cluster can be similar to the connection method shown in Figure 11 or Figure 12. The difference is that the memory 1003 of one or more storage nodes 1000 in this storage cluster can store the same instructions for executing the data replication method in the distributed file system.
[0292] In some possible implementations, the memory 1003 of one or more storage nodes 1000 in the storage cluster may also store partial instructions for executing data replication methods in a distributed file system. In other words, a combination of one or more storage nodes 1000 can jointly execute instructions for executing data replication methods in a distributed file system.
[0293] It should be noted that the memory 1003 in different storage nodes 1000 in the storage cluster can store different instructions, which are used to execute part of the functions of the data replication device in the distributed file system. That is, the instructions stored in the memory 1003 in different storage nodes 1000 can realize the functions of one or more of the aforementioned receiving module 801, generating module 802, sending module 803, receiving module 901, and playback module 902.
[0294] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions capable of running on a computing device or stored on any available medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to perform a data copying method in a distributed file system.
[0295] This application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store, or a data storage device such as a data center that includes one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to perform a data copying method in a distributed file system.
[0296] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the storage clusters and storage space involved in this application were obtained with full authorization.
[0297] Those skilled in the art will recognize that the method steps and units described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the steps and components of each embodiment have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0298] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be found in the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0299] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through some interfaces, apparatuses, or units, or they may be electrical, mechanical, or other forms of connection.
[0300] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0301] Furthermore, the units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or software.
[0302] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computing device (which may be a personal computer, server, or computing device, etc.) to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0303] In this application, the terms "first," "second," etc., are used to distinguish identical or similar items with substantially the same function. It should be understood that there is no logical or temporal dependency between "first," "second," and "nth," nor does it limit the quantity or execution order. It should also be understood that although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are merely used to distinguish one element from another. For example, without departing from the scope of various examples, a first storage cluster can be referred to as a second storage cluster, and similarly, a second storage cluster can be referred to as a first storage cluster. Both the first and second storage clusters can be storage clusters, and in some cases, they can be separate and different storage clusters.
[0304] In this application, the term "at least one" means one or more, and the term "multiple" means two or more. The terms "system" and "network" are often used interchangeably.
[0305] It should also be understood that the term "if" can be interpreted as meaning "when" or "upon" or "in response to determination" or "in response to detection." Similarly, depending on the context, the phrases "if determination..." or "if detection [the stated condition or event]" can be interpreted as meaning "when determination..." or "in response to determination..." or "when detection [the stated condition or event]" or "in response to detection [the stated condition or event]."
[0306] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0307] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. This computer program product includes one or more computer program instructions. When these computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
[0308] The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer program instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., digital video discs (DVDs), or semiconductor media (e.g., solid-state drives)).
[0309] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0310] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A data replication method in a distributed file system, characterized in that, The method, executed by a first storage cluster in a distributed file system, which further includes a second storage cluster, comprises: Receive a first operation request for a file, execute the first operation request based on the semantics of the first operation request, and obtain an execution result corresponding to the semantics of the first operation request. The execution result includes at least one of the execution result of the data of the file and the execution result of the metadata of the file. Based on the execution result and the execution order of the first operation request, an operation log for the first operation request is generated; The operation log is sent to the second storage cluster, and the operation log is used by the second storage cluster to replay the first operation request.
2. The method according to claim 1, characterized in that, Based on the semantics of the first operation request, the first operation request is executed to obtain an execution result corresponding to the semantics of the first operation request. Based on the execution result and the execution order of the first operation request, an operation log of the first operation request is generated, including: If the first operation request is an operation request on the metadata of the file, then the first operation request is executed to obtain the execution result of the metadata of the file. Based on the execution result of the metadata of the file, a metadata operation log is generated. Based on the metadata operation log and the execution order of the first operation request, an operation log of the first operation request is generated. If the first operation request is an operation request on the data of the file, then the first operation request is executed to obtain the execution result of the data of the file and the execution result of the metadata of the file. Based on the execution result of the data of the file, a data operation log is generated. Based on the execution result of the metadata of the file, a metadata operation log is generated. Based on the metadata operation log, the data operation log and the execution order of the first operation request, an operation log of the first operation request is generated.
3. The method according to claim 2, characterized in that, The first storage cluster includes a first network-attached storage (NAS) server cluster and a first metadata storage node; If the first operation request is an operation request on the metadata of the file, then the first operation request is executed to obtain the execution result on the metadata of the file. Based on the execution result on the metadata of the file, a metadata operation log is generated. Based on the metadata operation log and the execution order of the first operation request, an operation log for the first operation request is generated, including: If the first operation request is an operation request on the metadata of the file, the first NAS server cluster sends the metadata targeted by the first operation request to the first metadata storage node; The first metadata storage node assigns a logical sequence number to the first operation request, executes the first operation request on the metadata, obtains the execution result of the metadata of the file, and generates the metadata operation log based on the execution result of the metadata of the file. The logical sequence number is used to indicate the execution order of the first operation request. The first NAS server generates the operation log of the first operation request based on the metadata operation log and the execution order of the first operation request.
4. The method according to claim 2, characterized in that, The first storage cluster includes a first NAS server cluster, a first metadata storage node, and a first data storage cluster; If the first operation request is a data operation request for the file, then the first operation request is executed to obtain the execution result of the data operation on the file and the execution result of the metadata operation on the file. Based on the execution result of the data operation on the file, a data operation log is generated. Based on the execution result of the metadata operation on the file, a metadata operation log is generated. Based on the metadata operation log, the data operation log, and the execution order of the first operation request, an operation log for the first operation request is generated, including: If the first operation request is a data operation request for the file, then the first NAS server cluster sends the data carried in the first operation request to the first data storage cluster; The first data storage cluster executes the first operation request on the data carried by the first operation request, obtains the execution result of the data on the file, and generates the data operation log based on the execution result of the data on the file; The first NAS server cluster sends the metadata targeted by the first operation request to the first metadata storage node; The first metadata storage node assigns a logical sequence number to the first operation request, executes the first operation request on the metadata, obtains the execution result of the metadata of the file, and generates the metadata operation log based on the execution result of the metadata of the file. The logical sequence number is used to indicate the execution order of the first operation request. The first NAS server cluster generates the operation log of the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request.
5. The method according to claim 3, characterized in that, The first NAS server cluster includes at least one first NAS server and at least one second NAS server. The first NAS server is used to receive and forward operation requests, and the second NAS server is used to generate and forward operation logs of operation requests. The first NAS server cluster generates an operation log for the first operation request based on the metadata operation log and the logical sequence number of the first operation request, including any one of the following: The first NAS server receives the metadata operation log and the logical sequence number of the first operation request sent by the first metadata storage node, and sends the metadata operation log and the logical sequence number of the first operation request to the second NAS server. The second NAS server generates the operation log of the first operation request based on the metadata operation log and the logical sequence number of the first operation request. The second NAS server retrieves the metadata operation log from the first metadata storage node according to the logical sequence number of the first operation request, and generates the operation log of the first operation request based on the metadata operation log and the logical sequence number of the first operation request.
6. The method according to claim 4, characterized in that, The first NAS server cluster includes multiple first NAS servers and at least one second NAS server. The first NAS servers are used to receive and forward operation requests, and the second NAS servers are used to generate and forward operation logs of the operation requests. The first NAS server cluster generates an operation log for the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request, including any one of the following: The first NAS server receives the metadata operation log and the logical sequence number of the first operation request sent by the first metadata storage node, and sends the metadata operation log, the logical sequence number of the first operation request and the data carried by the first operation request to the second NAS server. The second NAS server generates the operation log of the first operation request based on the data operation log, the metadata operation log and the logical sequence number of the first operation request. The second NAS server obtains the metadata operation log from the first metadata storage node based on the logical sequence number of the first operation request, obtains the data carried by the first operation request from the first data storage cluster based on the metadata operation log, and generates the operation log of the first operation request based on the data operation log, the metadata operation log, and the logical sequence number of the first operation request.
7. The method according to claim 5 or 6, characterized in that, Each first NAS server is associated with a second NAS server. Each first NAS server is used to receive and forward a set of operation requests. The generation and forwarding of operation logs for the set of operation requests received and forwarded by each first NAS server are performed by the second NAS server associated with the first NAS server.
8. The method according to any one of claims 5 to 7, characterized in that, The first NAS server and the second NAS server are the same NAS server.
9. A data replication method in a distributed file system, characterized in that, The method, executed by a second storage cluster in a distributed file system, which further includes a first storage cluster, comprises: The system receives operation logs sent by the first storage cluster. The operation logs indicate the execution results corresponding to the semantics of the first operation request and the execution order of the first operation request. The first operation request is an operation request for a file. The execution results include at least one of the execution results of the data of the file and the execution results of the metadata of the file. Based on the operation log, the first operation request is replayed.
10. The method according to claim 9, characterized in that, The step of replaying the first operation request based on the operation log includes: The operation log is stored based on the execution order of the first operation request; Based on the operation log, the first operation request is replayed.
11. The method according to claim 10, characterized in that, The step of storing the operation log based on the execution order of the first operation request includes: If the operation log includes a metadata operation log, then the metadata operation log is stored based on the execution order of the first operation request, and the metadata operation log indicates the execution result of the metadata of the file; If the operation log includes a metadata operation log and a data operation log, then the metadata operation log and the data operation log are stored based on the execution order of the first operation request, and the data operation log indicates the execution result of the data of the file.
12. The method according to claim 11, characterized in that, The second storage cluster includes a second network attached storage (NAS) server cluster and a second data storage cluster. The second NAS server cluster is used to receive operation logs sent by the first storage cluster. The second data storage cluster includes multiple data storage nodes, and each data storage node includes at least one storage partition. If the operation log includes a metadata operation log, then based on the execution order of the first operation request, the metadata operation log is stored, and the metadata operation log indicates the execution result of the metadata of the file, including: If the operation log includes the metadata operation log, the second NAS server cluster determines the first storage partition in the second data storage cluster based on the logical sequence number of the first operation request in the operation log, and stores the metadata operation log in the first storage partition. The logical sequence number is used to indicate the execution order of the first operation request.
13. The method according to claim 11, characterized in that, The second storage cluster includes a second NAS server cluster and a second data storage cluster. The second NAS server cluster is used to receive operation logs sent by the first storage cluster. The second data storage cluster includes multiple data storage nodes, and each data storage node includes at least one storage partition. If the operation log includes a metadata operation log and a data operation log, then based on the execution order of the first operation request, the metadata operation log and the data operation log are stored. The data operation log indicates the execution result of the data in the file, including: If the operation log includes metadata operation log and data operation log, then the second NAS server cluster determines the second storage partition in the second data storage cluster based on the logical sequence number of the first operation request in the operation log, and stores the metadata operation log in the second storage partition. Based on the metadata operation log, the second NAS server cluster divides the data carried by the first operation request in the data operation log into at least one data block, determines at least one third storage partition in the second storage cluster, and stores one of the data blocks and the logical sequence number of the first operation request into one of the third storage partitions.
14. The method according to claim 12 or 13, characterized in that, The second NAS server cluster includes multiple third NAS servers, each of which is used to execute a stored procedure for the operation logs of a set of operation requests. The stored procedures for the operation logs of different sets of operation requests can be executed in parallel.
15. The method according to any one of claims 10 to 14, characterized in that, The step of replaying the first operation request based on the operation log includes: According to the execution order of the first operation requests, add the first operation requests to the waiting replay queue; If the first operation request is an operation request on the metadata of the file, when the first operation request meets the first concurrent replay condition, the first operation request is added to the metadata concurrent replay queue according to the execution order of the first operation requests, the metadata replay is performed on the first operation request, and the first operation request is added from the metadata replay queue to the replay queue. The first concurrent replay condition indicates that the first operation request and the operation requests in the metadata concurrent replay queue are not mutually exclusive, and multiple operation requests in the metadata concurrent replay queue can perform metadata replay concurrently. If the first operation request is a data operation request on the file, when the first operation request meets the second concurrent playback condition, the first operation request is added to the data concurrent playback queue according to the execution order of the first operation requests, and data playback is performed on the first operation request. When the first operation request meets the first concurrent playback condition, the first operation request is added from the data concurrent playback queue to the metadata concurrent playback queue according to the execution order of the first operation requests, and metadata playback is performed on the first operation request. The first operation request is added from the metadata playback queue to the already played-out queue. The second concurrent playback condition indicates that the first operation request and the operation requests in the data concurrent playback queue are not mutually exclusive, and multiple operation requests in the data concurrent playback queue can be played back concurrently.
16. The method according to claim 15, characterized in that, The second storage cluster also includes a second metadata storage node; The step of performing metadata replay on the first operation request includes: The second NAS server cluster obtains the metadata operation log from the second data storage cluster and sends the metadata operation log to the second metadata storage node; The second metadata storage node performs metadata replay of the first operation request based on the metadata operation log.
17. The method according to claim 15, characterized in that, The second storage cluster also includes a second metadata storage node. Each storage partition includes a data storage partition and a metadata storage partition. The data carried by the first operation request and the logical sequence number of the first operation request are stored in the metadata storage partition of the third storage partition. The step of replaying the data for the first operation request includes: The second NAS server cluster obtains the metadata operation log from the second data storage cluster and sends the metadata operation log to the second metadata storage node; Based on the metadata operation log, the second metadata storage node determines the plurality of third storage partitions in the second storage cluster, and writes the data carried by the first operation request in the third storage partition from the metadata storage partition of the third storage partition to the data storage partition of the third storage partition.
18. The method according to claim 16 or 17, characterized in that, The second NAS server cluster includes multiple fourth NAS servers, each of which is used to execute a replay process of a set of operation requests, and the replay processes of different sets of operation requests can be executed in parallel.
19. The method according to claim 18, characterized in that, The third NAS server and the fourth NAS server are the same NAS server.
20. A data replication method for a distributed file system, characterized in that, Applied to a distributed file system, the distributed file system comprising a first storage cluster and a second storage cluster, the method includes: The first storage cluster receives a first operation request for a file, executes the first operation request based on the semantics of the first operation request, and obtains an execution result corresponding to the semantics of the first operation request. The execution result includes at least one of the execution result of the data of the file and the execution result of the metadata of the file. The first storage cluster generates an operation log for the first operation request based on the execution result and the execution order of the first operation request; The first storage cluster sends the operation log to the second storage cluster; The second storage cluster receives the operation log sent by the first storage cluster and replays the first operation request based on the operation log.
21. A distributed file system, characterized in that, The distributed file system includes: The first storage cluster is used for: Receive a first operation request for a file, execute the first operation request based on the semantics of the first operation request, and obtain an execution result corresponding to the semantics of the first operation request. The execution result includes at least one of the execution result of the data of the file and the execution result of the metadata of the file. Based on the execution result and the execution order of the first operation request, an operation log for the first operation request is generated; Send the operation log to the second storage cluster; The second storage cluster is used for: The system receives the operation log sent by the first storage cluster and replays the first operation request based on the operation log.
22. A storage node, characterized in that, The storage node includes a memory and a processor; The processor is configured to execute instructions stored in the memory, causing the storage node to perform the data replication method in the distributed file system according to any one of claims 1 to 20.
23. A storage cluster, characterized in that, The storage cluster includes at least one storage node, and each storage node includes a memory and a processor; The processor of the at least one storage node is used to execute instructions stored in the memory of the at least one storage node, so that the storage cluster performs the data replication method in the distributed file system as described in any one of claims 1 to 20.
24. A computer program product containing instructions, characterized in that, When the instruction is executed by a storage node, the storage node performs the data replication method in a distributed file system as described in any one of claims 1 to 20; or, when the instruction is executed by a storage cluster, the storage cluster performs the data replication method in a distributed file system as described in any one of claims 1 to 20.
25. A computer-readable storage medium, characterized in that, The system includes computer program instructions, which, when executed by a storage node, enable the storage node to perform the log synchronization method as described in any one of claims 1 to 20; or, when executed by a storage cluster, enable the storage cluster to perform the data replication method in a distributed file system as described in any one of claims 1 to 20.
Citation Information
Patent Citations
Data synchronization method and data synchronization device based on log analysis
CN110297866A
Resynchronization to file system synchronous replication relationship endpoints
CN114127695A
Data processing system, data processing method and device and related equipment
CN117931831A
Redoing transaction log records in parallel
US20180144015A1