Data processing method and device and distributed storage system
By introducing coordinating points in the distributed storage system, processing data operation transactions distributed on multiple target nodes, and writing the results to the target node, the correctness and scalability of transaction processing in the distributed storage system is solved, and the system efficiency and reliability are achieved.
Patent Information
- Application Number
- CN202311825619.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
When a distributed storage system processes transactions and queries, if it does not support distributed transactions, it may lead to incorrectness and non-atomicity of concurrent operations, affecting the correctness of the system.
The data operation request is received by the coordination point, and it is determined that the data in the data operation range is distributed on multiple target nodes. These target nodes are called to perform data operation requests, and the data processing results of each target node are obtained, and the result is finally written to the corresponding target node.
Ensure the accuracy and scalability of the distributed storage system, and handle data operation transactions distributed on different target nodes through coordinating points to ensure the accuracy and consistency of data processing results.
Smart Images

Figure CN120215810A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and more particularly, to a data processing method, apparatus, and distributed storage system. Background Art
[0002] In a distributed storage system, transactions and queries may be executed on multiple nodes. If the distributed storage system does not support the execution of distributed transactions, the correctness and atomicity of concurrent operations cannot be guaranteed, thus affecting the correctness of the distributed storage system. Summary of the Invention
[0003] In view of this, embodiments of the present invention provide a data processing method, apparatus, and distributed storage system, which coordinate nodes to process data operation transactions distributed on different target nodes, and then write the data processing results to the corresponding target nodes respectively, so as to ensure the correctness and scalability of the distributed storage system.
[0004] In a first aspect, an embodiment of the present invention provides a data processing method, which is applied to a coordinator node in a distributed storage system. The distributed storage system includes multiple nodes, and each of the nodes stores data in a different range. The multiple nodes include a coordinator node and data nodes. The method includes:
[0005] Receiving a data operation request, where the data operation request includes a data operation range;
[0006] In response to the data in the data operation range being distributed on multiple target nodes, calling the corresponding data on the target nodes to execute the data operation request, and obtaining data processing results corresponding to each of the target nodes respectively;
[0007] Writing each of the data processing results to the corresponding target node.
[0008] In a second aspect, an embodiment of the present invention provides a distributed storage system. The distributed storage system includes multiple nodes, and each of the nodes stores data in a different range. The multiple nodes include:
[0009] A coordinator node, configured to execute the method as described above; and
[0010] Data nodes, configured to perform a commit operation on the data processing results sent by the coordinator node.
[0011] In a third aspect, an embodiment of the present invention provides a data processing apparatus. The apparatus includes:
[0012] A request receiving unit, configured to receive a data operation request, where the data operation request includes a data operation range;
[0013] A processing unit, configured to call corresponding data on the target nodes to execute the data operation request and obtain data processing results respectively corresponding to the target nodes in response to the data in the data operation range being distributed on multiple target nodes;
[0014] A writing unit, configured to write each of the data processing results into the corresponding target node.
[0015] In a fourth aspect, an embodiment of the present invention provides an electronic device, including a memory and a processor, where the memory is used to store one or more computer program instructions, and wherein the one or more computer program instructions are executed by the processor to implement the method as described above.
[0016] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method as described above is implemented.
[0017] In the technical solution of the embodiment of the present invention, a distributed storage system includes multiple nodes, and different ranges of data are stored in each node. The multiple nodes include a coordination node and data nodes. The coordination node is used to receive a data operation request. In response to the data in the data operation range corresponding to the data operation request being distributed on multiple target nodes, call the corresponding data on each target node to execute the data operation request, obtain data processing results respectively corresponding to the target nodes, and write each data processing result into the corresponding target node. Thus, in this embodiment, the coordination node processes data operation transactions distributed on different target nodes, and then writes the data processing results into the corresponding target nodes respectively, ensuring the correctness and scalability of the distributed storage system. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Through the following description of the embodiments of the present invention with reference to the drawings, the above and other objects, features, and advantages of the present invention will become clearer. In the drawings:
[0019] Figure 1 is a schematic diagram of a data processing system according to an embodiment of the present invention;
[0020] Figure 2 is a schematic diagram of a distributed storage system according to an embodiment of the present invention;
[0021] Figure 3 is a flowchart of a data processing method according to an embodiment of the present invention;
[0022] Figure 4 is a flowchart of a distributed transaction submission method according to an embodiment of the present invention;
[0023] Figure 5It is a flowchart of the distributed transaction commit phase in an embodiment of the present invention;
[0024] Figure 6 It is an interaction flowchart of the distributed transaction commit process in an embodiment of the present invention;
[0025] Figure 7 It is a schematic diagram of the load status determination process of the metadata tree in an embodiment of the present invention;
[0026] Figure 8 It is a schematic diagram of the process of the metadata tree migration method in an embodiment of the present invention;
[0027] Figure 9 It is a schematic diagram of the data processing device in an embodiment of the present invention;
[0028] Figure 10 It is a schematic diagram of the electronic device in an embodiment of the present invention. Detailed implementation manners
[0029] The following describes the present application based on embodiments, but the present application is not limited to these embodiments. In the following detailed description of the present application, some specific details are described in detail. Those skilled in the art can fully understand the present application without the description of these detail parts. In order to avoid obscuring the essence of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0030] In addition, those of ordinary skill in the art should understand that the accompanying drawings provided herein are for illustrative purposes only, and the drawings are not necessarily drawn to scale.
[0031] Unless the context clearly requires otherwise, words such as "including" and "comprising" in the entire application document should be interpreted as having an inclusive meaning rather than an exclusive or exhaustive meaning; that is, it is the meaning of "including but not limited to".
[0032] In the description of the present application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present application, unless otherwise specified, the meaning of "a plurality" is two or more.
[0033] For the solutions described in this specification and embodiments, if they involve personal information processing, they will be processed on the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for performing a contract, etc.), and will only be processed within the specified or agreed scope. If the user refuses to process personal information other than the necessary information required for basic functions, it will not affect the user's use of basic functions.
[0034] Figure 1It is a schematic diagram of the data processing system according to an embodiment of the present invention. Figure 2 It is a schematic diagram of the distributed storage system according to an embodiment of the present invention. As Figure 1 shown, the data processing system of this embodiment includes a service layer 11, a server 12, and a distributed storage system 13. Among them, the service layer 11 may include but is not limited to various computers, laptops, tablets, smartphones, Internet of Things devices, or wearable devices, etc. The server 12 may be an independent server or a server cluster composed of multiple servers.
[0035] Optionally, the distributed storage system 13 may perform read-write semantic fusion on various storage services, so that each service layer 11 can call an appropriate storage interface to read data from the underlying storage. Among them, the distributed storage system 13 may include a storage interface service module with multiple storage interfaces. Optionally, the storage interface service module may include interface services such as Posix file interface, HDFS SDK, CSI, S3 interface, S2 interface, and image processing interface, etc., to support multiple storage services for storage semantic fusion. It should be understood that this embodiment is not limited to the above storage interfaces, and other storage interfaces for implementing storage services can also be integrated into this embodiment.
[0036] In this embodiment, the service layer 11 communicates with the server 12 through a network, and the distributed storage system 13 can store the data that the server 12 needs to process. Further, the distributed storage system 13 can be integrated on the server 12, or placed on the cloud or other servers. The server 12 can receive a data operation request sent by a terminal in the service layer 11 and forward the data operation request to the distributed storage system 13. The distributed storage system 13 may include multiple nodes, and at least one coordination node and at least one data node are included in the multiple nodes. Among them, the coordination node is used to manage some or all of the nodes in the distributed storage system 13, and at the same time manage or process the received data operation requests, that is, distribute the transactions corresponding to the data operation requests to the relevant nodes for execution.
[0037] Taking the distributed storage system 13 as a metadata service system (Meta Data Server, MDS) as an example, in order to facilitate the client to access the metadata of various files in the distributed storage system 13, each metadata server will store the corresponding metadata tree in the local storage module or the remote storage module. As the number of files continues to grow, the scale of the metadata tree is also growing, and there are differences in the load status information such as the access volume of the metadata tree, which makes the load of each node storing the metadata tree in the distributed storage system 13 different, thus causing the phenomenon of load imbalance. Therefore, this embodiment enables the distributed storage system 13 to support subtree partitioning to achieve load balancing, thereby reducing system latency and increasing data throughput.
[0038] Further, the sub-tree partitioning also splits the metadata tree with a high load into multiple sub-trees and distributes these sub-trees to different data nodes, so that the distributed storage system 13 achieves load balancing. In this case, the distributed storage system 13 needs to support distributed transactions across sub-trees of multiple nodes to ensure the correctness and low latency of data operations.
[0039] Among them, a transaction is a series of operations executed as a single logical unit of work, either executed completely or not executed at all. Transaction processing can ensure that data-oriented resources are not permanently updated unless all operations within the transactional unit are successfully completed. By combining a set of related operations into a unit that either succeeds or fails completely, error recovery can be simplified and the application can be made more reliable. A logical unit of work must satisfy the ACID (atomicity, consistency, isolation, and durability) properties to become a transaction. A distributed transaction means that the participants in the transaction execution are located on different nodes.
[0040] Further, as Figure 2 shown, the distributed storage system 13 of this embodiment includes at least one coordinator node 21 and at least one data node 22.
[0041] In an alternative implementation, the coordinator node 21 can only act as a coordinator and does not store the corresponding data. In this case, the coordinator node 21 can store configuration information, which includes the data ranges stored by each data node 22, so as to send a data operation request to the corresponding data node 22 for processing, or call the data required to execute the data operation request from the corresponding data node 22 based on the configuration information, execute the data operation request, and write the data processing result of the data operation request into the corresponding data node 22. In another alternative implementation, the coordinator node 21 is a coordinator determined from the data nodes, and it also stores data within a certain data range, and has the roles of both a coordinator and a data node, and executes the corresponding transactions.
[0042] In this embodiment, as Figure 2 shown, the specific description is mainly given by taking the example that the coordinator node 21 also stores data within a certain data range. It should be understood that this embodiment is not limited thereto.
[0043] Optionally, the coordination node 21 may represent a single node or a cluster of nodes. Optionally, when the coordination node 21 represents a cluster of nodes, the coordination node 21 includes a coordination master node and at least one coordination slave node. Among them, the coordination master node (i.e., the Leader node in the coordination node cluster) is determined through election, responsible for managing the entire node cluster, synchronizing data to the coordination slave nodes (i.e., the Follower nodes in the coordination node cluster), and ensuring the global consistency of the data. Further, the coordination slave nodes can provide a data query function, and they can send the received data operation requests to the coordination master node. At the same time, the coordination master node and the coordination slave nodes can achieve linear consistent reads of the data, thus ensuring data consistency in the distributed system.
[0044] Similarly, data nodes such as the data node 22 in this embodiment can also represent a single node or a cluster of nodes. When the data node 22 represents a cluster of nodes, it includes a data master node and at least one data slave node. Among them, the data master node (i.e., the Leader node in the data node cluster) is determined through election, responsible for managing the entire node cluster, synchronizing data to the data slave nodes (i.e., the Follower nodes in the data node cluster), and ensuring the global consistency of the data. Further, the data slave nodes can provide a data query function, and they can send the received data operation requests to the data master node. At the same time, the data master node and the data slave nodes can achieve linear consistent reads of the data, thus ensuring data consistency in the distributed system.
[0045] In this embodiment, each coordination node 21 and each data node 22 stores data within different data ranges. Taking metadata as an example, among them, metadata is usually managed and stored in the form of a metadata tree. Further optionally, different metadata trees are stored in the coordination node 21 and the data node 22 of this embodiment.
[0046] As described above, when the load or storage resource consumption of the metadata tree on a single node is too large, the metadata tree can be split to form new metadata subtrees, which are distributed on different nodes to achieve load balancing in the distributed storage system.
[0047] In an optional implementation manner, the metadata tree stored in the data node can be determined based on the splitting of the metadata tree of the coordination node. Further optionally, the configuration information of each metadata tree includes the data range of the corresponding metadata tree and the identifier of the split metadata tree (i.e., the metadata subtree) or the data range of the split metadata tree. Further optionally, the metadata tree can have a corresponding identifier, or it can be represented by the identifier of the data node where it is located.
[0048] For example, assume that the load of the metadata tree Subtree0 on node A is too large, and a metadata subtree Subtree1 is split off and distributed on node B. Among them, the data range of the split metadata tree Subtree0 will change, and its configuration information includes the current data range of the split metadata tree Subtree0 and the data range of the split metadata subtree Subtree1 or the identifier of the metadata subtree Subtree1.
[0049] Based on this, in this embodiment, the node A where the metadata tree Subtree0 is located can be determined as the coordination node, and the nodes B where other metadata trees (i.e., metadata subtrees) are located can be determined as data nodes, so as to further facilitate the management and execution of data operation requests and reduce system latency. It should be understood that if node A represents a node cluster, the cluster represented by node A can be used as the coordination node (i.e., the coordination node represents a cluster), or the Leader node in the cluster represented by node A can be used as the coordination node (i.e., the coordination node represents a single node).
[0050] In this embodiment, the coordination node 21 receives a data operation request, determines the transaction type based on the data distribution in the data operation range corresponding to the data operation request, and processes the transaction corresponding to the data operation request based on the transaction type. Optionally, the transaction corresponding to the data operation request in this embodiment can be a single-machine transaction or a distributed transaction. Among them, a single-machine transaction means that the data in the data operation range corresponding to the request is distributed on a target node (data node or coordination node), and a distributed transaction means that the data in the data operation range corresponding to the request is distributed on multiple target nodes. That is, the coordination node needs to coordinate multiple target nodes to execute the distributed transaction. Among them, the multiple target nodes can include multiple data nodes or can also include the coordination node and data nodes at the same time.
[0051] Figure 3 is a flowchart of the data processing method of the embodiment of the present invention. Further, as Figure 3 shown, the coordination node 21 of this embodiment can execute the following steps:
[0052] Step S110, receive a data operation request. Among them, the data operation request includes a data operation range. Optionally, the operations corresponding to the data operation request may include file - type operations and directory - type operations. Among them, file - type operations may include Write operation, Read operation, Truncate operation, Fallocate operation, etc. Directory - type operations may include Rename operation, Link operation, and Symlink operation, etc. It should be understood that this embodiment is not limited to the above - exemplified operation requests, and it may cover all requests that a distributed storage system can support. Further, the data operation request range represents the files or directories involved in this data operation request. Taking the Write operation request in file - type operations as an example, the data operation range may include one or more files corresponding to the Write operation request.
[0053] Among them, the Write operation is an operation to write data into a file. The Read operation is an operation to read data from a file. The Truncate operation is used to truncate a file, that is, to delete some or all of the data in the file. The Fallocate operation is an operation to pre - allocate a file, that is, to pre - allocate a specified - size disk space to the file. The Rename operation is usually used to rename a file. The Link operation refers to creating a hard link to an existing file path. The newly created hard - link file and the existing original file have the same index and data blocks. The Symlink operation refers to creating a new symbolic link pointing to an existing file or directory, that is, the created symbolic link contains a reference to another file or directory. When accessing the symbolic link, the file or directory it references is accessed.
[0054] Step S120, in response to the data in the data operation range corresponding to the data operation request being distributed on multiple target nodes, call the corresponding data on each target node to execute the data operation request, and obtain the data - processing results corresponding to each target node respectively.
[0055] In an optional implementation manner, the configuration information of the coordination node includes the identifiers of at least one data node or the stored data range. Optionally, the configuration information of the coordination node also includes its own stored data range.
[0056] Further, at least two nodes in the distributed storage system in this embodiment have an association relationship, so that the coordination node can obtain the data ranges on each data node through the association relationship between the nodes. Optionally, the association relationship in this embodiment represents that the identifiers or stored data ranges of each data node are stored in the configuration information of other data nodes or the coordination node.
[0057] Taking the distributed storage system of this embodiment as the metadata storage system as an example, assume that the coordination node stores the metadata tree Subtree0, the data node A1 and the data node B1 store the metadata trees Subtree1 and subtree2 respectively, and the data node C1 stores the metadata tree Subtree3. Among them, the metadata trees Subtree1 and subtree2 are subtrees split from the metadata tree Subtree0, and the metadata tree Subtree3 is a subtree split from the metadata tree Subtree1. At this time, the configuration information of the coordination node includes the current data range of the metadata tree Subtree0, the identifiers (which can be subtree identifiers or the identifiers of the data nodes A1 and B1) or data ranges of the split subtrees Subtree1 and subtree2. The configuration information of the data node A1 includes the data range of the metadata tree Subtree1 and the identifier or data range of the split subtree Subtree3.
[0058] In other alternative implementation manners, after a new subtree is split from the metadata tree, each data node can also report to the coordination node in real time the data ranges currently stored by each data node, so that the coordination node can determine the data nodes where the data in the data operation range corresponding to the data operation request is distributed according to its own configuration information. This embodiment does not limit the setting of the configuration information of each data node and the manner in which the coordination node invokes other data nodes based on the configuration information.
[0059] That is to say, in this embodiment, if the operation corresponding to the current data operation request is a distributed transaction, the coordination node invokes the corresponding data in each target node through the configuration information to execute the data operation request, and obtains and caches the data processing result.
[0060] Step S130, write each data processing result into the corresponding target node.
[0061] In an alternative implementation manner, this embodiment realizes multiversion concurrency control (MVCC) by assigning a corresponding timestamp to each transaction, that is, allowing multiple transactions to read, write, and modify the same data simultaneously without interfering with each other, improving the data processing efficiency while ensuring the correctness of data processing.
[0062] The distributed storage system of this embodiment further includes a global time synchronization server (Timestamp Oracle, TSO) to provide monotonically increasing timestamps. Among them, the TSO server is a system or mechanism that maintains global timestamps. Through strictly monotonically increasing timestamps, it ensures that there is a strict order among all timestamps and realizes snapshot isolation. TSO is a way to implement MVCC. It uses the global time synchronization server to assign a unique timestamp to each transaction and uses this timestamp to record the execution process and commit time of the transaction.
[0063] The TSO timestamp is a 64-bit integer value, consisting of a physical part and a logical part. Among them, the high 48 bits are the physical part, which is the millisecond time of unixtime (that is, the millisecond time elapsed since a certain fixed time xx year xx month xx day xx hour xx minute xx second xx millisecond), and the low 18 bits are the logical part, which is a numerical counter. In theory, 262,144,000 (that is, 2 18 *1000) TSO timestamps can be generated per second.
[0064] Further optionally, in this embodiment, the global timestamp is allocated and persisted by a coordination node (such as a coordination master node). When the transaction corresponding to the data operation request starts, a TSO is obtained from the TSO server as the start timestamp (that is, read_ts). When the transaction executes a read operation, it will obtain the start timestamp read_ts to determine which data versions are visible according to this start timestamp read_ts. If the data version is not visible, the read data operation needs to wait until the data version to be read becomes visible. Another TSO is obtained from the TSO server as the commit timestamp (that is, commit_ts) when the transaction is committed. This embodiment can splice the start timestamp read_ts and / or the commit timestamp commit_ts behind the field of the data processing result to form the final data processing result. That is to say, when the transaction executes a modification operation (such as add, delete, or modify operations), the start timestamp read_ts and / or the commit timestamp commit_ts will be written into the database together with the modification content. This can enable each transaction to have its own data version, and these data versions maintain visibility and consistency under the strict monotonically increasing order of the global timestamp. When the data needs to be accessed, it can be judged which transaction's modification should be applied according to the order of the global timestamps. Thus, this embodiment can set the start timestamp and commit timestamp of each transaction through the TSO server, realizing multiversion concurrency control (MVCC).
[0065] It should be understood that in this embodiment, other methods can also be used to implement MVCC, such as pessimistic concurrency control, optimistic concurrency control, lock-based concurrency control, etc. This embodiment does not limit this. Among them, pessimistic concurrency control assumes the worst case and believes that other transactions will modify the data at any time, so a lock is added when reading the data to prevent other transactions from modifying it simultaneously. Optimistic concurrency control assumes that other transactions do not often modify the data, so no lock is added when reading the data, and only when updating the data, it is checked whether other transactions have modified the data simultaneously. Lock-based concurrency control uses read-write locks to implement concurrency control. Read locks can be held by multiple transactions simultaneously, and write locks can only be held by one transaction.
[0066] In this embodiment, when dealing with distributed transactions, the process of distributed transaction submission is described by taking the method of implementing MVCC using TSO as an example. Further, the transaction submission process of this embodiment can adopt the two-phase commit (Two-Phase Commit, 2PC) method. It should be understood that this embodiment does not limit the way of transaction submission. For example, the three-phase commit (Three-Phase Commit, 3PC) method, TCC, Saga, local pre-submission and other distributed transaction submission methods can also be used, as long as they can ensure the correct submission of the transaction.
[0067] Figure 4 is the flowchart of the distributed transaction submission method of the embodiment of the present invention. As Figure 4 shown, the coordinator node submits the distributed transaction through the following steps:
[0068] Step S131, send a pre-write instruction to each target node. In this embodiment, the coordinator node sends a pre-write instruction (i.e., Prewrite instruction) to each target node in the preparation submission stage to inquire whether all target nodes can perform the submission operation. Optionally, in this embodiment, the coordinator node can also send each data processing result to the corresponding target node when sending the pre-write instruction.
[0069] Step S132, in response to receiving the pre-write success messages feedback by each of the target nodes, obtain the transaction submission timestamp.
[0070] In this embodiment, if the coordinator node receives the pre-write success messages feedback by all target nodes, that is, the messages indicating that it can be submitted, it obtains a TSO from the global time-granting server as the transaction submission timestamp.
[0071] Step S133, send a commit instruction (i.e., commit instruction) to each target node based on the transaction submission timestamp to submit the transaction corresponding to the data operation request.
[0072] In this embodiment, the coordination node sends a commit instruction to each target node according to the transaction commit timestamp to notify each target node to commit the corresponding transaction. It should be understood that the distributed transaction corresponding to the data operation request in this embodiment includes sub-transactions corresponding to each target node, and the successful commit of each target node's transaction indicates the successful commit of the corresponding entire distributed transaction.
[0073] In an alternative implementation, in response to a pre-write failure message fed back by one or more target nodes, that is, a message indicating that the current commit is not possible, the coordination node notifies each target node to perform a rollback operation.
[0074] In an alternative implementation, the coordination node in this embodiment can send a commit instruction to each target node simultaneously, so that each target node performs a transaction commit operation. In response to receiving the commit success messages fed back by all target nodes, it is determined that the distributed transaction corresponding to the data operation request is committed successfully, and a feedback message indicating successful data operation is sent to the corresponding client in the service layer (or a feedback message indicating successful data operation is sent to the client through the server). In this way, if one or more target nodes feed back a commit failure message, the coordination node notifies each target node to perform a rollback operation.
[0075] In an alternative implementation, if the transaction rollback reaches a predetermined time or a predetermined number of times, the coordination node determines that the data operation fails, and sends a feedback message indicating data operation failure to the corresponding client in the service layer (or a feedback message indicating data operation failure is sent to the client through the server).
[0076] In the commit stage of the above commit method, the coordination node needs to receive the commit success messages fed back by all target nodes before it can send a feedback message to the corresponding client, which causes a certain delay for the client. Therefore, in another alternative implementation, in order to further reduce the delay, the coordination node in this embodiment selects one target node as the main target node (i.e., the primary node), determines the other non-target nodes as secondary target nodes (i.e., the secondary nodes), and makes the secondary nodes point to the primary node. Further optionally, the coordination node can randomly select one target node as the main target node or specify a certain target node as the main target node, and this embodiment does not limit this.
[0077] Figure 5 is the flowchart of the distributed transaction commit stage of the embodiment of the present invention. Specifically, as Figure 5 shown, in this embodiment, step S133 can specifically execute the commit stage through the following steps:
[0078] Step S133A: Send a first commit instruction to the primary target node based on the transaction commit timestamp. That is, after obtaining the transaction commit timestamp from the global time server, the coordinator sends a first commit instruction to the primary target node according to the transaction commit timestamp to notify the primary target node to perform the corresponding transaction commit.
[0079] Step S133B: In response to receiving the commit success message feedback from the primary target node, send a data operation success message to the client.
[0080] Since in this embodiment, each target node is determined as a primary target node and a secondary target node, and the secondary target node points to the primary target node, therefore, after the transaction commit of the primary target node is successful, the coordinator can determine that the corresponding distributed transaction commit is successful. At this time, it can directly send a data operation success message to the client. Thus, there is no need to wait for the feedback of other secondary target nodes, further reducing the latency. Further, in response to receiving the commit failure message feedback from the primary target node, the coordinator controls the transaction to perform a rollback operation.
[0081] Step S133C: Send a second commit instruction to each secondary target node to enable the secondary target node to commit the transaction corresponding to the data operation request. In an optional implementation manner, the coordinator of this embodiment asynchronously sends a second commit instruction to each secondary target node to enable each secondary target node to asynchronously and concurrently execute the transaction commit operation, thereby further reducing the latency and improving the data throughput.
[0082] Figure 6 It is an interaction flowchart of the distributed transaction commit process of an embodiment of the present invention. As Figure 6 shown, in this embodiment, it is described by taking the primary target node and the secondary target node not being the coordinator node and taking one secondary target node as an example. It should be understood that if there are multiple secondary target nodes, they are processed in a similar manner.
[0083] After determining to enter the distributed transaction commit stage, the coordinator 61 sends a Prewrite instruction w1 to the primary target node 62 and a Prewrite instruction w2 to the secondary target node 63 to inquire whether all target nodes can perform the commit operation. After the primary target node 62 confirms that it can perform the commit operation of this distributed transaction, it sends a feedback message r1 to the coordinator 61 to inform the coordinator that the commit operation can be performed. After the secondary target node 63 confirms that it can perform the commit operation of this distributed transaction, it sends a feedback message r1 to the coordinator 61 to inform the coordinator that the commit operation can be performed.
[0084] After receiving the feedback messages r1 and r2 that can be submitted from all target nodes (i.e., the primary target node 62 and the secondary target nodes 63), the coordination node 61 requests a TSO from the global time synchronization server as the transaction commit timestamp commit_ts, and sends a commit instruction c1 to the primary target node 62 based on the transaction commit timestamp commit_ts to notify the primary target node 62 to perform the commit operation. After receiving the commit instruction c1, the primary target node 62 performs the transaction commit operation. After the transaction commit is successful, the primary target node 62 sends a feedback message r3 indicating the successful transaction commit to the coordination node 61. After receiving the feedback message r3 indicating the successful transaction commit, the coordination node 61 sends a data operation success message to the corresponding client.
[0085] Furthermore, the coordination node 61 sends a commit instruction c2 to the secondary target node 63 to notify the secondary target node 63 to perform the commit operation. After receiving the commit instruction c2, the secondary target node 63 performs the transaction commit operation. After the transaction commit is successful, the secondary target node 63 sends a feedback message r4 indicating the successful transaction commit to the coordination node 61. At this time, the coordination node 61 determines that the current distributed transaction has been completed. Optionally, if there are multiple secondary target nodes 63, the coordination node 61 can also asynchronously send a second commit instruction to each secondary target node to enable each secondary target node to asynchronously and concurrently perform the transaction commit operation, thereby further reducing the latency and improving the data throughput.
[0086] In this embodiment, the version-based distributed transaction processing is implemented through the TSO server, and the data processing result can be the corresponding KV value. Among them, KV is divided into two parts, namely the version column and the data column. The version column is used to store the latest version of KV, and the data column is used to store the values of different versions of this KV. Further, KV can be as shown in Table (1) below.
[0087] Table (1)
[0088]
[0089]
[0090] In a distributed transaction, the data processing results obtained by the coordinator node executing the data operation request, that is, the KV results, involve multiple different target nodes. Therefore, based on the above submission method, the KV results cached by the coordinator node can be written to the corresponding multiple target nodes respectively to ensure atomic writing to multiple target nodes. Specifically, when the main target node executes the transaction submission operation, that is, writes the KV version corresponding to the main target node to the version column, that is, writes the corresponding Key_commit_ts--->read_ts, and after the writing is successful, feedback a message indicating successful data operation to the corresponding client. Similarly, when the subsequent slave target nodes asynchronously execute the transaction submission operation, that is, write the KV version corresponding to the slave target node to the corresponding version column, that is, write the corresponding Key_commit_ts--->read_ts.
[0091] Further, after the distributed transaction is successfully committed, if it is necessary to read a certain KV, the latest read_ts of the corresponding KV can be obtained from the version column first, and then the data (that is, value) corresponding to the read_ts can be obtained through the data column.
[0092] In summary, in this embodiment, after the main target node feedbacks a message indicating successful submission, a message indicating successful data operation can be sent to the client, and then the submission phases of other slave target nodes are executed asynchronously and concurrently, which further reduces the latency and improves the data throughput for the client.
[0093] The distributed storage system according to the embodiment of the present invention sets a coordinator node, and enables the coordinator node to receive a data operation request. In response to the data distribution in the data operation range corresponding to the data operation request on multiple target nodes, the corresponding data on each target node is called to execute the data operation request, and the data processing results corresponding to each target node are obtained, and each data processing result is written to the corresponding target node. Thus, in this embodiment, the coordinator node processes data operation transactions distributed on different target nodes, and then writes the data processing results to the corresponding target nodes respectively, ensuring the correctness and scalability of the distributed storage system.
[0094] Further, the data operation request in this embodiment can also be a single-machine transaction. In this regard, in response to the data distribution in the data operation range corresponding to the data operation request on one target node, the target node is controlled to execute the data operation request and submit the corresponding data processing result. Optionally, in this embodiment, the single-machine transactions on each node can achieve data consistency for concurrent processing through the Raft protocol, which will not be elaborated here.
[0095] In an alternative implementation, after receiving a data operation request from the service layer 11, the server 12 sends the data operation request to the coordination node in the distributed storage system 13, so that the coordination node determines whether the transaction corresponding to the data operation request is a single-machine transaction or a distributed transaction. If the transaction corresponding to the data operation request is a single-machine transaction, the coordination node sends the data operation request to the corresponding target node, so that the target node executes the data operation request and submits the corresponding data processing result. If the transaction corresponding to the data operation request is a distributed transaction, the coordination node invokes the corresponding data on each target node through the configuration information to execute the data operation request, and caches the data processing results corresponding to each target node respectively.
[0096] In another alternative implementation, if the server 12 maintains configuration information, where the configuration information stores the data ranges stored by each node in the distributed storage system 13, the server 12 can determine whether the data operation request is a single-machine transaction or a distributed transaction through the configuration information, send the data operation request corresponding to the distributed transaction and the data operation request corresponding to the single-machine transaction with the coordination node as the target node to the coordination node, and send the data operation request corresponding to the single-machine transaction with the data node as the target node to the corresponding data node.
[0097] In yet another alternative implementation, since file operations are basically single-machine transactions and directory operations include single-machine transactions and distributed transactions, the coordination node or the server 12 can also first determine whether the operation corresponding to the data operation request is a file operation or a directory operation. If the data operation request is a file operation, it can be directly determined as a single-machine transaction, and the node where the file to be operated is located is determined as the target node. If the data operation request is a directory operation, the target node is determined through the data distribution in the data operation range, which can further improve the data processing efficiency and reduce latency.
[0098] In an alternative implementation, this embodiment adopts the HybridTxn model to implement the combination of single-machine transactions and distributed transactions in the embodiment to ensure low latency, high scalability, and high throughput. Among them, the HybridTxn model is a transaction processing model that combines the characteristics of OLTP (Online Transaction Processing) and OLAP (Online Analytical Processing). While maintaining the transaction processing ability of the OLTP system, it provides the analysis ability of the OLAP system, and can provide efficient, flexible, and scalable transaction processing and analysis capabilities.
[0099] On-Line Transaction Processing (OLTP) is a technology for processing business operations in a real-time environment. In an OLTP system, each business operation is regarded as a transaction, and this transaction must satisfy the ACID properties: atomicity, consistency, isolation, and durability. This means that in a transaction, all operations either succeed or fail completely, ensuring data integrity and consistency.
[0100] On-Line Analytical Processing (OLAP) is a data analysis and processing technology that can be used to quickly perform complex analysis and queries on large-scale data. It is mainly used in data warehouses and intelligent systems. Through multidimensional data analysis, it can explore and analyze data from multiple dimensions, enabling multi-angle and multi-dimensional analysis of data, thus better supporting decision-making analysis.
[0101] In an optional implementation manner, to maintain the concurrency of data operations, each node (coordination node or data node) in the distributed storage system of the embodiments of the present invention also maintains a CoordinateLock pessimistic lock. Transactions of different metadata trees within a single node will be serially executed with locking at the master node, and transactions in different nodes can be executed concurrently.
[0102] The embodiments of the present invention adopt a combination of single-machine transactions and distributed transactions, such that metadata operations within a single data tree fall on the same target node, and metadata operations across metadata trees are distributed on different target nodes. While ensuring data correctness and high system scalability, low latency is also guaranteed.
[0103] In an optional implementation manner, as described above, when the load and / or storage resource consumption of the metadata tree on each node in the distributed storage system of this embodiment is too large, the distributed storage system can be expanded by splitting the metadata tree to obtain corresponding metadata subtrees, that is, new metadata trees, thereby ensuring data balance in the distributed storage system.
[0104] Further optionally, in this embodiment, when it is detected that the load amount and / or storage resource consumption amount of the metadata tree on a certain node in the distributed storage system is greater than a predetermined load threshold or storage resource consumption threshold, the target migration amount is determined based on the current load amount and / or storage resource consumption amount of the metadata tree, and the metadata subtree of the split tree is determined based on the migration amount, and the subtree is migrated to a node with a smaller load amount and storage resource consumption amount (that is, the corresponding data operation delay change amount is less than a predetermined value after accommodating the target migration amount), or can also be migrated to a newly allocated node (that is, a node that does not store the metadata tree). Thereby, the data balance of the distributed storage system can be ensured.
[0105] In an alternative implementation, in this embodiment, the load status information of the metadata tree can be determined by the storage resource consumption amount and / or operation request parameters in the node, and the target metadata tree to be split is determined according to the load status information of the metadata tree. Among them, the storage resource consumption amount is used to represent the size of the storage space occupied by the metadata tree, and the operation request parameters can be determined according to the number of requests of the data operation requests corresponding to the metadata tree, and / or information such as the data operation time or delay corresponding to the data operation requests.
[0106] Further optionally, when it is determined that the metadata tree is in an overloaded state based on the load status information of the metadata tree in this embodiment, it is determined as the target metadata tree to be separated. Optionally, in this embodiment, multiple load status information of the metadata tree within a predetermined time period can be obtained, and the load score corresponding to each load status information is determined. In response to the existence of a predetermined number of load scores greater than the load threshold, the metadata tree is determined to be in an overloaded state, that is, determined as the target metadata tree to be split. Thereby, this embodiment can avoid the problem of resource waste caused by the metadata tree being split due to accidental overload.
[0107] In an alternative implementation, in this embodiment, the storage resource consumption amount of the metadata tree can be converted into a corresponding score. For example, the value or normalized value of the storage resource consumption amount or the calculated value with a predetermined constant is used as the score of the storage resource consumption amount, or the ratio of the storage resource consumption amount to the total storage space of the node where the metadata tree is located, or the normalized value of the ratio or the calculated value of the ratio with a predetermined constant is used as the score of the storage resource consumption amount. This embodiment does not limit this.
[0108] In an alternative implementation, in this embodiment, the operation request scores corresponding to different data operation requests are determined based on the different degrees of load imposed on the nodes by different data operation requests, and the scores corresponding to the operation request parameters are determined based on the quantity and type of each data operation request corresponding to the metadata. For example, for operation request categories such as StatFS, Open, Close, Access, GetAttr, Resolve, ReadLink, GetXattr, ListXattr, and RemoveXAttr, the corresponding operation request scores can be set to 1 point; for operation request categories such as Lookup, Read, NextSlice, and NextINode, the corresponding operation request scores can be set to 2 points; for operation request categories such as Write, Create, MKnod, Mkdir, Rename, SetAttr, Rmdir, Unlink, Truncate, Fallocate, Flock, SetLlk, Link, Symlink, AppendFile, CommitCompact, NewSession, and SessionHeartbeat, the corresponding operation request scores can be set to 6 points; for operation request categories such as ReadDir and CommitAppend, the corresponding operation request scores can be set to 10 points. It should be understood that the above scoring criteria are merely exemplary, and this embodiment does not limit this, and it can be specifically set based on the operating conditions of the distributed storage system.
[0109] In an alternative implementation, in this embodiment, the score of the storage resource consumption can be used as the load score of the corresponding metadata tree, or the score corresponding to the operation request parameter can be used as the load score of the metadata tree, or the weighted value of the score of the storage resource consumption and the score corresponding to the operation request parameter can be used as the load score. This embodiment does not limit this. Further optionally, when determining whether the metadata tree is overloaded according to the combination of the storage resource consumption and the operation request parameter, this embodiment can first determine whether there are a predetermined number of scores of the storage resource consumption exceeding the corresponding first threshold (the first threshold can be determined according to the total storage space size corresponding to the node) within a predetermined time period. If it exceeds, it is determined that the metadata tree is overloaded. If it does not exceed, it can further determine whether there are a predetermined number of scores of the operation request parameter exceeding the corresponding second threshold (the second threshold can be determined according to the score of the operation request parameter when the total delay of the node does not exceed the delay threshold) within a predetermined time period. If it exceeds, it is determined that the metadata tree is overloaded. Further optionally, if the load score is the weighted value of the score of the storage resource consumption and the score corresponding to the operation request parameter, it can directly determine whether there are a predetermined number of load scores exceeding the load threshold within a predetermined time period. If it exceeds, it is determined that the metadata tree is overloaded. It should be understood that in this embodiment, the weights of the score of the storage resource consumption and the score corresponding to the operation request parameter are not limited and can be adjusted based on the actual application situation of the distributed storage system.
[0110] Optionally, this embodiment does not limit the above-mentioned predetermined time period and the predetermined number of load scores exceeding the load threshold within the predetermined time period, which can be determined according to the actual load situation of the distributed storage system. For example, if the metadata tree update frequency is high and the server cluster in the distributed storage system needs to perform load balancing every 60 minutes, the predetermined time period can be set to 30 minutes, 45 minutes, etc.; if the number of load scores in the predetermined time period is 30, the predetermined number can be set to 10, 15, etc.
[0111] Figure 7 It is a schematic diagram of the method for determining the load status of the metadata tree according to the embodiment of the present invention. In an alternative implementation, in this embodiment, the sliding window algorithm (Rolling Window) is used to determine whether the metadata tree is in an overloaded state. Specifically, as Figure 7 shown, in this embodiment, n load scores of the metadata tree M, that is, load score s1 - load score sn (where n is greater than or equal to 1), are stored in the sliding window queue 70 in the order of acquisition. Each window in the sliding window queue 70 stores load scores s1 - sn.
[0112] Furthermore, this embodiment sequentially determines whether the load scores in the sliding window queue 70 exceed the load threshold to update the sliding window queue 70. AsFigure 7 As shown, when it is determined that the load score s1 is less than the load threshold, the load score s1 is removed from the sliding window queue 70, and the (n + 1)-th load score sn+1 is obtained to form a new sliding window queue 70', so as to continue to determine the load status of the metadata tree M based on the new sliding window queue 70'.
[0113] Further, in this embodiment, the load scores in the sliding window queue are counted separately. For example, if the i-th (1 ≤ i ≤ n) load score exceeds the load threshold, the hot degree of this metadata tree is incremented by 1, and the counter value of this server group is determined to be 0; if the i-th load score does not exceed the load threshold, the hot degree of this metadata tree is decremented by 1, and the counter value of this metadata tree is incremented by 1. When the counter value of this metadata tree is lower than the minimum counter parameter, that is, a predetermined number of load scores among the multiple load scores corresponding to this metadata tree exceed the load threshold, that is, when the hot degree of this metadata tree is not lower than the minimum hot degree parameter, it is determined that this metadata subtree is in an overloaded state. When the counter value of this metadata tree is not lower than the minimum counter parameter, that is, more than a certain number of load scores among the multiple load scores corresponding to this metadata tree do not exceed the load threshold, it is determined that this metadata tree is not overloaded.
[0114] The management server can determine that the first load status of this metadata tree is not overloaded; when the hot degree of this metadata tree is not lower than the minimum hot degree parameter, the management server can determine that the first load status of this metadata tree is overloaded. Wherein, n is the number of load status information of this server group obtained in a predetermined time period.
[0115] In an alternative implementation, when the multi-metadata tree is in a non-overloaded state, this embodiment can also determine whether the node corresponding to this metadata tree can be migrated into the metadata subtree according to the specific value of the load status information corresponding to this metadata tree, and how much is the migration amount (that is, the load amount that can be migrated). Furthermore, this embodiment can match the migration amount of the target subtree to be migrated out and the migration amount of the nodes that can be migrated into the metadata subtree, so as to perform the migration in and out of the metadata subtree. Further optionally, if the load of the metadata subtree to be migrated out of the metadata tree is large and no eligible nodes for migration are found, new nodes can be allocated by the server to migrate into the split metadata subtree. In another alternative implementation, for a metadata tree with a small migration amount, in order to save resources and improve the overall load balance of the system, it can also be selected not to migrate out.
[0116] In this embodiment, after determining the load scores of the same metadata tree at different moments within a predetermined time period, the metadata trees with a predetermined number of load scores greater than the load threshold within the predetermined time period are determined as the target metadata trees that need to be split, avoiding the large impact of large short-term load fluctuations on the load status evaluation of the metadata tree, and further ensuring the load balance of the distributed storage system.
[0117] In an alternative implementation, after determining the target metadata tree to be split and the nodes of the metadata subtree to be migrated, this embodiment can split the target metadata tree, that is, update the configuration information of the target metadata tree and generate the configuration information of the newly split metadata tree. Among them, the configuration information of the metadata tree includes the data range stored in the metadata tree and the node identifier where it is located. Optionally, the configuration information of the target metadata tree may further include the data range stored in the newly split metadata tree and / or the node identifier where the newly split metadata tree is located.
[0118] In an alternative implementation, this embodiment can use any metadata tree migration method to implement the migration of the metadata tree. For example, it can pause the data operation requests of the new metadata tree, and after successfully migrating the new metadata tree to the corresponding node, receive and process the corresponding data operation requests again.
[0119] Figure 8 It is a schematic diagram of the process of the metadata tree migration method of the embodiment of the present invention. Further optionally, in order to avoid the latency problem caused by the inability to access data during the metadata tree migration process, this embodiment adopts the metadata tree migration method as shown in Figure 8 to perform the migration of the metadata tree. As shown in Figure 8 , this embodiment describes the example of splitting the metadata tree Subtree0 stored in the coordination node 81 into a new metadata tree Subtree1 and migrating the new metadata tree Subtree1 to the data node 82. It should be understood that the migration of the metadata tree can also occur between each data node, and this embodiment does not limit this.
[0120] As shown in Figure 8 , the coordination node 81 represents a node cluster, including a coordination master node 811, a coordination slave node 812, and a coordination slave node 813. Among them, before the migration of the metadata tree Subtree1, the master nodes of both the metadata tree Subtree0 and the metadata tree Subtree1 can be the coordination master node 811.
[0121] Further, after receiving the metadata tree splitting instruction (SplitSubTree), the coordination master node 811 splits the target metadata tree Subtree0, updates the configuration information of the target metadata tree Subtree0, and obtains a new metadata tree Subtree1. In this embodiment, after the coordination master node 811 splits out the new metadata tree Subtree1, it sends a data node acquisition request to the server to obtain the identifier of the migration node 82 corresponding to the new metadata tree Subtree1. Among them, the migration node 82 is a data node cluster, including node 821, node 822, and node 823.
[0122] Further, after obtaining the identifier of the migration node 82 fed back by the server, the coordination master node 811 generates the configuration information of the new metadata tree Subtree1. Optionally, after generating the configuration information of the new metadata tree Subtree1, the coordination master node 811 also sends the updated metadata tree Subtree0 and the configuration information of the new metadata tree Subtree1 to the configuration module of the distributed storage system (specifically, it can be the root server RootServer) so that the configuration module updates the configuration information of the corresponding metadata tree.
[0123] In this embodiment, the update of the metadata tree Subtree0 includes updating the node change information in the metadata tree and the node range information of the metadata tree Subtree0 (that is, updating the data range of the metadata). Further optionally, the configuration information after the change of the metadata tree Subtree0 may also include the node range information of the split new metadata tree Subtree1 and / or the identifier of the migration node 82, so as to facilitate the data call when the coordination node executes the distributed transaction subsequently. Similarly, the configuration information of the new metadata tree Subtree1 includes the node range information of the new metadata tree Subtree1 (that is, updating the data range of the metadata) and the identifier of the data node 82 where it is located.
[0124] It can be seen that at this time, the distributed storage system has only completed the splitting of the target metadata tree Subtree0, but the master node of the new metadata Subtree1 is still the coordination master node 811, that is, at this time, the data operation requests corresponding to the new metadata Subtree1 are still executed and submitted by the coordination master node 811.
[0125] Further, the server of this embodiment sends a first migration instruction for the coordinated slave node 812 to the coordinated master node 811 (or directly sends the first migration instruction to the coordinated slave node 812) to migrate the replicated subtree Subtree1' of the new metadata tree Subtree1 in the coordinated slave node 812 to a node 822 in the incoming node cluster 82. It should be understood that in this embodiment, the first migration instruction can be first sent to any one of the coordinated slave nodes in the coordinated node cluster 81 to migrate the corresponding replicated subtree Subtree1' to any one of the nodes in the incoming node cluster 82. This embodiment does not limit the first migrated coordinated slave node and the first incoming node.
[0126] Further, after the coordinated slave node 812 migrates the corresponding replicated subtree Subtree1' to the node 822 based on the first migration instruction, the new metadata tree Subtree1 in the coordinated slave node 812 is deleted. Specifically, the server sends a Subtree RemovePeer instruction to the coordinated slave node 812 and a Subtree AddPeer instruction to the incoming node cluster 82 or the node 822 in the incoming node cluster 82 to create the new metadata tree Subtree1 in the node 822 and remove the new metadata tree Subtree1 in the coordinated slave node 812.
[0127] Further, during the migration of the coordinated slave node 812, the coordinated master node 811 still acts as the master node of the new metadata tree Subtree1 to execute relevant data operation requests, updates the metadata tree Subtree1, and synchronizes the data of the metadata tree Subtree1 with the coordinated slave node 813.
[0128] In an alternative implementation, the master node transfer operation of the metadata tree Subtree1 can be performed at this time, or the metadata tree Subtree1 in the coordinated slave node 813 can be first migrated to a node in the incoming node cluster 82. This embodiment does not limit this. The following takes the migration of the metadata tree Subtree1 in the coordinated slave node 813 as an example for specific description.
[0129] The coordinated slave node 813 migrates the replicated subtree Subtree1” corresponding to the metadata tree Subtree1 to a node 821 in the incoming node cluster 82 based on the second migration instruction issued by the server, and deletes the new metadata tree Subtree1 in the coordinated slave node 813.
[0130] At this time, the server sends a master node transfer instruction (i.e., the TransferLeader instruction) of the new metadata tree Subtree1 to the coordination master node 811. The coordination master node 811 sends a master node transfer request to the current corresponding child nodes of the new metadata tree Subtree1 (i.e., node 821 and node 822). After receiving the master node transfer request of the new metadata tree Subtree1, node 821 and node 822 send a feedback message carrying the identifier of the newly elected master node to the coordination master node 811. Assuming that the identifier of the new master node is node 821, after receiving the feedback messages of all the slave nodes of the new metadata tree Subtree1, the coordination master node 811 marks the master node status of its new metadata tree Subtree1 as the transferred state, that is, the coordination master node 811 is converted into a slave node of the new metadata tree Subtree1. After receiving the message that it is elected as the new master node, node 821 marks its own status as the master node and starts to receive and process the data processing operations corresponding to the new metadata tree Subtree1.
[0131] After that, based on the third migration instruction sent by the server, the coordination master node 811 migrates the replication subtree Subtree1”' corresponding to the metadata tree Subtree1 to a node 823 in the incoming node cluster 82, and deletes the new metadata tree Subtree1 in the coordination master node 811.
[0132] At this time, the migration operation of the new metadata tree Subtree1 is completed. Among them, node 821 is the master node (i.e., the Leader node) of the metadata tree Subtree1, and node 822 and node 823 are the slave nodes (i.e., the Follower nodes) of the metadata tree Subtree1. After that, data synchronization is performed among node 821, node 822 and node 823 to ensure the data consistency of the metadata tree Subtree1.
[0133] In summary, the metadata tree migration method of this embodiment can first generate the configuration information of the new metadata tree in the outgoing node cluster, first migrate the replication subtree corresponding to the slave node of the new metadata tree in the outgoing node cluster, and perform the master node transfer operation of the new metadata tree. Thus, in the metadata tree migration process of this embodiment, only the execution of the data operation request of the new metadata tree is paused when performing the master node transfer operation, which greatly reduces the data processing delay caused by the data migration process, improves the data throughput of the system, and at the same time ensures the high scalability of the system.
[0134] In the distributed storage system of this embodiment, the target metadata tree to be split is determined based on the load status information corresponding to the metadata tree through the sliding window algorithm, and the migration amount of the target metadata tree and the migration amount of the nodes that can be migrated are determined based on the load status information corresponding to each metadata tree, so as to determine the migration nodes. Then, the configuration information of the new metadata tree is generated in the migration node cluster first, and then the replication subtree corresponding to the slave node of the new metadata tree in the migration node cluster is migrated out, and then the master node transfer operation of the new metadata tree is executed, so that during the metadata tree migration process, only the execution of the data operation request of the new metadata tree is paused when the master node transfer operation is executed, which greatly reduces the data processing delay caused by the data migration process, improves the data throughput of the system, and at the same time ensures the high scalability of the system. In the multi-node distributed storage system formed after the above data migration, a combination of single-machine transactions and distributed transactions is adopted, so that the metadata operations within a single data tree fall on the same target node, and the metadata operations across metadata trees are distributed on different target nodes, ensuring both data correctness and high scalability of the system, while ensuring low latency.
[0135] Figure 9 It is a schematic diagram of the data processing device according to an embodiment of the present invention. The data processing device 9 according to an embodiment of the present invention includes a request receiving unit 91, a processing unit 92, and a writing unit 93.
[0136] The request receiving unit 91 is configured to receive a data operation request, and the data operation request includes a data operation range. The processing unit 92 is configured to, in response to the data in the data operation range being distributed on multiple target nodes, call the corresponding data on the target nodes to execute the data operation request, and obtain the data processing results respectively corresponding to the target nodes. The writing unit 93 is configured to write each of the data processing results to the corresponding target node.
[0137] In an alternative implementation, the distributed storage system is a distributed metadata storage system, and different metadata trees are stored in each of the nodes.
[0138] In an alternative implementation, the metadata tree stored in the data node is determined based on the split of the metadata tree of the coordination node.
[0139] In an alternative implementation, the configuration information of each metadata tree includes the data range of the corresponding metadata tree and the identifier of the split metadata tree.
[0140] In an alternative implementation, the writing unit 93 includes a pre-writing subunit, a timestamp obtaining subunit, and a submission subunit.
[0141] The pre-write subunit is configured to send pre-write instructions to each of the target nodes. The timestamp acquisition subunit is configured to acquire a transaction commit timestamp in response to receiving pre-write success messages fed back by each of the target nodes. The commit subunit is configured to send commit instructions to each of the target nodes based on the transaction commit timestamp to commit the transaction corresponding to the data operation request.
[0142] In an alternative implementation, the data processing device 9 further includes a node type determination unit configured to select one of the target nodes as the main target node and determine the non-target nodes as slave target nodes.
[0143] In an alternative implementation, the commit subunit is further configured to perform:
[0144] Send a first commit instruction to the main target node based on the transaction commit timestamp;
[0145] In response to receiving a commit success message fed back by the main target node, feed back a data operation success message to the client;
[0146] Send a second commit instruction to each of the slave target nodes to cause the slave target nodes to commit the transaction corresponding to the data operation request.
[0147] In an alternative implementation, the commit subunit is further configured to perform:
[0148] Asynchronously send a second commit instruction to each of the slave target nodes.
[0149] In an alternative implementation, the data processing device 9 further includes a single-machine transaction processing unit configured to control the target node to execute the data operation request in response to the data in the data operation range being distributed on one target node.
[0150] In the technical solution of the embodiment of the present invention, the distributed storage system includes multiple nodes, each node stores data in a different range, the multiple nodes include a coordination node and data nodes, the coordination node is used to receive a data operation request, and in response to the data in the data operation range corresponding to the data operation request being distributed on multiple target nodes, call the corresponding data on each target node to execute the data operation request, obtain the data processing results corresponding to each target node respectively, and write each data processing result to the corresponding target node. Thus, in this embodiment, the coordination node processes data operation transactions distributed on different target nodes, and then writes the data processing results to the corresponding target nodes respectively, ensuring the correctness and scalability of the distributed storage system.
[0151] Figure 10 It is a schematic diagram of the electronic device of the embodiment of the present invention. AsFigure 10 As shown, the electronic device 100 is a general data processing device, which includes a general computer hardware structure, and at least includes a processor 101 and a memory 102. The processor 101 and the memory 102 are connected by a bus 103. The memory 102 is adapted to store instructions or programs executable by the processor 101. The processor 101 can be an independent microprocessor or a set of one or more microprocessors. Thus, by executing the instructions stored in the memory 102, the processor 101 executes the method flow of the embodiment of the present invention as described above to implement data processing and control of other devices. The bus 103 connects the above-mentioned multiple components together, and at the same time connects the above-mentioned components to a display controller 104, a display device, and an input / output (I / O) device 105. The input / output (I / O) device 105 can be a mouse, a keyboard, a modem, a network interface, a touch input device, a body sensing input device, a printer, and other devices well known in the art. Typically, the input / output device 105 is connected to the system through an input / output (I / O) controller 106.
[0152] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device (equipment), or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be implemented as a computer program product on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0153] The present application is described with reference to the flowcharts of methods, devices (equipment), and computer program products according to the embodiments of the present application. It should be understood that each process in the flowchart can be implemented by computer program instructions.
[0154] These computer program instructions can be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device, and the instruction device implements the process Figure 1 the functions specified in one process or multiple processes.
[0155] These computer program instructions can also be provided to the processor of a general computer, a special computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices produce a device for implementing the functions specified in Figure 1 one process or multiple processes.
[0156] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, which is used for a computer to execute the above-mentioned partial or all method embodiments.
[0157] That is, those skilled in the art can understand that all or part of the steps in implementing the above-mentioned embodiment methods can be completed by specifying relevant hardware through a program. This program is stored in a storage medium and includes several instructions to enable a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0158] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A data processing method, characterized in that, Applied to a coordination node in a distributed storage system, the distributed storage system includes multiple nodes, each of the nodes stores data in a different range, the multiple nodes include a coordination node and data nodes, and the method includes: Receiving a data operation request, the data operation request including a data operation range; In response to the data in the data operation range being distributed on multiple target nodes, calling the corresponding data on the target nodes to execute the data operation request, and obtaining data processing results respectively corresponding to each of the target nodes; Writing each of the data processing results to the corresponding target node.
2. The method according to claim 1, wherein The distributed storage system is a distributed metadata storage system, and each of the nodes stores a different metadata tree.
3. The method according to claim 2, wherein The metadata tree stored in the data node is determined based on the splitting of the metadata tree of the coordination node.
4. The method according to claim 2 or 3, characterized in that, The configuration information of each of the metadata trees includes the data range of the corresponding metadata tree and the identifier of the split metadata tree.
5. The method according to claim 1, wherein The writing each of the data processing results to the corresponding target node includes: Sending a pre-write instruction to each of the target nodes; In response to receiving pre-write success messages feedback from each of the target nodes, obtaining a transaction commit timestamp; Sending a commit instruction to each of the target nodes based on the transaction commit timestamp to commit the transaction corresponding to the data operation request.
6. The method according to claim 5, characterized in that, The method includes: Selecting one target node as the main target node and determining the non-target nodes as slave target nodes.
7. The method according to claim 6, wherein The sending a commit instruction to each of the target nodes based on the transaction commit timestamp includes: Sending a first commit instruction to the main target node based on the transaction commit timestamp; In response to receiving a commit success message feedback from the main target node, sending a data operation success message to the client; Sending a second commit instruction to each of the slave target nodes to cause the slave target nodes to commit the transaction corresponding to the data operation request.
8. The method according to claim 7, characterized in that The sending a second commit instruction to each of the slave target nodes includes: Asynchronously sending a second commit instruction to each of the slave target nodes.
9. The method according to claim 1, characterized in that, The method further includes: In response to the data in the data operation range being distributed on one target node, controlling the target node to execute the data operation request.
10. A distributed storage system, characterized in that, The distributed storage system includes multiple nodes, each of the nodes stores data in a different range, and the multiple nodes include: A coordination node configured to execute the method according to any one of claims 1-9; and Data nodes configured to perform a commit operation on the data processing results sent by the coordination node.
11. The system according to claim 10, wherein The data nodes are further configured to execute the data operation requests sent by the coordination node and perform a commit operation on the obtained data processing results.
12. The system according to claim 10 or 11, characterized in that The distributed storage system is a distributed metadata storage system, each of the nodes stores a different metadata tree, the metadata tree stored in the data node is determined based on the splitting of the metadata tree of the coordination node, and the configuration information of each of the metadata trees includes the data range of the corresponding metadata tree and the identifier of the split metadata tree.
13. A data processing device, characterized in that, The apparatus includes: A request receiving unit configured to receive a data operation request, the data operation request including a data operation range; A processing unit, configured to call corresponding data on the target nodes to execute the data operation request and obtain data processing results respectively corresponding to the target nodes in response to the data in the data operation range being distributed on multiple target nodes; A writing unit, configured to write the data processing results into corresponding target nodes.
14. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method according to any one of claims 1-9.
15. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method according to any one of claims 1-9 is implemented.