A method for concurrent data writing and a distributed system for concurrent data writing
By generating a unique and incremental version number for the target data in the distributed file storage system and carrying the version number in the write request, the problem of data inconsistent when multiple clients write concurrently to the same target data is solved, data consistency and reliability are achieved, and the system's concurrent writing capabilities are improved.
Patent Information
- Application Number
- CN202011431452.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-09
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-12-09
AI Technical Summary
In a distributed file storage system, when multiple clients write concurrently to the same target data, data inconsistency may result in data inconsistency, affecting the reliability of the data, and limiting the system's capacity in concurrent writes.
By generating a unique and incremental version number for the target data at the control node and sending it with the storage node address information to the client, the client carries the version number when sending a write request, so that the storage node can judge the validity and sequence of the write operation based on the version number.
It realizes that when multiple clients write concurrently to the same target data in a distributed file storage system, data consistency and reliability are ensured, data inconsistency is solved, and the system's concurrent writing capabilities are improved.
Smart Images

Figure CN112486932B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method for concurrent data writing and a distributed system for concurrent data writing. Background Art
[0002] A distributed file storage system refers to a file system where the physical storage resources managed are not necessarily directly connected to the local node, but are connected to the node (which can be simply understood as a computer) through a computer network; or it is a complete hierarchical file system formed by combining several different logical disk partitions or volume labels. The distributed file storage system disperses a large amount of data to different nodes for storage, greatly reducing the risk of data loss.
[0003] The existing method for a client to perform a single data write is as follows: The client sends a write request for the target data to the control node; the control node receives and verifies the write request. When the verification is successful, it returns the address information of the storage node corresponding to the data to be written to the client; the client receives the address information and sends a data write request containing the data to be written to the storage node corresponding to the data to be written; the storage node receives and executes the write operation corresponding to the data write request, and then returns write result information indicating whether the write is successful or failed to the client; the client receives the write result information returned by the storage node and sends the write result information to the control node; when the control node determines that the write result is successful, it updates the metadata information corresponding to the target data.
[0004] In practical applications, to ensure data reliability, multiple copies of file data are generally saved. However, when using the above data write method to write data, when multiple clients perform write operations on the same target data, due to network latency and the actual business conditions of the storage nodes where each copy is located, it may lead to inconsistent data in multiple copies corresponding to the target data. For example, client A and client B simultaneously perform write operations on the same target data. For copy 1 of the target data, client A performs the write operation first and client B performs the write operation later. For copy 2 of the target data, client B performs the write operation first and client A performs the write operation later, resulting in inconsistent data in copy 1 and copy 2 of the same target data, thereby affecting data reliability. In a distributed file storage system, only serial writing is allowed for the same file, and concurrent writing is not allowed. Summary of the Invention
[0005] The purpose of the embodiments of the present invention is to provide a method for concurrent data writing and a distributed system for concurrent data writing, so as to solve the problem of inconsistent data caused by multiple clients performing write operations on the same file, which affects data reliability. The specific technical solutions are as follows:
[0006] In a first aspect, an embodiment of the present invention provides a method for concurrent data writing, which is applied to a control node in a distributed file storage system. The distributed file storage system includes: a control node and a storage node. The method includes:
[0007] Verify a first write request for target data sent by a client, and generate a version number corresponding to the first write request to obtain a first version number. The first version number increases as the number of times the target data is written increases. The first write request corresponds to corresponding data to be written;
[0008] When the verification of the first write request is successful, select a preset number of storage nodes for the data to be written to obtain address information corresponding to the preset number of storage nodes;
[0009] Send the address information corresponding to the preset number of storage nodes and the first version number to the client, so that the client sends a second write request to the preset number of storage nodes. The preset number of storage nodes respectively perform corresponding write operations according to the second write request and return write result information to the client; wherein, the second write request includes the data to be written and the first version number;
[0010] Update the metadata information corresponding to the target data based on the write result information returned by the client; wherein, the metadata information corresponding to the target data includes the size of the target data and information about the storage nodes to which the target data is written.
[0011] Optionally, the step of generating a version number corresponding to the first write request to obtain a first version number includes:
[0012] Generate a first version number that is higher than the version number corresponding to the most recent first write request for the target data.
[0013] Optionally, the step of selecting a preset number of storage nodes for the data to be written includes:
[0014] Select a preset number of storage nodes for the data to be written according to the size of the data to be written corresponding to the first write request and the remaining space size of each storage node.
[0015] Optionally, the method further includes:
[0016] Receive the self-status information sent by the storage node. The self-status information of the storage node includes: the node status of the storage node and the remaining space size of the storage node.
[0017] Second aspect, an embodiment of the present invention provides a method for concurrent data writing, which is applied to a storage node in a distributed file storage system. The distributed file storage system includes: a control node and a storage node. The method includes:
[0018] Receiving a second write request sent by a client, where the second write request includes a first version number corresponding to a first write request for target data, and the first version number increases as the number of times the target data is written increases;
[0019] Determining whether the first version number is higher than the local version number of the target data, where the local version number of the target data is the version number carried in the second write request during the last data write of the target data locally;
[0020] If the first version number is higher than the local version number of the target data, performing the write operation corresponding to the second write request, and returning write result information indicating successful writing to the client, so that the client sends the write result information to the control node, and the control node updates the metadata information corresponding to the target data.
[0021] Optionally, the method further includes:
[0022] If the first version number is not higher than the local version number of the target data, rejecting the write operation corresponding to the second write request, and returning write result information indicating failed writing to the client, so that the client sends the write result information to the control node.
[0023] Optionally, the second write request further includes the data to be written corresponding to the first write request for the target data; the step of performing the write operation corresponding to the second write request includes:
[0024] Updating the local data of the target data using the data to be written;
[0025] Updating the local version number of the target data using the first version number.
[0026] Optionally, the method further includes:
[0027] Sending its own status information to the control node, where the own status information includes: the node status and the remaining space size.
[0028] Third aspect, an embodiment of the present invention provides a distributed data concurrent writing system, the distributed data concurrent writing system includes: a control node and a storage node;
[0029] The control node is used to verify the first write request for the target data sent by the client and generate a first version number corresponding to the first write request. The first write request corresponds to the data to be written. When the verification of the first write request is successful, a preset number of storage nodes are selected for the data to be written, and the address information corresponding to the preset number of storage nodes is obtained. The address information corresponding to the preset number of storage nodes and the first version number are sent to the client, and the metadata information corresponding to the target data is updated based on the write result information returned by the client. Among them, the first version number increases as the number of times the target data is written increases. The metadata information corresponding to the target data includes the size of the target data and the information of the storage nodes where the target data is written.
[0030] The storage node is used to receive the second write request sent by the client, determine whether the first version number included in the second write request is higher than the local version number of the target data. When it is determined that the first version number is higher than the local version number of the target data, the write operation corresponding to the second write request is executed, and the write result information indicating successful writing is returned to the client. Among them, the second write request includes the first version number corresponding to the first write request for the target data and the data to be written. The local version number of the target data is the version number carried in the second write request when the target data was last written locally.
[0031] Optionally, the system further includes a client.
[0032] The client is used to send a first write request for the target data to the control node, and receive the address information corresponding to a preset number of storage nodes and the first version number returned by the control node; and send the second write request to the preset number of storage nodes, receive the write result information returned by the preset number of storage nodes, and send the write result information to the control node.
[0033] Optionally, the control node is specifically used for:
[0034] Generate a first version number higher than the version number corresponding to the most recent historical first write request for the target data.
[0035] Optionally, the storage node is further used for:
[0036] When it is determined that the first version number is not higher than the local version number of the target data, reject the write operation corresponding to the second write request, and return the write result information indicating write failure to the client.
[0037] Optionally, the storage node is specifically configured to:
[0038] Use the data to be written to update the local data of the target data;
[0039] Use the first version number to update the local version number of the target data.
[0040] Optionally, the storage node is further configured to:
[0041] Send its own status information to the control node, where the own status information includes: the node status and the remaining space size.
[0042] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method steps of a data concurrent writing method described in the first aspect above are implemented.
[0043] In a fifth aspect, an embodiment of the present invention provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the method steps of a data concurrent writing method described in the second aspect above are implemented.
[0044] Beneficial effects of the embodiments of the present invention:
[0045] A data concurrent writing method and a distributed data concurrent writing system provided by the embodiments of the present invention. Since the control node can generate a corresponding first version number for the first writing request for the target data sent by the client, and the first version number increases as the number of times of writing the target data increases, that is, the version number for the same target data is unique and increasing, so that when the storage node receives the second writing request, it can determine whether to execute the current data writing operation according to the first version number, ensuring the consistency of the data of multiple copy files when multiple clients concurrently write to the same target data, solving the problem of inconsistent data of multiple copy files when multiple clients concurrently write to the same target data, and improving the reliability of the data. Further, in a distributed file storage system, it is possible to allow multiple clients to concurrently write to the same target data.
[0046] Of course, it is not necessary for any product or method implementing the present invention to simultaneously achieve all the above-mentioned advantages. Description of the Drawings
[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings.
[0048] Figure 1 It is a schematic flow chart of a data concurrent writing method provided by an embodiment of the present invention;
[0049] Figure 2 It is a schematic flow chart of another data concurrent writing method provided by an embodiment of the present invention;
[0050] Figure 3 It is a schematic flow chart of yet another data concurrent writing method provided by an embodiment of the present invention;
[0051] Figure 4 It is a schematic flow chart of a data writing implementation manner provided by an embodiment of the present invention;
[0052] Figure 5 It is an interaction schematic diagram of a data concurrent writing method provided by an embodiment of the present invention;
[0053] Figure 6 It is a schematic structural diagram of a data concurrent writing system provided by an embodiment of the present invention. Specific embodiments
[0054] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.
[0055] Distributed file storage systems have redundancy. The failure of some nodes does not affect the normal operation of the whole. Moreover, even if the data stored in a faulty computer is damaged, it can be recovered by other nodes. A typical distributed file storage system generally consists of a control node (or metadata server) and storage nodes (or data servers). The main functions of the control node are: managing the directory tree, managing the metadata information of the files uploaded by clients (or users) (such as file directories, file names, upload times of files, file sizes, which storage nodes the file data is distributed on, etc.), and managing all storage nodes in the distributed file storage system. The main functions of the storage nodes are: storing the file data uploaded by clients.
[0056] To solve the problem that when multiple clients perform write operations on the same target data, it may cause data inconsistency in multiple copies corresponding to the target data, an embodiment of the present invention provides a data concurrent writing method, which is applied to a control node in a distributed file storage system. The distributed file storage system includes: a control node and a storage node. The method may include:
[0057] Verify the first write request for the target data sent by the client, and generate a version number corresponding to the first write request to obtain a first version number. The first version number increases as the number of times the target data is written increases. The first write request corresponds to the data to be written;
[0058] When the verification of the first write request is successful, select a preset number of storage nodes for the data to be written to obtain the address information corresponding to the preset number of storage nodes;
[0059] Send the address information corresponding to the preset number of storage nodes and the first version number to the client, so that the client sends a second write request to the preset number of storage nodes. The preset number of storage nodes respectively perform corresponding write operations according to the second write request and return the write result information to the client; wherein, the second write request includes the data to be written and the first version number;
[0060] Update the metadata information corresponding to the target data based on the write result information returned by the client; wherein, the metadata information corresponding to the target data includes the size of the target data and the information of the storage nodes where the target data is written.
[0061] In the data concurrent writing method provided by the embodiment of the present invention, since the control node can generate a corresponding first version number for the first write request for the target data sent by the client, and the first version number increases as the number of times the target data is written increases, that is, the version number for the same target data is unique and increasing. When the storage node receives the second write request, it can determine whether to perform the current data write operation according to the first version number, which ensures the consistency of the data in multiple copy files when multiple clients concurrently write the same target data, solves the problem of data inconsistency in multiple copy files when multiple clients concurrently write the same target data, and improves the reliability of the data. Further, it realizes that in a distributed file storage system, multiple clients can concurrently write the same target data.
[0062] The following provides a detailed introduction to the data concurrent writing method provided by the embodiment of the present invention:
[0063] In an embodiment of the present invention, the data concurrent writing method is applied to a distributed file storage system, which may include a control node and multiple storage nodes. Exemplarily, the control node and the multiple storage nodes may be data storage modules respectively arranged on different servers, etc.
[0064] As Figure 1 shown, an embodiment of the present invention provides a data concurrent writing method, which is applied to a control node in a distributed file storage system. The method may include the following steps:
[0065] S101, verify a first write request for target data sent by a client, and generate a version number corresponding to the first write request to obtain a first version number.
[0066] The first write request may be a write request for target data sent by any client. In an embodiment of the present invention, multiple clients may concurrently perform write operations on the target data at the same time; the control node may verify the write requests for the same target data sent by one or more received clients, and at the same time generate corresponding version numbers for the write requests to obtain the first version number corresponding to the first write request.
[0067] Among them, the first version number increases as the number of times the target data is written increases, and the first write request corresponds to corresponding data to be written. Exemplarily, when the control node receives the third write request for the target data, the generated corresponding version number may be expressed as V3, and when the control node receives the fifth write request for the target data, the generated corresponding version number may be expressed as V5, and so on.
[0068] In an embodiment of the present invention, for the same target data, the version number corresponding to the data write is unique and strictly increasing, so that each time a write operation is performed on the same target data, the version number corresponding to the target data increases, avoiding the problem of data disorder of the target data when multiple clients concurrently write to the same target data.
[0069] As an optional implementation manner of an embodiment of the present invention, the first write request for target data sent by a client may include: the identifier of the client, information such as the size of the data volume corresponding to the first write request for the target data, etc. Furthermore, when the control node verifies the first write request for target data sent by the client, it may be to verify the permissions of the client. For example, when receiving the first write request for target data sent by the client, it may be determined whether the client has the permission to perform the write operation on the target data, etc. When the control node verifies the first write request for target data sent by the client, it may also be to verify the data volume of the data to be written corresponding to the first write request, etc.
[0070] As an optional implementation manner of an embodiment of the present invention, the step of generating a version number corresponding to the first write request to obtain the first version number may include:
[0071] Generate a first version number that is higher than the version number corresponding to the most recent first write request for the target data in history.
[0072] In an embodiment of the present invention, when the control node receives a first write request for the target data sent by the client, it may generate a first version number that is higher than the version number corresponding to the most recent first write request for the target data in history, so that for the same target data, the version numbers corresponding to data writing are unique and strictly increasing.
[0073] S102. When the verification of the first write request is successful, select a preset number of storage nodes for the data to be written to obtain the address information corresponding to the preset number of storage nodes.
[0074] In an embodiment of the present invention, the control node may store a management directory tree of different files, metadata information of files, storage node information, etc. The metadata information may include: file directory, file name, upload time of the file, size of the file, which storage nodes the file data is distributed on, etc. The storage node information may include the remaining space size of the storage node or the available utilization rate of the remaining space of the storage node, etc.
[0075] When the control node successfully verifies the first write request, it may select a preset number of storage nodes for the data to be written to the target data to obtain the address information corresponding to the preset number of storage nodes. Those skilled in the art can set the preset number according to actual needs. Exemplarily, to ensure data reliability, the preset number can be set to 3. Then, when the client needs to perform a write operation on the target data, 3 storage nodes can be selected for the data to be written.
[0076] As an optional implementation manner of an embodiment of the present invention, the step of selecting a preset number of storage nodes for the data to be written may include:
[0077] Select a preset number of storage nodes for the data to be written according to the size of the data to be written corresponding to the first write request and the remaining space size of each storage node.
[0078] In an embodiment of the present invention, the control node may sort the storage nodes capable of storing the data to be written according to the size of the data volume of the data to be written corresponding to the first write request and the remaining space size of each storage node to obtain a storage node sequence, and then select a preset number of storage nodes from the storage node sequence.
[0079] In S103, send the address information corresponding to a preset number of storage nodes and the first version number to the client, so that the client sends a second write request to the preset number of storage nodes. The preset number of storage nodes respectively execute corresponding write operations according to the second write request and return the write result information to the client.
[0080] The control node sends the address information corresponding to a preset number of storage nodes and the first version number to the client, so that the client sends a second write request including the data to be written and the first version number to the preset number of storage nodes. Then, the preset number of storage nodes respectively execute corresponding write operations according to the second write request and return the write result information to the client, and the client then forwards the write result information to the control node.
[0081] In S104, update the metadata information corresponding to the target data based on the write result information returned by the client.
[0082] The control node receives the write result information returned by the client, and determines whether the write result information indicates a successful write or a failed write. If the write result information indicates a successful write, update the metadata information corresponding to the target data, where the metadata information corresponding to the target data includes the size of the target data and the information of the storage node to which the target data is written. If the write result information indicates a failed write, do not update the metadata information corresponding to the target data.
[0083] In the data concurrent write method provided by the embodiment of the present invention, since the control node can generate a corresponding first version number for the first write request for the target data sent by the client, and this first version number increases as the number of times the target data is written increases, that is, the version number for the same target data is unique and increasing. This enables the storage node to determine whether to execute the current data write operation according to the first version number when receiving the second write request, ensuring the consistency of the data of multiple copy files when multiple clients concurrently write to the same target data, solving the problem of inconsistent data of multiple copy files when multiple clients concurrently write to the same target data, and improving the reliability of the data. Further, it realizes that in a distributed file storage system, multiple clients can concurrently write to the same target data.
[0084] As an optional implementation manner of the embodiment of the present invention, the control node can also receive the self-state information sent by the storage node, where the self-state information of the storage node can include: the node state of the storage node and the remaining space size of the storage node. Exemplarily, the state of the storage node can include information such as whether the storage node is working properly and the normal speed of data writing, and the remaining space size of the storage node is the size of the space that the storage node can still utilize currently.
[0085] In an embodiment of the present invention, the control node may receive the status information of the storage node sent by the storage node, and then update the status information of the storage node stored in the control node, so as to more accurately select the storage node to be written for the data to be written.
[0086] As Figure 2 shown, an embodiment of the present invention provides another method for concurrent data writing, which is applied to a storage node in a distributed file storage system. The storage node may be one of a preset number of storage nodes selected by the control node for the data to be written. The method may include the following steps:
[0087] S201, Receive a second write request sent by the client.
[0088] Among them, the second write request may include the first version number corresponding to the first write request for the target data, and the first version number increases as the number of times the target data is written increases.
[0089] S202, Determine whether the first version number is higher than the local version number of the target data.
[0090] The storage node receives the second write request sent by the client, which includes the first version number corresponding to the first write request for the target data, and then determines the relationship between the first version number and the local version number of the target data. The local version number of the target data may be: the version number carried in the second write request corresponding to the last data write of the target data locally. If it is determined that the first version number is higher than the local version number of the target data, it means that the second write request for the target data is the latest data write request for the target data, so the steps of S203 are executed.
[0091] S203, If the first version number is higher than the local version number of the target data, perform the write operation corresponding to the second write request, and return the write result information indicating successful writing to the client, so that the client sends the write result information to the control node, and the control node updates the metadata information corresponding to the target data.
[0092] If it is determined that the first version number is higher than the local version number of the target data, perform the write operation corresponding to the second write request, update the local target data, and return the write result information indicating successful writing to the client, so that the client sends the write result information to the control node, and the control node updates the metadata information corresponding to the target data.
[0093] A data concurrent writing method provided by an embodiment of the present invention. Since the control node can generate a corresponding first version number for the first writing request for the target data sent by the client, and this first version number increases as the number of times the target data is written increases, that is, for the same target data, the version number is unique and increasing. This enables the storage node to determine whether to execute the current data writing operation according to the first version number when receiving the second writing request, ensuring the consistency of the data of multiple replica files when multiple clients concurrently write to the same target data, solving the problem of inconsistent data of multiple replica files when multiple clients concurrently write to the same target data, and improving the reliability of the data. Further, in a distributed file storage system, it is realized that multiple clients can concurrently write to the same target data.
[0094] Based on the embodiment shown in Figure 2 as shown in Figure 3 Another data concurrent writing method is provided by an embodiment of the present invention, which is applied to a storage node in a distributed file storage system. The method may include the following steps:
[0095] S301. Receive a second writing request sent by the client.
[0096] Wherein, the second writing request may include the first version number corresponding to the first writing request for the target data, and this first version number increases as the number of times the target data is written increases.
[0097] S302. Determine whether the first version number is higher than the local version number of the target data.
[0098] Wherein, the local version number of the target data may be: the version number carried in the second writing request corresponding to the last data writing of the target data locally.
[0099] S303. If the first version number is higher than the local version number of the target data, execute the write operation corresponding to the second writing request, and return the write result information indicating successful writing to the client, so that the client sends the write result information to the control node, and the control node updates the metadata information corresponding to the target data.
[0100] Wherein, the implementation process of steps S301 - S303 may be the same as that of steps S201 - S203, and the embodiments of the present invention will not elaborate herein.
[0101] S304. If the first version number is not higher than the local version number of the target data, reject the write operation corresponding to the second writing request, and return the write result information indicating failed writing to the client, so that the client sends the write result information to the control node.
[0102] If it is determined that the first version number is not higher than the local version number of the target data, it indicates that the target data stored locally is already the latest file. Therefore, the update of the target data is rejected, and further, the write operation corresponding to the second write request is rejected, and the write result information indicating write failure is returned to the client, so that the client can send the write result information to the control node.
[0103] In the embodiments of the present invention, when it is determined that the first version number is not higher than the local version number of the target data, the write operation corresponding to the second write request is rejected, which ensures the consistency of the data of multiple replica files when multiple clients concurrently write to the same target data, solves the problem of inconsistent data of multiple replica files caused by multiple clients concurrently writing to the same target data, and improves the reliability of the data.
[0104] As an optional implementation manner of the embodiments of the present invention, as Figure 4 shown, the second write request received by the storage node may further include the data to be written corresponding to the first write request for the target data. Correspondingly, when it is determined that the first version number is higher than the local version number of the target data, the steps of performing the write operation corresponding to the second write request may include:
[0105] S3031, updating the local data of the target data using the data to be written.
[0106] S3032, updating the local version number of the target data using the first version number.
[0107] When the storage node determines that the first version number is higher than the local version number of the target data, it updates the local data of the target data using the data to be written included in the received second write request, and updates the local version number of the target data using the first version number included in the received second write request. That is, it uses the data to be written corresponding to the first write request of the client for the target data and the first version number generated by the control node for the first write request to overwrite the local data and local version number corresponding to the target data stored locally, updates the target data, and completes the write operation corresponding to the second write request.
[0108] As an optional implementation manner of the embodiments of the present invention, the storage node may further send its own status information to the control node, where the own status information of the storage node may include: node status and remaining space size. Exemplarily, the status of the node may include information such as whether the storage node is working properly and the normal speed of data writing, and the remaining space size is the size of the space that the storage node can still utilize currently.
[0109] In an embodiment of the present invention, the storage node may send its own status information to the control node regularly or periodically, so that the control node can update the status information of the storage node it stores, and more accurately select the storage node to which the data to be written is to be written.
[0110] Exemplarily, when there are two clients A and B performing concurrent write operations on the target data at the same time, the control node receives the first write requests for the same target data sent by clients A and B, and can respectively verify the write requests sent by clients A and B. At the same time, a corresponding version number is generated for the write request of client A, denoted as V1, and a corresponding version number is generated for the write request of client B, denoted as V2, and V2 is higher than V1. When the control node successfully verifies the write requests sent by clients A and B, three storage nodes C, D, and E are respectively selected for the data to be written corresponding to the write requests of clients A and B, and the address information corresponding to the storage nodes C, D, and E and the version number V1 are sent to client A, and the address information corresponding to the storage nodes C, D, and E and the version number V2 are sent to client B.
[0111] Client A receives the address information corresponding to the storage nodes C, D, and E and the version number V1 returned by the control node, and sends second write requests to the storage nodes C, D, and E respectively. The second write request carries the version number V1 and the corresponding data to be written.
[0112] Client B receives the address information corresponding to the storage nodes C, D, and E and the version number V2 returned by the control node, and sends second write requests to the storage nodes C, D, and E respectively. The second write request carries the version number V2 and the corresponding data to be written.
[0113] Taking the storage node as C and the local version number of the target data as V0 for illustration, the storage node C first receives the second write request sent by client A. At this time, it is judged that the version number V1 carried in the second write request is higher than the local version number V0, and the local data is updated to the data to be written corresponding to the second write request sent by client A, and the local version number V0 is updated to V1. The storage node C then receives the second write request sent by client B. At this time, it is judged that the version number V2 carried in the second write request is higher than the local version number V1, and the local data is updated to the data to be written corresponding to the second write request sent by client B, and the local version number V1 is updated to V2.
[0114] Alternatively, storage node C first receives the second write request sent by client B. At this time, it is determined that the version number V2 carried in the second write request is higher than the local version number V0. The local data is updated to the data to be written corresponding to the second write request sent by client B, and the local version number V0 is updated to V2. When storage node C then receives the second write request sent by client A, it is determined that the version number V1 carried in the second write request is not higher than the local version number V2, so the local data is not updated, and the local version number of the local data is maintained as V2.
[0115] Storage nodes D and E perform the same operations, so that regardless of whether storage nodes C, D, and E first receive the second write request sent by client A or the second write request sent by client B, the version number corresponding to the data finally saved in storage nodes C, D, and E is V2, ensuring the consistency of the data in multiple replica files when multiple clients concurrently write to the same target data, solving the problem of inconsistent data in multiple replica files caused by multiple clients concurrently writing to the same target data, and improving the reliability of the data.
[0116] As Figure 5 shown, Figure 5 is an interaction schematic diagram of a data concurrent writing method provided by an embodiment of the present invention.
[0117] The client sends a first write request for the target data to the control node.
[0118] The control node verifies the first write request for the target data sent by the client and generates a first version number corresponding to the first write request. The first write request corresponds to the corresponding data to be written. When the verification of the first write request is successful, a preset number of storage nodes are selected for the data to be written to obtain the address information corresponding to the preset number of storage nodes, and the address information corresponding to the preset number of storage nodes and the first version number are sent to the client.
[0119] The client receives the address information corresponding to the preset number of storage nodes and the first version number sent by the control node, and sends the second write request to the preset number of storage nodes.
[0120] The storage node receives the second write request sent by the client, determines whether the first version number included in the second write request is higher than the local version number of the target data, and responds to the write operation corresponding to the second write request, and returns the write result information to the client.
[0121] The client receives the write result information returned by the storage node and returns the write result information to the control node.
[0122] Based on the write result information returned by the client, the control node updates the metadata information corresponding to the target data.
[0123] Corresponding to the above method embodiments, the embodiments of the present invention also provide corresponding system embodiments.
[0124] As Figure 6 shown, the embodiments of the present invention provide a distributed data concurrent write system, and the distributed data concurrent write system includes: a control node and a storage node;
[0125] The control node 401 is configured to verify a first write request for target data sent by the client 402, and generate a first version number corresponding to the first write request. The first write request corresponds to the corresponding data to be written. When the verification of the first write request is successful, a preset number of storage nodes 403 are selected for the data to be written, and the address information corresponding to the preset number of storage nodes is obtained. The address information corresponding to the preset number of storage nodes and the first version number are sent to the client 402, and the metadata information corresponding to the target data is updated based on the write result information returned by the client 402; wherein, the first version number increases as the number of times the target data is written increases, and the metadata information corresponding to the target data includes the size of the target data and the information of the storage nodes where the target data is written.
[0126] The storage node 403 is configured to receive a second write request sent by the client 402, determine whether the first version number included in the second write request is higher than the local version number of the target data, and when it is determined that the first version number is higher than the local version number of the target data, perform the write operation corresponding to the second write request, and return the write result information indicating successful writing to the client 402; wherein, the second write request includes the first version number corresponding to the first write request for the target data and the data to be written, and the local version number of the target data is the version number carried in the second write request when the target data was last written locally.
[0127] For the distributed data concurrent write system provided by the embodiments of the present invention, since the control node can generate a corresponding first version number for the first write request for the target data sent by the client, and the first version number increases as the number of times the target data is written increases, that is, the version number for the same target data is unique and increasing, so that when the storage node receives the second write request, it can determine whether to perform the current data write operation according to the first version number, ensuring the consistency of the data of multiple copy files when multiple clients concurrently write to the same target data, solving the problem of inconsistent data of multiple copy files caused by multiple clients concurrently writing to the same target data, and improving the reliability of the data. Further, it realizes that in a distributed file storage system, multiple clients can concurrently write to the same target data.
[0128] Optionally, the system further includes a client 402:
[0129] The client 402 is configured to send a first write request for target data to the control node 401, and receive address information corresponding to a preset number of storage nodes and a first version number returned by the control node 401; and send a second write request to the preset number of storage nodes 403, receive write result information returned by the preset number of storage nodes, and send the write result information to the control node 401.
[0130] Optionally, the control node 401 is specifically configured to:
[0131] Generate a first version number that is higher than the version number corresponding to the most recent first write request for the target data in history.
[0132] Optionally, the storage node 403 is further configured to:
[0133] When it is determined that the first version number is not higher than the local version number of the target data, reject the write operation corresponding to the second write request, and return write result information indicating write failure to the client 402.
[0134] Optionally, the storage node 403 is specifically configured to:
[0135] Use the data to be written to update the local data of the target data;
[0136] Use the first version number to update the local version number of the target data.
[0137] Optionally, the storage node 403 is further configured to:
[0138] Send its own status information to the control node 401, where its own status information includes: node status and remaining space size.
[0139] In another embodiment provided by the present invention, there is also provided a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above data concurrent writing methods are implemented to achieve the same effect.
[0140] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.
[0141] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.
[0142] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included within the protection scope of the present invention.
Claims
1. A data concurrent writing method, characterized in that, Applied to the control node in a distributed file storage system, the distributed file storage system includes: a control node and storage nodes, and the method includes: Verify the first write requests for target data sent by multiple clients, and generate a corresponding version number for each first write request to obtain a first version number, where the first version number increases as the number of times the target data is written increases, and the first write request corresponds to corresponding data to be written; For each of the first write requests, when the verification of the first write request is successful, select a preset number of storage nodes for the data to be written corresponding to the first write request to obtain the address information corresponding to the preset number of storage nodes; among them, the address information corresponding to the preset number of storage nodes selected for the first write requests of different clients is the same; Send the address information corresponding to the preset number of storage nodes and the first version number to the client corresponding to the first write request, so that the client sends a second write request to the preset number of storage nodes, and the preset number of storage nodes respectively execute corresponding write operations according to the second write request and return the write result information to the corresponding client; among them, the second write request includes the data to be written and the first version number; Update the metadata information corresponding to the target data based on the write result information returned by each client; among them, the metadata information corresponding to the target data includes the size of the target data and the information of the storage nodes where the target data is written.
2. The method according to claim 1, wherein The step of generating a corresponding version number for each first write request to obtain a first version number includes: Generate a first version number that is higher than the version number corresponding to the most recent first write request for the target data in history.
3. The method according to any one of claims 1 or 2, characterized in that, The step of selecting a preset number of storage nodes for the data to be written corresponding to the first write request includes: Select a preset number of storage nodes for the data to be written according to the size of the data to be written corresponding to the first write request and the remaining space size of each storage node.
4. The method according to claim 1, wherein The method further includes: Receive the self-status information sent by the storage node, and the self-status information of the storage node includes: the node status of the storage node and the remaining space size of the storage node.
5. A method for concurrent data writing, characterized in that, Applied to the storage node in a distributed file storage system, the distributed file storage system includes: a control node and storage nodes, and the method includes: Receive second write requests sent by multiple clients, where the second write requests contain the first version number corresponding to the first write request for the target data, and the first version number increases as the number of times the target data is written increases; the second write request is: after the control node verifies the first write requests for the target data sent by multiple clients and generates a corresponding version number for each first write request to obtain the first version number, and for each of the first write requests, when the verification of the first write request is successful, select a preset number of storage nodes for the data to be written corresponding to the first write request, and send the address information corresponding to the preset number of storage nodes and the first version number to the corresponding client, the client sends it; the address information corresponding to the preset number of storage nodes selected by the control node for the first write requests of different clients is the same; For each of the second write requests, determine whether the first version number is higher than the local version number of the target data, where the local version number of the target data is the version number carried in the second write request when the target data was last written locally; If the first version number is higher than the local version number of the target data, then perform the write operation corresponding to the second write request, and return write result information indicating successful writing to the client, so that the client sends the write result information to the control node, and the control node updates the metadata information corresponding to the target data.
6. The method according to claim 5, wherein The method further includes: If the first version number is not higher than the local version number of the target data, then reject the write operation corresponding to the second write request, and return write result information indicating write failure to the client, so that the client sends the write result information to the control node.
7. The method according to any one of claims 5 or 6, characterized in that The second write request also contains the data to be written corresponding to the first write request for the target data; the step of performing the write operation corresponding to the second write request includes: Using the data to be written to update the local data of the target data; Using the first version number to update the local version number of the target data.
8. The method according to claim 5, wherein The method further includes: Sending its own status information to the control node, where the own status information includes: the node status and the remaining space size.
9. A distributed data concurrent writing system, characterized in that, The distributed data concurrent write system includes: a control node and storage nodes; The control node is used to verify the first write requests for the target data sent by multiple clients, and generate a corresponding first version number for each first write request. The first write request corresponds to the data to be written. For each first write request, when the verification of the first write request is successful, a preset number of storage nodes are selected for the data to be written corresponding to the first write request, and the address information corresponding to the preset number of storage nodes is obtained. The address information corresponding to the preset number of storage nodes and the first version number are sent to the client corresponding to the first write request, and the metadata information corresponding to the target data is updated based on the write result information returned by each client. Among them, the first version number increases as the number of times the target data is written increases. The address information corresponding to the preset number of storage nodes selected for the first write requests of different clients is the same. The metadata information corresponding to the target data includes the size of the target data and the information of the storage nodes where the target data is written. The storage node is used to receive the second write requests sent by multiple clients. For each second write request, it is judged whether the first version number included in the second write request is higher than the local version number of the target data. When it is judged that the first version number is higher than the local version number of the target data, the write operation corresponding to the second write request is executed, and the write result information indicating successful writing is returned to the client. Among them, the second write request includes the first version number corresponding to the first write request for the target data and the data to be written. The local version number of the target data is the version number carried in the second write request when the target data was last written locally.
10. The system according to claim 9, wherein The control node is specifically used for: Generating a first version number higher than the version number corresponding to the most recent first write request for the target data in history.
11. The system according to claim 9, characterized in that, The storage node is further used for: When it is judged that the first version number is not higher than the local version number of the target data, rejecting the write operation corresponding to the second write request, and returning the write result information indicating write failure to the client.
12. The system according to claim 9, wherein The storage node is specifically used for: Using the data to be written to update the local data of the target data. Using the first version number to update the local version number of the target data.
13. The system according to claim 9, wherein The storage node is further used for: Sending its own status information to the control node. The own status information includes: the node status and the remaining space size.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method steps described in any one of claims 1-4.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method steps described in any one of claims 5-8.
Citation Information
Patent Citations
Method and device for updating data in distributed storage system
CN103294675A