File operation method, distributed storage system, electronic device, and storage medium
By introducing file operation methods and consensus mechanisms in distributed storage systems, the problem of inconsistency in storage content in high-availability files is solved, and strong file consistency and high system reliability are achieved.
Patent Information
- Application Number
- CN202111589942.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-23
AI Technical Summary
Existing highly available file distributed storage systems cannot guarantee the consistency of stored content, especially in the case of failure and parallel storage, where there may be inconsistencies between multiple copies of the same data.
By introducing file operation methods in the distributed storage system, the client's file operation request is obtained, and it is sent to the associated storage node in accordance with the preset organization mode, a consensus message is obtained, the operation log in the local storage is synchronized into the status indicated by the consensus message, the matching between the file and the operation log is detected, and the files are synchronized when they do not match to ensure consistency.
It ensures strong consistency of files in a distributed storage system, solves the problem of inconsistent storage content, and improves the reliability and availability of the system.
Smart Images

Figure CN114218169B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data storage, and particularly to a file operation method, a distributed storage system, an electronic device, and a storage medium. Background Art
[0002] A distributed storage system is a system that disperses data storage across multiple independent devices. As the number of servers increases, the probability of server failures also increases. To ensure the system remains available in the event of a server failure, a single piece of data is usually stored in multiple copies on different servers.
[0003] In the related art, a distributed storage system is provided. This distributed storage system is provided with a name service node, a task dispatching node, and a data service node. The task dispatching node receives and parses a data read / write command sent by a client, obtains the access information corresponding to the data read / write command, and sends the received data read / write command, access information, and the global path information of the file to the name service node; the name service node obtains the access information corresponding to the data service node from the corresponding tree structure according to the access information and the global path information, and sends a read / write operation command to the data service node; the data service node reads or writes the file data stored on the corresponding data service node according to the read / write operation command, and returns the execution result data to the task dispatching node, and the task dispatching node returns the execution result data to the corresponding client.
[0004] This distributed storage system uses a centralized scheduling node and still follows the traditional distributed storage mode, which cannot meet the requirements of high-availability distributed storage of files. A high-availability distributed storage system for files needs to consider consistency factors, that is, due to the existence of failures and parallel storage, etc., there may be inconsistent situations among multiple copies of the same data. Currently, most high-availability distributed storage systems for files adopt a master-slave backup mode. When the master node fails to provide services, the slave node provides services. In this master-slave backup mode, node downtime or network failures will both lead to inconsistent stored content, and the consistency of the stored content cannot be guaranteed.
[0005] Regarding the problem that the high-availability file distributed storage system in the related art cannot guarantee the consistency of stored content, no effective solution has been proposed yet. Summary of the Invention
[0006] In this embodiment, a file operation method, a distributed storage system, an electronic device, and a storage medium are provided to solve the problem that the high-availability file distributed storage system in the related art cannot guarantee the consistency of stored content.
[0007] In a first aspect, in the present embodiment, a file operation method is provided, which is applied to a first storage node in a distributed storage system. The distributed storage system further includes at least one storage node other than the first storage node. The method includes:
[0008] Obtain a file operation request initiated by a client, obtain a target file to be operated according to the file operation request, and perform corresponding operations on the target file;
[0009] Send the file operation request to a second storage node associated with the first storage node according to a preset organization mode to instruct the second storage node to perform corresponding operations on the target file;
[0010] Obtain a consensus message, and synchronize the first operation log in the local storage to the state indicated by the consensus message, where the first operation log stores the operation process performed on the target file;
[0011] Detect whether the file in the local storage matches the file operation indicated in the first operation log, and in the case where it is detected that the file in the local storage does not match the file operation indicated in the first operation log, synchronize the file in the local storage to a state that is at least consistent with the file stored in the second storage node.
[0012] In some embodiments, detecting whether the file in the local storage matches the file operation indicated in the first operation log includes:
[0013] Judge whether there is a situation of file loss or file redundancy in the local storage according to the file operation indicated in the first operation log;
[0014] In the case where it is detected that there is file loss or file redundancy in the local storage, it is determined that the file in the local storage does not match the file operation indicated in the first operation log.
[0015] In some embodiments, synchronizing the file in the local storage to a state that is at least consistent with the file stored in the second storage node includes:
[0016] Synchronize the file in the local storage to a state consistent with the file stored in the second storage node; or,
[0017] Query target storage nodes that meet the file operation indicated in the first operation log in sequence according to the preset organization mode, and synchronize the file in the local storage to a state consistent with the file stored in the target storage node.
[0018] In some of these embodiments, the file operation request carries the target file and information indicating the addition of the target file. Obtain the target file to be operated according to the file operation request, perform corresponding operations on the target file, and send the file operation request to a second storage node associated with the first storage node according to a preset organization mode to instruct the second storage node to perform corresponding operations on the target file, including:
[0019] Write the target file to local storage, and during the process of writing the target file to local storage, send the target file to the second storage node to instruct the second storage node to add the target file.
[0020] In some of these embodiments, writing the target file to local storage and instructing the second storage node to add the target file includes:
[0021] In local storage, split the target file into multiple file blocks, assign file block IDs to each of the file blocks, and store the multiple file blocks in the structure of a MerkleTree. Among them, the root node of the MerkleTree stores the file ID of the target file, the leaf nodes of the MerkleTree store the file content and file block IDs of each of the file blocks, and the non-leaf nodes of the MerkleTree store links pointing to the corresponding leaf nodes;
[0022] Instruct the second storage node to perform the same file addition operation as the local storage.
[0023] In some of these embodiments, after writing the target file to local storage and instructing the second storage node to add the target file, the method further includes:
[0024] Receive a file reading request initiated by the client, where the file reading request carries the file ID of the target file;
[0025] Determine multiple target storage nodes according to the file reading request, and allocate reading tasks among the multiple target storage nodes, where the file block IDs carried in each of the reading tasks are different;
[0026] Parallelly read different file blocks in each of the target storage nodes, and splice the multiple read file blocks according to the file block IDs to obtain the target file;
[0027] Return the target file to the client.
[0028] In some of these embodiments, the file operation request carries the file ID of the target file and information indicating deletion of the target file. Obtaining the target file to be operated according to the file operation request, performing corresponding operations on the target file, and sending the file operation request to a second storage node associated with the first storage node according to a preset organization mode to instruct the second storage node to perform corresponding operations on the target file includes:
[0029] Deleting the target file in the local storage according to the file ID, and instructing the second storage node to delete the target file.
[0030] In some of these embodiments, after obtaining a file operation request initiated by a client, obtaining the target file to be operated according to the file operation request, and performing corresponding operations on the target file, the method further includes:
[0031] Writing the operation process performed on the target file into the first operation log in sequence according to the execution order.
[0032] In some of these embodiments, after sending the file operation request to a second storage node associated with the first storage node according to a preset organization mode, the method further includes:
[0033] Receiving an execution result returned by the second storage node, where the execution result carries the node IDs of the storage nodes through which the file operation request has flowed and the execution status of each storage node;
[0034] Determining the storage nodes in the distributed storage system that have performed corresponding operations on the target file and have succeeded, and determining whether the number of storage nodes that have succeeded reaches a preset threshold;
[0035] In the case where it is determined that the number of storage nodes that have succeeded reaches the preset threshold, returning a successful execution response message to the client; and,
[0036] In the case where it is determined that the number of storage nodes that have succeeded does not reach the preset threshold, returning a failed execution response message to the client.
[0037] In some of these embodiments, after writing the operation process performed on the target file into the first operation log in sequence according to the execution order, the method further includes:
[0038] Setting the first operation log to a non-consensus state;
[0039] In the case where the number of storage nodes with successful execution does not reach the preset threshold, the status of the first operation log remains the state of not reaching consensus, where the operation log set to the state of not reaching consensus is allowed to be changed; and,
[0040] In the case where the number of storage nodes with successful execution reaches the preset threshold, the first operation log is set to the state of reaching consensus, where the operation log set to the state of reaching consensus is not allowed to be changed.
[0041] In some of the embodiments, the consensus message is generated by the Leader node elected by each storage node. In the case where the number of storage nodes with successful execution does not reach the preset threshold, the method further includes:
[0042] Participating in re-election of the Leader node in the distributed storage system;
[0043] Receiving the new consensus message sent by the re-elected Leader node, synchronizing the first operation log to the state indicated by the new consensus message, and performing corresponding operations on the files stored locally according to the updated first operation log.
[0044] In a second aspect, in the present embodiment, a distributed storage system is provided, and the system includes: a plurality of storage nodes, and each storage node is sequentially associated according to a preset organization mode. When the first storage node receives a file operation request initiated by a client, the first storage node is configured to execute the file operation method described in the first aspect above.
[0045] In some of the embodiments, each storage node includes: a file reading and writing module, a consensus module, a command execution module, and a file storage module; where,
[0046] The file reading and writing module is configured to transfer files between each storage node;
[0047] Each storage node stores an operation log, and the consensus module is configured to maintain the consistency of the operation log between each storage node, where the operation log stores the operation process performed on the file;
[0048] The command execution module is configured to maintain the consistency of the files between each storage node;
[0049] The file storage module is configured to store files.
[0050] In a third aspect, in the present embodiment, an electronic device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the file operation method described in the first aspect above is implemented.
[0051] Fourthly, in this embodiment, a storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the file operation method described in the first aspect above is implemented.
[0052] Compared with the related art, in the file operation method, distributed storage system, electronic device, and storage medium provided in this embodiment, by obtaining a file operation request initiated by a client, obtaining a target file to be operated according to the file operation request, and performing corresponding operations on the target file; sending the file operation request to a second storage node associated with the first storage node according to a preset organization mode to instruct the second storage node to perform corresponding operations on the target file; obtaining a consensus message and synchronizing the first operation log in the local storage to the state indicated by the consensus message, where the first operation log stores the operation process performed on the target file; detecting whether the file in the local storage matches the file operation indicated in the first operation log, and in the case where it is detected that the file in the local storage does not match the file operation indicated in the first operation log, synchronizing the file in the local storage to a state that is at least consistent with the file stored in the second storage node, the problem that the high-availability file distributed storage system in the related art cannot guarantee the consistency of stored content is solved, and the strong consistency of the stored content in the high-availability file distributed storage system is ensured.
[0053] The details of one or more embodiments of the present application are set forth in the following drawings and description to make the other features, objects, and advantages of the present application more concise and understandable. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0055] Figure 1 is a hardware structure block diagram of a terminal of the file operation method according to an embodiment of the present application;
[0056] Figure 2 is a flowchart of the file operation method according to an embodiment of the present application;
[0057] Figure 3 is a structure block diagram of a distributed storage system according to an embodiment of the present application;
[0058] Figure 4 is a schematic diagram of the file upload process of a distributed storage system according to an embodiment of the present application;
[0059] Figure 5 is a schematic diagram of the consensus message transmission process of a distributed storage system according to an embodiment of the present application;
[0060] Figure 6 It is a schematic diagram of the file storage structure of a distributed storage system according to an embodiment of the present application;
[0061] Figure 7 It is a flowchart of uploading a file in a distributed storage system according to an embodiment of the present application;
[0062] Figure 8 It is a flowchart of deleting a file in a distributed storage system according to an embodiment of the present application. Detailed implementation manners
[0063] To understand the purpose, technical solution and advantages of the present application more clearly, the present application is described and illustrated below with reference to the accompanying drawings and embodiments.
[0064] Unless otherwise defined, the technical terms or scientific terms involved in the present application shall have the general meanings understood by those with ordinary skills in the technical field to which the present application belongs. In the present application, words such as "a", "one", "a kind of", "the", "these" and the like do not indicate a limitation in quantity, and they can be singular or plural. The terms "including", "comprising", "having" and any variants thereof involved in the present application are intended to cover non-exclusive inclusion; for example, a process, method, system, product or device including a series of steps or modules (units) is not limited to the listed steps or modules (units), but may include unlisted steps or modules (units), or may include other steps or modules (units) inherent in these processes, methods, products or devices. The terms "connected", "coupled" and the like involved in the present application are not limited to physical or mechanical connections, but may include electrical connections, whether directly or indirectly connected. The term "plurality" involved in the present application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may represent: A exists alone, A and B exist simultaneously, and B exists alone. Usually, the character " / " indicates that the objects associated before and after are in an "or" relationship. The terms "first", "second", "third" and the like involved in the present application are only used to distinguish similar objects and do not represent a specific sorting of the objects.
[0065] The method embodiments provided in this embodiment can be executed on a terminal, a computer or a similar computing device. For example, running on a terminal, Figure 1 It is a hardware structure block diagram of a terminal of a file operation method according to an embodiment of the present application. As Figure 1 shown, the terminal may include one or more ( Figure 1Only one processor 102 and a memory 104 for storing data are shown. Among them, the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA. The above terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 The structure shown is only schematic and does not limit the structure of the above terminal. For example, the terminal may further include more or fewer components than Figure 1 shown in, or have a different configuration from Figure 1 that shown.
[0066] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the file operation method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0067] The transmission device 106 is used to receive or send data via a network. The above network includes a wireless network provided by the communication provider of the terminal. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0068] In this embodiment, a file operation method is provided, which is applied to a first storage node in a distributed storage system. The distributed storage system further includes at least one storage node other than the first storage node. Figure 2 is a flowchart of the file operation method according to an embodiment of the present application. As Figure 2 shown, the process includes the following steps:
[0069] Step S201, obtain a file operation request initiated by a client, obtain a target file to be operated according to the file operation request, and perform corresponding operations on the target file.
[0070] The client can arbitrarily select a storage node in the distributed storage system to initiate a file operation request. For example, the first storage node can be either an ordinary Follower node or a Leader node determined through election. If the file operation request is to indicate adding a file uploaded by the client, the first storage node obtains the target file from the client and writes the target file to local storage; if the file operation request is to indicate deleting a file already stored in the distributed storage system, the first storage node obtains the target file from local storage and deletes the target file from local storage.
[0071] Step S202: Send the file operation request to a second storage node associated with the first storage node according to a preset organization mode, so as to instruct the second storage node to perform corresponding operations on the target file.
[0072] In this embodiment, the second storage node is configured to send the file operation request to a third storage node, and the third storage node is a storage node associated with the second storage node according to a preset organization mode. The preset organization mode means that data is transmitted through pipelines in the distributed storage system, where each storage node is sequentially associated. For example, the next storage node associated with the first storage node is the second storage node, and the next storage node associated with the second storage node is the third storage node. The first storage node receives data from the client while transmitting data to the second storage node. The second storage node receives data from the first storage node while transmitting data to the third storage node, and so on, until all storage nodes in the entire distributed storage system have completed data reception.
[0073] Step S203: Obtain a consensus message and synchronize the first operation log in local storage to the state indicated by the consensus message, where the first operation log stores the operation process performed on the target file.
[0074] The consensus message originates from the Leader node. Since the Leader node is determined through election, it can be ensured that the operation process in the Leader node is the most complete and can be used as a reference for the current entire distributed storage system. If the previous storage node of the first storage node happens to be the Leader node, the first storage node can directly receive the consensus message sent by the Leader node. If the previous storage node of the first storage node is not the Leader node, the first storage node can indirectly receive the consensus message through other storage nodes.
[0075] In some embodiments, the first storage node sequentially writes the operation process performed on the target file into the first operation log in the execution order, that is, the operation process can be a set of operation commands arranged in the execution order. For example, if the client requests the distributed storage system to add file file1 and file file2 in sequence and then deletes file file1, the first storage node sequentially writes Add file1, Add file2, and Del file1 into the first operation log.
[0076] Similarly, the Leader node also writes its own set of operation commands into the consensus message, sends the consensus message to the Follower nodes, and instructs the Follower nodes to synchronize the content of their own operation logs to a state consistent with this set of operation commands. In this embodiment, the Raft protocol can be used among the storage nodes to implement the transmission of the consensus message.
[0077] Step S204, detect whether the files in the local storage match the file operations indicated in the first operation log, and in the case where it is detected that the files in the local storage do not match the file operations indicated in the first operation log, synchronize the files in the local storage to a state that is at least consistent with the files stored in the second storage node.
[0078] Although the first storage node has written the operation process of the target file into the first operation log, however, in the case where the first storage node crashes or restarts, the file operations of the first storage node before crashing or restarting may become invalid, and there is still a possibility that files are lost or not deleted in the local storage, resulting in the inconsistency between the files stored in the first storage node and the files stored in the other storage nodes. For this reason, the first storage node will detect whether the files in the local storage are lost or redundant according to the first operation log. If it is detected that there are file losses or file redundancies in the local storage, the first storage node will synchronize the lost files in the local storage according to the other storage nodes, or delete the redundant files in the local storage according to the other storage nodes. Specifically, during implementation, the first storage node will maintain a file list based on the files in the local storage, write the file IDs of the files in the local storage into the file list, and detect whether the files in the local storage are lost or redundant according to the file list and the first operation log.
[0079] In the above steps S201 to S204, each storage node stores a complete file backup, eliminating the centralized scheduling of storage nodes. Each storage node can provide operations such as file addition, deletion, modification, and reading to the outside. Although the files of all storage nodes are not consistent in real time, the consistency of the operation logs in the entire distributed storage system can be ensured through message consensus. Further, by synchronizing the files stored locally through the operation logs after consensus, the consistency of the files in the entire distributed storage system can be ensured. At the same time, all file operation requests are processed by the first storage node. Therefore, from the perspective of the client, the distributed storage system is strongly consistent. Through the above steps, the problem in the related technology that a highly available file distributed storage system cannot guarantee the consistency of stored content is solved, and the strong consistency of the stored content in the highly available file distributed storage system is ensured.
[0080] In an embodiment of the present application, when the file stored in the second storage node matches the file operation indicated in the first operation log, the file stored locally can be synchronized to a state consistent with the file stored in the second storage node.
[0081] In an embodiment of the present application, when the file stored in the second storage node does not match the file operation indicated in the first operation log, the target storage node that meets the file operation indicated in the first operation log is sequentially queried according to the preset organization mode, and the file stored locally is synchronized to a state consistent with the file stored in the target storage node.
[0082] With such a setting, when the first storage node fails or restarts, other storage nodes for reference can always be found in the entire distributed storage system to synchronize files.
[0083] In an embodiment of the present application, the file operation request carries the target file and information indicating the addition of the target file. When the first storage node obtains the target file to be operated according to the file operation request, performs the corresponding operation on the target file, and sends the file operation request to the second storage node associated with the first storage node according to the preset organization mode to instruct the second storage node to perform the corresponding operation on the target file, the first storage node writes the target file into local storage, and during the process of writing the target file into local storage, the target file is sent to the second storage node to instruct the second storage node to add the target file.
[0084] In an embodiment of the present application, writing the target file to the local storage and instructing the second storage node to add the target file can be achieved by the following method: In the local storage, the first storage node splits the target file into multiple file blocks, assigns a file block ID to each file block, and stores the multiple file blocks in the structure of a Merkle Tree. Among them, the root node of the Merkle Tree stores the file ID of the target file, the leaf nodes of the Merkle Tree store the file content and file block ID of each file block, and the non-leaf nodes of the Merkle Tree store the links pointing to the corresponding leaf nodes; instruct the second storage node to perform the same file addition operation as the local storage.
[0085] In this embodiment, the target file is not stored as a complete entity. The target file can be evenly split into several file blocks of 256 KB in size, and the multiple file blocks are stored in the structure of a Merkle DAG (Merkle Directed Acyclic Graph).
[0086] In an embodiment of the present application, after writing the target file to the local storage and instructing the second storage node to add the target file, the first storage node will also receive a file reading request initiated by the client. Among them, the file reading request carries the file ID of the target file; determine multiple target storage nodes according to the file reading request, and allocate reading tasks among the multiple target storage nodes. Among them, the file block IDs carried in each reading task are different; read different file blocks in parallel in each target storage node, and splice the multiple read file blocks according to the file block ID to obtain the target file; return the target file to the client.
[0087] With such a setting, the first storage node can read the target file in parallel from multiple storage nodes, improving the file reading efficiency.
[0088] In an embodiment of the present application, the file operation request carries the file ID of the target file and information indicating the deletion of the target file. When the first storage node obtains the target file to be operated according to the file operation request, performs the corresponding operation on the target file, and sends the file operation request to the second storage node associated with the first storage node according to the preset organization mode to instruct the second storage node to perform the corresponding operation on the target file, the first storage node deletes the target file in the local storage according to the file ID, and instructs the second storage node to delete the target file.
[0089] In one embodiment of the present application, after the first storage node sends a file operation request to a second storage node associated with the first storage node according to a preset organization mode, the first storage node will receive the execution result returned by the second storage node. The execution result carries the node IDs of the storage nodes through which the file operation request has passed and the execution status of each storage node. Determine the storage nodes in the distributed storage system that have successfully executed the corresponding operation on the target file according to the execution result, and determine whether the number of storage nodes that have executed successfully reaches a preset threshold. In the case where it is determined that the number of storage nodes that have executed successfully reaches the preset threshold, return a success response message to the client. And, in the case where it is determined that the number of storage nodes that have executed successfully does not reach the preset threshold, return a failure response message to the client.
[0090] In this embodiment, the preset threshold can be determined according to the total number of storage nodes in the distributed storage system. By way of example and not limitation, the preset threshold can be set to one-half of the total number of storage nodes. That is to say, when half or more of the storage nodes have executed successfully, the first storage node will return a success response message to the client. Otherwise, the first storage node will return a failure response message to the client.
[0091] In a traditional distributed storage system, a success response message will be returned to the client only when all storage nodes have executed successfully, resulting in low performance of the distributed storage system. Compared with the traditional distributed storage system, in this embodiment, only most storage nodes need to execute successfully, and slow processing of a small number of storage nodes will not delay the overall system operation, ensuring that the distributed storage system has high performance.
[0092] In one embodiment of the present application, after the first storage node sequentially writes the operation process of the target file into the first operation log according to the execution order, the first storage node also sets the first operation log to an unconsensus state. In the case where it is determined that the number of storage nodes that have executed successfully does not reach the preset threshold, keep the state of the first operation log still in the unconsensus state. The operation log set to the unconsensus state allows changes. And, in the case where it is determined that the number of storage nodes that have executed successfully reaches the preset threshold, set the first operation log to a consensus state. The operation log set to the consensus state does not allow changes.
[0093] When the first storage node writes the operation process into the first operation log, it writes the operation process performed on the target file into the first operation log strictly in the execution order. After writing the operation process into the first operation log, without consensus, the currently written operation process will be automatically set to the unconsensus state, and only after consensus, the state of the operation process will be changed. By way of example and not limitation, when half or more of the storage nodes execute successfully, the first storage node sets the currently written operation process to the consensus reached state. In this embodiment, the operation log set to the consensus reached state is not allowed to be changed. Therefore, when the client receives a successful execution response message, it means that the file in the system will no longer change. Even if less than half of the storage nodes fail, the file will not be lost, ensuring the high reliability of the distributed storage system.
[0094] In addition, the distributed storage system is divided into multiple networks, which may cause some storage nodes to be unable to communicate with each other. If the number of storage nodes in the network where the first storage node is located does not reach half, the first storage node will return a failed execution response message to the client, so that the client knows that the file operation request has not been successfully requested, and the first storage node does not confirm the state of the first operation log, so as to facilitate subsequent modification of the first operation log.
[0095] In an embodiment of the present application, the consensus message is generated by the Leader node elected by each storage node. When it is determined that the number of storage nodes that execute successfully does not reach the preset threshold, the first storage node will participate in the re-election of the Leader node in the distributed storage system; receive the new consensus message sent by the re-elected Leader node, synchronize the first operation log to the state indicated by the new consensus message, and perform corresponding operations on the files stored locally according to the updated first operation log.
[0096] In this embodiment, the number of storage nodes in the network where the old Leader node is located does not reach half. After the distributed storage system restores the network connection, the old Leader node is automatically demoted to a Follower node, and the first storage node will enter the process of re-electing the Leader node, follow the re-elected Leader node to update the first operation log (the first operation log is still in the unconsensus state), and perform corresponding operations on the files stored locally according to the updated first operation log. With such a setting, even if the Leader node fails, the system will spontaneously elect a new Leader node without manual intervention, and the unavailable time is extremely short. Less than half of the storage nodes failing or network anomalies will not affect the availability of the system, ensuring the high availability of the distributed storage system.
[0097] Combined with the file operation method of the above embodiments, the present embodiment further provides a distributed storage system, which includes: a plurality of storage nodes, and each storage node is sequentially associated according to a preset organization mode. When the first storage node among them receives a file operation request initiated by a client, the first storage node is used to execute the file operation method of any of the above embodiments.
[0098] In this embodiment, the distributed storage system is composed of a plurality of storage nodes. Each storage node stores a complete file backup, and each storage node can provide operations such as file addition, deletion, modification, and reading to the outside.
[0099] Figure 3 It is a structural block diagram of the distributed storage system according to an embodiment of the present application. As Figure 3 shown, the system includes a plurality of storage nodes, and each storage node includes a file reading and writing module, a consensus module, a command execution module, and a file storage module. Among them, the file reading and writing module is used to transfer files between each storage node; each storage node stores an operation log, and the consensus module is used to maintain the consistency of the operation log between each storage node. Among them, the operation log stores the operation process performed on the file; the command execution module is used to maintain the consistency of the file between each storage node; the file storage module is used to store files.
[0100] In this embodiment, between every two storage nodes, file data messages can be transmitted through the file reading and writing module to transfer files, consensus messages can be transmitted through the consensus module to maintain the consistency of the operation log, and file synchronization messages can be transmitted through the command execution module to maintain the consistency of the file.
[0101] The following will introduce each module in each storage node through preferred embodiments.
[0102] (1) File reading and writing module
[0103] Figure 4 It is a schematic diagram of the file upload process of the distributed storage system according to an embodiment of the present application. As Figure 4As shown in the figure, the file read / write module is responsible for file data upload, download, and file data transmission between different storage nodes. The upload speed of the file read / write module directly affects the performance of the entire system. To minimize the data upload latency as much as possible, the file is uploaded to all storage nodes in the system through a pipeline method during upload. For example, the client Client selects any one storage node, sends the file to it, and then this storage node forwards the file to the next storage node while receiving the file from the Client, and so on, thus completing the transmission of the file from the Client to all storage nodes in the system. This method is similar to a data flow pipeline, which can make full use of the bandwidth of each storage node, and the file content can be written into the file storage module in parallel, reducing the upload latency. When the file is written into the file storage module, a file ID representing the file will be returned, and this file ID can be used to read the file content from the file storage module. The response message contains the file upload error information of the current storage node and the unique file ID, and the last storage node in the upload path returns in the reverse order, and finally the upload result is returned to the Client storage node. Each storage node in the distributed storage system has a full backup of the file, and the Client can read the file content from any storage node and can read in parallel from multiple storage nodes, with extremely high reading efficiency.
[0104] By way of example and not limitation, the distributed storage system of this embodiment defines a file upload message and a file upload response message.
[0105] Definition of file upload message:
[0106] Field Meaning FileStream File data stream NextPeer Next storage node address information
[0107] Definition of file upload response message:
[0108] Field Meaning Error Upload file error information. If the upload is successful, the Error field is empty FileID File unique ID information
[0109] (2) Consensus module
[0110] Figure 5 It is a schematic diagram of the consensus message transmission process of the distributed storage system according to an embodiment of the present application. As Figure 5 shown, each storage node maintains an operation log through the consensus module. The operation log stores file operation commands, and the file operation commands are stored in the execution order and cannot be modified once confirmed by consensus. The command execution module ensures the consistency of the files stored in the distributed storage system by executing the file operation commands in the operation log.
[0111] By way of example and not limitation, the distributed storage system of this embodiment defines two types of consensus commands, namely a file addition command and a file deletion command.
[0112] File addition command:
[0113] Field Meaning Name File name FileID File ID ExpireAt File expiration time. To save storage space, expired files can be deleted
[0114] File deletion command:
[0115]
[0116]
[0117] (3) Command execution module
[0118] The command execution module is responsible for executing the file operation commands in the operation log. The consensus module ensures the consistency of the operation logs in each storage node, and the command execution module ensures the consistency of the files in each storage node. The command execution module reads and executes the file operation commands in the operation log. The command execution module and the consensus module can be executed asynchronously, and the time-consuming operations executed by the command execution module will not affect the efficiency of the consensus module.
[0119] In this embodiment, the functions of the command execution module include: adding or deleting files according to the operation log; maintaining the current file list stored, and regularly detecting whether the files required in the file list and the operation log are consistent; synchronizing the files missing locally from other storage nodes; regularly deleting expired files to optimize the storage space.
[0120] (4) File storage module
[0121] The file storage module is responsible for storing the specific file content. The file is not stored as a complete individual. The file is sliced into several file blocks of 256KB in size by the average segmentation method, and the file blocks are stored in the structure of MerkleDAG. Each file corresponds to a unique storage root ID (file ID), and the file content is stored in the leaf nodes of MerkleDAG. Figure 6 It is a schematic diagram of the file storage structure of a distributed storage system according to an embodiment of the present application, as Figure 6 shown:
[0122] The root node is the root node, and the root node stores the file ID.
[0123] Node3, node4, node5, and node6 are all leaf nodes. Node3, node4, node5, and node6 store the file blocks and the corresponding file block IDs respectively.
[0124] Node1 is the parent storage node of node3 and node4, and node1 stores the links pointing to node3 and node4; node2 is the parent storage node of node5 and node6, and node2 stores the links pointing to node5 and node6.
[0125] Figure 7 It is a flowchart of uploading a file in a distributed storage system according to an embodiment of the present application. As Figure 7 shown, the process includes the following steps:
[0126] Step S701: Receive a file addition request initiated by an application party.
[0127] Step S702: The file read / write module is responsible for transmitting the target file to the file storage modules of each storage node in a pipelined manner.
[0128] Step S703: The file storage module stores the target file in blocks and returns a file ID.
[0129] Step S704: Call the consensus module to consensus on operations related to adding a file.
[0130] Step S705: The command execution module monitors the file ID, detects the files stored locally. If there are missing or redundant files, synchronize the file to other storage nodes in the system or delete the file.
[0131] Step S706: Return the file addition result to the application party.
[0132] Figure 8 It is a flowchart of deleting a file in a distributed storage system according to an embodiment of the present application. As Figure 8 shown, the process includes the following steps:
[0133] Step S801: Receive a file deletion request initiated by an application party.
[0134] Step S802: Call the consensus module to consensus on operations related to deleting a file.
[0135] Step S803: The command execution module monitors the operation of deleting a file ID.
[0136] Step S804: The command execution module monitors the file ID, detects the files stored locally, and detects that the target file has been deleted.
[0137] Step S805: Return the file deletion result to the application party.
[0138] In this embodiment, an electronic device is further provided, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0139] Optionally, the above electronic device may further include a transmission device and an input / output device. Among them, the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0140] Optionally, in this embodiment, the above-mentioned processor may be configured to execute the following steps through a computer program:
[0141] S1. Obtain a file operation request initiated by a client, obtain a target file to be operated according to the file operation request, and perform corresponding operations on the target file;
[0142] S2. Send the file operation request to a second storage node associated with the first storage node according to a preset organization mode to instruct the second storage node to perform corresponding operations on the target file;
[0143] S3. Obtain a consensus message, and synchronize the first operation log in the local storage to the state indicated by the consensus message, where the first operation log stores the operation process performed on the target file;
[0144] S4. Detect whether the file in the local storage matches the file operation indicated in the first operation log, and in the case where it is detected that the file in the local storage does not match the file operation indicated in the first operation log, synchronize the file in the local storage to a state that is at least consistent with the file stored in the second storage node.
[0145] It should be noted that specific examples in this embodiment may refer to the examples described in the above-mentioned embodiment and optional implementation manners, and will not be elaborated in this embodiment.
[0146] In addition, in combination with the file operation method provided in the above-mentioned embodiment, a storage medium may also be provided in this embodiment to implement it. A computer program is stored on the storage medium; when the computer program is executed by a processor, any one of the file operation methods in the above-mentioned embodiment is implemented.
[0147] It should be understood that the specific embodiments described here are only used to explain this application, rather than to limit it. According to the embodiments provided in this application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the protection scope of this application.
[0148] Obviously, the accompanying drawings are only some examples or embodiments of this application. For those of ordinary skill in the art, this application can also be applied to other similar situations according to these drawings without creative work. In addition, it can be understood that although the work done during the development process here may be complex and time-consuming, for those of ordinary skill in the art, certain design, manufacturing, or production changes made according to the technical content disclosed in this application are only conventional technical means and should not be regarded as insufficient disclosure of this application.
[0149] As used in this application, the term "embodiment" means that the specific features, structures or characteristics described in connection with an embodiment may be included in at least one embodiment of this application. The phrase appears in various positions in the specification and does not necessarily mean the same embodiment, nor does it mean that it is independent or alternative to other embodiments and mutually exclusive. Those of ordinary skill in the art can clearly or implicitly understand that the embodiments described in this application can be combined with other embodiments without conflict.
[0150] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of patent protection. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all fall within the protection scope of this application. Therefore, the protection scope of this application shall be subject to the appended claims.
Claims
1. A file operation method, applied to a first storage node in a distributed storage system, the distributed storage system further including at least one storage node other than the first storage node, characterized in that, the method includes: obtaining a file operation request initiated by a client, obtaining a target file to be operated according to the file operation request, and performing a corresponding operation on the target file; sending the file operation request to a second storage node associated with the first storage node according to a preset organization mode to instruct the second storage node to perform a corresponding operation on the target file, so that each of the storage nodes in the distributed storage system performs a corresponding operation on the target file, where the preset organization mode means that the storage nodes are associated in sequence; obtaining a consensus message, and synchronizing a first operation log in local storage to the state indicated by the consensus message, where the first operation log stores the operation process performed on the target file; detecting whether the file in local storage matches the file operation indicated in the first operation log, and in the case where it is detected that the file in local storage does not match the file operation indicated in the first operation log, synchronizing the file in local storage to a state that is at least consistent with the file stored in the second storage node; wherein, detecting whether the file in local storage matches the file operation indicated in the first operation log includes: judging whether there is a situation of file loss or file redundancy in local storage according to the file operation indicated in the first operation log; in the case where it is detected that there is file loss or file redundancy in local storage, determining that the file in local storage does not match the file operation indicated in the first operation log.
2. The file operation method according to claim 1, characterized in that, synchronizing the file in local storage to a state that is at least consistent with the file stored in the second storage node includes: synchronizing the file in local storage to a state that is consistent with the file stored in the second storage node; or, sequentially querying target storage nodes that meet the file operation indicated in the first operation log according to the preset organization mode, and synchronizing the file in local storage to a state that is consistent with the file stored in the target storage node.
3. The file operation method according to claim 1, characterized in that, the file operation request carries the target file and information indicating adding the target file, obtaining the target file to be operated according to the file operation request, performing a corresponding operation on the target file, and sending the file operation request to a second storage node associated with the first storage node according to a preset organization mode to instruct the second storage node to perform a corresponding operation on the target file includes: writing the target file into local storage, and during the process of writing the target file into local storage, sending the target file to the second storage node to instruct the second storage node to add the target file.
4. The file operation method according to claim 3, characterized in that, Writing the target file to the local storage and instructing the second storage node to add the target file includes: In the local storage, splitting the target file into multiple file blocks, assigning a file block ID to each file block, and storing the multiple file blocks in the structure of a MerkleTree. Among them, the root node of the MerkleTree stores the file ID of the target file, the leaf nodes of the MerkleTree store the file content and file block ID of each file block, and the non-leaf nodes of the MerkleTree store links pointing to the corresponding leaf nodes; Instructing the second storage node to perform the same file addition operation as the local storage.
5. The file operation method according to claim 4, wherein, After writing the target file to the local storage and instructing the second storage node to add the target file, the method further includes: Receiving a file reading request initiated by the client, wherein the file reading request carries the file ID of the target file; Determining multiple target storage nodes according to the file reading request, and allocating reading tasks among the multiple target storage nodes, wherein the file block IDs carried in each reading task are different; Parallelly reading different file blocks in each target storage node, and splicing the multiple read file blocks according to the file block ID to obtain the target file; Returning the target file to the client.
6. The file operation method according to claim 1, wherein, The file operation request carries the file ID of the target file and information indicating the deletion of the target file. Obtaining the target file to be operated according to the file operation request, performing corresponding operations on the target file, and sending the file operation request to the second storage node associated with the first storage node according to a preset organization mode to instruct the second storage node to perform corresponding operations on the target file includes: Deleting the target file in the local storage according to the file ID, and instructing the second storage node to delete the target file.
7. The file operation method according to claim 1, wherein, After obtaining a file operation request initiated by the client, obtaining the target file to be operated according to the file operation request, and performing corresponding operations on the target file, the method further includes: Sequentially writing the operation process performed on the target file into the first operation log according to the execution order.
8. The file operation method according to claim 7, wherein, After sending the file operation request to the second storage node associated with the first storage node according to a preset organization mode, the method further includes: Receiving the execution result returned by the second storage node, wherein the execution result carries the node IDs of each storage node through which the file operation request flows and the execution status of each storage node; Determine the storage nodes in the distributed storage system that perform the corresponding operations on the target file and succeed, and determine whether the number of storage nodes that succeed in execution reaches a preset threshold; In the case where it is determined that the number of storage nodes that succeed in execution reaches the preset threshold, return a success response message to the client; and, In the case where it is determined that the number of storage nodes that succeed in execution does not reach the preset threshold, return a failure response message to the client.
9. The file operation method according to claim 8, wherein, after sequentially writing the operation process performed on the target file into the first operation log according to the execution order, the method further includes: Set the first operation log to an unconsensus state; In the case where it is determined that the number of storage nodes that succeed in execution does not reach the preset threshold, keep the state of the first operation log still in the unconsensus state, wherein the operation log set to the unconsensus state allows modification; and, In the case where it is determined that the number of storage nodes that succeed in execution reaches the preset threshold, set the first operation log to a consensus state, wherein the operation log set to the consensus state is not allowed to be modified.
10. The file operation method according to claim 8, wherein, The consensus message is generated by the Leader node elected by each storage node. In the case where it is determined that the number of storage nodes that succeed in execution does not reach the preset threshold, the method further includes: Participate in re-election of the Leader node in the distributed storage system; Receive the new consensus message sent by the re-elected Leader node, synchronize the first operation log to the state indicated by the new consensus message, and perform the corresponding operation on the file stored locally according to the updated first operation log.
11. A distributed storage system, wherein, including: Multiple storage nodes, each storage node is sequentially associated according to a preset organization mode. When the first storage node receives a file operation request initiated by the client, the first storage node is used to execute the file operation method according to any one of claims 1 to 10 above.
12. The distributed storage system according to claim 11, wherein, Each storage node includes: a file reading and writing module, a consensus module, a command execution module, and a file storage module; wherein, The file reading and writing module is used to transfer files between storage nodes; Each storage node stores an operation log, and the consensus module is used to maintain the consistency of the operation log between storage nodes, wherein the operation log stores the operation process performed on the file; The command execution module is used to maintain the consistency of files between storage nodes; The file storage module is used to store files.
13. An electronic device, including a memory and a processor, wherein, A computer program is stored in the memory, and the processor is configured to run the computer program to execute the file operation method according to any one of claims 1 to 10.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, the steps of the file operation method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Distributed file synchronization system and method
CN105956110A
Method and device for operating data object, computing equipment and storage medium
CN113742050A