Client and Distributed File System
By using programmable switches to manage write operations of multiple replicas in the client of a distributed file system, the problems of write latency and consistency management are solved, achieving high-performance and reliable data storage.
Patent Information
- Application Number
- CN202310314877.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-28
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-03-28
AI Technical Summary
In a distributed file system, multiple replicas are maintained to improve throughput and reduce read latency, but this increases the performance overhead of write latency and data consistency management.
By configuring a programmable switch on the client, sending it to the master server after receiving the write request, the master server generates multiple write requests and distributes them to the node server for copy writing. When the programmable switch receives the reply message for successful writes, it feedbacks the write success information to the user, and manages and rewrites temporary data copies if necessary.
By reducing the number of forwarding hops of write requests and reducing write latency, system performance and reliability are improved, while effectively managing write latency and data consistency in multiple replica systems.
Smart Images

Figure CN116303328B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a client and a distributed file system. Background Art
[0002] Computers manage and store data through file systems. In the era of information explosion, the amount of data that people can store has increased exponentially. Modern network services such as social networks, search engines, and e-commerce usually have strict throughput and latency requirements. To meet these needs, a high-performance distributed file system is essential, which can effectively solve the problems of data storage and management, and can also improve data reliability and access efficiency.
[0003] The file system realizes data stability and reliability by maintaining multiple copies of files, and improves throughput by dispersing data read requests to different copies, which requires maintaining multiple copies of the same data. Maintaining multiple copies is beneficial to improving system throughput and reducing read latency. By maintaining copies of data on multiple different servers, the throughput of the system can be linearly extended. If multiple servers can complete read operations, the tail latency of read requests will be greatly reduced.
[0004] However, maintaining multiple copies is not conducive to maintaining data consistency and will increase the write latency of the system. Since data is distributed on different servers, maintaining its consistency will generate additional system performance overhead, thus increasing the write latency of the system. Summary of the Invention
[0005] To solve at least one of the above technical problems, the present disclosure provides a client and a distributed file system.
[0006] A first aspect of the present disclosure provides a client of a distributed file system. The client is configured with a programmable switch, and the programmable switch is respectively communicatively connected to a master server and a node server of the distributed file system. The programmable switch is configured to perform the following operations: in response to a received first write request, send the first write request to the master server, so that the master server generates n second write requests according to the first write request, where n>1; receive the n second write requests sent by the master server and respectively forward them to corresponding first node servers for replica writing, where each second write request includes data to be written, and the number of the second write requests is determined according to the first write request; receive first reply information fed back by the first node server, where the first reply information indicates successful writing; when the number of first node servers that feedback the first reply information reaches m, feedback a writing success message to the user, where 1≤m<n.
[0007] According to an embodiment of the present disclosure, m = n - 1.
[0008] According to an embodiment of the present disclosure, when the number of node servers that feed back the first reply information reaches m, the programmable switch is configured to perform the following operations: send a reminder message to a second node server in the first node servers that has not fed back the first reply information, so that the second node server starts timing according to the reminder message; when the reply information fed back by the second node server is second reply information, perform replica writing to a third node server, where the second reply information indicates a write failure due to the failure to complete writing within a preset duration.
[0009] According to an embodiment of the present disclosure, the step of the programmable switch performing replica writing to the third node server includes: in response to the received second reply information, send a second allocation request to the master server, so that the master server generates and sends a third write request according to the second allocation request; in response to the third write request sent by the master server, perform replica writing to the third node server according to the temporary data replica of the data to be written stored in itself.
[0010] According to an embodiment of the present disclosure, the temporary data replica is established in the following manner: after receiving the second write request, the programmable switch generates and stores a temporary data replica of the data to be written.
[0011] According to an embodiment of the present disclosure, the data packet format used by the client to communicate with the master server and the node server includes a first packet header and a second packet header. The first packet header contains system control information, and the second packet header contains the data required by the request. The number of the second packet headers is determined according to the system control information.
[0012] According to an embodiment of the present disclosure, the first packet header includes the following fields: a type field, used to identify the type of the request information; a length field, used to identify the length of the data to be written and the number of the second packet headers; a request ID field, used to identify the unique ID information of the request information; an address field, used to identify the address of the request data; a node ID field, used to identify the ID information of the node server; an offset field, used to identify the offset information of the request data. The content of the node ID field and the offset field in the second write request is obtained by the master server through the address field.
[0013] According to an embodiment of the present disclosure, the programmable switch is further configured to perform the following operations: in response to a received first read request, if the data replicas corresponding to the first read request have completed all replica writes as indicated by the corresponding first write request, send the first read request to the primary server, so that the primary server generates and sends a second read request based on the first read request; receive the second read request sent by the primary server and forward it to the corresponding fourth node server for replica reading; receive the data replicas sent by the fourth node server.
[0014] According to an embodiment of the present disclosure, when the programmable switch receives a first read request, if the data replicas corresponding to the first read request have not completed all replica writes as indicated by the corresponding first write request, and the number of node servers that feedback the first reply information reaches m, the programmable switch reads the temporary data replicas of the data to be written stored in itself and feedbacks them to the user.
[0015] A second aspect of the present disclosure provides a distributed file system, including: a client, a primary server, and multiple node servers, where the client is the client in any of the above embodiments; when the programmable switch of the client receives a first write request, in response to the received first write request, send the first write request to the primary server; the primary server generates n second write requests based on the first write request and sends them to the programmable switch, where each second write request includes data to be written, and the number of the second write requests is determined according to the first write request, and n>1; the programmable switch receives the n second write requests and forwards them to the corresponding first node servers for replica writing respectively; the first node server sends a first reply information to the programmable switch when the writing is successful; the programmable switch receives the first reply information fed back by the first node server, and when the number of the first node servers that feedback the first reply information reaches m, feedbacks a write success information to the user, where 1≤m<n. Description of the Drawings
[0016] The drawings illustrate exemplary embodiments of the present disclosure and, together with the description thereof, are used to explain the principles of the present disclosure. These drawings are included to provide a further understanding of the present disclosure and are included in this specification and form a part of this specification.
[0017] Figure 1 is a schematic diagram of a distributed file system adopting a hardware implementation manner of a processing system according to an embodiment of the present disclosure.
[0018] Figure 2 is a schematic diagram of the workflow of a distributed file system for processing write requests according to an embodiment of the present disclosure.
[0019] Figure 3 It is a schematic diagram of the data packet format of a distributed file system according to an embodiment of the present disclosure.
[0020] Figure 4 It is a schematic diagram of the first header format of a distributed file system data packet according to an embodiment of the present disclosure.
[0021] Figure 5 It is a schematic diagram of the workflow for a distributed file system to process a read request according to an embodiment of the present disclosure.
[0022] Figure 6 It is a schematic diagram of the topological structure of a distributed file system test platform according to an embodiment of the present disclosure.
[0023] Figure 7 It is a CDF test result curve graph of the response latency of two distributed file systems.
[0024] Figure 8 It is a bar graph of the latency test results of two distributed file systems at the 90th percentile.
[0025] Figure 9 It is a pie chart of time decomposition of a distributed file system according to an embodiment of the present disclosure. Detailed Embodiments
[0026] The present disclosure will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the relevant content and do not limit the present disclosure. Additionally, it should be noted that only parts related to the present disclosure are shown in the drawings for ease of description.
[0027] It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other. The technical solutions of the present disclosure will be described in detail below with reference to the accompanying drawings and embodiments.
[0028] Unless otherwise specified, the exemplary embodiments / examples shown are understood to provide exemplary features of various details of some ways that can implement the technical concept of the present disclosure in practice. Therefore, unless otherwise specified, without departing from the technical concept of the present disclosure, the features of various embodiments / examples can be additionally combined, separated, interchanged, and / or rearranged.
[0029] The terms used in this specification are for the purpose of describing particular embodiments and are not limiting. As used herein, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are also intended to include the plural forms. In addition, when the terms "comprising" and / or "including" and their variants are used in this specification, it is stated that there are the stated features, integers, steps, operations, components, assemblies and / or groups thereof, but does not exclude the presence or addition of one or more other features, integers, steps, operations, components, assemblies and / or groups thereof. It should also be noted that, as used herein, the terms "substantially", "about" and other similar terms are used as approximate terms and not as terms of degree, so they are used to explain the inherent deviations of measured values, calculated values and / or provided values that would be recognized by a person of ordinary skill in the art.
[0030] A large-scale distributed file system may include a master server, a client, and multiple node servers. For example, Google's GFS system (Google File System) is a scalable distributed file system for large-scale, distributed applications that access large amounts of data. The node servers are used to store, manage, and maintain data blocks in the distributed file system and respond to read and write operations requested by the client.
[0031] For a multi-replica distributed file system, the files in the system are divided into blocks of a fixed size. For each file block, the system will copy it into multiple file replicas and store these replicas in different node servers according to a certain strategy.
[0032] The master server maintains the metadata of all files, including access control information, the mapping from files to blocks, the current locations of the blocks, etc. For the data operated by the user, whether it is adding, deleting, or modifying data, it needs to be synchronized to multiple replicas in different node servers. Therefore, when the user sends a write request in the distributed file system, multiple replicas need to be updated synchronously.
[0033] The client, distributed file system, and data reading and writing method of the present disclosure will be described below with reference to the accompanying drawings.
[0034] Figure 1 is a schematic structural diagram of a distributed file system adopting a hardware implementation manner of a processing system according to an embodiment of the present disclosure. Refer to Figure 1 , the distributed file system may adopt the system architecture of Gecko. Gecko is a tool designed according to the multi-replica writing strategy of the distributed file system to accelerate the writing operation of the distributed file system using a programmable switch.
[0035] In this embodiment, a distributed file system (such as Hadoop) may include 1 client, 1 master server, and 4 node servers (chunk servers). Among them, the master server is represented by master, and the 4 node servers are represented by chunk1 to chunk4 respectively. The client accesses the network. The client is configured with a programmable switch, which is represented by switch. The programmable switch switch is communicatively connected to the master server and the node servers chunk1 to chunk4 of the distributed file system respectively. When a request arrives at the distributed file system from the network, it needs to be processed by the client and the master server in sequence, and finally reaches the specific node server chunkserver to execute the request.
[0036] It can be understood that the numbers of the client, the master server, and the node servers included in the distributed file system can also be set to other numbers, and the present disclosure does not limit this.
[0037] The programmable switch switch is configured to perform the following operations S100 to S400.
[0038] S100, in response to the received first write request, send the first write request to the master server so that the master server generates n second write requests according to the first write request, where n>1.
[0039] S200, receive the n second write requests sent by the master server and forward them to the corresponding first node servers for replica writing respectively. Each of the second write requests contains data to be written, and the number of the second write requests is determined according to the first write request.
[0040] S300, receive the first reply information fed back by the first node server. The first reply information indicates that the writing is successful.
[0041] S400, when the number of the first node servers that feed back the first reply information reaches m, feed back the writing success information to the user. Where 1≤m<n.
[0042] Figure 2 It is a schematic diagram of the workflow for the distributed file system to process write requests according to an embodiment of the present disclosure. Refer to Figure 2 , the switch switch of the client receives the first write request issued by the user through the network, corresponding to Figure 2Phase ①. The first write request contains the data to be written into the node server, i.e., the data to be written. At this time, before sending the first write request to the master server, the switch can first process the first write request. The processing process may include: allocating a globally unique ID for the first write request and adding the ID information to the corresponding field content of the first write request.
[0043] The way to allocate a globally unique ID can be implemented through a monotonically increasing register configured on the switch. Whenever a request arrives at the switch from the network, the switch will allocate a unique ID for it, and the value of the register will increase by 1.
[0044] After allocating the unique ID, in order to know which node servers need to be written with the data to be written (replicas), the switch sends the processed first write request to the master server, corresponding to Figure 2 Phase ②.
[0045] In Gecko, when the first write request arrives at the switch from the network, by directly completing the processing of the first write request on the switch, the number of forwarding hops is reduced, so that both write requests and read requests are accelerated.
[0046] Exemplarily, the packet format adopted when the client communicates with the master server and the node server may include a first header header1 and a second header header2. Among them, the first header header1 contains system control information, and the second header header2 contains the data required by the request. The number of the second headers header2 is determined according to the above system control information.
[0047] Figure 3 is a schematic diagram of the packet format of a distributed file system according to an embodiment of the present disclosure. Refer to Figure 3 , in order to carry the data required by the system, two dedicated headers designed for Gecko can be adopted. The first header of Gecko is the control header, called GECKO, which contains some system control information and is carried by the TCP header. The second header is a dedicated header called TAIL, which is responsible for carrying the data required by the request. TAIL follows immediately after the previous header, and its number k is specified by the control information.
[0048] Figure 4 is a schematic diagram of the first header format of a distributed file system packet according to an embodiment of the present disclosure. Refer to Figure 4, the first packet header may include a type field, a length field, a request ID field, an address field, a node ID field, and an offset field. The first packet header may also include a data carrying field.
[0049] The type field (OP field) is used to identify the type of the request information, such as identifying whether the request is a read request or a write request, and the status of the current request. The length fields (LENGTH field and SUCCESS_LEN field) are used to identify the length of the data to be written and the number of the second packet headers. The data carrying field (HAVE_DATA field) is used to identify whether there is a second packet header TAIL after the first packet header GECKO. The request ID field (REQUEST_ID field) is used to identify the globally unique ID of the request information, which is given by the client. The address field (ADDRESS field) is used to identify the address of the request data. The node ID field (CHUNK_ID field) is used to identify the ID information of the node server. The offset field (OFFSET field) is used to identify the offset information of the request data, which is obtained by the master server through conversion based on the address field. Among them, when the switch receives the first write request and processes it, it will calculate the CHUNK_ID and OFFSET according to the ADDRESS and fill them into the corresponding fields. When the master server receives the first write request and processes it, it will find the corresponding CHUNK_ID and make corresponding modifications to the field content.
[0050] The second packet header TAIL may include a single 32-bit field for storing data.
[0051] Let N represent the number of node servers included in the distributed file system. In this embodiment, N = 4. When the master server receives the first write request (including the data to be written) sent by the switch, it can map the first write request to a specific node server. Let n represent the number of node servers that need to create data replicas. Generally, the value range of n is: 1 < n < N.
[0052] Suppose the distributed file system needs to write three copies of the same data. At this time, n = 3. Then the master server can determine 3 node servers among all the node servers. The determined node servers are called the first node servers. Taking the 3 determined first node servers as chunk1, chunk2, and chunk3 as an example, the master server will generate three second write requests corresponding to the three copies respectively, so as to map the first write request to the first node servers chunk1, chunk2, and chunk3. The mapping process mainly includes: determining the specific storage locations of n copies through a read-write address provided by the user. In this embodiment, the specific storage locations are the first node servers chunk1, chunk2, and chunk3.
[0053] After the master server generates three second write requests, it forwards them to the three first node servers chunk1, chunk2, and chunk3 through the switch switch, corresponding to Figure 2 phase ③. Each of the 3 first node servers will receive a corresponding second write request, and then create a data copy according to the data to be written contained therein.
[0054] After the first node server completes the creation of the data replica, it will feedback a first reply message to the switch switch to indicate that the data replica has been successfully created. The first reply message can be a success reply message, and this message may not contain the second packet header TAIL header. There will be a sequence of times for the first node servers to complete the creation of the data replicas, and there will also be a sequence of times for the switch switch to receive the first reply messages from each first node server.
[0055] When the programmable switch receives the first reply message, it determines the number of node servers that feedback the first reply message by counting the ID information in the first reply message. Specifically, in Figure 2 phase ③, when the switch switch receives the second write request sent by the master server, it can allocate a counter for the write request. Then, when the switch switch receives the first first reply message, assuming the first first reply message is sent by the first node server chunk1, corresponding to Figure 2Phase ④. At this time, the switch can obtain the request ID in the first reply message and use this counter to make statistics, so as to calculate the number of times the replicas are successfully created in each first node server. It can be understood that whenever a first reply message is received, the number of node servers that have currently sent the first reply message can be determined by counting the request ID. In addition, the number of times the replicas are successfully created can also be counted by other means.
[0056] m is a preset value, which is used to determine whether the number of first node servers that have successfully created replicas has reached a value that can report to the user in advance that all replicas have been successfully written. m is set to a value less than n, and m is usually set to a value close to n, that is, the difference between the two is small. In this embodiment, m = n - 1 = 2.
[0057] At Figure 2 Phase ④, only the first node server chunk1 sends the first reply message. At this time, the number of first node servers that have successfully created replicas counted by the counter is 1, which has not reached m. After counting the request ID of the first reply message of chunk1, the switch can discard the content of the first reply message.
[0058] At Figure 2 Phase ⑤, the first node server chunk2 sends the first reply message. At this time, the number of first node servers that have successfully created replicas is 2, which has reached m. At this time, the first node server chunk3 has not sent the first reply message yet. However, since only one first node server is still in the process of creating replicas and most of the first node servers have successfully created data replicas, at this time, it can be reported to the user that all replicas have been successfully written, that is, the write request has been completed. Compared with replying to the user that the write is successful after all replicas have been written, it avoids the long-tail delay of the write operation caused by transient server failures and improves the system performance.
[0059] A client of a distributed file system proposed according to an embodiment of the present disclosure utilizes the feature that a programmable switch can process data packets at wire speed, places the programmable switch on the data forwarding path of write requests in the distributed file system, thereby accelerating the write requests of large distributed storage systems in the network. At the same time, taking advantage of the significant advantages of the programmable switch in terms of latency and throughput, the functions of the client in the distributed file system are deployed on the programmable switch rather than on a traditional server, thereby reducing the forwarding hops of write requests in the distributed file system and reducing the long-tail latency of write requests. And when only a small number of replicas among multiple replicas have not been completed writing to the node server, report to the user that all replicas have been written successfully, and continue to independently process the remaining unfinished replicas in the background to reduce the tail latency. In this way, it can report the write success to the user in advance, reduce the write latency, accelerate the processing speed of write requests, and has better reliability; while ensuring the consistency of multiple replicas in the distributed file system, it can better apply to the elimination of write latency of multiple file replicas in a file system with a server as the carrier.
[0060] Exemplarily, when the number of node servers that have feedback the first reply information reaches m, the programmable switch is configured to further perform the following operations S500 to S700.
[0061] S500, send a reminder message to the second node server in the first node server that has not feedback the first reply information, so that the second node server can start timing according to the reminder message.
[0062] S600, when the reply information feedback by the second node server is the first reply information, the programmable switch erases the data to be written stored in itself.
[0063] S700, when the reply information feedback by the second node server is the second reply information, perform replica writing to the third node server. Wherein, the second reply information indicates a write failure due to not being completed within a preset duration.
[0064] Continuing with Figure 1 and Figure 2 as an example. After both the first node servers chunk1 and chunk2 have feedback the information of successful writing, the switch switch can report to the user that all replicas have been written successfully, and at the same time send a reminder message to the second node server chunk3 that has not yet feedback the successful writing. It can be understood that the second node server chunk3 belongs to the first node server. In addition, if m is set to be greater than 1, it means that there are multiple first node servers that have not yet feedback the successful writing at this time, so reminder messages will be sent to these multiple first node servers.
[0065] After the second node server chunk3 receives the reminder message, it can start timing. At the same time, chunk3 is also configured with a preset duration t0, which can be carried in the reminder message or pre-configured in chunk3. The timing value and the preset duration t0 are jointly used to determine whether the writing of chunk3 times out.
[0066] It can be understood that the timing method of the second node server can adopt forward timing or reverse timing.
[0067] If the second node server chunk3 completes data writing before the timing value reaches the preset duration t0, then chunk3 sends a first reply message to the switch switch, corresponding to Figure 2 phase ⑥. After receiving the first reply message sent by chunk3, the switch switch determines that all the first node servers have completed the creation of data replicas. Therefore, the switch switch can erase the temporary data replicas stored in itself and can release the count value in the counter.
[0068] Exemplarily, the establishment method of the temporary data replica can be: after receiving the second write request, the programmable switch generates a temporary data replica of the data to be written and stores it. Specifically, in Figure 2 phase ③, when the switch switch forwards the second write request to the first node server, when the switch switch identifies the data packet of the second write request, it can generate a temporary data replica of the data to be written according to the data to be written included in the second write request.
[0069] If the second node server chunk3 fails to complete data writing before the timing value reaches the preset duration t0, then chunk3 sends a second reply message to the switch switch, corresponding to Figure 2 phase ⑦. After receiving the second reply message sent by chunk3, the switch switch determines that there is a second node server that fails to complete the creation of data replicas on time and is regarded as a writing failure. Therefore, the switch switch needs to write replicas to other node servers, and the other node servers are other node servers except the first node servers.
[0070] By creating a temporary data replica on the switch switch, even if the second node server fails to write, it can ensure that successful writing can be carried out again on other servers, maintaining the validity of all replica writes successfully reported to the user before.
[0071] Due to the limited capacity of the programmable switch, it may be difficult to map the entire storage space to the switch, and maintaining a free space linked list for space allocation is relatively complex for the switch. Therefore, multiple hash functions can be used to map the address space. Whenever a new copy needs to be written to the switch, three different hash functions can be called to allocate available space. When any one of the three hash functions determines a legal free space, the switch can write the copy to the corresponding address. By using three different hash functions, the collision probability can be reduced.
[0072] In a distributed file system, the length of each piece of data is usually different. Therefore, all data can be divided into a preset minimum length unit, such as 4B. The data is written to the programmable switch according to the minimum length unit. Specifically, the written data is divided into several data units according to the minimum length unit, and each data unit is stored in a TAIL packet header. The length field of the GECKO packet header specifies the number of TAIL packet headers and the length of the data in the copy. On the switch, the length is determined by matching the length field of the GECKO packet header in each storage unit. If the length of the data is greater than the length specified by the storage unit, the data is written to the storage unit, thereby realizing the writing of variable-length data to the switch.
[0073] Exemplarily, in step S700, the steps for the programmable switch to write a copy to the third node server include S710 and S720.
[0074] S710, in response to the received second reply message, send a second allocation request to the primary server, so that the primary server can generate and send a third write request based on the second allocation request.
[0075] S720, in response to the third write request sent by the primary server, write a copy to the third node server based on the temporary data copy of the data to be written stored in itself.
[0076] In Figure 2 phase ⑦, when the switch determines that there is a second node server with a write failure through the received second reply message, it will generate a second allocation request and send it to the primary server master, so as to apply to the primary server master for reallocating a node server to write the copy.
[0077] After the master server receives the second allocation request, it generates a third write request and sends it to the switch. The third write request contains the ID of the third node server as the new allocation target. After receiving the third write request, the switch sends the temporary data copy of the data to be written that it created before to the third node server corresponding to the third write request, corresponding to Figure 2 in stage ⑧. In this embodiment, the third node server is chunk4.
[0078] Compared with reading data from the node server where the copy has been successfully written and writing it to the third node server, the method of reading data from the switch and writing it to the third node server reduces the communication overhead in the system and further accelerates the write request.
[0079] When the third node server chunk4 finishes writing the data, it sends a first reply message to the switch indicating that the writing is successful, corresponding to Figure 2 in stage ⑨. When the switch receives the first reply message sent by chunk4, it counts the request IDs and determines that 3 node servers have completed the creation of data copies, reaching the target of n = 3. At this time, the switch can erase the temporary data copy stored in itself and can release the count value in the counter.
[0080] Exemplarily, the programmable switch is further configured to perform the following operations S810 to S830.
[0081] S810, in response to the received first read request, if the data copy corresponding to the first read request has completed all copy writes as indicated by the corresponding first write request, send the first read request to the master server so that the master server generates and sends a second read request based on the first read request.
[0082] S820, receive the second read request sent by the master server and forward it to the corresponding fourth node server for copy reading.
[0083] S830, receive the data copy sent by the fourth node server.
[0084] Figure 5 is a schematic diagram of the workflow for a distributed file system to process read requests according to an embodiment of the present disclosure. Refer to Figure 5 , when the client receives the first read request (corresponding to Figure 5 in stage ①), it forwards it to the master server through the switch for data location query (corresponding to Figure 5 in stage ②).
[0085] It can be understood that before the switch sends the first read request to the master server, it can first process the first read request. The processing process may include: calculating the node ID and offset based on the content of the address field, and writing the node ID and offset into the corresponding fields in the request data packet respectively.
[0086] When the master server receives the first read request forwarded by the switch, it determines the node server storing the data to be read according to the first read request, and then selects a node server therefrom. The selected node server is called the fourth node server. Then, the data replica of the fourth node server is read. In this embodiment, the fourth node server is chunk1 (corresponding to Figure 5 phase ③). The fourth node server sends the data replica to the switch to implement data reading.
[0087] Exemplarily, when the programmable switch receives the first read request, if the data replica corresponding to the first read request has not completed all replica writes according to the indication of the corresponding first write request, and the number of node servers that feedback the first reply message reaches m, the programmable switch reads the temporary data replica of the data to be written stored by itself and feedbacks it to the user.
[0088] If the data replica to be read is being written to the node server, that is to say, the same data is to be read before all replicas are created. At this time, if at least m node servers have completed replica creation, the switch directly uses the temporary data replica of the data to be read (which is also the data to be written) generated by itself before as the data replica to be read and feedbacks it to the user (corresponding to Figure 5 phase ④). If the number of node servers that have completed replica creation has not reached m, the read request will be blocked on the switch until the number of node servers that have completed replica creation reaches m.
[0089] The following is the test process of Gecko of the distributed file system in this embodiment.
[0090] The Gecko prototype of this embodiment is implemented on the Barefoot Tofino switch and server. The user accesses the data in Gecko using the virtual linear storage space, and Gecko cuts the data into blocks of the same size internally. Each piece of data will be copied into 3 replicas and stored on the blocks across different servers.
[0091] In the multi-copy structure as a comparative example, requests are processed by a server-based client instead of a switch-based client. Additionally, in the multi-copy structure, the client only returns success to the user after all three copies of the write request have been successfully written.
[0092] Figure 6 FIG. is a schematic diagram of the topology of a distributed file system test platform according to an embodiment of the present disclosure. Refer to Figure 6 , experiments were conducted on a Barefoot Tofino switch and six servers. Each server machine was equipped with 2 Intel Xeon Silver 4210R CPUs, 128GB of memory, and a Mellanox ConnectX-5 network card. Using Figure 6 the test platform topology shown to implement the Gecko prototype and the multi-copy structure. The Gecko prototype was built with a programmable switch-based client, a user, a master server, and four node servers chunk1-chunk4. When implementing the multi-copy structure, a server was used as the client, and a normal switch was used to connect a user, a master server, a client, and three node servers chun1-chunk3 as a distributed file system.
[0093] A write workload was used for evaluation. The workload was generated by concurrent clients with arrival intervals. The latency of Gecko and the multi-copy system can be tested and analyzed, that is, the time interval from when the user issues a request to when the user confirms that the request has been successfully written. Additionally, the long-tail latency (i.e., one of the three copies has a large latency) can be manually increased for 10% of the write requests to simulate an abnormal jitter scenario.
[0094] Figure 7 FIG. is a CDF (Cumulative Distribution Function) test result curve graph of the response latency of two distributed file systems. Refer to Figure 7 , in terms of abnormal jitter, there is an abnormal surge in 10% of the response latency in curve C2 of the traditional multi-copy system (Multi-replica). This is due to the long-tail latency in the scenario of abnormal jitter. The write request can only be completed when the slowest request among all copies is completed. Gecko of the present embodiment can tolerate abnormal jitter, as shown in curve C1. By placing the abnormally jittered copies in the background, the front-end traffic can flow normally while ensuring data consistency. This greatly reduces the impact of long-tail latency on performance.
[0095] Figure 8 It is a bar chart showing the latency test results of two distributed file systems at the 90th percentile. Refer to Figure 8 , in terms of latency, Figure 8 shows the 90th percentile response latency of the traditional multi-copy system and Gecko of this embodiment. It can be seen that Gecko of this embodiment significantly reduces the latency by 60 times. In addition, the latency of Gecko is lower than that of the multi-copy system in all cases. Because the number of copies to be completed reported by Gecko is always one less than that of the traditional multi-copy system, which makes Gecko have a faster completion time even under normal circumstances. And it also changes the processing flow of multi-copy in the traditional distributed system. By deploying the client function on the switch, the system greatly reduces the forwarding hops of write requests.
[0096] Figure 9 It is a pie chart showing the time decomposition of a distributed file system according to an embodiment of the present disclosure. Refer to Figure 9 , in terms of time decomposition, it shows the time decomposition of Gecko. The transmission of data in the network occupies most of the processing time in Gecko. Since the data in the experiment is transmitted using the TCP protocol, the data packets need to pass through the kernel to reach the network card, which increases the latency of write requests to a certain extent.
[0097] Through strict analysis and a large number of experiments on the Barefoot Tofino switch, it shows that Gecko not only greatly reduces the long-tail write latency caused by server transient failures, but also maintains the same reliability level as the classic triple-copy write operation.
[0098] Refer to Figure 1 , the present disclosure also provides a distributed file system, including: a client, a master server, and multiple node servers chunk1-chunk4, and the client is the client of any of the above embodiments.
[0099] When the programmable switch of the client receives the first write request, in response to the received first write request, it processes the first write request and sends the processed first write request to the master server.
[0100] The master server generates n second write requests based on the processed first write request and sends them to the programmable switch, where each second write request contains the data to be written, and the number of second write requests is determined according to the first write request, and n>1.
[0101] The programmable switch receives n second write requests and forwards them to the corresponding first node servers for replica writing respectively.
[0102] When the first node server writes successfully, it sends a first reply message to the programmable switch.
[0103] The programmable switch receives the first reply message fed back by the first node server, and when the number of first node servers that feed back the first reply message reaches m, it feeds back a write success message to the user, where 1 ≤ m < n.
[0104] It should be noted that for the details not disclosed in the distributed file system of this embodiment, reference can be made to the details disclosed in the client of the distributed file system of the above embodiment proposed in this disclosure, which will not be elaborated here.
[0105] It should be understood that each part of the present disclosure can be implemented by hardware, software, or a combination thereof. In the above embodiment, multiple steps or methods can be implemented by software stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following technologies well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0106] Those of ordinary skill in the art of this technology can understand that all or part of the steps for implementing the above embodiment method can be completed by a program instructing relevant hardware. The program can be stored in a readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0107] In addition, in each embodiment of the present disclosure, each functional unit can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a readable storage medium. The storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0108] In the description of this specification, the descriptions referring to terms such as "one embodiment / way", "some embodiments / ways", "specific examples", or "some examples", etc., mean that the specific features, structures, materials, or characteristics described in connection with the embodiment / way or example are included in at least one embodiment / way or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment / way or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments / ways or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments / ways or examples described in this specification and the features of different embodiments / ways or examples.
[0109] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0110] Those skilled in the art should understand that the above embodiments are only for clearly explaining the present disclosure and are not intended to limit the scope of the present disclosure. For those skilled in the art, other changes or modifications can be made on the basis of the above disclosure, and these changes or modifications are still within the scope of the present disclosure.
Claims
1. A client of a distributed file system, characterized in that The client is configured with a programmable switch, which is communicatively connected to the master server and the node servers of the distributed file system respectively, and the programmable switch is configured to perform the following operations: In response to a received first write request, send the first write request to the master server so that the master server generates n second write requests based on the first write request, where n>1; Receive the n second write requests sent by the master server and forward them to the corresponding first node servers for replica writing respectively. Each second write request contains data to be written, and the number of the second write requests is determined according to the first write request; Receive the first reply information fed back by the first node server, and the first reply information indicates successful writing; When the number of first node servers that have fed back the first reply information reaches m, feedback a write success message to the user, where 1≤m<n; When the number of node servers that have fed back the first reply information reaches m, the programmable switch is configured to perform the following operations: Send a reminder message to the second node servers among the first node servers that have not fed back the first reply information, so that the second node servers start timing according to the reminder message; When the reply information fed back by the second node server is the second reply information, perform replica writing on the third node server, and the second reply information indicates write failure due to failure to complete writing within a preset time; The steps for the programmable switch to perform replica writing on the third node server include: In response to the received second reply information, send a second allocation request to the master server so that the master server generates and sends a third write request based on the second allocation request; In response to the third write request sent by the master server, perform replica writing on the third node server according to the temporary data replica of the data to be written stored by itself.
2. The client according to claim 1, wherein m=n-1.
3. The client according to claim 1, characterized in that, The temporary data replica is established in the following way: after receiving the second write request, the programmable switch generates and stores a temporary data replica of the data to be written.
4. The client according to claim 1, wherein The data packet format used by the client to communicate with the master server and the node servers includes a first packet header and a second packet header. The first packet header contains system control information, and the second packet header contains the data required by the request. The number of the second packet headers is determined according to the system control information.
5. The client according to claim 4, wherein The first packet header includes the following fields: A type field, used to identify the type of request information; A length field, used to identify the length of the data to be written and the number of the second packet headers; A request ID field, used to identify the unique ID information of the request information; An address field, used to identify the address of the request data; A node ID field, used to identify the ID information of the node server; An offset field, used to identify the offset information of the request data. Among them, the content of the node ID field and the offset field in the second write request is obtained by the master server through the address field conversion.
6. The client according to claim 1, wherein The programmable switch is also configured to perform the following operations: In response to the received first read request, if the data replicas corresponding to the first read request have completed all replica writes as instructed by the corresponding first write request, the first read request is sent to the primary server so that the primary server generates and sends a second read request based on the first read request; Receive the second read request sent by the primary server and forward it to the corresponding fourth node server for replica reading; Receive the data replicas sent by the fourth node server.
7. The client according to claim 6, characterized in that, When the programmable switch receives a first read request, if the data replicas corresponding to the first read request have not completed all replica writes as instructed by the corresponding first write request, and the number of node servers that feedback the first reply information reaches m, the programmable switch reads the temporary data replicas of the data to be written stored in itself and feedbacks them to the user.
8. A distributed file system, characterized in that, Comprising: A client, a primary server, and multiple node servers, where the client is the client described in any one of claims 1-7; When the programmable switch of the client receives a first write request, in response to the received first write request, the first write request is sent to the primary server; The primary server generates n second write requests based on the first write request and sends them to the programmable switch, where each second write request contains the data to be written, and the number of the second write requests is determined based on the first write request, n>1; The programmable switch receives the n second write requests and forwards them to the corresponding first node servers for replica writing respectively; When the first node server writes successfully, it sends a first reply information to the programmable switch; The programmable switch receives the first reply information fed back by the first node server, and when the number of first node servers that feedback the first reply information reaches m, it feedbacks a write success message to the user, 1≤m<n.
Citation Information
Patent Citations
Multistage cache distributed key value storage system based on programmable switch
CN114844846A
Transaction commit system and method, and related device
WO2021103036A1