Data Processing Method, System, Device, Electronic Device and Computer Storage Medium

Through the coordinated work of the client and the management node, the data inconsistency problem caused by the write exception of data nodes in the client-base distributed storage system is solved, and the data consistency guarantee is achieved.

CN113420035BActive Publication Date: 2025-07-11ALIBABA GROUP HOLDING LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110164604.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-05
Publication Date
2025-07-11
Estimated Expiration
2041-02-05

AI Technical Summary

Technical Problem

In client-base distributed storage system, when data nodes write abnormalities, the prior art can easily cause request glitch problems and cannot guarantee data consistency.

Method used

The client determines the incomplete data node and sends the write length to the management node. The management node guides the data node to increment or complete the entire amount according to the write length.

Benefits of technology

Through the coordinated work of the client and management node, data consistency in the distributed data storage system is ensured and request glitches are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113420035B_ABST
    Figure CN113420035B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a data processing method, system, device, electronic device, and computer storage medium. In a distributed data storage system, a client determines the storage progress of each data node. When there is a first data node with incomplete written data, the data is retried for writing, and after the retry writing fails, the coordination management node is coordinated to asynchronously complete the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular, to a data processing method, system, device, electronic device, and computer storage medium. Background Art

[0002] Distributed storage systems have currently been widely applied to various data storage scenarios. In a distributed storage system, it is necessary to maintain the consistency of data on each data node.

[0003] In the currently commonly used client-base distributed storage system, multiple data nodes are peer-to-peer, and the client coordinates the read / write consistency and availability among multiple data nodes. If there is an exception when writing data to a certain data node, the client needs to keep retrying or continue after the client successfully replicates all the data, which is likely to cause request jitter problems. In addition, since the client itself is not necessarily reliable, even if the client replicates all the data, it may not necessarily ensure complete consistency.

[0004] Based on this, a more reliable data processing solution for guaranteeing data consistency is needed. Summary of the Invention

[0005] In view of this, the embodiments of the present application provide a reliable data processing solution for guaranteeing data consistency to at least partially solve the above problems.

[0006] According to the first aspect of the embodiments of the present application, a data processing method is provided, which is applied to a client of a client-base storage system. The method includes: after determining that the retry of writing data to a first data node fails, determining a first write length in the first data node and a second write length in a second data node, where the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; sending the first write length and the second write length to a management node, so that the management node guides the first data node to supplement the data according to the first write length and the second write length.

[0007] According to a second aspect of the embodiments of the present application, another data processing method is provided, which is applied to a management node of a client-base storage system. The method includes: receiving a first write length in a first data node and a second write length in a second data node sent by a client, where the first write length and the second write length are determined by the client when writing data to the data node, the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; determining a filling position of the first data node according to the first write length and the second write length; guiding the first data node to obtain data corresponding to the second write length from the second data node and perform incremental filling starting from the filling position.

[0008] According to a third aspect of the embodiments of the present application, a data processing system is provided, including a client, a data node, and a management node. In the system, the client, after determining that the retry of writing data to the first data node fails, determines a first write length in the first data node and a second write length in the second data node, where the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; sends the first write length and the second write length to the management node; the management node is configured to receive the first write length in the first data node and the second write length in the second data node sent by the client, determine the filling position of the first data node according to the first write length and the second write length; guide the first data node to obtain data corresponding to the second write length from the second data node and perform incremental filling starting from the filling position.

[0009] According to a fourth aspect of the embodiments of the present application, a data processing device is provided, which is applied to a client of a client-base storage system. The device includes: a determination module, after determining that the retry of writing data to the first data node fails, determines a first write length in the first data node and a second write length in the second data node, where the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; a sending module, sends the first write length and the second write length to the management node so that the management node guides the first data node to fill the data according to the first write length and the second write length.

[0010] According to a fifth aspect of the embodiments of the present application, another data processing device is provided, which is applied to a management node of a client-base storage system. The device includes: a receiving module, configured to receive a first write length in a first data node and a second write length in a second data node sent by a client, where the first write length and the second write length are determined by the client when writing data to the data node, the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; a position determination module, configured to determine a filling position of the first data node according to the first write length and the second write length; a guiding module, configured to guide the first data node to obtain data corresponding to the second write length from the second data node and perform incremental filling starting from the filling position.

[0011] According to a sixth aspect of the embodiments of the present application, an electronic device is provided, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations corresponding to the method described in the first aspect or the second aspect.

[0012] According to a seventh aspect of the embodiments of the present application, a computer storage medium is provided, on which a computer program is stored. When the program is executed by a processor, it implements the data processing method described in the first aspect or the second aspect.

[0013] According to the solution provided by the embodiments of the present application, in a client-base distributed data storage system, the client determines the storage progress of each data node. When there is a first data node with incomplete written data, retry writing the data, and after the retry writing fails, coordinate with the management node to asynchronously fill the data, thereby ensuring data consistency in the distributed data storage system. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0015] Figure 1 It is a schematic diagram of the architecture of a system related to the present application;

[0016] Figure 2 It is a schematic flowchart of the steps of a data processing method provided by the embodiments of the present application;

[0017] Figure 3a Schematic diagram of a client retrying to write data provided by an embodiment of the present application;

[0018] Figure 3b Schematic diagram of a management node guiding data completion provided by an embodiment of the present application;

[0019] Figure 3c Schematic diagram of data completion when a client is abnormal provided by an embodiment of the present application;

[0020] Figure 4 Schematic flowchart of another data processing method provided by an embodiment of the present application;

[0021] Figure 5 Schematic diagram of the structure of a data processing device provided by an embodiment of the present application;

[0022] Figure 6 Schematic diagram of the structure of another data processing device provided by an embodiment of the present application

[0023] Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0024] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art shall fall within the protection scope of the embodiments of the present application.

[0025] The Client-base storage system is a distributed storage system. In this system, it includes a client, multiple data nodes, and a management node. As Figure 1 shown, Figure 1 Schematic diagram of the architecture of a client-base storage system related to the present application. The client writes the same piece of data into multiple different data nodes. Currently, if an exception occurs when writing data to a certain data node, usually the client needs to perform a full copy of the data to continue writing data, which is likely to cause request jitter problems. Based on this, the present application provides a more reliable data processing solution for ensuring data consistency.

[0026] The following further illustrates the specific implementation of the embodiments of the present application with reference to the accompanying drawings of the embodiments of the present application. Specifically, it includes two aspects: the client and the management node.

[0027] For the first aspect of the present application, as Figure 2 shown, Figure 2 is a schematic flowchart of the steps of a data processing method provided by an embodiment of the present application, which is applied to the client of a client-base storage system and includes:

[0028] S201, after determining that the retry of writing data to the first data node fails, determine the first write length in the first data node and the second write length in the second data node, where the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data.

[0029] As the data writer, the client can know the write status information of each data node while writing data, including the write length of the data, the result of writing the data, and whether the communication status between itself and each data node is normal, etc.

[0030] In practical applications, due to various reasons such as network communication speed and machine performance (i.e., there may be slow nodes or bad nodes in the data nodes), the possible write lengths of the same data on different data nodes may be different. For example, assume that the complete data length is 100, the write lengths in data nodes 1 and 2 are 100 (i.e., complete writing), while the write length in data node 3 is 80 (i.e., incomplete writing).

[0031] At this time, the data node with incomplete written data can be confirmed as the first data node, and the corresponding write length is the first write length; while the data node with complete written data is the second data node, and the corresponding write length is the second write length.

[0032] In the previous example, data node 3 is the first data node, and data nodes 1 and 2 are the second data nodes. Obviously, the second write length in the second data node is greater than the first write length in the first data node.

[0033] In the present application, the client can adopt the most synchronous write method when writing data to each data node. That is, as long as the client confirms that the number of data nodes with complete writing meets the preset conditions of the most synchronous write (for example, the number of data nodes with complete writing exceeds a certain value, or the proportion of the number of data nodes with complete writing exceeds a certain proportion, usually more than 50%), it is considered that the data writing is successful, thus effectively avoiding slow nodes and bad nodes. Then at this time, other data nodes with incomplete written data need to supplement the data to ensure data consistency in the distributed system.

[0034] While writing data, the client can monitor the status information in real time to determine whether there is a first data node in the system. If it is determined that there is a first data node, then the data can be retried to be written into the first data node first. As Figure 3a shown Figure 3a is a schematic diagram of a client retrying to write data provided by an embodiment of the present application.

[0035] This retry write is usually used to handle a short-term exception that occurs in the first data node, including situations such as slow disk data writing, slow network communication speed, and temporary network disconnection, etc. At this time, the data to be written is still saved in the memory of the client. Therefore, the client can try to perform incremental complementation to the first data node.

[0036] The way of incremental complementation is to take the position corresponding to the first length as the starting point of the complementation position, and the data position corresponding to the second length as the end point of the complementation position, and perform data complementation from the starting point to the end point to achieve complementing the write length of the first data node to the write length of the second data node. Continuing the previous example, for data node 3, the client will try to complement the data segment corresponding to the length from 81 to 100 to it.

[0037] As mentioned above, this retry complementation is for data nodes with short-term exceptions. Therefore, it usually cannot last for too long to avoid excessive load caused by continuous client retries. Specifically, the client can perform a specified number of times (for example, 5 times) of retrying to write the data into the first data node, or perform a specified duration (for example, retry for 1 minute) of retrying to write the data into the first data node. Exceeding the specified number of times or specified duration is considered a failure to retry writing the data.

[0038] At this time, the client will terminate the retry write, consider that the data node may have a continuous exception (such as crashing or being offline for a long time), and determine the first data node and its first write length, and determine the second data node (since there are multiple second data nodes, any one of the second data nodes can be selected randomly) and its second write length.

[0039] S203. Send the first write length and the second write length to the management node so that the management node can guide the first data node to complement the data according to the first write length and the second write length.

[0040] After the client sends the first write length and the second write length to the management node, the management node can then guide the first data node to complete the data based on this. For example, determine the filling position of the first data node according to the first write length and the second write length; guide the first data node to obtain the data corresponding to the second write length from the second data node, and start incremental filling from the filling position, etc.

[0041] Through this embodiment, in a client - base distributed data storage system, the client determines the storage progress of each data node. When there is a first data node with incomplete written data, retry writing the data, and after the retry write fails, coordinate with the management node to asynchronously complete the data, thereby ensuring data consistency in the distributed data storage system.

[0042] The foregoing part has described the workflow on the client side. In the second aspect of the embodiments of the present application, as Figure 4 shown, Figure 4 is a schematic flowchart of another data processing method provided by the embodiments of the present application, which is applied to the management node of a client - base storage system. The method includes:

[0043] S401, receive the first write length in the first data node and the second write length in the second data node sent by the client, where the first write length and the second write length are determined by the client when writing data to the data node, the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data.

[0044] In a client - base distributed storage system, the management node stores the relevant meta - information of each data node, such as the status information of the stored file, the location information of the data block (chunk), etc. The management node can also perform system member liveness detection and member change, etc.

[0045] By receiving the first write length in the first data node and the second write length in the second data node, the management node can know which data nodes have writing exceptions and which data nodes have complete writes. When the client notifies the management node of the lag in writing of the first data node, it can also apply to the management node for a new data block to perform subsequent data writing to the first data node.

[0046] S402, determine the filling position of the first data node according to the first write length and the second write length.

[0047] The method for determining the filling position in the first data node is similar to the determination method in the client, and will not be elaborated here.

[0048] S403, guide the first data node to obtain the data corresponding to the second write length from the second data node, and perform incremental padding starting from the padding position.

[0049] The management node can then notify the first data node and inform the first data node of the write lengths of other second data nodes at this time. If the first data node has relevant position information of the second data node (including the second data node identifier, address, file storage location, etc.), then the first data node can pull the corresponding data from the second data node, and determine the padding position locally according to the second write length to perform incremental padding.

[0050] If the first data node does not have the relevant information of the second data node, then the management node will send the relevant information of the second data node and the second write length to the first data node, so that the first data node can perform incremental padding according to the relevant information and the second write length.

[0051] According to the solution provided by the embodiment of the present application, in a client-base distributed data storage system, the client determines the storage progress of each data node. When there is a first data node with incomplete written data, retry writing the data, and after the retry write fails, send the relevant information to the management node. The management node determines the padding position of the first data node according to the information sent by the client, including the first write length of the first data node and the second write length of the second data node, and then guides the first data node to asynchronously pad the data, thereby ensuring data consistency in the distributed data storage system.

[0052] In one embodiment, the management node also needs to know the padding status of the first data node, that is, it needs to receive the incremental padding result of the incremental padding returned by the first data node, including padding success, padding failure, or no response message.

[0053] Whether it is a padding failure or a no response message, the management node can consider that there is a problem with the hardware or network of the first data node, resulting in the inability to pad the data, which will lead to the inability to meet the data consistency required in the distribution.

[0054] Therefore, in one embodiment, when the padding result received by the management node indicates a padding failure (including padding failure or no response message), the management node can also select another data node from the system and guide the other data node to perform full padding on the data. The selected other node is different from the previously determined first data node and second data node.

[0055] Specifically, it is possible to guide another data node to perform full replenishment from the client or from the aforementioned second data node. When performing full replenishment from the client, the management node can guide the other data node to establish communication with the client, so that the client can write full data into the other data node. When performing full replenishment from the second data node, the management node can send relevant information of the second node to the other data node, so that the other data node can pull and store full data from the second data node. As Figure 3b shown, Figure 3b is a schematic diagram of a management node guiding data replenishment provided by an embodiment of the present application. In this schematic diagram, data nodes 1 and 2 are identified as second data nodes, and the replenishment steps are carried out in the order shown by the numbers in the figure.

[0056] As mentioned above, the management node can also probe the liveness of the members in the system in real time. Here, the members include each data node and the client. For example, the management node establishes a heartbeat monitoring mechanism with the client and each data node to ensure the effectiveness of the connection. When the heartbeat data packet of the client cannot be received, it can be known that the client is in an abnormal working state.

[0057] As mentioned above, the management node itself also stores the meta-information of each data node. Therefore, in one embodiment, when the management node determines that the client is not working properly, the management node can also query the length of the data written in each data node from the meta-information, so as to determine the first data node and the second data node.

[0058] In one embodiment, if the most synchronous write method is adopted in the system, the writing of data still needs to meet the preset conditions of the most synchronous write: that is, complete data is written on a certain number or a certain proportion of data nodes. If this condition is not met, then the current data write is considered a write failure, and each data node will roll back to the state before the write.

[0059] After determining that the write is successful, it is obvious that more than half of the data nodes are second data nodes at this time, and the corresponding write length is the second write length. Other lengths lower than the second write length are the first write lengths, and the corresponding nodes are the first data nodes. Furthermore, the replenishment position can be determined based on the first write length and the second write length as described above, and the first data node can be guided to obtain data from the second data node for incremental or full data replenishment.

[0060] In another embodiment, the management node may also compare the write lengths on each data node, determine the maximum write length as the second write length, and the corresponding data node is the second data node, and the remaining data nodes are the first data nodes. Further, the first data nodes may be guided to complete the data according to the written data in the second data node. As Figure 3c shown Figure 3c is a schematic diagram of data completion when a client is abnormal provided by an embodiment of the present application. In this actual diagram, the management node determines that data nodes 1 and 2 are the second data nodes, and determines one first data node, and the completion steps are carried out in the order shown by the numbers in the figure. After the incremental completion of the first data node fails, another node may be selected for full completion as described above.

[0061] In a third aspect of the embodiments of the present application, a data processing system is further provided. The data processing system adopts a client-base storage method. The data processing system includes a client, data nodes, and a management node. The architecture of the system is as Figure 1 shown. In the system:

[0062] The client, after determining that the retry of writing data to the first data node fails, determines the first write length in the first data node and the second write length in the second data node, where the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; sends the first write length and the second write length to the management node;

[0063] The management node is configured to receive the first write length in the first data node and the second write length in the second data node sent by the client, determine the completion position of the first data node according to the first write length and the second write length; guide the first data node to obtain the data corresponding to the second write length from the second data node, and start incremental completion from the completion position.

[0064] Optionally, the client is further configured to write data to multiple data nodes before determining that the retry of writing data to the first data node fails; determine the status information of each data node for the written data, where the status information is used to characterize the write length and write result of the client writing data to multiple data nodes; determine whether there are first data nodes and second data nodes from the multiple data nodes according to the status information; when there is a first data node, retry writing the data to the first data node.

[0065] In this embodiment, in a client-base based distributed data storage system, the management node determines the data filling position of the first data node according to the first writing length in the first data node and the second writing length in the second data node sent by the client, and guides the first data node to fill the data at the filling position with the data of the second writing length obtained from the second data node, thereby ensuring data consistency in the distributed data storage system.

[0066] In the fourth aspect of the embodiments of the present application, a data processing device is further provided, as Figure 5 shown Figure 5 is a schematic structural diagram of a data processing device provided by the embodiments of the present application, which is applied to the client of the client-base storage system. The device includes:

[0067] A determination module 501, after determining that the retry writing of data to the first data node fails, determines the first writing length in the first data node and the second writing length in the second data node, where the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data;

[0068] A sending module 503 sends the first writing length and the second writing length to the management node, so that the management node guides the first data node to fill the data according to the first writing length and the second writing length.

[0069] Optionally, the device further includes a retry writing module 505, which is used to write data to multiple data nodes before determining that the retry writing of data to the first data node fails; determine the status information of each data node for the written data, where the status information is used to characterize the writing length and writing result of the client writing data to multiple data nodes; determine whether there are a first data node and a second data node among the multiple data nodes according to the status information; when there is a first data node, retry writing the data to the first data node.

[0070] Optionally, the retry writing module 505 retries writing the data to the first data node for a specified number of times, or retries writing the data to the first data node for a specified duration.

[0071] The data processing device in this embodiment is used to implement the corresponding data processing method of the management node in the foregoing method embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the data processing device in this embodiment can be referred to the description of the corresponding part in the foregoing method embodiment, which will not be elaborated here either.

[0072] In the fifth aspect of the embodiments of the present application, another data processing device is further provided. As Figure 6 shown, Figure 6 FIG. is a schematic structural diagram of another data processing device provided by the embodiments of the present application, which is applied to a management node of a client-base storage system. The device includes:

[0073] A receiving module 601, configured to receive a first write length in a first data node and a second write length in a second data node sent by a client, where the first write length and the second write length are determined by the client when writing data into the data node. The first data node is a data node with incomplete written data, and the second data node is a data node with complete written data;

[0074] A position determining module 603, configured to determine a filling position of the first data node according to the first write length and the second write length;

[0075] A guiding module 605, configured to guide the first data node to obtain data corresponding to the second write length from the second data node and perform incremental filling starting from the filling position.

[0076] Optionally, the receiving module 601 receives an incremental filling result of the incremental filling returned by the first data node.

[0077] Optionally, when the filling result indicates filling failure, the guiding module 605 is further configured to select another data node and guide the other data node to perform full filling on the data.

[0078] Optionally, the guiding module 605 guides the other data node to perform full filling on the data from the second data node or from the client.

[0079] Optionally, the device further includes a monitoring module 607, configured to determine whether the client is working properly; when it is determined that the client is not working properly, query the lengths of the written data in each data node; determine the first data node and the second data node according to the lengths of the written data in each data node; correspondingly, the guiding module 605 guides the first data node to perform filling according to the written data in the second data node.

[0080] The data processing device in this embodiment is used to implement the corresponding data processing method of the client in the foregoing method embodiment, and has the beneficial effects of the corresponding method embodiment, which will not be elaborated here. In addition, the function implementation of each module in the data processing device in this embodiment can refer to the description of the corresponding part in the foregoing method embodiment, which will not be elaborated here either.

[0081] In the sixth aspect of the embodiments of the present application, an electronic device is further provided. Refer to Figure 7 , Figure 7 which is a schematic structural diagram of an electronic device provided by the embodiments of the present application. The specific implementation of the electronic device is not limited in the specific embodiments of the present invention.

[0082] As Figure 7 shown, the electronic device may include: a processor 702, a communications interface 704, a memory 706, and a communication bus 708.

[0083] Among them:

[0084] The processor 702, the communications interface 704, and the memory 706 communicate with each other through the communication bus 708.

[0085] The communications interface 704 is used to communicate with other electronic devices or servers.

[0086] The processor 702 is used to execute the program 710, and specifically can execute the relevant steps in the above-mentioned data processing method embodiments.

[0087] Specifically, the program 710 may include program code, and the program code includes computer operation instructions.

[0088] The processor 702 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application. One or more processors included in the intelligent device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.

[0089] The memory 706 is used to store the program 710. The memory 706 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0090] The program 710 is specifically used to enable the processor 702 to perform the operations of the aforementioned data processing methods of the client or the management node.

[0091] For the specific implementation of each step in Program 710, reference may be made to the corresponding steps and descriptions in the corresponding units in the foregoing embodiments of the data processing method, which will not be elaborated herein. Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the devices and modules described above can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated herein again.

[0092] In a seventh aspect of the embodiments of the present application, a computer storage medium is further provided, on which a computer program is stored, and when the program is executed by a processor, it implements the data processing method as described in any one of Figure 2 or Figure 4 the above.

[0093] It should be noted that, according to the needs of implementation, each component / step described in the embodiments of the present application can be split into more components / steps, or two or more components / steps or partial operations of components / steps can be combined into new components / steps to achieve the purpose of the embodiments of the present application.

[0094] The method according to the embodiments of the present application can be implemented in hardware, firmware, or be implemented as software or computer code that can be stored in a recording medium (such as a CD ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or be implemented as computer code originally stored in a remote recording medium or a non-transitory machine-readable medium and downloaded through a network and to be stored in a local recording medium, so that the method described herein can be processed by such software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component (such as a RAM, a ROM, a flash memory, etc.) that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, it implements the data processing method described herein. In addition, when a general-purpose computer accesses the code for implementing the data processing method shown herein, the execution of the code converts the general-purpose computer into a dedicated computer for executing the data processing method shown herein.

[0095] Those of ordinary skill in the art can realize that the units and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.

[0096] The above embodiments are only used to illustrate the embodiments of the present application, rather than to limit the embodiments of the present application. Those of ordinary skill in the relevant technical field can also make various changes and modifications without departing from the spirit and scope of the embodiments of the present application. Therefore, all equivalent technical solutions also belong to the scope of the embodiments of the present application. The patent protection scope of the embodiments of the present application shall be defined by the claims.

Claims

1. A data processing method, applied to the client of a client - base storage system, the method comprising: After determining that the retry of writing data to the first data node fails, determining a first write length in the first data node and a second write length in the second data node, where the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; Sending the first write length and the second write length to a management node, so that the management node guides the first data node to complete the data filling according to the first write length and the second write length.

2. The method according to claim 1, wherein Before determining that the retry of writing data to the first data node fails, the method further comprises: Writing data to multiple data nodes; Determining status information of each data node for the written data, where the status information is used to characterize the write length and write result of the client writing data to multiple data nodes; Determining whether there are a first data node and a second data node among the multiple data nodes according to the status information; When there is a first data node, retrying to write the data to the first data node.

3. The method according to claim 2, wherein Retrying to write the data to the first data node includes: Retrying to write the data to the first data node for a specified number of times, or retrying to write the data to the first data node for a specified duration.

4. A data processing method, applied to the management node of a client - base storage system, the method comprising: Receiving the first write length in the first data node and the second write length in the second data node sent by the client, where the first write length and the second write length are determined by the client when writing data to the data node, the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; Determining the filling position of the first data node according to the first write length and the second write length; Guiding the first data node to obtain the data corresponding to the second write length from the second data node and perform incremental filling starting from the filling position.

5. The method according to claim 4, wherein, The method further comprises: Receiving the incremental filling result of the incremental filling returned by the first data node.

6. The method according to claim 5, wherein When the filling result indicates filling failure, the method further comprises: Selecting another data node and guiding the other data node to perform full - volume filling of the data.

7. The method according to claim 6, wherein Guiding the other data node to perform full - volume filling of the data includes: Guiding the other data node to perform full - volume filling of the data from the second data node or from the client.

8. The method according to claim 4, wherein The method further comprises: Determining whether the client is working properly; When determining that the client is not working properly, querying the lengths of the written data in each data node; Determining a first data node and a second data node according to the lengths of the written data in each data node; Guiding the first data node to complete filling according to the written data in the second data node.

9. A data processing system that adopts a client - base storage method, the data processing system includes a client, data nodes, and a management node. In the system, After determining that the retry of writing data to the first data node fails, the client determines the first write length in the first data node and the second write length in the second data node, where The first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; Send the first write length and the second write length to the management node; The management node is configured to receive the first write length in the first data node and the second write length in the second data node sent by the client, determine the filling position of the first data node according to the first write length and the second write length; guide the first data node to obtain the data corresponding to the second write length from the second data node, and start incremental filling from the filling position.

10. The system according to claim 9, wherein, In the system, The client is further configured to write data to multiple data nodes before determining that the retry of writing data to the first data node fails; determine the status information of each data node for the written data, and the status information is used to represent the write length and write result of the client writing data to multiple data nodes; Determine whether there are a first data node and a second data node among the multiple data nodes according to the status information; When there is a first data node, retry writing the data to the first data node.

11. A data processing device applied to the client of a client - base storage system, the device includes: A determination module, after determining that the retry of writing data to the first data node fails, determine the first write length in the first data node and the second write length in the second data node, where the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; A sending module, send the first write length and the second write length to the management node so that the management node guides the first data node to fill the data according to the first write length and the second write length.

12. A data processing device applied to the management node of a client - base storage system, the device includes: A receiving module, receive the first write length in the first data node and the second write length in the second data node sent by the client, where the first write length and the second write length are determined by the client when writing data to the data node, the first data node is a data node with incomplete written data, and the second data node is a data node with complete written data; A position determination module, determine the filling position of the first data node according to the first write length and the second write length; A guiding module, guide the first data node to obtain the data corresponding to the second write length from the second data node, and start incremental filling from the filling position.

13. An electronic device, comprising: A processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to the method described in any one of claims 1-8.

14. A computer storage medium having a computer program stored thereon, and when the program is executed by a processor, it implements the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Data writing method and device, computer equipment and storage medium

    CN111625601A