Data processing method, computing device, storage medium and computer program product

By identifying the target data storage node and obtaining data description information in a distributed storage system, the problem of complete data replication after file trimming is solved, enabling accurate reconstruction of missing information and ensuring the security and efficiency of data storage.

CN120929438APending Publication Date: 2025-11-11ALIBABA CLOUD COMPUTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410568543.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-08
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

In a distributed storage system, how can we ensure the security and accuracy of data storage by completely replicating a data copy with gaps after pruning the data file?

Method used

By identifying the target data storage node in the distributed storage system, obtaining data description information, and replicating the target data based on this information to generate and store a data copy, the integrity and accuracy of the void information are ensured.

Benefits of technology

It enables complete reconstruction of holes after data pruning, ensuring the integrity and security of data copies, avoiding wasted space during copy reconstruction, and reducing storage costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929438A_ABST
    Figure CN120929438A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data processing method, computing equipment, a storage medium and a computer program product, and the data processing method comprises the steps that the data processing method is applied to a distributed storage system, and the distributed storage system comprises a plurality of data storage nodes, the method comprises the steps that in response to a data reconstruction instruction for target data, a target data storage node is determined from the multiple data storage nodes, the target data is data obtained after initial data is cut and stored in a first data storage node, and the target data is stored in a second data storage node; the first data storage node is any data storage node except the target data storage node in the plurality of data storage nodes; obtaining data description information of the target data, and copying the target data according to the data description information to obtain a data copy of the target data; and storing the data copy of the target data to the target data storage node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments in this specification relate to the field of computer technology, and in particular to data processing methods, computing devices, storage media, and computer program products. Background Technology

[0002] In distributed storage systems, data can be stored across multiple independent data storage nodes and managed via network services. Typically, to ensure data security and redundancy, multiple copies of the data are stored on different data storage nodes. However, after processing the data using file trimming techniques, a segment of data may be marked as invalid and deleted, leaving a blank, non-physical space. How to replicate this blank data, creating a complete copy with this blank space, to ensure data security and accuracy is a pressing problem that needs to be solved. Summary of the Invention

[0003] In view of this, embodiments of this specification provide two data processing methods. One or more embodiments of this specification also relate to two data processing apparatuses, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.

[0004] According to a first aspect of the embodiments of this specification, a data processing method is provided, applied to a distributed storage system, the distributed storage system including multiple data storage nodes, the method comprising:

[0005] In response to a data reconstruction instruction for target data, a target data storage node is determined from the plurality of data storage nodes, wherein the target data is data obtained by pruning the initial data stored in a first data storage node, and the first data storage node is any data storage node other than the target data storage node among the plurality of data storage nodes;

[0006] Obtain the data description information of the target data, and copy the target data according to the data description information to obtain a data copy of the target data;

[0007] A data copy of the target data is stored in the target data storage node.

[0008] According to a second aspect of the embodiments of this specification, a data processing apparatus is provided, applied to a distributed storage system, the distributed storage system including multiple data storage nodes, the apparatus comprising:

[0009] The determination module is configured to determine a target data storage node from the plurality of data storage nodes in response to a data reconstruction instruction for target data, wherein the target data is data obtained by pruning the initial data stored in a first data storage node, and the first data storage node is any one of the plurality of data storage nodes other than the target data storage node;

[0010] The copying module is configured to acquire data description information of the target data and copy the target data according to the data description information to obtain a data copy of the target data;

[0011] The storage module is configured to store a data copy of the target data to the target data storage node.

[0012] According to a third aspect of the embodiments of this specification, a data processing method is provided, applied to a target data storage node, wherein the target data storage node is any one of a plurality of data storage nodes included in a distributed storage system, the method comprising:

[0013] In response to a data reconstruction instruction for target data, a first data storage node for storing the target data is determined, wherein the target data is data obtained after cropping the initial data, and the first data storage node is any one of the plurality of data storage nodes other than the target data storage node;

[0014] Obtain the target data from the first data storage node, and obtain the data description information of the target data;

[0015] Based on the data description information, the target data is copied to obtain a data copy of the target data and stored.

[0016] According to a fourth aspect of the embodiments of this specification, a data processing apparatus is provided, applied to a target data storage node, wherein the target data storage node is any one of a plurality of data storage nodes included in a distributed storage system, the apparatus comprising:

[0017] The determination module is configured to determine a first data storage node for storing the target data in response to a data reconstruction instruction for the target data, wherein the target data is data obtained after cropping the initial data, and the first data storage node is any one of the plurality of data storage nodes other than the target data storage node;

[0018] The acquisition module is configured to acquire the target data from the first data storage node and acquire the data description information of the target data;

[0019] The copying module is configured to copy the target data according to the data description information, obtain a data copy of the target data, and store it.

[0020] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising:

[0021] Memory and processor;

[0022] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above-described data processing method.

[0023] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the data processing method described above.

[0024] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the data processing method described above.

[0025] This specification provides a data processing method in one embodiment, applied to a distributed storage system, the distributed storage system including multiple data storage nodes. The method includes: responding to a data reconstruction instruction for target data, determining a target data storage node from the multiple data storage nodes, wherein the target data is data stored in a first data storage node obtained after trimming initial data, and the first data storage node is any one of the multiple data storage nodes other than the target data storage node; obtaining data description information of the target data, and copying the target data according to the data description information to obtain a data copy of the target data; and storing the data copy of the target data in the target data storage node.

[0026] In the above method, after responding to a data reconstruction instruction for the target data, the target data storage node can be determined from multiple data storage nodes in the distributed storage system. This target data storage node is then used as the node for reconstructing and storing the data copy. Furthermore, the data description information of the target data is obtained, and the target data is copied based on this information to obtain a data copy, ensuring the integrity and accuracy of the copied data copy. For the target data obtained after trimming the initial data, the information about any gaps in the target data can be determined based on the data description information. This ensures that any gaps in the target data can be completely and accurately reconstructed during the data copy process, guaranteeing the safety of both the target data and the trimmed information. Attached Figure Description

[0027] Figure 1 This is a schematic diagram illustrating an application scenario of a data processing method provided in one embodiment of this specification;

[0028] Figure 2 This is a flowchart illustrating a data processing method provided in one embodiment of this specification;

[0029] Figure 3 This is a schematic diagram of a data storage node in a data processing method provided in one embodiment of this specification;

[0030] Figure 4 This is a schematic diagram of the storage of initial data in a data processing method provided in one embodiment of this specification;

[0031] Figure 5 This is a schematic diagram of initial data and target data in a data processing method provided in one embodiment of this specification;

[0032] Figure 6 This is a schematic diagram of data description information in a data processing method provided in one embodiment of this specification;

[0033] Figure 7 This is a flowchart of the first processing procedure of a data processing method provided in one embodiment of this specification;

[0034] Figure 8 This is a flowchart of the second processing step of a data processing method provided in one embodiment of this specification;

[0035] Figure 9 This is a schematic diagram of the structure of a data processing apparatus provided in one embodiment of this specification;

[0036] Figure 10 This is a flowchart of another data processing method provided in one embodiment of this specification;

[0037] Figure 11 This is a schematic diagram of the structure of another data processing device provided in one embodiment of this specification;

[0038] Figure 12 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation

[0039] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0040] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0041] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."

[0042] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0043] This specification provides two data processing methods, and also relates to two data processing devices, a computing device, a computer-readable storage medium, and a computer program product, which are described in detail in the following embodiments.

[0044] See Figure 1 , Figure 1 A schematic diagram illustrating an application scenario of a data processing method provided according to an embodiment of this specification is shown.

[0045] Figure 1 It includes a distributed storage system 100, which includes multiple data storage nodes 102, 104, 106, 108, etc.

[0046] In practical implementation, for user-uploaded file data, data copies of that file data can be stored in data storage nodes 102, 104, and 106 respectively to achieve multi-copy redundancy. If the distributed storage system determines that communication with data storage node 106 is interrupted, resulting in the inability to access the data copies in data storage node 106, then data storage node 108 is selected as the target data storage node. A data copy is obtained from data storage node 102, and based on the data description information of the file data, this data copy is replicated to obtain a new data copy. This new data copy is then stored in data storage node 108, ensuring that the number of data copies accessible to the user meets data security requirements.

[0047] See Figure 2 , Figure 2 A flowchart of a data processing method according to an embodiment of this specification is shown, applied to a distributed storage system, the distributed storage system including multiple data storage nodes, specifically including the following steps.

[0048] Step 202: In response to the data reconstruction instruction for the target data, determine the target data storage node from the plurality of data storage nodes.

[0049] The target data is data obtained by pruning the initial data and storing it in the first data storage node. The first data storage node is any one of the plurality of data storage nodes other than the target data storage node.

[0050] The target data can be understood as the data stored in the distributed storage system, such as file data uploaded by users. This file data can be text files, video files, audio files, etc. Understandably, for file data uploaded by users, the distributed storage system can allocate multiple data storage nodes, replicating the file data into multiple copies and storing them on each data storage node to ensure data security. For example, for video files uploaded by users, the distributed storage system can replicate the video data to obtain multiple copies, storing each copy on a different data storage node, achieving multi-copy redundancy to ensure user data security. However, if communication between a data storage node containing a data copy and the distributed storage system is interrupted, making that data copy inaccessible, a data reconstruction instruction can be generated to achieve multi-copy redundancy. In other words, a data reconstruction instruction can be understood as an instruction to rebuild a data copy of the target data; data reconstruction can be understood as data copying.

[0051] In practical applications, data storage nodes can be implemented using zoned storage devices. Zoned storage devices, also known as block storage devices, divide their address space into multiple blocks with specific write constraints. Each block must be written sequentially, i.e., there is a sequential write restriction. This design simplifies hard drive architecture and reduces the need for complex data management. The advantages of zoned storage devices include increased storage capacity, reduced costs, improved service quality, and increased device durability. It is suitable for various storage needs, such as cold data storage, archive storage, and big data storage and analysis scenarios. Figure 3 As shown, Figure 3 A schematic diagram of a data storage node is shown in a data processing method according to an embodiment of this specification. For example... Figure 3 As shown, zoned storage devices are divided into multiple fixed-size, contiguous zone spaces. Each zone space supports strict sequential writing, meaning that target data is written to the zone space in order, facilitating subsequent searching and copying. Furthermore, it enables deletion at the zone space level. In other words, target data can be written sequentially to zone space 1, zone space 2, and so on, and the data stored in a particular zone space can be deleted.

[0052] Trimming, or file pruning, is a technique used in distributed storage systems to free up storage space. Typically, file data, as the smallest unit of resource management, requires complete deletion to free up storage space. However, file pruning allows you to delete a specific segment of data within a file without deleting the entire file, thus freeing up storage space. For example, in a 512KB file, pruning can delete segments from 1KB to 128KB, retaining the rest. This pruning process can be performed on zoned storage devices. After pruning, gaps may appear in the file data; these gaps are the pruned data.

[0053] In specific implementation, before determining the target data storage node from the plurality of data storage nodes in response to the data reconstruction instruction for the target data, the method further includes:

[0054] In response to a data upload request, the system receives the initial data carried in the data upload request.

[0055] In response to a cropping request for the initial data, the initial data is cropped to obtain the target data.

[0056] In this context, initial data can be understood as the file data before trimming, and target data can be understood as the file data after trimming. A data upload request can be understood as a data upload request sent by a user to the distributed storage system through a client. Initial data can be understood as the file data that the user wants to upload to the distributed storage system, such as text files, audio files, and video files. A trimming request can be understood as a data processing request sent by the user to the distributed storage system through a client. This trimming request can be used to trim the initial data, marking a segment of data in the initial data as invalid data. The data storage node storing this initial data can check the initial data at preset time intervals, and delete the data content marked as invalid data based on the detection results, thereby obtaining the target data containing the trimmed data (i.e., the gaps).

[0057] Based on this, in response to a data upload request sent by a user through a client, the system can receive the initial data carried in the data upload request, and receive a trimming request sent by the user through a client for the initial data. The system can then trim the initial data, mark a segment of the initial data as invalid data, and obtain the trimmed target data.

[0058] In addition, after receiving the initial data carried in the data upload request, the initial data can be copied to obtain multiple data copies of the initial data, and these multiple data copies can be stored in different data storage nodes.

[0059] In practical applications, see Figure 4 , Figure 4 A schematic diagram illustrating the storage of initial data in a data processing method according to an embodiment of this specification is shown. Figure 4 As shown, the initial data can be file data, which can be divided into data blocks 402, 404, 406, and 408. Taking data block 404 as an example, data block 404 can be copied to obtain data copies 4042, 4044, and 4046. Data copy 4042 is stored on data storage node 1, data copy 4044 is stored on data storage node 2, and data copy 4046 is stored on data storage node 3. Understandably, a similar distributed storage operation is performed on the target data obtained after trimming.

[0060] See Figure 5 , Figure 5 A schematic diagram of initial data and target data is shown in a data processing method according to an embodiment of this specification. Figure 5As shown, the initial data 502 is trimmed using file trimming technology, resulting in target data 504 that includes trimmed data (i.e., holes). Further, the target data 504 can be divided into multiple data blocks. A data copy is created for each data block and stored separately to ensure efficient file storage. For the data block containing trimmed data, it can be copied to obtain data copies 5042, 5044, and 5046, which are then stored in data storage nodes 1, 2, and 3, respectively.

[0061] In other words, in one embodiment of this specification, the target data can be further understood as data blocks obtained by dividing file data.

[0062] In summary, by setting up multiple data replicas for separate storage, multi-replica redundancy is achieved, further ensuring the security of data storage.

[0063] Furthermore, after obtaining the target data, and before determining the target data storage node from the plurality of data storage nodes in response to the data reconstruction instruction for the target data, the method further includes:

[0064] The target data is copied to obtain multiple data copies;

[0065] Each of the multiple data replicas is stored in one of the multiple data storage nodes.

[0066] Specifically, after obtaining the target data but before responding to a data reconstruction instruction for it, the target data can be copied to obtain multiple data copies, which are then stored on multiple data storage nodes. This process is similar to the initial data storage process described above and will not be repeated here.

[0067] In practice, the distributed storage system is communicatively connected to each data storage node;

[0068] Prior to the data reconstruction instruction for the target data, the method further includes:

[0069] If communication with the second data storage node is interrupted, a data reconstruction instruction for the target data is generated.

[0070] The second data storage node is any one of the plurality of data storage nodes that stores a data copy of the target data.

[0071] Specifically, if communication with a second data storage node that stores a copy of the target data is interrupted, it indicates that a copy has been lost and the data copy stored in the second data storage node is inaccessible. In order to ensure that the number of data copies accessible to users meets the data security requirements, a data reconstruction instruction for the target data can be generated. Based on the data reconstruction instruction, a new data storage node (i.e., the target data storage node) is reallocated to rebuild a new data copy.

[0072] In one embodiment of this specification, a data reconstruction instruction for the target data can be generated if it is determined that the communication interruption with the second data storage node exceeds a preset time period.

[0073] In summary, by performing communication checks on data storage nodes that store data copies, the number of data copies accessible to users can be guaranteed to meet data security requirements.

[0074] In specific implementation, after cropping the initial data to obtain the target data, the process further includes:

[0075] Based on the initial data and the target data, generate data description information for the target data;

[0076] The data description information is written into the indexing system so that the data description information can be retrieved through the indexing system in response to the copy rebuild instruction.

[0077] Data description information can be understood as information describing data attributes. It can include data status description, data location description, and data length description. An indexing system can be understood as the in-memory data structure of a data storage node, which can be used to record data description information of file data stored in the data storage node.

[0078] In practical applications, data status description information can include a norm field and a delete field. The norm field indicates that the data is valid, while the delete field indicates that the data is clipped (i.e., has holes). Data location description information can be offset information, which indicates the starting position of the data within the target data. Data length description information can indicate the valid length of the data.

[0079] See Figure 6 , Figure 6 A schematic diagram of data description information is shown in a data processing method according to an embodiment of this specification. For example... Figure 6As shown, for the first segment of the target data, the starting position of the first segment is 0, the length is 1M, and it is valid data. For the second segment of the target data, the starting position of the second segment is 1M, the length is 1M, and it is clipped data (i.e., holes). For the third segment of the target data, the starting position of the third segment is 2M, the length is 1M, and it is valid data.

[0080] Based on this, data description information of the target data can be generated by comparing the initial data and the target data, or by the pruning process from the initial data to the target data. This data description information can then be written into the index system to ensure that the data description information can be obtained through the index system when a replica reconstruction task is subsequently performed.

[0081] In summary, by defining the data description information, the target data is interpreted, and the truncated information is fully preserved, ensuring the correctness of subsequent data copying.

[0082] Furthermore, the data description information can be defined using the pb format, enabling it to have compatible information expansion capabilities. Moreover, the data description information can also include the data pruning time and the original data's checksum. During data reconstruction, the accuracy of data copy replication is further enhanced by verifying the checksum and determining the pruning time.

[0083] Step 204: Obtain the data description information of the target data, and copy the target data according to the data description information to obtain a data copy of the target data.

[0084] Specifically, data description information of the target data can be obtained from the indexing system, and the target data can be copied based on the data description information to obtain a data copy of the target data.

[0085] In specific implementation, the step of copying the target data according to the data description information to obtain a data copy of the target data includes:

[0086] If, based on the data description information, it is determined that the target data includes cropped data, then the cropping information of the cropped data is determined based on the data description information.

[0087] Identify the other data in the target data besides the cropped data;

[0088] Based on the cropping information and the other data, a data copy of the target data is determined.

[0089] In this context, the target data, excluding the cropped data, can be understood as valid data that has not been cropped. Cropped data can be understood as invalid data that has been deleted due to cropping of the target data, i.e., gaps. Cropping information can be understood as information about the gaps in the target data, such as information describing which segment of the target data was cropped.

[0090] Based on this, if the target data includes cropped data according to the data description information, the cropping information of the cropped data can be determined according to the data description information, and the valid data in the target data other than the cropped data can be determined. Based on the cropping information and the valid data, a data copy of the target data can be determined.

[0091] In practice, the presence or absence of clipped data in the target data can be determined based on the data status description information. For example, if the data status description information includes the "deleted" field, it indicates that the target data contains clipped data; if the data status description information does not include the "deleted" field, it indicates that the target data does not contain clipped data.

[0092] In practical applications, target data may include multiple data segments, each with its own data description information. If, based on this description, a segment within the target data is determined to be truncated, then the target data contains truncated data. In this case, the truncated information for that segment can be determined based on its description and written to the indexing system. The remaining data segments from the target data, excluding the truncated segment, are then written to a new data copy. In this new copy, the position corresponding to the truncated segment is left blank, treated as a void, thus obtaining a copy of the target data. Specifically, the process described above can be used to sequentially determine whether each data segment is truncated, thereby completing the copying of the target data.

[0093] In summary, by preserving the pruning information (i.e., the hole information), the correctness and integrity of data copy replication are ensured. Furthermore, there is no need to fill the holes, avoiding the waste of space caused by filling holes during copy reconstruction and reducing storage costs.

[0094] Further, determining the cropping information of the cropped data based on the data description information includes:

[0095] The data status of the cropped data is determined based on the data status description information in the data description information;

[0096] The data location of the cropped data is determined based on the data location description information in the data description information;

[0097] The data length of the cropped data is determined based on the data length description information in the data description information;

[0098] The cropping information of the cropped data is determined based on the data status, the data position, and the data length;

[0099] After determining the cropping information of the cropping data, the method further includes:

[0100] The cropping information is written into the indexing system.

[0101] Specifically, for each data segment in the target data, the data state of the data segment can be determined as truncated data based on the data state description information, the data position of the data segment in the target data can be determined based on the data position description information, and the data length of the data segment in the target data can be determined based on the data length description information. Thus, the data state, data position, and data length are used as the truncating information for the data segment.

[0102] Furthermore, the step of copying the target data according to the data description information to obtain a data copy of the target data further includes:

[0103] If, based on the data description information, it is determined that the target data does not include cropped data, the target data is copied to obtain a data copy of the target data.

[0104] Specifically, if the target data does not contain vacant data based on the data description information, the target data can be directly copied to obtain a data copy of the target data.

[0105] In specific implementation, the step of copying the target data according to the data description information to obtain a data copy of the target data includes:

[0106] A second data storage node is determined from the plurality of data storage nodes, wherein the second data storage node is any one of the plurality of data storage nodes that stores a data copy of the target data;

[0107] Obtain a data copy of the target data from the second data storage node;

[0108] Based on the data description information, a data copy of the target data is copied to obtain a data copy of the target data.

[0109] Specifically, during the replica reconstruction task, a data copy of the target data stored in the second data storage node can be made. The data description information at this point can then be obtained from the indexing system corresponding to that second data storage node.

[0110] For example, second data storage nodes 1, 2, and 3 each store a copy of the target data. In the event of a communication interruption at second data storage node 1, the distributed storage system can allocate a target data storage node 4. Target data storage node 4 can obtain a copy of the target data from second data storage nodes 2 or 3, and correspondingly retrieve the data description information of the target data from the index system corresponding to second data storage nodes 2 or 3. Based on this data description information, it copies the target data copy to obtain a new data copy, which is then stored in target data storage node 4.

[0111] Step 206: Store a data copy of the target data to the target data storage node.

[0112] Specifically, after obtaining a copy of the target data, this copy can be stored on the target data storage node to ensure multi-copy redundancy.

[0113] In summary, the above method, upon responding to a data reconstruction instruction for the target data, can determine the target data storage node from multiple data storage nodes in the distributed storage system. This target data storage node is then used as the node for reconstructing and storing the data copy. Furthermore, the data description information of the target data is obtained, and the target data is copied based on this information to obtain a data copy, ensuring the integrity and accuracy of the copied data copy. For the target data obtained after pruning the initial data, the information about any gaps in the target data can be determined based on the data description information. This ensures that any gaps in the target data can be reconstructed completely and accurately during the data copying process, guaranteeing the safety of both the target data and the pruning information.

[0114] The following is in conjunction with the appendix Figure 7 and attached Figure 8 Taking the application of the data processing method provided in this specification in a distributed storage system as an example, the data processing method will be further explained. Figure 7 A flowchart of the first processing procedure of a data processing method provided in one embodiment of this specification is shown, which specifically includes the following steps.

[0115] Step 702: The distributed storage system initiates a replica reconstruction task.

[0116] Specifically, a distributed storage system includes multiple data storage nodes. For user-uploaded file data, this data can be divided into multiple data blocks. Taking one data block as an example, this data block is replicated to obtain three copies, which are stored in different data storage nodes. For instance, the data copies of the data block can be stored in data storage node 1, data storage node 2, and data storage node 3, respectively. If the distributed storage system determines that communication with data storage node 1 is interrupted, the data copy stored in data storage node 1 cannot be accessed, and the data copy is lost. To ensure redundancy, the distributed storage system can initiate a copy reconstruction task for this data block (i.e., the aforementioned copy reconstruction instruction).

[0117] Step 704: The distributed storage system allocates the target data storage node.

[0118] Specifically, a distributed storage system can allocate a data storage node that does not store a copy of the data block from multiple data storage nodes as the target data storage node, and send the copy reconstruction task to the target data storage node.

[0119] Step 706: The target data storage node receives and executes the replica reconstruction task.

[0120] Specifically, the process described in steps 802 to 810 below is the process of the target data storage node performing a replica reconstruction task.

[0121] Figure 8 A flowchart of a second processing procedure of a data processing method provided in one embodiment of this specification is shown, which specifically includes the following steps.

[0122] Step 802: Obtain a data copy of the data block from the source data storage node.

[0123] In this context, the source data storage node is a data storage node that stores a copy of the data block. Continuing with the previous example, the source data storage node can be data storage node 2 or data storage node 3.

[0124] Step 804: Determine whether there are holes in the data block based on the data description information of the data block. If yes, proceed to step 806; otherwise, proceed to step 808.

[0125] If the source data storage node is data storage node 2, then the data description information of the data block can be obtained from the index system corresponding to data storage node 2.

[0126] Specifically, a data block can include multiple data segments, such as data segment 1, data segment 2, and data segment 3. Based on the data description information of data segment 1, it can be determined whether data segment 1 is a hole. If yes, proceed to step 806; otherwise, proceed to step 808. Continue determining whether data segment 2 is a hole. If yes, proceed to step 806; otherwise, proceed to step 808. Continue determining whether data segment 3 is a hole. If yes, proceed to step 806; otherwise, proceed to step 808.

[0127] Step 806: Record the hole information into the indexing system.

[0128] Specifically, taking data segment 1 as an example, if data segment 1 is a hole, the hole information is recorded in the index system corresponding to the target data storage node, so as to generate data description information of the data copy stored in the target data storage node based on the hole information.

[0129] Step 808: Directly write the data block to the new data copy.

[0130] Specifically, taking data segment 1 as an example, if data segment 1 is not empty, then data segment 1 is directly written into a new data copy.

[0131] Step 810: Determine whether the new data copy has been rebuilt. If yes, end the process; otherwise, continue to step 802.

[0132] Following the previous example, after data segment 1, data segment 2, and data segment 3 have all undergone the processing steps 804 to 808, it indicates that the new data copy has been reconstructed.

[0133] In summary, the above method, upon responding to a data reconstruction instruction for the target data, can determine the target data storage node from multiple data storage nodes in the distributed storage system. This target data storage node is then used as the node for reconstructing and storing the data copy. Furthermore, the data description information of the target data is obtained, and the target data is copied based on this information to obtain a data copy, ensuring the integrity and accuracy of the copied data copy. For the target data obtained after pruning the initial data, the information about any gaps in the target data can be determined based on the data description information. This ensures that any gaps in the target data can be reconstructed completely and accurately during the data copying process, guaranteeing the safety of both the target data and the pruning information.

[0134] Corresponding to the above method embodiments, this specification also provides embodiments of a data processing apparatus applied to a distributed storage system, wherein the distributed storage system includes multiple data storage nodes. Figure 9 A schematic diagram of the structure of a data processing apparatus according to one embodiment of this specification is shown. Figure 9 As shown, the device includes:

[0135] The determining module 902 is configured to determine a target data storage node from the plurality of data storage nodes in response to a data reconstruction instruction for target data, wherein the target data is data obtained by pruning the initial data stored in a first data storage node, and the first data storage node is any data storage node other than the target data storage node among the plurality of data storage nodes;

[0136] The copying module 904 is configured to acquire data description information of the target data and copy the target data according to the data description information to obtain a data copy of the target data;

[0137] Storage module 906 is configured to store a data copy of the target data to the target data storage node.

[0138] In an optional embodiment, the copying module 904 is further configured to:

[0139] If, based on the data description information, it is determined that the target data includes cropped data, then the cropping information of the cropped data is determined based on the data description information.

[0140] Identify the other data in the target data besides the cropped data;

[0141] Based on the cropping information and the other data, a data copy of the target data is determined.

[0142] In an optional embodiment, the copying module 904 is further configured to:

[0143] The data status of the cropped data is determined based on the data status description information in the data description information;

[0144] The data location of the cropped data is determined based on the data location description information in the data description information;

[0145] The data length of the cropped data is determined based on the data length description information in the data description information;

[0146] The cropping information of the cropped data is determined based on the data status, the data position, and the data length;

[0147] After determining the cropping information of the cropping data, the method further includes:

[0148] The cropping information is written into the indexing system.

[0149] In an optional embodiment, the copying module 904 is further configured to:

[0150] If, based on the data description information, it is determined that the target data does not include cropped data, the target data is copied to obtain a data copy of the target data.

[0151] In an optional embodiment, the device further includes a trimming module configured to:

[0152] In response to a data upload request, the system receives the initial data carried in the data upload request.

[0153] In response to a cropping request for the initial data, the initial data is cropped to obtain the target data.

[0154] In an optional embodiment, the copying module 904 is further configured to:

[0155] The target data is copied to obtain multiple data copies;

[0156] Each of the multiple data replicas is stored in one of the multiple data storage nodes.

[0157] In an optional embodiment, the copying module 904 is further configured to:

[0158] A second data storage node is determined from the plurality of data storage nodes, wherein the second data storage node is any one of the plurality of data storage nodes that stores a data copy of the target data;

[0159] Obtain a data copy of the target data from the second data storage node;

[0160] Based on the data description information, a data copy of the target data is copied to obtain a data copy of the target data.

[0161] In an optional embodiment, the apparatus further includes a generation module configured to:

[0162] Based on the initial data and the target data, generate data description information for the target data;

[0163] The data description information is written into the indexing system so that the data description information can be retrieved through the indexing system in response to the copy rebuild instruction.

[0164] In one optional embodiment, the distributed storage system is communicatively connected to each data storage node;

[0165] The generation module is further configured as follows:

[0166] If communication with the second data storage node is interrupted, a data reconstruction instruction for the target data is generated.

[0167] In summary, in the above-described apparatus, upon responding to a data reconstruction command for the target data, the target data storage node can be determined from multiple data storage nodes in the distributed storage system. This target data storage node is then used as the node for reconstructing and storing the data copy. Furthermore, the data description information of the target data is obtained, and the target data is copied based on this information to obtain a data copy, ensuring the integrity and accuracy of the copied data copy. For the target data obtained after trimming the initial data, the information about any gaps in the target data can be determined based on the data description information. This ensures that any gaps in the target data can be completely and accurately reconstructed during the data copy process, guaranteeing the security of both the target data and the trimmed information.

[0168] The above is an illustrative scheme of a data processing apparatus according to this embodiment. It should be noted that the technical solution of this data processing apparatus and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing apparatus, please refer to the description of the technical solution of the data processing method described above.

[0169] See Figure 10 , Figure 10 A flowchart of another data processing method according to an embodiment of this specification is shown, which is applied to a target data storage node, wherein the target data storage node is any one of a plurality of data storage nodes included in a distributed storage system, and the specific steps are as follows.

[0170] Step 1002: In response to the data reconstruction instruction for the target data, determine the first data storage node storing the target data.

[0171] Wherein, the target data is the data obtained after cropping the initial data, and the first data storage node is any one of the multiple data storage nodes except the target data storage node;

[0172] Step 1004: Obtain the target data from the first data storage node, and obtain the data description information of the target data;

[0173] Step 1006: Based on the data description information, copy the target data to obtain a data copy of the target data and store it.

[0174] In summary, in the above method, after responding to a data reconstruction instruction for the target data, the target data storage node can determine the first data storage node storing the target data from multiple data storage nodes in the distributed storage system, obtain the target data from the first data storage node, and acquire the data description information of the target data. Based on the data description information, the target data is copied to obtain a data copy, ensuring the integrity and accuracy of the copied data copy. For the target data obtained after trimming the initial data, the information on any gaps in the target data can be determined based on the data description information, thereby ensuring that any gaps in the target data can be completely and accurately reconstructed when copying the data copy, thus guaranteeing the safety of the target data and the trimming information.

[0175] Corresponding to the above method embodiments, this specification also provides a data processing apparatus embodiment, applied to a target data storage node, wherein the target data storage node is any one of a plurality of data storage nodes included in a distributed storage system. Figure 11 A schematic diagram of another data processing apparatus provided in one embodiment of this specification is shown. Figure 11 As shown, the device includes:

[0176] The determining module 1102 is configured to determine a first data storage node storing the target data in response to a data reconstruction instruction for the target data, wherein the target data is data obtained after cropping the initial data, and the first data storage node is any one of the plurality of data storage nodes other than the target data storage node;

[0177] The acquisition module 1104 is configured to acquire the target data from the first data storage node and acquire the data description information of the target data;

[0178] The copying module 1106 is configured to copy the target data according to the data description information, obtain a data copy of the target data, and store it.

[0179] In summary, in the aforementioned apparatus, after responding to a data reconstruction command for the target data, the target data storage node can determine the first data storage node storing the target data from multiple data storage nodes in the distributed storage system, obtain the target data from the first data storage node, acquire the data description information of the target data, and copy the target data according to the data description information to obtain a data copy of the target data, ensuring the integrity and accuracy of the copied data copy. For the target data obtained after trimming the initial data, the information on any gaps in the target data can be determined based on the data description information, thereby ensuring that any gaps in the target data can be completely and accurately reconstructed when copying the data copy, thus guaranteeing the security of the target data and the trimming information.

[0180] The above is an illustrative scheme of another data processing device according to this embodiment. It should be noted that the technical solution of this data processing device and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the data processing device, please refer to the description of the technical solution of the data processing method described above.

[0181] Figure 12 A structural block diagram of a computing device 1200 according to an embodiment of this specification is shown. The components of the computing device 1200 include, but are not limited to, a memory 1210 and a processor 1220. The processor 1220 is connected to the memory 1210 via a bus 1230, and a database 1250 is used to store data.

[0182] The computing device 1200 also includes an access device 1240, which enables the computing device 1200 to communicate via one or more networks 1260. Examples of such networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. Access device 1240 may include one or more of any type of wired or wireless network interface (e.g., network interface card (NIC)), such as IEEE 802.11 Wireless Local Area Network (WLAN) interface, Wi-MAX (Worldwide Interoperability for Microwave Access) interface, Ethernet interface, Universal Serial Bus (USB) interface, cellular network interface, Bluetooth interface, Near Field Communication (NFC) interface, and so on.

[0183] In one embodiment of this application, the aforementioned components of the computing device 1200 and Figure 12 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 12 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0184] The computing device 1200 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1200 can also be a mobile or stationary server.

[0185] The processor 1220 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above-described data processing method.

[0186] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.

[0187] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.

[0188] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the data processing method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the data processing method embodiments.

[0189] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described data processing method.

[0190] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the data processing method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the data processing method described above.

[0191] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0192] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.

[0193] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.

[0194] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0195] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. A data processing method applied to a distributed storage system, the distributed storage system comprising multiple data storage nodes, the method comprising: In response to a data reconstruction instruction for target data, a target data storage node is determined from the plurality of data storage nodes, wherein the target data is data obtained by pruning the initial data stored in a first data storage node, and the first data storage node is any data storage node other than the target data storage node among the plurality of data storage nodes; Obtain the data description information of the target data, and copy the target data according to the data description information to obtain a data copy of the target data; A data copy of the target data is stored in the target data storage node.

2. The data processing method according to claim 1, wherein copying the target data according to the data description information to obtain a data copy of the target data includes: If, based on the data description information, it is determined that the target data includes cropped data, then the cropping information of the cropped data is determined based on the data description information. Identify the other data in the target data besides the cropped data; Based on the cropping information and the other data, a data copy of the target data is determined.

3. The data processing method according to claim 2, wherein determining the cropping information of the cropped data based on the data description information includes: The data status of the cropped data is determined based on the data status description information in the data description information; The data location of the cropped data is determined based on the data location description information in the data description information; The data length of the cropped data is determined based on the data length description information in the data description information; The cropping information of the cropped data is determined based on the data status, the data position, and the data length; After determining the cropping information of the cropping data, the method further includes: The cropping information is written into the indexing system.

4. The data processing method according to claim 2, wherein copying the target data according to the data description information to obtain a data copy of the target data further includes: If, based on the data description information, it is determined that the target data does not include the cropped data, the target data is copied to obtain a data copy of the target data.

5. The data processing method according to claim 1, further comprising, before determining the target data storage node from the plurality of data storage nodes in response to a data reconstruction instruction for the target data: In response to a data upload request, the system receives the initial data carried in the data upload request. In response to a cropping request for the initial data, the initial data is cropped to obtain the target data.

6. The data processing method according to claim 5, further comprising, after obtaining the target data and before determining the target data storage node from the plurality of data storage nodes in response to a data reconstruction instruction for the target data: The target data is copied to obtain multiple data copies; Each of the multiple data replicas is stored in one of the multiple data storage nodes.

7. The data processing method according to claim 6, wherein copying the target data according to the data description information to obtain a data copy of the target data comprises: A second data storage node is determined from the plurality of data storage nodes, wherein the second data storage node is any one of the plurality of data storage nodes that stores a data copy of the target data; Obtain a data copy of the target data from the second data storage node; Based on the data description information, a data copy of the target data is copied to obtain a data copy of the target data.

8. The data processing method according to claim 7, wherein the distributed storage system is communicatively connected to each data storage node; Prior to the data reconstruction instruction for the target data, the method further includes: If communication with the second data storage node is interrupted, a data reconstruction instruction for the target data is generated.

9. The data processing method according to claim 5, after cropping the initial data to obtain the target data, further comprising: Based on the initial data and the target data, generate data description information for the target data; The data description information is written into the indexing system so that the data description information can be retrieved through the indexing system in response to the copy rebuild instruction.

10. A data processing method applied to a target data storage node, wherein the target data storage node is any one of a plurality of data storage nodes included in a distributed storage system, the method comprising: In response to a data reconstruction instruction for target data, a first data storage node for storing the target data is determined, wherein the target data is data obtained after cropping the initial data, and the first data storage node is any one of the plurality of data storage nodes other than the target data storage node; Obtain the target data from the first data storage node, and obtain the data description information of the target data; Based on the data description information, the target data is copied to obtain a data copy of the target data and stored.

11. A computing device, comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the data processing method according to any one of claims 1 to 10.

12. A computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the data processing method according to any one of claims 1 to 10.

13. A computer program product comprising a computer program / instructions that, when executed by a processor, implement the steps of the data processing method according to any one of claims 1 to 10.