Data processing method and device, computer equipment and cluster
By processing the data to be stored in parallel and adjusting the parallelism degree, the problem of high calculation time and low redeletion rate caused by inappropriate sliding window length is solved, and more efficient point-cutting acquisition and data redeletion are achieved.
Patent Information
- Application Number
- CN202311670969.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-06
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2043-12-06
AI Technical Summary
Due to the inappropriate length of the selected sliding window, the calculation time is high and the deletion rate is low.
By using multiple sliding windows to process the data to be stored in parallel, a first set of tangents including at least one tangent is obtained, and the parallelism is adjusted according to the calculation time and cache occupancy of the tangents to optimize the number and length of the sliding windows.
It reduces the calculation time required by computing devices to obtain tangent points, improves the accuracy of tangent points, and enhances the data deletion rate.
Smart Images

Figure CN120104040A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of storage technology, and in particular to a data processing method, device, computer equipment and cluster. Background Art
[0002] Data deduplication is a data reduction technology that divides a data stream into multiple data blocks and deletes data blocks in the data stream that are identical to the stored data to reduce the storage capacity occupied by redundant data in the storage device. The data stream is divided according to a variable-length block algorithm. For example, a single sliding window is used to traverse the data stream to obtain multiple tangent points, and multiple data blocks are obtained based on the multiple tangent points. However, if the length of the sliding window is too long, the tangent point cannot be accurately obtained, resulting in a low deduplication rate; if the length of the sliding window is too short, the calculation of the tangent point is time-consuming. Summary of the invention
[0003] The present application provides a data processing method, apparatus, computer equipment and cluster, which solve the problems of high computation time and low deduplication rate caused by inappropriately selected sliding window length.
[0004] In a first aspect, the present application provides a data processing method. The method includes: a computing device obtains a first tangent point set of a first segment of data in the data to be stored according to a first degree of parallelism, divides the first segment of data according to at least one tangent point in the first tangent point set, and obtains a first data set. The first degree of parallelism is used to indicate M sliding windows. The computing device adjusts the first degree of parallelism according to at least one of the calculation time of the first tangent point set and the cache occupancy rate to obtain a second degree of parallelism. The first cache occupancy rate is used to indicate the cache occupancy rate of performing data deduplication on the first segment of data according to the first tangent point set. The second degree of parallelism is used to indicate N sliding windows. M is not equal to N, and M and N are integers greater than or equal to 2. The computing device obtains a second tangent point set of a second segment of data in the data to be stored according to the second degree of parallelism, divides the second segment of data according to at least one tangent point in the second tangent point set, and obtains a second data set. And the computing device performs data deduplication on the first data set and the second data set. The first data set includes at least one data block, and the second data set includes at least one data block.
[0005] Compared with the computing device using a single sliding window to determine the tangent point of the data to be stored. In the present application, the computing device uses multiple sliding windows to process the data to be stored in parallel to obtain a first tangent point set including at least one tangent point. Multiple sliding windows can obtain at least one tangent point by processing the data to be stored at a time, which reduces the calculation time required for the computing device to obtain the tangent point. If the number of sliding windows is too large, multiple sliding windows slide in parallel to find the tangent point of the data to be stored, the storage space of the cache may not be enough, and the cache occupancy rate is high, or multiple sliding windows slide in parallel to determine multiple tangent points, and the calculation time of the tangent point may also be long. Then the computing device adjusts the parallelism of processing the data to be stored according to at least one of the calculation time and cache occupancy rate of the first tangent point set to adjust the number of sliding windows, so as to minimize the calculation time of the tangent point and reduce the cache occupancy, so as to find the reasonable tangent point of the data to be stored as soon as possible, adjust the length of the data block selected by the multiple sliding windows, improve the accuracy of the computing device in obtaining the tangent point, improve the acquisition of the data to be stored that is repeated with the stored data, and improve the data deduplication rate.
[0006] In a possible implementation, M sliding windows slide at least once on the data to be stored, and the computing device obtains a first tangent point set of the first segment of data.
[0007] In this way, the computing device uses M sliding windows to slide on the data to be stored multiple times, and then adjusts the parallelism of processing the data to be stored, thereby avoiding the computing and storage resources consumed by frequently adjusting the parallelism.
[0008] In another possible implementation, the computing device divides the first segment of data according to M sliding windows to obtain a first initial data set, and the computing device obtains a first tangent point set according to a tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution feature.
[0009] In another possible implementation, the computing device obtains the first tangent point set according to a tangent point determination rule associated with a length interval determined by a length of a data block in the first initial data set and a maximum value, a minimum value, and an expected value of the length.
[0010] In this way, the computing device uses different cut-point determination rules to determine the cut-points of data blocks of different lengths, thereby improving the accuracy of the computing device in acquiring the cut-points.
[0011] In another possible implementation, the first segment of data is divided by M sliding windows with an interval of at least 1 byte, and the computing device obtains a first initial cut-point set and a first initial data set.
[0012] In this way, the M sliding windows divide the first segment of data at intervals of at least 1 byte, so that each sliding window selects a different position of the first segment of data, thereby improving the efficiency of the sliding window traversing the first segment of data.
[0013] In another possible implementation, the lengths of the M sliding windows are consistent.
[0014] In this way, the lengths of the M sliding windows are consistent, so that the lengths of the data blocks selected by each sliding window frame are the same, thereby improving the accuracy of the tangent point of the computing device.
[0015] In another possible implementation, the computation time of the first tangent point set includes at least one of a first duration, a second duration, and a sum of the first duration and the second duration. The first duration is used to indicate the duration required for M sliding windows to divide the first segment of data into the first initial data set. The second duration is used to indicate the duration required for the computing device to obtain the first tangent point set based on a tangent point determination rule associated with the length of a data block in the first initial data set and a length distribution feature.
[0016] In this way, the computing time is obtained according to the time required for the computing device to obtain the first initial data set and the time required for the first tangent point set. In the case where the time required for the computing device to obtain the first initial data set is longer than the time required to obtain the first tangent point set, the computing device reduces the time required to obtain the first initial data set. And in the case where the time required for the computing device to obtain the first initial data set is shorter than the time required to obtain the first tangent point set, the computing device increases the time required to obtain the first initial data set. Thus, the time required to obtain the first initial data set and the time required for the first tangent point set are balanced.
[0017] In another possible implementation, when at least one of the computation time and the cache occupancy rate of the first tangent point set is greater than the respective upper limit values, the computing device reduces the M sliding windows to N sliding windows.
[0018] In this way, when the computation time of the first tangent point set is greater than the corresponding upper limit value and / or the cache occupancy rate is greater than the corresponding upper limit value, the resources that the computing device can provide cannot support the computing device to use M sliding windows to process the data to be stored. The computing device reduces the number of sliding windows so that the resources that it can provide match the resources required for the sliding windows to process the data to be stored.
[0019] In another possible implementation, when either the computation time and the cache occupancy rate of the first tangent point set is less than the respective lower limit value, the computing device increases the M sliding windows to N sliding windows.
[0020] In this way, when the computation time of the first tangent point set is less than the corresponding upper limit value and / or the cache occupancy rate is less than the corresponding upper limit value, the computing device cannot effectively utilize the resources that the computing device can provide by using M sliding windows to process the data to be stored. The computing device increases the number of sliding windows so that the resources that it can provide match the resources required for the sliding windows to process the data to be stored.
[0021] In another possible implementation, the computing device deletes a data block that matches the stored data from at least one data block included in the first data set, and stores a data block that does not match the stored data from at least one data block included in the first data set. And the computing device deletes a data block that matches the stored data from at least one data block included in the second data set, and stores a data block that does not match the stored data from at least one data block included in the second data set.
[0022] In this way, the computing device deletes the matching data blocks and stores the non-matching data blocks, thereby reducing the storage capacity required by the computing device to store the data to be stored and improving the deduplication rate.
[0023] In another possible implementation, the stored data includes a second data block and a third data block. At least one data block included in the first data set matches the second data block. The computing device divides the third segment of the data to be stored using the length of the third data block to obtain a fourth data block.
[0024] In this way, the computing device avoids using a sliding window to select and divide the data to be stored byte by byte to obtain data blocks, which reduces the time required to obtain the cut point and improves the efficiency of data deduplication.
[0025] In another possible implementation, the first data set includes a fifth data block, a sixth data block, and a seventh data block. The length and maximum length of the fifth data block and the sixth data block are the same. The third segment of the data to be stored is divided by the first tangent point to obtain the fifth data block. The fourth segment of the data to be stored is divided by the second tangent point to obtain the seventh data block. The computing device generates a tangent point feature table for the fifth data block and the sixth data block. The tangent point feature table includes: a tangent point feature and a tangent point distance. The tangent point feature is used to indicate: the fifth data block and the sixth data block. The tangent point distance is used to indicate the spacing between the second tangent point and the first tangent point.
[0026] In this way, the computing device generates a cut-point feature table for multiple cut-points, saving storage resources required by the computing device to store the cut-point feature table.
[0027] In a second aspect, the present application provides a data processing device. The device includes a cut-point determination module, an adjustment module, and a processing module. The cut-point determination module is used to: obtain a first cut-point set of a first segment of data in the data to be stored according to a first parallelism, and divide the first segment of data according to at least one cut-point in the first cut-point set to obtain a first data set. The first parallelism is used to indicate M sliding windows. The adjustment module is used to: adjust the first parallelism according to at least one of the calculation time consumption and the cache occupancy of the first cut-point set to obtain a second parallelism. The first cache occupancy is used to indicate the cache occupancy of performing data deduplication on the first segment of data according to the first cut-point set. The second parallelism is used to indicate N sliding windows, M is not equal to N, and M and N are integers greater than or equal to 2. The cut-point determination module is also used to: obtain a second cut-point set of a second segment of data in the data to be stored according to the second parallelism, and divide the second segment of data according to at least one cut-point in the second cut-point set to obtain a second data set. The processing module is also used to: perform data deduplication on the first data set and the second data set. The first data set includes at least one data block, and the second data set includes at least one data block.
[0028] In a possible scenario, the tangent point determination module is specifically configured to: slide M sliding windows on the data to be stored at least once to obtain a first tangent point set for the first segment of data.
[0029] In another possible scenario, the tangent point determination module is specifically used to divide the first segment of data according to M sliding windows to obtain a first initial data set. And the tangent point determination module is also specifically used to obtain a first tangent point set according to a tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution feature.
[0030] In another possible scenario, the tangent point determination module is specifically used to obtain a first tangent point set according to a tangent point determination rule associated with a length interval determined by a length of a data block in a first initial data set and a maximum value, a minimum value and an expected value of the length.
[0031] In another possible scenario, the tangent point determination module is specifically used to: divide the first segment of data into M sliding windows with an interval of at least 1 byte to obtain a first initial data set.
[0032] In another possible scenario, the lengths of the M sliding windows are the same.
[0033] In another possible scenario, the calculation time of the first tangent point set includes: one of: a first duration, a second duration, or the first duration and the second duration.
[0034] The first duration is used to indicate the duration required to divide the first segment of data into M sliding windows to obtain the first initial data set. The second duration is used to indicate the duration required to obtain the first tangent point set according to the tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution characteristics.
[0035] In another possible scenario, the adjustment module is specifically configured to reduce the M sliding windows to N sliding windows when at least one of the calculation time consumption and the cache occupancy rate of the first tangent point set is greater than the respective upper limits, wherein M is greater than N.
[0036] In another possible scenario, the adjustment module is specifically configured to: when at least one of the calculation time consumption and the cache occupancy rate of the first tangent point set is less than the respective lower limit value, increase the M sliding windows to N sliding windows, where M is less than N.
[0037] In another possible scenario, the processing module is specifically used to: delete the data block that matches the stored data in the at least one data block included in the first data set, and store the data block that does not match the stored data in the at least one data block included in the first data set. And the processing module is also specifically used to: delete the data block that matches the stored data in the at least one data block included in the second data set, and store the data block that does not match the stored data in the at least one data block included in the second data set.
[0038] In another possible scenario, the stored data includes a second data block and a third data block. At least one data block included in the first data set matches the second data block. The cut-point partitioning module is further used to: divide the third segment of the data to be stored by using the length of the third data block to obtain a fourth data block.
[0039] In another possible scenario, the first data set includes a fifth data block, a sixth data block, and a seventh data block. The lengths of the fifth data block and the sixth data block are the same as the maximum lengths. The first tangent point divides the data to be stored to obtain the fifth data block, and the second tangent point divides the data to be stored to obtain the seventh data block. The processing module is also used to: generate a tangent point feature table for the fifth data block and the sixth data block. The tangent point feature table includes: a tangent point feature and a tangent point distance. The tangent point feature is used to indicate the fifth data block and the sixth data block. The tangent point distance is used to indicate the spacing between the second tangent point and the first tangent point.
[0040] In a third aspect, the present application provides a computer device. The computer device includes a memory and a processor, and the memory is used to store a set of computer instructions. When the processor executes the set of computer instructions, the processor is used to execute the operation steps of the data processing method in the first aspect or any possible design of the first aspect.
[0041] In a fourth aspect, the present application provides a cluster. The cluster includes a computing node and a storage node. The computing node performs the operation steps of the data processing method in the first aspect or any possible design of the first aspect, and the storage node is used to store data after the computing node performs data deduplication on the data to be stored.
[0042] In a fifth aspect, the present application provides a computer-readable storage medium, including: computer software instructions; when the computer software instructions are executed in a computing device, the computing device executes the operation steps of the method described in the first aspect or any possible implementation of the first aspect.
[0043] In a sixth aspect, the present application provides a computer program product comprising instructions. When the computer program product is run on a cluster, the cluster is caused to execute the operation steps of the method described in the first aspect or any possible implementation of the first aspect.
[0044] The beneficial effects of the third to sixth aspects above can be referred to the description of any implementation in the first or second aspect, and will not be repeated here. Based on the implementation provided by the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A schematic diagram of the structure of a data processing system provided for this application;
[0046] Figure 2 A schematic diagram of a fixed-length division method provided in this application;
[0047] Figure 3 A flowchart for obtaining a data block provided by this application;
[0048] Figure 4 A schematic diagram of a sliding window to obtain data blocks provided by this application;
[0049] Figure 5 A schematic diagram of a variable length partitioning method provided in this application;
[0050] Figure 6 A schematic diagram of a combination of fixed-length and variable-length partitioning methods provided in this application;
[0051] Figure 7 A flowchart of a data processing method provided in this application;
[0052] Figure 8 A schematic diagram of a process for obtaining a cut-point set provided in this application;
[0053] Fig. 9 A schematic diagram of data block length distribution provided for this application;
[0054] Fig.10 A schematic diagram of the relationship between the probability of a cut point determination and the length of a data block provided in this application;
[0055] Fig.11A corresponding diagram of a cut point determination rule and a length interval provided in this application;
[0056] Fig.12 A schematic diagram of obtaining a cut point provided in this application;
[0057] Fig.13 A schematic diagram of generating a cut point feature table provided in this application;
[0058] Fig.14 A schematic diagram of a process for obtaining a cut point provided in this application;
[0059] Fig.15 A schematic diagram of adjusting the degree of parallelism provided for this application;
[0060] Fig.16 A schematic diagram of the relationship between cache hit rate and parallelism provided for this application;
[0061] Fig.17 A data processing block diagram provided for this application;
[0062] Fig.18 A schematic diagram of the structure of a data processing device provided in this application;
[0063] Fig.19 A schematic diagram of the structure of a computing device provided for this application;
[0064] Fig. 20 A schematic diagram of a cluster provided for this application;
[0065] Fig.21 A connection diagram of a computing device provided for this application. DETAILED DESCRIPTION
[0066] In order to make the description of the following embodiments clear and concise, a brief introduction to the related technology is first given.
[0067] Figure 1 A structural diagram of a data processing system provided for this application, such as Figure 1 As shown, the data processing system 100 includes: a client 110 and a computing device 120. Figure 1 In the application scenario shown, users access data through applications. The computers running these applications can be called "clients".
[0068] In one possible example, the client 110 may send data to be stored to the computing device 120 through the network 130. For example, the network 130 may include a switch.
[0069] In another possible example, the client 110 may also communicate with the computing device 120 via a wired connection. For example, the client 110 communicates with the computing device 120 via a universal serial bus (USB) or a peripheral component interconnect express (PCIe) bus.
[0070] Figure 1 The computing device 120 shown can be used to perform a data deduplication operation on the data to be stored, and store the data after the deduplication operation is performed. The computing device 120 can be a centralized data processing system or a distributed data processing system, which will be described below respectively.
[0071] In a first possible scenario, computing device 120 is a centralized data processing system.
[0072] In this case, the computing device 120 has a unified entrance, and all data from external devices must pass through this entrance, which can also be called engine 121.
[0073] like Figure 1 As shown, the engine 121 may have one or more controllers. Figure 1 The example of engine 121 including one controller is used for explanation. In a possible example, if engine 121 has multiple controllers, any two controllers may have a mirror channel to achieve the function of any two controllers backing up each other, thereby avoiding hardware failures that cause the entire computing device 120 to be unavailable. It should be understood that if engine 121 includes multiple controllers, engine 121 may also be referred to as an array controller of computing device 120.
[0074] The engine 121 also includes a front-end interface 1211 and a back-end interface 1214. The front-end interface 1211 is used to communicate with the client 110, thereby providing data access services for the client 110. The back-end interface 1214 is used to communicate with the hard disk to expand the capacity of the computing device 120. Through the back-end interface 1214, the engine 121 can connect more hard disks to form a storage resource pool with a very large storage capacity.
[0075] like Figure 1 As shown, each controller in the engine 121 includes at least: a processor 1212 and a memory 1213 .
[0076] The processor 1212 is a central processing unit (CPU) for processing data access requests from outside the computing device 120 (server or other data processing system), and also for processing requests generated inside the computing device 120. Exemplarily, when the processor 1212 receives write data requests sent by the client 110 through the front-end interface 1211, the data in these write data requests are temporarily saved in the memory 1213. When the total amount of data in the memory 1213 reaches a certain threshold, the processor 1212 sends the data stored in the memory 1213 to at least one of the mechanical hard disk 1221, the mechanical hard disk 1222, the solid state drive (SSD) 1223 or other hard disks 1224 through the back-end port for persistent storage.
[0077] The memory 1213 refers to an internal memory that directly exchanges data with the processor 1212. It can read and write data at any time and at a very fast speed, and serves as a temporary data storage for the operating system or other running programs. The memory 1213 includes at least two types of memories, for example, the memory can be either a random access memory or a read-only memory (ROM). For example, the random access memory is DRAM or SCM. DRAM is a semiconductor memory, and like most random access memories (RAM), it is a volatile memory device. However, DRAM and SCM are only exemplary descriptions in this embodiment, and the memory can also include other random access memories, such as static random access memory (SRAM), etc. As for the read-only memory, for example, it can be a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), etc.
[0078] In addition, the memory 1213 may also be a dual in-line memory module or a dual in-line memory module (DIMM), that is, a module composed of a dynamic random access memory (DRAM), or an SSD. In practical applications, multiple memories 1213 and different types of memories 1213 may be configured in the controller. This embodiment does not limit the number and type of memory 1213. In addition, the memory 1213 may be configured to have a power-saving function. The power-saving function means that when the system loses power and then powers on again, the data stored in the memory 1213 will not be lost. A memory with a power-saving function is called a non-volatile memory.
[0079] The memory 1213 stores software programs, and the processor 1212 runs the software programs in the memory 1213 to manage the hard disk. For example, the hard disk is abstracted into a storage resource pool, and the storage resource pool is provided to the server in the form of a logical unit number (LUN). The LUN here is actually the hard disk seen on the server. Of course, some centralized data processing systems are also file servers themselves, which can provide shared file services for the server.
[0080] like Figure 1 As shown, in the data processing system 100, the engine 121 may not have a hard disk slot. In this case, the hard disk needs to be placed in the hard disk frame 122, and the back-end interface 1214 communicates with the hard disk frame 122. The back-end interface 1214 exists in the engine 121 in the form of an adapter card, and two or more back-end interfaces 1214 can be used simultaneously on one engine 121 to connect multiple hard disk frames. Alternatively, the adapter card can also be integrated on the motherboard, and the adapter card can communicate with the processor 1212 via the PCIe bus.
[0081] It should be noted that Figure 1 Only one engine 121 is shown in the figure. However, in actual applications, the data processing system may include two or more engines 121, and redundancy or load balancing is performed between the multiple engines 121.
[0082] The hard disk enclosure 122 includes: a control unit 1225 and a plurality of hard disks. The control unit 1225 can have various forms. In one case, the hard disk enclosure 122 is an intelligent disk enclosure, such as Figure 1As shown, the control unit 1225 includes a CPU and a memory. The CPU is used to perform operations such as address conversion and reading and writing data. The memory is used to temporarily store data to be written to the hard disk, or read from the hard disk to send data to the controller. In another case, the control unit 1225 is a programmable electronic component, such as a data processing unit (DPU). The DPU has the versatility and programmability of the CPU, but is more specialized and can run efficiently on network data packets, storage requests or analysis requests. The DPU is distinguished from the CPU by a large degree of parallelism (needing to process a large number of requests). Optionally, the DPU here can also be replaced by a processing chip such as a graphics processing unit (GPU) and an embedded neural network processor (NPU). Generally, the number of control units 1225 can be one, or two or more. The functions of the control unit 1225 can be offloaded to the network card 1226. In other words, in this embodiment, the hard disk frame 122 does not have a control unit 1225 inside, but the network card 1226 completes data reading and writing, address conversion and other computing functions. At this time, the network card 1226 is an intelligent network card. It can include a CPU and a memory. The CPU is used to perform operations such as address conversion and reading and writing data. The memory is used to temporarily store data to be written to the hard disk, or data read from the hard disk to be sent to the controller. It can also be a programmable electronic component, such as a DPU. There is no ownership relationship between the network card 1226 and the hard disk in the hard disk frame 122. The network card 1226 can access any hard disk in the hard disk frame 122 (such as Figure 1 The mechanical hard disk 1221, mechanical hard disk 1222, solid state disk 1223 and other hard disks 1224 are shown, so it is more convenient to expand the hard disk when the storage space is insufficient.
[0083] According to the type of communication protocol between the engine 121 and the hard disk frame 122, the hard disk frame 122 may be a serial attached small computer system interface (SAS) hard disk frame, or it may be an NVMe (Non-Volatile Memory express) hard disk frame and other types of hard disk frames. The SAS hard disk frame adopts the SAS3.0 protocol, and each frame supports 25 SAS hard disks. The engine 121 is connected to the hard disk frame 122 through the onboard SAS interface or the SAS interface module. The NVMe hard disk frame is more like a complete computer system, and the NVMe hard disk is inserted in the NVMe hard disk frame. The NVMe hard disk frame is then connected to the engine 121 through the RDMA port.
[0084] For example, the computing device 120 may refer to a storage array, such as an all-flash storage array in which all storage media are flash memories.
[0085] In an optional implementation, the computing device 120 may be a centralized data processing system with integrated disk and controller. The computing device 120 does not have the above-mentioned hard disk frame 122, and the engine 121 is used to manage multiple hard disks connected through the hard disk slots. The function of the hard disk slots may be implemented by the back-end interface 1214.
[0086] In a second possible scenario, computing device 120 is a distributed data processing system.
[0087] In this case, the distributed data processing system includes: a computing device cluster and a storage device cluster.
[0088] The computing device cluster includes one or more computing devices, and each computing device can communicate with each other. The computing device can be, but is not limited to: a server, a desktop computer or a controller of a storage array. In terms of hardware, the computing device may include a processor, a memory, a network card, etc. Among them, the processor is a CPU, which is used to process data access requests from outside the computing device, or requests generated inside the computing device. Exemplarily, when the processor receives a write data request sent by a user, the data in these write data requests is temporarily stored in the memory. When the total amount of data in the memory reaches a certain threshold, the processor sends the data stored in the memory to the storage device for persistent storage. In addition, the processor is also used for data calculation or processing, such as deduplication, data compression, garbage cleaning, data replication, virtualized storage space, and address conversion. In an example, any computing device can access any storage device in the storage device cluster through a network. The storage device cluster includes multiple storage devices. A storage device includes one or more controllers, a network card and multiple hard disks, and the network card is used to communicate with the computing device.
[0089] When the data processing system stores data, data reduction technology can be used to delete data blocks in the data to be stored that are the same as the stored data, so as to reduce the storage capacity occupied by redundant data in the storage device. The above process can be specifically described as follows: First, the computing device divides the data to be stored to obtain multiple data blocks. Secondly, the computing device generates fingerprint features corresponding to each data block. Finally, the computing device compares the fingerprint features of the stored data with the fingerprint features of each data block, deletes the data blocks in the data to be stored that match the fingerprint features of the stored data, and stores the data blocks in the data to be stored that do not match the fingerprint features of the stored data.
[0090] In the above process, the computing device may divide the data to be stored into multiple data blocks and store the data to be stored in multiple ways. For example, the computing device may divide the data to be stored into multiple data blocks by using a fixed-length block division algorithm, a variable-length block division algorithm, or a combination of a fixed-length block division algorithm and a variable-length block division algorithm. The above three ways are described below.
[0091] In a first approach, a computing device uses a fixed-length block algorithm to divide the data to be stored into a plurality of data blocks, and stores the data to be stored.
[0092] The data to be stored below are data 1 and data 2, and the computing device divides the data to be stored into data blocks containing 4 bytes for storage as an example to illustrate the process.
[0093] Figure 2 A schematic diagram of a fixed-length division method provided in this application, such as Figure 2 As shown, for data 1 stored for the first time: the computing device divides data 1 to obtain data blocks 11 to 16, and stores data blocks 11 to 16. Among them, data blocks 11 to 16 each contain 4 bytes. For data not stored for the first time: the computing device divides data 2 to obtain data blocks 21 to 27. The computing device compares data blocks 11 to 16 with data blocks 21 to 27, and determines that among data blocks 21 to 27, only data block 21 is the same as data block 11. The computing device deletes data block 21 in data 2, and stores data blocks 22 to 27 in data 2. In this way, when the stored data is partially changed, the fixed-length block cutting algorithm is used to process the data to be stored, which has the problem of low deduplication rate.
[0094] In a second method, the computing device uses a variable-length block algorithm to divide the data to be stored into multiple data blocks, and stores the data to be stored.
[0095] In this case, the computing device may use a sliding window to implement variable-length slicing of the data to be stored. For example, the length of the sliding window is L bytes.
[0096] Figure 3 A flowchart for obtaining a data block provided by this application, such as Figure 3 As shown, the process includes the following steps ① to ③.
[0097] In step ①, the computing device uses a sliding window to select L bytes of data to be stored, generates a tangent point feature based on the L bytes, and determines whether the tangent point feature meets the conditions.
[0098] Exemplarily, the computing device uses a sliding window to select data corresponding to the 1st byte to the Lth byte of the data to be stored. The computing device performs a hash algorithm on the data corresponding to the 1st byte to the Lth byte of the data to be stored, and obtains the tangent feature f1 corresponding to the Lth byte of the data to be stored. The computing device performs f1 modulo D to obtain a calculation value. And the computing device determines whether the calculation value is within the range indicated by r. Wherein D is a fixed value, r is greater than 0 and less than or equal to D.
[0099] In step ②, if the cut-off feature meets the conditions, the computing device determines that the Lth character of the data to be stored is a cut-off point, and the computing device obtains the data from the 1st byte to the Lth byte of the data to be stored as a data block. The sliding window steps L bytes, and the computing device continues to execute step ①.
[0100] For example, Figure 4 A schematic diagram of a sliding window to obtain data blocks provided by this application, such as Figure 4 As shown, if the operation value is within the range indicated by r, the tangent feature meets the condition, and the computing device obtains the data corresponding to the 1st byte to the Lth byte of the data to be stored, and obtains a data block containing L bytes. Sliding window 1 steps L bytes, and the L+1th byte to the 2Lth byte of the data to be stored are used as the content selected by sliding window 1 next time, and the computing device continues to execute step ①.
[0101] In step ③, if the cut-point feature does not meet the conditions and the sliding window advances by 1 byte, the computing device continues to execute step ①.
[0102] Exemplarily, if the calculated value is within the range indicated by r, the sliding window 1 steps forward by 1 byte, and the 2nd byte to the L+1th byte of the data to be stored are used as the content selected by the next sliding window 1, and the computing device continues to execute step ①. The computing device repeats the above steps ① to ③ until the sliding window 1 traverses the data 2 and obtains multiple data blocks.
[0103] Based on the above detailed description of the method in which the computing device uses a sliding window to obtain data blocks, the process of storing data using a variable-length block algorithm by the computing device is described by still taking the computing device storing data 1 and data 2 as an example.
[0104] Figure 5 A schematic diagram of a fixed-length division method provided in this application, such as Figure 5As shown, for data 1 stored for the first time: the computing device adopts the sliding window data partitioning method described above to partition data 1 to obtain data blocks 11 to 16, and stores data blocks 11 to 16. For data 2: the computing device adopts the sliding window data partitioning method described above to partition data 2 to obtain data blocks 21 to 26. The computing device compares data blocks 11 to 16 with data blocks 21 to 26, and determines that data blocks 21 to 26 are different only in data block 22. The computing device stores data block 22, and deletes data blocks 21 and data blocks 23 to 26. Although a higher deduplication rate can be obtained by adopting a block cutting algorithm, in the process of the computing device using a sliding window to obtain data blocks, each time the sliding window steps one byte, the computing device needs to calculate the cut-point feature once. For example, for data to be stored with a data volume of 1GB, the computing device needs to calculate about 10 9 There are problems of large amount of calculation and low efficiency in the process of using the tangent point feature.
[0105] Method three: the computing device divides the data to be stored into multiple data blocks by combining a fixed-length block cutting algorithm and a variable-length cut-point algorithm, and stores the data to be stored.
[0106] To address the low efficiency problem in method 2, the computing device can first execute a fixed-length block algorithm to divide the data to be stored into multiple data segments of equal length, such as Figure 6 As shown, Figure 6 A schematic diagram of a combination of a fixed-length and variable-length partitioning method provided for the present application. And a block cutting algorithm is executed on each data segment to divide the data block to obtain multiple data blocks. For the specific description of the process of a computing device dividing the data to be stored to obtain multiple data segments, please refer to the specific description of method one, and for the process of a computing device dividing the data segment to obtain multiple data blocks, please refer to the specific description of method two, which will not be repeated here. Although this method can improve efficiency, in the case where some of the stored data increases or decreases, the data segments after the changed data segment are changed, resulting in a reduction in the number of data blocks in the divided data blocks that match the stored data, a low deduplication rate, and a tangent feature that still needs to be performed once for the sliding window to step one byte in the process of obtaining a data block, which requires a large amount of calculation.
[0107] In order to solve the problem of high computational time consumption and low data deduplication rate of the computing device in obtaining the tangent point, the present application provides a data processing method. In the method, the computing device uses M sliding windows to divide the data to be stored to obtain a first data set, the computing device adjusts the M sliding windows to N sliding windows according to at least one of the computational time consumption and cache occupancy rate of obtaining the first tangent point set, and the computing device uses N sliding windows to divide the data to be stored to obtain a second data set. The computing device performs data deduplication on the first data set and the second data set. In this way, multiple sliding windows can obtain at least one tangent point by processing the data to be stored at a time, which reduces the computational time required for the computing device to obtain the tangent point. The computing device adjusts the parallelism of processing the data to be stored according to at least one of the computational time consumption and cache occupancy rate of the first tangent point set, so as to adjust the length of the data to be stored selected by the multiple sliding windows. Improve the accuracy of the computing device in obtaining the tangent point, increase the data to be stored that is repeated with the stored data, and improve the data deduplication rate.
[0108] As a functional feature of distributed storage systems, data deduplication can cover the current primary storage system and backup storage system scenarios that support data deduplication (such as cloud storage, distributed storage, all-flash high-end storage, etc.), including storage systems managed based on open source distributed storage software (such as Ceph, HDFS), etc.
[0109] Figure 7 A flow chart of a data processing method provided in this application is as follows: Figure 7 As shown, this method can be Figure 1 The described computing device is executed, and the method includes the following S710 to S740.
[0110] S710: The computing device obtains a first tangent point set of a first segment of data in the data to be stored according to a first degree of parallelism, and divides the first segment of data according to at least one tangent point in the first tangent point set to obtain a first data set.
[0111] The process of the computing device acquiring the first data set may include the following steps ① and ②.
[0112] In step ①, the computing device uses M sliding windows to obtain the first tangent point set of the first segment of data in the data to be stored.
[0113] The computing device may adopt a method of sliding M sliding windows on the data to be stored to obtain the tangent points in the first tangent point set. The computing device may also adopt a method of selecting tangent points to obtain the tangent points in the first tangent point set.
[0114] In a first possible scenario, the computing device uses M sliding windows to slide on the data to be stored to obtain the tangent points in the first tangent point set.
[0115] The computing device may use M sliding windows to slide at least once on the data to be stored to obtain a first tangent point set of the first segment of data. Figure 8 A schematic diagram of a process for obtaining a cut-point set provided in this application, such as Figure 8 As shown, the process includes the following S11 and S12.
[0116] S11, the computing device uses M sliding windows to divide the first segment of data to obtain a first initial data set.
[0117] The first segment of data is not pre-divided, and the first segment of data may refer to data covered when the M sliding windows are not sliding, or may refer to data covered when the M sliding windows are sliding.
[0118] In a possible scenario, the computing device may divide the first segment of data using M sliding windows with an interval of at least 1 byte to obtain a first initial data set, such as an interval of a bytes, where a is a positive integer greater than or equal to 1.
[0119] Depending on whether it is the first time for the M sliding windows to slide on the data to be stored and the division of the data to be stored, the process of obtaining the first initial data set is also different, which are described below respectively.
[0120] (1) M sliding windows slide on the data to be stored for the first time, and divide the data to be stored.
[0121] For example, the data to be stored is data 1, the M sliding windows are sliding windows 1 to sliding window m, the M sliding windows divide data 1 at intervals of a bytes, and sliding window 1 divides data 1 starting from the i-th byte of data 1.
[0122] Sliding window 1 selects the 1st byte to the Lth byte of data 1, sliding window 2 selects the ath byte to the (L+a-1)th byte of data 1, and sliding window j selects the j*ath byte to the (L+j*a-1)th byte of data 1. In this case, the first segment of data may refer to the data segment corresponding to the 1st byte to the (L+j*a-1)th byte in data 1.
[0123] Sliding window 1 divides the first segment of data from the 1st byte to the Lth byte to obtain a data block 11 with a length of L bytes. Sliding window 2 divides the first segment of data from the ath byte to the L+a-1th byte to obtain a data block 12 with a length of L bytes. Sliding window j divides the first segment of data from the j*ath byte to the (L+j*a-1)th byte to obtain a data block 1j with a length of L bytes. Wherein, i is a positive integer greater than or equal to 1, and j is a positive integer greater than or equal to 3 and less than or equal to m. The computing device obtains a first initial data set containing m data blocks according to data block 11, data block 12, and data block 1j. The following takes the example of M sliding windows dividing the first segment of data at an interval of 1 byte to illustrate the present application.
[0124] (2) The M sliding windows slide on the data to be stored non-for the first time, and the data to be stored is divided.
[0125] The computing device slides the M sliding windows on the data to be stored for the second time and divides the data to be stored. Depending on whether the position of the last byte of the data block in the first initial data set is a tangent point, the computing device divides the data to be stored using the M sliding windows in different processes.
[0126] 1. The situation where the last byte of the data block in the first initial data set is located at the tangent point.
[0127] In this case, the sliding window that divides the data to be stored to obtain data blocks steps L bytes, and the computing device uses the position after stepping L bytes as the starting position of the sliding window to re-divide the data to be stored to obtain new data blocks.
[0128] Exemplarily, the sliding window k is divided to obtain data block 1k, and data block 1k contains data from the kth byte to the (L+k-1)th byte of data 1. The position of the last byte of data block 1k is the tangent point. The sliding window k steps L bytes, and the sliding window k after stepping L bytes takes the (L+k)th byte of data 1 as the starting position and is re-divided to obtain data block k1 with a length of L bytes. Data block k1 contains data from the (L+k)th byte to the (2L+k-1)th byte of data 1.
[0129] 2. The situation where the position of the last byte of the data block in the first initial data set is not the tangent point.
[0130] In this case, the sliding window stepping of the data blocks obtained by dividing the data to be stored is 1 byte, and the data to be stored is continuously selected to obtain data blocks with a length of (L+1) bytes.
[0131] Exemplarily, sliding window k divides data 1 to obtain data block 1k, and data block 1k contains data from the kth byte to the (L+k-1)th byte of data 1. The position where the last byte of data block 1k is located is not the tangent point. Sliding window k steps 1 byte, and after stepping 1 byte, sliding window k takes the (k+1)th byte of data 1 as the starting position and re-divides to obtain data block k1 with a length of (L+1) bytes. Data block k1 contains data from the (k+1)th byte to the (L+k)th byte of data 1.
[0132] In this case, the first segment of data may refer to data covered after some of the M sliding windows step by 1 byte and another part of the sliding windows step by L bytes.
[0133] The above text takes the second sliding of M sliding windows on the data to be stored and the division of the data to be stored as an example to illustrate the process of using M sliding windows to divide the data to be stored to obtain the first initial data set. After the M sliding windows have slid twice, the computing device can obtain the first initial data set containing data blocks with a length of L bytes or (L+1) bytes. Similarly, the computing device can use the method described in (2) to use M sliding windows to slide and divide the data to be stored for the third, fourth, and so on times. For related content, please refer to the description in (2), which will not be repeated here.
[0134] The computing device may determine the length of the data block in the first initial data set according to the length of the sliding window and the number of times the sliding window slides. After obtaining the first initial data set containing at least one data block, the computing device may use the method described in S12 to determine whether the length of the data block in the first initial data set is within the length interval determined by the maximum and minimum values of the length. If not, the computing device may continue to slide the sliding window to increase the length of the data block in the first initial data set. If it is, the computing device may use the tangent point determination rule corresponding to the length to determine whether the position where the last byte of the data block is located is the tangent point.
[0135] S12, the computing device obtains a first tangent point set according to a tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution feature.
[0136] The distribution characteristics may refer to the distribution of various indicators of a set of data. The distribution characteristics may be described by the central tendency of the distribution, which may reflect the degree to which each data converges or aggregates to its central value. In this application, the central value may also be referred to as the expectation. The distribution characteristics of the length may refer to the maximum value, minimum value, and expected value of the data block length.
[0137] In one possible scenario, the computing device obtains the first tangent point set according to a tangent point determination rule associated with the length of the data block in the first initial data set and the length interval determined by the maximum value, minimum value and expected value of the length. The length interval determined by the maximum value, minimum value and expected value of the length can be set by the user according to actual needs.
[0138] When the content of the data to be stored is random, the cut-off feature f generated by the data block selected by the sliding window is also random. In this case, from a statistical point of view, it can be considered that the probability of the position where the last byte of the data block selected by the sliding window each time is determined as the cut-off point is the same. Let this probability be p, then the probability of the last byte of the data block selected by the sliding window for the nth time being determined as the cut-off point is shown in formula (1).
[0139] P = p × (1-p) (n-1) Formula (1)
[0140] Among them, p is the probability that the position of the last byte of the data block is determined as the cut point, (1-p) (n-1) is the probability that the position of the last byte of the (n-1) data blocks obtained by the previous (n-1) sliding of the sliding window is not determined as the tangent point. From formula (1), it can be seen that the probability that the position of the last byte of the data block obtained by each sliding of the sliding window is determined as the tangent point is an exponential function. Based on this property, the maximum value, minimum value and expected value of the given length are combined to establish the following Fig. 9 The data block length distribution diagram shown in FIG. Fig. 9 A schematic diagram of data block length distribution provided for this application.
[0141] Depending on the different relationships between the length of the data block in the first initial data set and the maximum value, minimum value and expected value of the length, the process of the computing device obtaining the first tangent point set is also different, which is described below in different cases.
[0142] Case A is for a data block in the first initial data set whose length is less than the minimum length.
[0143] In this case, the computing device does not determine whether the position of the last byte of the data block is the tangent point. The computing device makes the sliding window of the data block stepped by 1 byte. After the sliding window has stepped by 1 byte, it continues to divide the data to be stored by the method described in S11 to obtain a new data block.
[0144] For the new data block, the computing device obtains the length of the new data block, and again determines the relationship between the length of the new data block and the maximum, minimum and expected values of the length. If the length of the new data block is less than the minimum value of the length, the new data block is processed using the method described in scenario A. If the length of the new data block is within the length interval determined by the minimum and maximum values of the length, the new data block is processed using the method described in scenario B. If the length of the new data block is equal to the maximum value of the length, and the computing device determines that the last byte of the data block is not a tangent point using the method described in scenario B, the computing device processes the new data block using the method described in scenario C to obtain a tangent point in the first tangent point set.
[0145] Exemplarily, sliding window 1 divides data 1 to obtain data block 11, and the length of data block 11 is l 1 The minimum length is Lmin. If l 1 <Lmin, the computing device does not determine whether the position of the last byte of data block 11 in data 1 is the tangent point, and the sliding window 1 steps forward by 1 byte. After stepping forward by 1 byte, the sliding window 1 divides data 1 to obtain a length of l 1 +1 byte of data block 21.
[0146] For data block 11, the computing device obtains the length of data block 11 (ie, (l 1 +1) bytes), and again determine the relationship between the length of data block 11 and the maximum value (i.e., Lmax), the minimum value (i.e., Lmin) and the expected value of the length. 1 +1<Lmin, the computing device processes data block 11 using the method described in case A. 1 +1≤Lmax, the computing device processes data block 11 using the method described in scenario B. 1 +1=Lmin, and the computing device uses the method described in situation B to determine that the position of the last byte of the data block is not a tangent point, and the computing device uses the method described in situation C to process data block 11 to obtain a tangent point in the first tangent point set.
[0147] Case B: For a data block in the first initial data set, the length of the data block is within a length interval determined by a minimum length and a maximum length.
[0148] In one possible scenario, the computing device may divide the length interval determined by the maximum length value and the minimum length value into a plurality of sub-length intervals, and determine a different tangent point determination rule for each sub-length interval. Fig.10 A schematic diagram of the relationship between the probability of cutting point determination and the length of the data block provided in this application is as follows: Fig.10 As shown in the figure, when the length of the data block changes from the minimum value to the expected value, the probability that the position where the last byte of the data block is located is determined as the tangent point gradually increases, and when the length of the data block changes from the expected value to the maximum value, the probability that the position where the last byte of the data block is located is determined as the tangent point gradually increases. In this way, the length of the data block obtained by the computing device using the sliding window division can be as long as possible without being less than the expected value.
[0149] For example, Fig.11 A corresponding diagram of a cut point determination rule and a length interval provided in this application, such as Fig.11As shown, the computing device divides the length interval determined by the maximum value of the length and the minimum value of the length into sub-length interval 1 to sub-length interval i, and sub-length interval (i+1) to sub-length interval n. Sub-length interval 1 to sub-length interval n correspond to tangent point determination rules 1 to tangent point determination rules n respectively. Sub-length interval i is the interval where the expected value is located. The probability that the position where the last byte of the data block is determined to be the tangent point by tangent point determination rules 1 to tangent point determination rules n gradually increases. The computing device can increase the probability that the position where the last byte of the data block is determined to be the tangent point by the tangent point determination rule by reducing the number of judgment conditions shown in formula (2). As shown in Table 1, if the length of the data block is within the length range of sub-length interval 1 close to the minimum value of the length, the computing device determines that the position where the last byte of the data block is the tangent point only when n judgment conditions shown in formula (2) are met at the same time. If the length of the data block is within the length range of sub-length interval i where the expected value is located, the computing device only needs to meet i judgment conditions shown in formula (2) at the same time to determine that the position where the last byte of the data block is the tangent point. And if the length of the data block is within the length range of the sub-length interval n close to the maximum value of the length, only one determination condition shown in formula (2) needs to be satisfied, and the computing device determines that the position where the last byte of the data block is located is the tangent point. In this way, the length of the data block obtained by dividing the data to be stored by the tangent point is adjusted so that the length of the data block obtained by the tangent point division is as long as possible without being less than the expected value.
[0150] f mod D = r Formula (2)
[0151] Among them, f is the cut-point feature generated by the data block selected by sliding window, D is a fixed value, and r is greater than zero and less than or equal to D.
[0152] Table 1
[0153]
[0154]
[0155] The computing device may divide the length interval into multiple sub-length intervals in a variety of ways, and several possible examples are given below.
[0156] Example 11: The computing device divides the length interval into multiple sub-length intervals of different lengths. For example, if the length of the length interval is 15 bytes, the computing device divides the length interval into sub-length interval 1 to sub-length interval 5. The lengths of sub-length interval 1 to sub-length interval 5 are 1 byte to 5 bytes, respectively.
[0157] Example 12: The computing device divides the length interval into multiple sub-length intervals, some of which are equal in length and some of which are different in length. For example, if the length of the length interval is 15 bytes, the computing device divides the length interval into sub-length interval 1 to sub-length interval 10. Among them, the lengths of sub-length interval 1 to sub-length interval 5 are all 1 byte, and the lengths of sub-length interval 6 to sub-length interval 10 are all 2 bytes.
[0158] Example 13: The computing device divides the length interval into multiple sub-length intervals of the same length, so that the length of each sub-length interval is the same. For example, if the length of the length interval is 20 bytes, the computing device evenly divides the length interval into sub-length interval 1 to sub-length interval 10. Among them, the lengths of sub-length intervals 1 to sub-length interval 10 are all 2 bytes.
[0159] After determining the cut-point determination rule corresponding to the length interval, the computing device may use the following B11 to B13 to determine whether the position where the last byte of the data block is located is a cut-point.
[0160] B11, the computing device obtains the data of the sliding window frame selected by dividing the data blocks.
[0161] Exemplarily, sliding window k divides the kth byte to the (k+x-1)th byte of data 1 to obtain a data block 1k of length x bytes. In this case, the computing device obtains the content from the ((k+x-1)-L)th byte to the (k+x-1)th byte of data 1 selected by sliding window k.
[0162] B12, the computing device generates a tangent point feature based on the data selected by the sliding window.
[0163] Exemplarily, the computing device performs a hash operation on the contents from the ((k+x-1)-L)th byte to the (k+x-1)th byte of data 1 to obtain the cut-point feature f corresponding to data block 1k k .
[0164] B13, the computing device determines whether the cut-point feature complies with the cut-point determination rule associated with the length interval in which the data block length is located.
[0165] The computing device determines the sub-length interval where the length of the data block is located according to the length of the data block, and determines whether the position where the last byte of the data block is located is the tangent point according to the tangent point determination rule corresponding to the sub-length interval. If so, the computing device obtains a tangent point in the first tangent point set, and the computing device makes the sliding window k step L bytes, and uses the method described in S11 above to continue dividing the data to be stored to obtain a new data block. If not, the computing device makes the sliding window k step 1 byte, and uses the method described in S11 above to continue dividing the data to be stored to obtain a new data block.
[0166] Exemplarily, the length of the data block 1k is in the sub-length interval n. The computing device uses the determination condition f corresponding to the row with sequence number n in Table 1. k mod D k =r k , determine whether the position of the last byte of the data block 1k is a tangent point. If so, the computing device obtains a tangent point, and the computing device makes the sliding window k step L bytes, and uses the method described in S11 above to continue dividing the data to be stored to obtain a new data block. If not, the computing device makes the sliding window k step 1 byte, and uses the method described in S11 above to continue dividing the data to be stored to obtain a new data block.
[0167] In case C, the length of the data block in the first initial data set is equal to the maximum length, and the computing device uses the method described in case B to determine that the position of the last byte of the data block is not the data block of the tangent point.
[0168] In this case, the computing device directly determines the position of the last byte of the data block in the data to be stored as the tangent point, and obtains a tangent point in the first tangent point set. The computing device makes the sliding window step L bytes, and continues to divide the data to be stored to obtain a new data block using the method described in S11 above.
[0169] Exemplarily, sliding window 1 divides data 1 to obtain data block 11, and the length of data block 11 is l 1 The maximum length is L max , l 1 Equal to L max And L max > L. The computing device uses the method described in Situation B to determine that the position where the last byte of data block 1 is located is not a tangent point. In this case, the computing device directly determines that the position where the last byte of data block 11 in data 1 is located is a tangent point, and obtains a tangent point in the first tangent point set. The computing device makes sliding window 1 step L bytes, and uses the method described in S11 above to continue dividing the data to be stored to obtain new data blocks.
[0170] In a second possible scenario, the computing device obtains the tangent points in the first tangent point set by selecting tangent points.
[0171] In this case, the computing device can also generate a tangent feature table for the stored data, and generate a tangent feature for each tangent point in the first tangent point set. The computing device compares the tangent features of each tangent point in the first tangent point set with the tangent features in the tangent feature table. If the tangent feature of the first tangent point in the first tangent point set matches the tangent feature in the tangent feature table, the computing device starts from the position of the first tangent point and directly shifts the length of a specific byte to obtain another tangent point. The length of a specific byte may refer to the tangent distance corresponding to the tangent feature in the tangent feature table. In this way, in the case where the data to be stored is partially modified relative to the stored data, the computing device does not need to use a sliding window to slide byte by byte on the data to be stored to obtain the tangent points of the data blocks in the first data set. This saves the amount of computation required by the computing device to calculate the tangent points and improves the efficiency of data deduplication.
[0172] For example, Fig.12 A schematic diagram of obtaining a cut point provided for this application, such as Fig.12 As shown, the stored data includes data blocks 1 to data blocks n. The cut-point feature of data block 1 is cut-point feature 1, and the cut-point distance corresponding to cut-point feature 1 is L 1 . Cut point 1 divides data 1 to obtain data block 11, which includes data from the 1st byte to the i-th byte of data 1. The position of cut point 1 is the i-th byte of data block 11. The cut point feature of data block 11 is cut point feature 11. When cut point feature 1 matches cut point feature 11, the computing device obtains the cut point distance L corresponding to cut point feature 1. 1 . And the displacement L between the position of the tangent point 1 of the computing device (i.e. the i-th byte of data 1) 1 bytes, and get the cut point 2 (that is, the Lth 1 +i bytes).
[0173] The cut-point feature table includes: cut-point features and cut-point distances. The cut-point features are used to indicate the contents of the sliding window selection of the data block obtained by dividing. The cut-point distance is used to indicate: the distance between the first cut-point and the second cut-point, the first cut-point is used to indicate the position of the last byte of the data block, and the second cut-point is used to indicate the subsequent adjacent cut-point of the first cut-point. The cut-point distance is the same as the length of the subsequent adjacent data block of the data block. Table 2 is a cut-point feature table provided by the present application.
[0174] Table 2
[0175] Data Block Cut point feature (key) Tangent distance (value) Data Block Tangent point feature Tangent distance
[0176] Exemplarily, the stored data includes data block 1 and data block 2, and data block 2 is the subsequent adjacent data block of data block 1. Data block 1 is obtained by dividing data 1 at cut point 1, and data block 2 is obtained by dividing data 1 at cut point 2. The computing device may generate a cut point feature table for data block 1 using the following process, which is specifically as follows: ① The computing device generates a cut point feature 1 corresponding to data block 1. ② The computing device obtains the distance 1 between cut point 2 and cut point 1. ③ The computing device generates a cut point feature table for data block 1 based on the cut point feature 1 and distance 1.
[0177] The computing device may generate a tangent point feature table in a variety of ways, and two possible ways are given below.
[0178] In method a, the computing device generates a corresponding tangent point feature table for each tangent point.
[0179] This method is applicable to the tangent points obtained by the computing device using the method described in the above scenario B. Assume that the stored data includes data blocks 1 to data blocks K, and the computing device stores the tangent point feature table corresponding to data blocks 1 to 1K. Take the data to be stored as data 1, and the tangent point 1 divides data 1 to obtain data block 11 as an example to illustrate the process of generating the tangent point feature table.
[0180] In step ①, the computing device obtains the divided data 1 to obtain the content 11 of the sliding window selection of the data block 11, and generates the tangent point feature 11 according to the content 11.
[0181] Exemplarily, the computing device may execute a hash algorithm on the content 11 to generate a cut-point feature 11 of the content 11 .
[0182] In step ②, the computing device compares the cut-point feature 11 with the cut-point features in the cut-point feature table. If the cut-point feature 11 matches any one in the cut-point feature table, the computing device executes step ③. If the cut-point feature 11 does not match any of the cut-point features in the cut-point feature table, the computing device executes step ④.
[0183] Exemplarily, the computing device compares the cut-point feature 11 with the cut-point features corresponding to data blocks 1 to data blocks K. If the cut-point feature 11 is the same as the cut-point feature corresponding to data block 1, it is considered that the cut-point feature 11 matches the cut-point feature corresponding to data block 1. If the cut-point feature 11 is different from the cut-point features corresponding to data blocks 1 to data blocks K, it is considered that data block 11 does not match any of the cut-point features in the cut-point feature table.
[0184] In step ③, the computing device obtains the tangent distance corresponding to the tangent feature of data block 1, and the computing device shifts the tangent distance corresponding to the tangent feature of data block 1 from the position where tangent 1 is located, to obtain tangent 2. The computing device obtains the data between tangent 1 and tangent 2 in data 1, to obtain data block 21. And the computing device continues to use the method described in steps ① and ② to determine whether the tangent feature of data block 21 is the same as the tangent feature in the stored tangent feature table.
[0185] In step ④, the computing device stores the data block 11, generates a tangent point feature table 11 of the data block 11, and adds the tangent point feature table 11 to the stored tangent point feature table to obtain a new tangent point feature table including the tangent point feature table 11.
[0186] Exemplarily, the computing device generates a tangent feature table 11 according to the tangent feature 11 and the tangent distance 11. Specifically, the computing device generates the tangent feature 11 using the method in step ①, the computing device obtains the tangent distance 2 corresponding to the tangent feature of the data block 2 of the stored data as the tangent distance 11, and updates the tangent distance 1 corresponding to the tangent feature of the data block 1 in the stored tangent feature table as the length of the data block 11. According to the needs of the actual application, the computing device can also use other methods to determine the tangent distance 11, which is not limited in this application.
[0187] In method b, the computing device generates a tangent point feature table corresponding to the multiple tangent points.
[0188] This scenario is applicable to the tangent points obtained by the computing device using the method described in scenario C above.
[0189] In one possible scenario, the first data set includes adjacent fifth and sixth data blocks, and the fifth and sixth data blocks are both data blocks obtained by dividing the fifth and sixth data blocks using the tangent point determined by the computing device using the method described in scenario C. The length and maximum length of the fifth and sixth data blocks are the same, and the computing device generates and stores a tangent point feature table corresponding to the fifth and sixth data blocks. The tangent point features in the tangent point feature table are jointly generated by the fifth and sixth data blocks. The tangent point distance in the tangent point feature table is the distance between the third tangent point and the fourth tangent point. The third tangent point is used to indicate the tangent point for dividing the third segment of data to be stored to obtain the fifth data block, and the fourth tangent point is used to indicate the tangent point for dividing the fourth segment of data to be stored to obtain the sixth data block. Fig.13 A schematic diagram of generating a cut point feature table provided in this application, such as Fig.13As shown, when there are multiple (such as A) adjacent cut points determined by the method described in situation C in the first data set, the computing device generates a cut point feature table for the multiple cut points. Compared with the computing device generating a cut point feature table for each of the multiple cut points, the computing device saves time in calculating the cut point features corresponding to (A-1) cut points, and saves the storage space occupied by the computing device to store the cut point feature table corresponding to the (A-1) cut points.
[0190] Fig.14 A schematic diagram of a process for obtaining a cut point provided in this application, such as Fig.14 As shown, the computing device uses a differentiated tangent point determination method (i.e., different tangent point determination rules are used for data blocks of different lengths) to determine the tangent point. The computing device generates a tangent point feature corresponding to the tangent point, and generates a tangent point feature table. The computing device compares the tangent point feature with the tangent point feature in the tangent point feature table of the stored data, and when the tangent point feature is the same as a tangent point feature in the tangent point feature table, the computing device starts to shift the first distance from the position of the tangent point to obtain the next tangent point, and the first distance may refer to: the tangent point distance corresponding to the tangent point feature in the tangent point feature table. The computing device determines whether the tangent point feature of the next tangent point is the same as the tangent point feature in the tangent point feature table of the stored data (i.e., it satisfies deduplication locality). Deduplication locality may mean that there is a small part of the data modified between the two data, and most of the content is the same.
[0191] In step ②, the computing device uses at least one tangent point in the first tangent point set to divide the first segment of data to obtain a first data set.
[0192] Exemplarily, the data to be stored is data 1, and the first cut point set includes cut points 1 to a. Cut point 1 indicates the xth byte of data 1, and cut point a indicates the yth byte of data 1. In this case, the first segment of data may refer to the 1st byte to the yth byte of data 1. Cut points 1 to a divide the 1st byte to the yth byte of data 1 respectively to obtain a first data set. Among them, the first data set includes data blocks 11 to 1x.
[0193] Optionally, for the case where the computing device directly shifts the length of a specific byte on the current tangent point to obtain another tangent point in the first tangent point set, the computing device may use the following process to divide the data to be stored to obtain data blocks in the first data set. The stored data includes a second data block and a third data block. At least one data block included in the first data set matches the second data block. The computing device divides the third segment of the data to be stored by the length of the third data block to obtain a fourth data block. The third data block may refer to the subsequent adjacent data block of the second data block. In this way, in the case where the data to be stored is only slightly modified compared to the stored data, compared to the computing device using a sliding window to slide byte by byte on the data to be stored to obtain the initial data block, and determining the tangent point according to the tangent point determination rule associated with the length of the initial data block and the length distribution feature, and using the tangent point to divide the data to obtain the data block. In the present application, the computing device uses a sliding window to directly slide the length of the third data block to obtain the data block, which reduces the number of calculations of the tangent point feature, reduces the amount of calculation, and improves the efficiency of data deduplication.
[0194] Exemplarily, the stored data includes data block 1, data block 2 to data block K. The length of data block 2 is x bytes. The first data set includes data block 3. Data block 3 contains data from the i-th byte to the y-th byte of data 1. If data block 3 matches data block 1, the computing device starts from the (y+1)th byte of data 1, selects the (y+1)th byte to the (y+x)th byte of data 1, and obtains data block 4 with a length of x bytes.
[0195] S720: The computing device adjusts the first degree of parallelism according to at least one of the computation time and the cache occupancy rate of the first tangent point set to obtain a second degree of parallelism.
[0196] The first cache occupancy rate is used to indicate the cache occupancy rate of performing data deduplication on the first segment of data according to the first tangent point set. The second parallelism is used to indicate N sliding windows. M is not equal to N, and M and N are integers greater than or equal to 2.
[0197] In a possible scenario, the computing device may obtain the computing time consumption of the first tangent point set in multiple ways. Three possible ways are given below.
[0198] In method A, the computing device uses the first duration as the computing time of the first tangent point set.
[0199] The first duration is used to indicate: the duration required for M sliding windows to divide the first segment of data into a first initial data set.
[0200] Exemplarily, the M sliding windows include: sliding window 1 to sliding window m. The data to be stored is data 1. The first initial data set includes: data block 11 to data block 1m. In this case, the first duration is used to indicate: the duration required for dividing data 1 from sliding window 1 to sliding window m to obtain data blocks 11 to data blocks 1m. For the process of dividing data from the sliding window to obtain data blocks, please refer to the relevant description in S11 above, which will not be repeated here.
[0201] In method B, the computing device uses the second duration as the computing time of the first tangent point set.
[0202] The second duration is used to indicate: the duration required for obtaining the first tangent point set according to a tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution feature.
[0203] The second duration may include: the duration required for the computing device to generate the tangent point feature according to the content of the sliding window frame selected by dividing the data to be stored to obtain the data block, and the duration required for the computing device to determine whether the tangent point feature meets the tangent point determination rule. For the above process, please refer to the relevant description in S12 above, which will not be repeated here.
[0204] Exemplarily, the first initial data set includes: data block 11 to data block 1m. Sliding windows 1 to m respectively divide data 1 to obtain data blocks 11 to 1m. The second duration includes: the computing device generates a cut-off feature f according to the content 1 to content m selected by sliding windows 1 to m. 1 To the tangent point feature f m The time required and the feature f of the cut point 1 To the tangent point feature f m The time required to meet the cut-point determination rules.
[0205] In method C, the computing device uses the sum of the first duration and the second duration as the computing time of the first tangent point set.
[0206] For the relevant contents about the first duration and the second duration, please refer to the description of method A and method B above, which will not be repeated here.
[0207] According to the different relationships between the computation time consumption and the cache occupancy rate and their respective upper limits, the computing device may use different methods to adjust the first degree of parallelism to the second degree of parallelism, which is described below in different situations.
[0208] Fig.15 A schematic diagram of adjusting the degree of parallelism provided for this application, such as Fig.15 As shown, the computing device adjusts the degree of parallelism according to the computing time and the cache occupancy rate, and the computing device can adjust the degree of parallelism using formula (3).
[0209]
[0210] Where X is the degree of parallelism, The length of each parallel processed data block, is the time required for each parallel processing, i.e., the first time. k is the calculation time required for sliding the window by 1 byte. f(X) is the cache occupancy rate. g(X) is the time required for determining the cut point, i.e., the second time. The computing device uses the calculation time and the cache occupancy rate (i.e., To achieve performance monitoring and The optimization solution is obtained to obtain X that minimizes the value of the formula, and the parallelism is adjusted using X.
[0211] There is a relationship between cache occupancy and parallelism as described by formula (4).
[0212]
[0213] Among them, X m is the maximum parallelism, indicating that when this parallelism is adopted, the cache occupancy rate is 100%. Fig.16 A schematic diagram of the relationship between cache hit rate and parallelism provided for this application, such as Fig.16 As shown, as the parallelism increases, the cache occupancy rate also increases. When the cache occupancy rate reaches 100%, the cache hit rate gradually decreases.
[0214] In one possible scenario, when at least one of the computation time and the cache occupancy rate of the first tangent point set is greater than the respective upper limit values, the computing device reduces the M sliding windows to N sliding windows. This scenario may include the following three possible examples.
[0215] Example 1: When the computing time is greater than the upper limit of the computing time and the cache occupancy rate is greater than the upper limit of the cache occupancy rate, the computing device reduces the M sliding windows to N sliding windows.
[0216] Example 2: When the computation time is greater than the upper limit of the computation time and the cache occupancy rate is less than or equal to the upper limit of the cache occupancy rate, the computing device reduces the M sliding windows to N sliding windows.
[0217] Example 3: When the computation time is less than or equal to the upper limit of the computation time and the cache occupancy rate is greater than the upper limit of the cache occupancy rate, the computing device reduces the M sliding windows to N sliding windows, where M is greater than N.
[0218] In another possible scenario, when at least one of the computation time consumption and the cache occupancy rate of the first tangent point set is less than the respective lower limit value, the M sliding windows are increased to N sliding windows, where M is less than N.
[0219] The computing device may use a variety of methods to achieve mutual adjustment between the number of sliding windows M and N. Several possible examples are given below.
[0220] Example 1: The computing device adjusts the number of sliding windows according to a preset adjustment number. For example, the adjustment number is X. When the computing time is greater than the upper limit of the computing time and the cache occupancy rate is greater than the upper limit of the cache occupancy rate, the computing device reduces the M sliding windows by X to obtain N sliding windows.
[0221] Example 2: The computing device adjusts the number of sliding windows to a preset value according to the computing time. For example, when the computing time reaches the first time, the computing device adjusts the M sliding windows to the preset value 1. When the computing time reaches the second time, the computing device adjusts the M sliding windows to the preset value 2, and so on.
[0222] Example 3: The computing device adjusts the number of sliding windows to a preset value according to the cache occupancy rate. For example, when the cache occupancy rate reaches a first cache occupancy rate, the computing device adjusts the M sliding windows to a preset value of 1. When the cache occupancy rate reaches a second cache occupancy rate, the computing device adjusts the M sliding windows to a preset value of 2, and so on.
[0223] Three computing devices are provided above to realize the adjustment between M sliding windows and N sliding windows. According to the needs of actual applications, other methods can also be used to realize the adjustment between M sliding windows and N sliding windows, such as adjusting the number of sliding windows one by one, etc. The specific adjustment method of the sliding windows is not limited in this application.
[0224] S730: The computing device obtains a second tangent point set for the second segment of data in the data to be stored according to the second degree of parallelism, and divides the second segment of data into at least one tangent point in the second tangent point set to obtain a second data set.
[0225] The second parallelism is used to indicate N sliding windows.
[0226] In a possible scenario, the computing device may use N sliding windows to slide on the data to be stored at least once to obtain a first tangent point set of the first segment of data.
[0227] In one possible scenario, the computing device divides the first segment of data according to N sliding windows to obtain a first initial data set, and the computing device obtains a first tangent point set according to a tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution characteristics.
[0228] In one possible scenario, the computing device divides the first segment of data into N sliding windows with an interval of at least 1 byte to obtain a first initial data set.
[0229] In one possible scenario, the computing device obtains the first tangent point set according to a tangent point determination rule associated with a length interval determined by a length of a data block in a first initial data set and a maximum value, a minimum value, and an expected value of the length.
[0230] The computing device uses N sliding windows to obtain a second tangent point set for the second segment of data in the data to be stored, and at least one tangent point in the second tangent point set divides the second segment of data to obtain a process of obtaining a second data set. This is the same as the process described in S710 above, in which the computing device uses M sliding windows to obtain a first tangent point set for the first segment of data in the data to be stored, and at least one tangent point in the first tangent point set divides the first segment of data to obtain a first data set. For related content, please refer to the description of S710 above, which will not be repeated here.
[0231] S740: The computing device performs data deduplication on the first data set and the second data set.
[0232] The first data set includes at least one data block, and the second data set includes at least one data block.
[0233] In one possible scenario, the computing device performs data deduplication on the first data set, including: the computing device deletes a data block that matches the stored data in at least one data block included in the first data set, and stores a data block that does not match the stored data in at least one data block included in the first data set. The computing device performs data deduplication on the second data set, including: the computing device deletes a data block that matches the stored data in at least one data block included in the second data set, and stores a data block that does not match the stored data in at least one data block included in the second data set.
[0234] The computing device may receive data to be stored, divide the data to be stored into a data set (such as a first data set and a second data set) including at least one data block, and generate content features of each data block. After obtaining the content features of each data block, the computing device queries whether the data block generating the content features has been deduplicated according to the content features. The computing device includes an index subsystem, and the index subsystem includes a database (such as a teraDB database). Under normal circumstances, the teraDB database stores a mapping between the logical address and the physical address of the data. In the data deduplication scenario, the teraDB database stores a mapping between the logical address and the content features of the data. The computing device determines whether the data has been deduplicated using the different mapping relationships in different scenarios. The computing device accesses the teraDB database, and if the data block stores the content features, the data has been deduplicated. Otherwise, the data has not been deduplicated. The process may be performed in the global cache (Global Cache) of the computing device.
[0235] Fig.17 A data processing block diagram provided for this application, such as Fig.17 As shown, the computing device receives the data to be stored sent by the client, and writes the data to be stored to the cache. The computing device can use a data segmentation module to process the data to be stored to obtain at least one data block. The computing device calculates the content features of each data block, queries the mapping in the database of the index subsystem according to the content features of the data block, determines whether the data block generating the content features has been deduplicated, and sets a deduplication mark for the data that has been deduplicated. And the computing device stores the data blocks that do not match the stored data to the distributed storage cluster. The data segmentation module can include a multi-stream parallel computing module and a multi-strategy tangent point determination module. The multi-stream parallel computing module uses multiple parallelisms (such as the first parallelism) to divide the data to be stored to obtain a first initial data set. The multi-strategy tangent point determination module uses different tangent point determination rules to determine whether the position of the last byte of the data blocks of different lengths in the first initial data set is a tangent point, and obtains a first tangent point set. And the multi-strategy tangent point determination module compares: whether the tangent point features of each tangent point in the first tangent point set match the tangent point features in the stored tangent point feature table. If there is a match, the computing device directly divides the data to be stored to obtain data blocks with lengths equal to the cut-off distances in the stored cut-off feature table, without sliding the window byte by byte to divide the data to be stored to obtain new data blocks, thus achieving redundant calculation. The computing device also adjusts the degree of parallelism according to the calculation time and cache occupancy of the calculated cut-off points.
[0236] It is understandable that in order to implement the functions in the above embodiments, the computing device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and method steps of each example described in the embodiments disclosed in this application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.
[0237] Combined with the above Figures 1 to 17 , describes in detail the data processing method provided by this embodiment, and will now be combined with Fig.18 , describing the data processing device provided according to this embodiment.
[0238] Fig.18 The present invention provides a schematic diagram of the structure of a data processing device provided in the application. The data processing device can be used to implement the functions of the processor in the above method embodiment, and thus can also achieve the beneficial effects of the above method embodiment. Fig.18 As shown, the data processing device 1800 includes a cut point determination module 1810, an adjustment module 1820, a processing module 1830 and a storage module 1840. The data processing device 1800 is used to implement the above Figure 7The function of the computing device in the method embodiment shown in . The cut point determination module 1810 can be used to implement the functions of S710 and S730 in the above method embodiment, the adjustment module 1820 can be used to implement the function of S720 in the above method embodiment, and the storage module 1840 can be used to store the data after data deduplication is performed on the data to be stored.
[0239] The tangent point determination module 1810 is used to: obtain a first tangent point set of the first segment of data in the data to be stored according to the first parallelism, and divide the first segment of data according to at least one tangent point in the first tangent point set to obtain a first data set. The first parallelism is used to indicate M sliding windows. The adjustment module 1820 is used to: adjust the first parallelism according to at least one of the calculation time consumption and the cache occupancy rate of the first tangent point set to obtain a second parallelism. The first cache occupancy rate is used to indicate the cache occupancy rate of performing data deduplication on the first segment of data according to the first tangent point set. The second parallelism is used to indicate N sliding windows, M is not equal to N, and M and N are integers greater than or equal to 2. The tangent point determination module 1810 is also used to: obtain a second tangent point set of the second segment of data in the data to be stored according to the second parallelism, and divide the second segment of data according to at least one tangent point in the second tangent point set to obtain a second data set. The processing module 1830 is used to: perform data deduplication on the first data set and the second data set. The first data set includes at least one data block, and the second data set includes at least one data block.
[0240] In a possible scenario, the tangent point determination module 1810 is specifically configured to: slide M sliding windows on the data to be stored at least once to obtain a first tangent point set of the first segment of data.
[0241] In another possible scenario, the tangent point determination module 1810 is specifically used to divide the first segment of data according to M sliding windows to obtain a first initial data set. And the tangent point determination module 1810 is also specifically used to obtain a first tangent point set according to a tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution feature.
[0242] In another possible scenario, the tangent point determination module 1810 is specifically used to obtain a first tangent point set according to a tangent point determination rule associated with a length interval determined by a length of a data block in a first initial data set and a maximum value, a minimum value and an expected value of the length.
[0243] In another possible scenario, the tangent point determination module 1810 is specifically configured to: divide the first segment of data into M sliding windows with an interval of at least 1 byte to obtain a first initial data set.
[0244] In another possible scenario, the lengths of the M sliding windows are the same.
[0245] In another possible scenario, the calculation time of the first tangent point set includes: one of: a first duration, a second duration, or the first duration and the second duration.
[0246] The first duration is used to indicate the duration required to divide the first segment of data into M sliding windows to obtain the first initial data set. The second duration is used to indicate the duration required to obtain the first tangent point set according to the tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution characteristics.
[0247] In another possible scenario, the adjustment module 1820 is specifically configured to reduce the M sliding windows to N sliding windows when at least one of the calculation time consumption and the cache occupancy rate of the first tangent point set is greater than the respective upper limits.
[0248] In another possible scenario, the adjustment module 1820 is specifically configured to: increase the M sliding windows to N sliding windows when at least one of the calculation time consumption and the cache occupancy rate of the first tangent point set is less than the respective lower limit values, wherein M is less than N.
[0249] In another possible scenario, the processing module 1830 is specifically configured to: delete a data block that matches the stored data in the at least one data block included in the first data set, and store a data block that does not match the stored data in the at least one data block included in the first data set. And the processing module 1830 is further specifically configured to: delete a data block that matches the stored data in the at least one data block included in the second data set, and store a data block that does not match the stored data in the at least one data block included in the second data set.
[0250] In another possible scenario, the storage module 1840 is used to store stored data. The stored data includes a second data block and a third data block. At least one data block included in the first data set matches the second data block. The tangent point partitioning module is further used to: divide the third segment of the data to be stored by using the length of the third data block to obtain a fourth data block.
[0251] In another possible scenario, the first data set includes a fifth data block, a sixth data block, and a seventh data block. The lengths of the fifth data block and the sixth data block are the same as the maximum lengths. The first tangent point divides the data to be stored to obtain the fifth data block, and the second tangent point divides the data to be stored to obtain the seventh data block. Processing module 1830 is also used to: generate a tangent point feature table for the fifth data block and the sixth data block. The tangent point feature table includes: a tangent point feature and a tangent point distance. The tangent point feature is used to indicate the fifth data block and the sixth data block. The tangent point distance is used to indicate the spacing between the second tangent point and the first tangent point.
[0252] Among them, the cut-point determination module 1810, the adjustment module 1820, the processing module 1830 and the storage module 1840 can all be implemented by software, or can be implemented by hardware. Exemplarily, the implementation of the cut-point determination module 1810 is introduced below by taking the cut-point determination module 1810 as an example. Similarly, the implementation of the adjustment module 1820, the processing module 1830 and the storage module 1840 can refer to the implementation of the cut-point determination module 1810.
[0253] As an example of a software functional unit, the tangent point determination module 1810 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the tangent point determination module 1810 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code can be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code can be distributed in the same availability zone (AZ) or in different AZs, each AZ including one data center or multiple data centers with close geographical locations. Among them, usually a region can include multiple AZs.
[0254] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.
[0255] As an example of a hardware functional unit, the cut-point determination module 1810 may include at least one computing device, such as a server, etc. Alternatively, the cut-point determination module 1810 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0256] The multiple computing devices included in the cut-point determination module 1810 can be distributed in the same region or in different regions. The multiple computing devices included in the cut-point determination module 1810 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the cut-point determination module 1810 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0257] It should be noted that, in other embodiments, the cut-point determination module 1810 can be used to execute any step in the data processing method, the adjustment module 1820 can be used to execute any step in the data processing method, and the storage module 1840 can be used to execute any step in the data processing method. The steps that the cut-point determination module 1810, the adjustment module 1820, the processing module 1830 and the storage module 1840 are responsible for implementing can be specified as needed. The cut-point determination module 1810, the adjustment module 1820, the processing module 1830 and the storage module 1840 respectively implement different steps in the data processing method to realize the full functions of the data processing device.
[0258] Fig.19 A schematic diagram of a computing device provided in this application, such as Fig.19 As shown, the computing device 1900 includes: a processor 1910, a bus 1920, a memory 1930, and a communication interface 1940. The processor 1910, the memory 1930, and the communication interface 1940 communicate with each other through the bus 1920. The computing device 1900 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 1900.
[0259] The bus 1920 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.19 The bus 104 is represented by only one line, but does not mean that there is only one bus or one type of bus. The bus 104 may include a path for transmitting information between various components of the computing device 1900 (eg, the memory 1930, the processor 1910, and the communication interface 1940).
[0260] The processor 1910 may be used to obtain the first data set and the second data set. The processor 1910 may include any one or more of a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0261] The memory 1930 may include a volatile memory, such as a random access memory (RAM). The processor 1910 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0262] The memory 1930 can store executable program codes, and the processor 1910 executes the executable program codes to respectively implement the functions of the aforementioned processing module, the cut-point determination module, the adjustment module, and the storage module, thereby implementing the data processing method. That is, the memory 1930 stores instructions for executing the data processing method. The memory 1930 can also store data after data deduplication is performed.
[0263] Alternatively, the memory 1930 stores executable codes, and the processor 1910 executes the executable codes to respectively implement the functions of the aforementioned data processing device, thereby implementing the data processing method. That is, the memory 1930 stores instructions for executing the data processing method.
[0264] The communication interface 1940 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 1900 and other devices or communication networks, such as receiving data to be stored sent by a client.
[0265] The present application also provides a cluster. The cluster includes at least one computing device. The computing device may be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device may also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0266] Fig. 20 A schematic diagram of a cluster provided for this application, such as Fig. 20 As shown, the cluster includes at least one computing device 1900. The memory 1930 in one or more computing devices 1900 in the cluster may store the same instructions for executing the data processing method.
[0267] In some possible implementations, the memory 1930 of one or more computing devices 1900 in the cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 1900 may jointly execute instructions for executing the data processing method.
[0268] It should be noted that the memory 1930 in different computing devices 1900 in the cluster can store different instructions, which are respectively used to execute part of the functions of the data processing device. That is, the instructions stored in the memory 1930 in different computing devices 1900 can implement the functions of one or more modules among the processing module, the cut point determination module, the adjustment module and the storage module.
[0269] In some possible implementations, one or more computing devices in the cluster may be connected via a network, which may be a wide area network or a local area network. Fig.21 A possible implementation is shown. Fig.21 A connection diagram of a computing device provided in this application, such as Fig.21 As shown, two computing devices 1900A and 1900B are connected via a network. Specifically, the network is connected via a communication interface in each computing device. In this type of possible implementation, the memory 1930 in the computing device 1900A stores instructions for executing the functions of the processing module. At the same time, the memory 1930 in the computing device 1900B stores instructions for executing the functions of the tangent point determination module, the adjustment module, and the storage module.
[0270] Fig.21The connection mode between the clusters shown may be considered to be that the data processing method provided by the present application requires a large amount of partitioning of the data to be stored and calculation of the cut point set. Therefore, it is considered that the functions implemented by the adjustment module and the storage module are handed over to the computing device 1900B for execution.
[0271] It should be understood that Fig.21 The functions of the computing device 1900A shown in FIG. 1 may also be completed by multiple computing devices 1900. Similarly, the functions of the computing device 1900B may also be completed by multiple computing devices 1900.
[0272] The present application embodiment also provides another cluster. The connection relationship between the computing devices in the cluster can be similarly referred to as Fig. 20 and Fig.21 The connection mode of the cluster is different in that the memory 1930 in one or more computing devices 1900 in the cluster may store the same instructions for executing the data processing method.
[0273] In some possible implementations, the memory 1930 of one or more computing devices 1900 in the cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 1900 may jointly execute instructions for executing the data processing method.
[0274] The embodiment of the present application also provides a computer program product including instructions. The computer program product may be software or a program product including instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the at least one computing device executes a data processing method or a data processing method.
[0275] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state hard disk). The computer-readable storage medium includes instructions that instruct a computing device to execute a data processing method, or instructs a computing device to execute a data processing method.
[0276] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.
Claims
1. A data processing method, It is characterized in that The method comprises: Acquire a first tangent point set of a first segment of data in the data to be stored according to a first degree of parallelism, and divide the first segment of data according to at least one tangent point in the first tangent point set to obtain a first data set, wherein the first degree of parallelism is used to indicate M sliding windows; The first degree of parallelism is adjusted according to at least one of the computation time and the cache occupancy of the first tangent point set to obtain a second degree of parallelism, wherein the first cache occupancy is used to indicate the cache occupancy of performing data deduplication on the first segment of data according to the first tangent point set; the second degree of parallelism is used to indicate N sliding windows, where M is not equal to N, and M and N are integers greater than or equal to 2; Acquire a second tangent point set for a second segment of data in the data to be stored according to the second degree of parallelism, wherein at least one tangent point in the second tangent point set divides the second segment of data to obtain a second data set; Data deduplication is performed on the first data set and the second data set, the first data set includes at least one data block, and the second data set includes at least one data block.
2. The method according to claim 1, It is characterized in that The M sliding windows slide at least once on the data to be stored to obtain a first tangent point set of the first segment of data.
3. The method according to claim 1 or 2, It is characterized in that The step of obtaining a first tangent point set of a first segment of data in the data to be stored according to the first degree of parallelism includes: Divide the first segment of data according to the M sliding windows to obtain a first initial data set; A first tangent point set is obtained according to a tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution feature.
4. The method according to claim 3, It is characterized in that The first tangent point set is obtained according to the tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution feature, including: A first tangent point set is obtained according to a tangent point determination rule associated with a length interval determined by a length of a data block in the first initial data set and a maximum value, a minimum value and an expected value of the length.
5. The method according to claim 3 or 4, It is characterized in that The dividing the first segment of data according to the M sliding windows to obtain a first initial data set includes: The M sliding windows divide the first segment of data at intervals of at least 1 byte to obtain a first initial data set.
6. The method according to any one of claims 1 to 5, It is characterized in that The lengths of the M sliding windows are consistent.
7. The method according to any one of claims 3 to 6, It is characterized in that The time consumption of calculating the first tangent point set includes: First duration; or, the second duration; or, the sum of the first duration and the second duration; Among them, the first time length is used to indicate: the time length required for the M sliding windows to divide the first segment of data into a first initial data set; the second time length is used to indicate: the time length required for obtaining the first tangent point set based on a tangent point determination rule associated with the length of the data block in the first initial data set and the length distribution characteristics.
8. The method according to any one of claims 1 to 7, It is characterized in that The adjusting the first parallelism according to at least one of the computation time consumption and the cache occupancy rate of the first tangent point set to obtain the second parallelism includes: When at least one of the computation time and the cache occupancy rate of the first tangent point set is greater than the respective upper limit values, the M sliding windows are reduced to N sliding windows, where M is greater than N.
9. The method according to any one of claims 1 to 7, It is characterized in that The adjusting the first parallelism according to at least one of the computation time consumption and the cache occupancy rate of the first tangent point set to obtain the second parallelism includes: When at least one of the computation time and the cache occupancy rate of the first tangent point set is less than the respective lower limit value, the M sliding windows are increased to N sliding windows, where M is less than N.
10. The method according to any one of claims 1 to 9, It is characterized in that The performing data deduplication on the first data set and the second data set includes: deleting a data block that matches the stored data in at least one data block included in the first data set, and storing a data block that does not match the stored data in at least one data block included in the first data set; A data block matching the stored data in at least one data block included in the second data set is deleted, and a data block not matching the stored data in at least one data block included in the second data set is stored.
11. The method according to any one of claims 1 to 10, It is characterized in that The stored data includes a second data block and a third data block, at least one data block included in the first data set matches the second data block, The method further comprises: The third segment of the data to be stored is divided using the length of the third data block to obtain a fourth data block.
12. The method according to any one of claims 1 to 10, It is characterized in that The first data set includes a fifth data block, a sixth data block and a seventh data block, the lengths of the fifth data block and the sixth data block are the same and the maximum length is the same, the first tangent point divides the data to be stored to obtain the fifth data block, and the second tangent point divides the data to be stored to obtain the seventh data block, The method further comprises: Generate a tangent point feature table for the fifth data block and the sixth data block; the tangent point feature table includes: a tangent point feature and a tangent point distance, the tangent point feature is used to indicate the fifth data block and the sixth data block, and the tangent point distance is used to indicate the distance between the second tangent point and the first tangent point.
13. A data processing device, It is characterized in that The device comprises: a tangent point determination module, configured to: obtain a first tangent point set of a first segment of data in the data to be stored according to a first degree of parallelism, and divide the first segment of data according to at least one tangent point in the first tangent point set to obtain a first data set, wherein the first degree of parallelism is used to indicate M sliding windows; an adjustment module, configured to: adjust the first degree of parallelism according to at least one of a calculation time consumption and a cache occupancy rate of the first tangent point set to obtain a second degree of parallelism, wherein the first cache occupancy rate is used to indicate a cache occupancy rate of performing data deduplication on the first segment of data according to the first tangent point set; and the second degree of parallelism is used to indicate N sliding windows, where M is not equal to N, and M and N are integers greater than or equal to 2; The tangent point determination module is further used to: obtain a second tangent point set for a second segment of data in the data to be stored according to the second degree of parallelism, wherein at least one tangent point in the second tangent point set divides the second segment of data to obtain a second data set; The processing module is further used to: perform data deduplication on the first data set and the second data set, the first data set includes at least one data block, and the second data set includes at least one data block.
14. A computer device, It is characterized in that The method comprises a memory and a processor, wherein the memory is used to store a group of computer instructions; when the processor executes the group of computer instructions, the processor executes the method described in any one of claims 1 to 12.
15. A cluster, It is characterized in that The cluster includes a computing node and a storage node. The computing node executes the method according to any one of claims 1 to 12. The storage node is used to store data after the computing node performs data deduplication on the data to be stored.
Citation Information
Patent Citations
Method and system for concurrent blocking for data deduplication process
CN104361068A
Data processing method and device
CN115599591A
Hard disk scanning method and device
CN115857793A
Systems and methods for throttling packet transmission in a scalable memory system protocol
US20150350082A1
De-duplication system and method thereof
US20150356134A1
Cited By
Data processing method and device, medium and program product
CN120723170A
Data processing method, device, medium and program product
CN120723170B