Data processing method, electronic device and computer program product
By dynamically selecting the data processing mode based on the data segment size and compression level, the problem of low efficiency in a single mode in existing technologies is solved, and more efficient data processing and storage system performance improvement are achieved.
Patent Information
- Application Number
- CN202110090766.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-22
- Publication Date
- 2025-12-19
- Estimated Expiration
- 2041-06-21
AI Technical Summary
In existing technologies, a single data processing mode is inefficient in data processing with different deduplication rates, becoming a bottleneck in system performance and failing to effectively utilize the potential of the storage system.
The data processing mode is dynamically selected based on the size of the data segment and the compression level. Non-repeating data segments are identified through matching operations, and combined with compression operations, an appropriate mode is selected to reduce processing time.
By dynamically selecting data processing modes, data processing efficiency is improved, storage space utilization is maximized, and the performance of the storage system and user experience are enhanced.
Smart Images

Figure CN114780501B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the field of computers, and more specifically, to a data processing method, an electronic device, and a computer program product. BACKGROUND
[0002] In the era of big data, the contradiction between the existence of massive data and the computing cost of limited storage systems raises the demand for reducing the processing cost of data. It can be understood that in the process of processing data, for example, when data is backed up to a storage system, it is necessary to perform deduplication processing and compression processing on the data. There are different data processing modes, such as performing a deduplication operation first and then performing data compression, or performing deduplication and compression operations together. For different types of data and different data deduplication rates, applying different modes can require processing costs (such as time costs), and there is a large cost difference therebetween. SUMMARY
[0003] Embodiments of the present disclosure provide a scheme for data processing.
[0004] In a first aspect of the present disclosure, a data processing method is provided, the method comprising determining, based on sizes of a plurality of data segments included in to-be-processed data, a first time required for performing a matching operation on each data segment, the matching operation being used to determine non-repeated data segments; determining, based on the size of each data segment and a compression level for the to-be-processed data, a second time required for performing a compression operation on each data segment; and determining, based on the first time, the second time, and a deduplication rate for the to-be-processed data, a target mode for processing the plurality of data segments from among a first mode and a second mode, in the first mode, the compression operation is performed only on the non-repeated data segments among the plurality of data segments, and in the second mode, the compression operation is performed on each data segment among the plurality of data segments.
[0005] In a second aspect of the present disclosure, an electronic device is provided, comprising a processor; and a memory coupled with the processor, the memory having stored therein instructions which, when executed by the processor, cause the electronic device to perform actions, the actions comprising: determining, based on sizes of a plurality of data segments included in to-be-processed data, a first time required for performing a matching operation on each data segment, the matching operation being used to determine non-repeated data segments; determining, based on the size of each data segment and a compression level for the to-be-processed data, a second time required for performing a compression operation on each data segment; and determining, based on the first time, the second time, and a deduplication rate for the to-be-processed data, a target mode for processing the plurality of data segments from among a first mode and a second mode, in the first mode, the compression operation is performed only on the non-repeated data segments among the plurality of data segments, and in the second mode, the compression operation is performed on each data segment among the plurality of data segments.
[0006] In a third aspect of the present disclosure, a computer program product is provided, the computer program product being tangibly stored on a computer readable medium and comprising machine executable instructions that, when executed, cause a machine to perform any of the steps of the method according to the first aspect.
[0007] The summary is provided to introduce a selection of concepts, in a simplified form, that are further described below in the detailed description. This summary is not intended to identify key features or essential features of the disclosure, and is not intended to limit the scope of the disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0008] The above and other objects, features and advantages of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings in which like reference characters refer to like parts throughout the figures. In the drawings:
[0009] Figure 1 A schematic diagram illustrating an example environment according to embodiments of the present disclosure is shown;
[0010] Figure 2 A schematic diagram illustrating various charts of various metrics related to size of data segments, compression levels is shown;
[0011] Figure 3 A flowchart illustrating a process of data processing according to embodiments of the present disclosure is shown;
[0012] Figure 4 A flowchart illustrating a process of determining a target mode according to embodiments of the present disclosure is shown; and
[0013] Figure 5 A block diagram illustrating an example device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0014] The principles of the present disclosure will now be described, by way of example only, with reference to the drawings in which several example embodiments are illustrated by way of example. The following detailed description is therefore not to be taken in a limiting sense.
[0015] The terms "including", "includes", "included", "include" and variations thereof are used synonymously with the terms "comprising", "comprises", "comprised" and "comprise" and are open-ended, i.e. they mean "including, but not limited to". The term "or" means "and / or" unless expressly stated otherwise. The term "based on" means "based, at least in part, on". The terms "one example embodiment" and "an example embodiment" mean "one of a group of example embodiments". The term "another embodiment" means "at least one of a group of other example embodiments". The terms "first", "second", and the like, can refer to different or identical objects. Further below, other explicit and implicit definitions can be included.
[0016] As discussed above, various data processing modes exist in storage systems. In the first mode, all data segments are first matched, and then unique data is compressed. In the second mode, both matching and compression are performed on all data segments. The time required to perform matching and compression on a single data segment in the second mode can be considered equal to the compression time in the first mode. Current solutions employ a single processing mode for all data types. However, for data with low deduplication rates, the time required for the second mode is significantly less than that for the first mode. Conversely, the time required for the first mode is significantly less than that for the second. The inefficiency of using a single data processing mode becomes a bottleneck limiting system performance.
[0017] To at least partially address the aforementioned drawbacks, embodiments of this disclosure provide a data processing scheme. In this scheme, a first time required for a matching operation is first determined based on the size of data segments in the data to be processed; the matching operation is used to identify non-repeating data segments. Then, a second time required for a compression operation is determined based on this size and the compression level for the data to be processed. Finally, a target pattern for the data to be processed is determined from a first pattern and a second pattern based on the aforementioned first and second times and the deduplication rate for the data to be processed. Therefore, this scheme can select a suitable data processing pattern based on the characteristics of the data, thereby reducing data processing time and improving data processing efficiency.
[0018] Figure 1 A schematic diagram of an exemplary environment 100 according to an embodiment of the present disclosure is shown, in which devices and / or methods according to embodiments of the present disclosure may be implemented. Figure 1 As shown, Figure 1 As shown, the exemplary environment may include storage system 150. Storage system 150 may include computing device 105 for handling various operations for data storage, including but not limited to data compression and decompression, data deduplication, data storage, data backup and recovery.
[0019] Storage system 150 may include (but not shown) storage disks for storing data. Storage disks may be various types of storage devices, including but not limited to hard disks (HDDs), solid-state drives (SSDs), removable disks, any other magnetic storage devices, and any other optical storage devices, or any combination thereof.
[0020] The computing device 105 can be configured to obtain a data segment size included in the data to be processed 110, a deduplication rate for the data to be processed 110, a compression level for the data to be processed 110, and the like.
[0021] The computing device 105 can be configured to compress the data to be processed 110 to obtain the compressed data 130. The compressed data 130 can be stored in the storage disk to save storage space of the storage disk.
[0022] In some embodiments, the storage system 150 can be a storage system for data backup, which is configured with a deduplication (sometimes also referred to as de-duplication or data deduplication) device to remove the duplicated parts of the data and only store the non-duplicated parts, thereby achieving efficient use of storage space. In some embodiments, the storage system can use a suitable coprocessor to perform a matching operation on the data, such as a SHA1 operation on the data. The storage system can use the SHA1 operation to obtain a fingerprint of the data to be processed, and match the fingerprint with the fingerprints of the existing data in the storage space to determine the non-duplicated data. In some embodiments, the storage system can use various compression techniques to perform a compression operation on the data. In some embodiments, the storage system can be configured to use various compression levels provided by various compression techniques to perform a compression operation on the data.
[0023] One example of the above-mentioned suitable coprocessor is a Quick Assist Technology (QAT) card, which can be used to accelerate compute-intensive tasks such as compression and encryption. By adding a QAT card to the storage system, the running of the application program can be accelerated, and the performance and efficiency of the storage system can be improved. The functions provided by the QAT card can include symmetric encryption, authentication, asymmetric encryption, digital signature, public key encryption, lossless data compression, and the like. In some cases, the computing device 105 can process the data using a first mode 120. For example, in the first mode 120, the computing device 105 can separately perform a data matching operation (such as a SHA1 operation, or an HMAC SHA1 operation) and a compression operation on the data via a Quick Assist Technology (QAT) card. In other cases, the computing device 105 can process the data using a second mode 130. For example, in the second mode 130, the computing device 105 can perform a data matching operation (such as a SHA1 operation) and a compression operation on the data in combination via a Quick Assist Technology (QAT) card (which can be collectively referred to as a "SHA1-compression chained operation"). In the following description, we consider that the time required to perform the data matching operation and the compression operation in combination for the same data segment is equal to the time required to perform the compression operation separately for the same data segment (in reality, the difference is extremely small, i.e., less than a threshold value, which can be ignored in the operation of the storage system).
[0024] It should be noted that the QAT and SHA1 operations described above are merely exemplary, and other suitable processors and algorithms can also be applied, and the present disclosure is not limited in this regard.
[0025] It can be appreciated that in some cases, such as the 0th generation backup case, the deduplication rate of the data is 0. In this case, it is more suitable to apply a data processing mode that combines the matching operation and the compression operation. In other cases, such as when the deduplication rate of the data is greater than 90%, it is more suitable to first apply the matching operation for deduplication, and then apply the compression operation for data processing. Moreover, due to various characteristics of the data, such as the size of the data segment, the deduplication rate of the data, and the compression level of the data, the processing time of the data will also be affected. Therefore, the storage system 150 (e.g., the computing device 105 of the storage system) can dynamically determine which data processing mode to apply according to various characteristics of the data 110 to be processed.
[0026] The processes according to embodiments of the present disclosure will be described in detail below in conjunction with Figure 2 , Figure 3 and Figure 4 . For ease of understanding, the specific data mentioned in the following description are all exemplary and are not intended to limit the scope of protection of the present disclosure. It can be appreciated that the embodiments described below can also include additional actions not shown or can omit the actions shown, and the scope of the present disclosure is not limited in this regard.
[0027] Figure 2 A schematic diagram 200 showing various graphs of data processing time related to the compression level and / or the size of the data segment is shown. It should be noted that Figure 2 only graphs of various indicators corresponding to various compression levels provided by the QAT compression technology in one hardware configuration are shown. It can be appreciated that similar graphs can be obtained by those skilled in the art through testing of the storage system in the case of taking different other hardware configurations and / or taking other compression technologies.
[0028] The graph 210 shows the time required for the matching operation for different sizes of data segments according to embodiments of the present disclosure. The matching operation can be the HMAC SHA1 algorithm in the SHA1 algorithm in the storage system. The SHA1 algorithm and the values in the graph are merely exemplary, and suitable other matching operations can also be applied to determine the non-duplicate data segment.
[0029] The graph 220 shows the time required for the compression operation for different sizes of data segments at various compression levels according to embodiments of the present disclosure. Compression and / or decompression can be divided into two types, dynamic type and static type, which can refer to dynamic huffman data compression and / or decompression, and static huffman data compression and / or decompression, respectively.
[0030] For example, the QAT compression technique can provide dynamic compression level 1 to dynamic compression level 4 (sometimes also referred to herein as dynamic levels), and static compression level 1 to static compression level 4 (sometimes also referred to herein as static levels). At different compression levels, the compression time required is different. Additionally, the size of the data segment, for example, 1 KB, 4 KB, 8 KB, 16 KB, 64 KB, can also affect the throughput.
[0031] The chart 230 shows the ratio of the time required for the matching operation to the time required for the compression operation (time required for matching operation / time required for compression operation) for different sizes of data segments at various compression levels, according to an embodiment of the present disclosure. The time required for the matching operation and the time required for the compression operation can be derived from the chart 210 and the chart 220.
[0032] It can be appreciated that the exact values of various metrics similar to those shown in the above charts can vary, but the relationships between them are similar to those described above with reference to the above charts, in case different other hardware configurations are taken and / or other matching operations and compression operations are taken.
[0033] Figure 3 A flowchart of a process 300 of data processing is shown, according to an embodiment of the present disclosure. The process 300 can be implemented at the computing device 105 shown in Figure 1 FIG. 1.
[0034] At 310, the computing device 105 determines, based on the sizes of the plurality of data segments included in the data to be processed 110, a first time required to perform a matching operation for each data segment, the matching operation being used to determine non-duplicate data segments.
[0035] In particular, the data to be processed 110 is data that is expected to be processed with various operations (techniques or algorithms), for example. In some embodiments, the data to be processed can be data to be stored (e.g., to be backed up) in a storage system. In some embodiments, the data to be processed can also be data obtained after the data to be stored is processed with data deduplication, which can be stored into a storage disk after compression for subsequent retrieval.
[0036] Alternatively, in some embodiments, the stored data that has already been stored in a storage disk, for example, in case of performing data reclamation processing such as garbage collection, is also expected to be processed with various compression techniques or algorithms. In this case, the data to be processed can also be data to be reclaimed.
[0037] The data to be processed 110 can be in the form of a data stream, which can include a plurality of data segments, the size of the plurality of data segments can be obtained in various ways. In some embodiments, the computing device 105 can determine the size of the data segments by utilizing various monitors of the storage system 150. For example, a data segment size monitor can be utilized to monitor the size of the data segments in real time, and additionally or alternatively, such parameters can be utilized to compute the size of the plurality of data segments included in the data to be processed 110. Alternatively, in some embodiments, the computing device 105 can directly configure the size of the plurality of data segments included in the data to be processed 110.
[0038] The computing device 105 can perform a matching operation on each of the plurality of data segments to determine the non-duplicate data segments. For example, the computing device 105 can perform a fingerprint matching operation (e.g., utilizing the SHA1 algorithm) on the data segments to obtain a data fingerprint for each data segment, and then compare the data fingerprint with existing data in the storage to determine the non-duplicate data segments. This is merely exemplary, and various suitable matching operations can be applied to determine the non-duplicate data segments. It can be appreciated that there is a different time for the matching operation for different sizes of the data segments. For example, as shown in the graph 210 in Figure 2
[0039] In some embodiments, the computing device 105 can determine an average value of the size of the data segments processed in a historical time period, and then determine the first time based on the average value. For example, the computing device 105 can determine that the size of the data segments processed in the past 24 hours is 16KB by the data segment monitor. The computing device 105 can then determine the first time as 26μs by the graph 210.
[0040] Alternatively, in some embodiments, the computing device 105 can receive the size of the data segments inputted by a user through, for example, a user configuration interface, and then can determine the first time by the graph 210.
[0041] At 320, the computing device 105 determines a second time required to perform a compression operation on each data segment based on the size of each data segment and the compression level for the data to be processed 110. For example, the computing device 105 can perform a compression operation on the non-duplicate data segments after determining the non-duplicate data segments.
[0042] It can be appreciated that the better the compression level, the higher the compression ratio, and thus the less storage space required for the compressed data. However, the data processing by the storage system usually needs to meet certain time requirements. In some cases, for a large amount of data to be compressed per unit of time, using, for example, the best compression level can likely result in too long a time for the storage system to process the data, and thus fail to meet the predetermined time requirement. In other cases, for a small amount of data to be compressed per unit of time, using, for example, the worst compression level can likely result in unnecessary occupation of storage space, although it can meet the predetermined time requirement.
[0043] Therefore, in some embodiments, the computing device 105 can select an optimal compression level according to the amount of data 110 to be processed or the number of non-repeated data segments determined above, so that the predetermined time requirement can be met while the compression ratio of the compressed data is the highest.
[0044] Alternatively, in some embodiments, the computing device 105 can also determine the compression level for the data 110 to be processed according to the average compression level in the historical time period. Additionally or alternatively, in some embodiments, the computing device 105 can receive a compression level input by a user through, for example, a user configuration interface.
[0045] After the computing device 105 determines the size of the data segment and the compression level, the computing device 105 can determine a second time required for performing the compression operation for each data segment. In some embodiments, the computing device 105 can obtain a compression level mapping table including a plurality of compression operation times corresponding to a plurality of sizes of data segments and a plurality of compression levels. Then the computing device 105 can determine the second time from the plurality of compression operation times based on the compression level mapping table, the size of the data segment and the compression level.
[0046] Specifically, the computing device 105 can first obtain a graph 220 in Figure 2 The graph 220 is a compression level mapping table, which can be obtained from a local database or externally, or can also be dynamically determined by the computing device 105 according to historical data. As can be seen from the graph 220, for each size of data segment and each compression level, there is a compression time. For example, for a compression operation of static level 1 and a data segment of 4 KB size, a compression time of 24 μs is required. After obtaining the compression level mapping table, the computing device 105 can determine the second time for the compression operation of the data segment to be 130 μs from the plurality of compression times in the graph 220 according to the size of the data segment and the compression level determined above (e.g., 16 KB and dynamic level 3).
[0047] The compression level and the data segment size of the data are determined by various suitable methods, which can save the cost in time and storage resources in subsequent processing. At the same time, it lays the foundation for the dynamic selection of subsequent processing modes.
[0048] At 330, the computing device 105 determines a target mode for processing the plurality of data segments from the first mode 120 and the second mode 130 based on the first time, the second time, and the deduplication rate for the data to be processed 110, in which the compression operation is only performed on the non-duplicate data segments in the plurality of data segments in the first mode 120, and in which the compression operation is performed on each data segment in the plurality of data segments in the second mode 130.
[0049] Specifically, after determining the above-mentioned first time and second time, the computing device 105 can determine whether to apply the first mode 120 or the second mode 130 according to the deduplication rate of the data, i.e., the proportion of the data that already exists in the storage system in the data to be processed.
[0050] The determination of the mode and the processing of the data in different modes will be described in detail. Figure 4 The determination of the mode and the processing of the data in different modes will be described in detail. Figure 4 A flowchart of a process 400 of determining a target mode according to an embodiment of the present disclosure is shown.
[0051] Before introducing the embodiments, we first introduce the first mode 120 and the second mode 130 for processing data. First, some parameters are defined to facilitate subsequent description. The number of data segments of the data to be processed 110 is defined as N, the deduplication rate of the data to be processed 110 is defined as R, the above-mentioned first time is defined as S (for example, the time in the graph 210), and the above-mentioned second time is defined as C (for example, the time in the graph 220). In the first mode, the computing device 105 first performs a matching operation on all data segments to determine non-duplicate data segments, and then performs a compression operation on the non-duplicate data segments. The time A required for processing the data to be processed 110 using the first mode is:
[0052] A = S * N + (1 - R) * N * C Equation (1)
[0053] In the second mode, the computing device 105 performs a matching operation and a compression operation (as described above, which can be a chain operation and is equal to the compression operation in time) on all data segments. The time B required for processing the data to be processed 110 using the second mode is:
[0054] B = C * N Equation (2)
[0055] Comparing equation (1) and equation (2), it can be concluded that when A < B, i.e. R > S / C, the first mode is used to save time. When A > B, i.e. R < S / C, the second mode is used to save time. That is, the ratio of the first time and the second time is the threshold value for determining whether to use the first mode or the second mode.
[0056] At 410, the computing device 105 determines whether the deduplication rate is greater than the ratio of the first time and the second time. The method of obtaining the deduplication rate of the to-be-processed data 110 is similar to the method of obtaining the size of the data segment of the to-be-processed data 110 and the compression level, i.e. through historical data or receiving configuration data, which will not be described here. The computing device 105 can determine the first time and the second time according to the method of the above charts 210 and 220, and then compare the deduplication rate with the ratio.
[0057] In some embodiments, the computing device 105 can obtain a predetermined chart 230 from a database, in which the ratio of the first time and the second time S / C corresponding to each compression level and each data segment size is stored. After determining the deduplication rate, the size of the data segment and the compression level of the to-be-processed data 110, the size relationship between the deduplication rate and the response ratio can be directly determined according to the chart 230.
[0058] At 420, the computing device 105 determines that the deduplication rate is greater than the ratio of the first time and the second time, and determines that the target mode is the first mode. In the case where the computing device 105 determines that the target mode is the first mode. The computing device 105 can perform a matching operation on each data segment. Then the computing device 105 can determine a set of to-be-compressed data segments from the plurality of data segments based on the result of the above matching operation. For example, the non-duplicate data segment is determined as a set of to-be-compressed data segments. Then the computing device 105 performs a compression operation on the set of to-be-compressed data segments according to the compression level.
[0059] At 430, the computing device 105 determines that the deduplication rate is less than the ratio of the first time and the second time, and determines that the target mode is the second mode. In the case where the computing device 105 determines that the target mode is the second mode. The computing device 105 performs a matching operation and a compression operation on each data segment in the plurality of data segments according to the compression level, for example, the chain operation described above. It can be understood that in the case where the deduplication rate is equal to the above ratio. Either of the first mode and the second mode can be applied.
[0060] According to the scheme of the present disclosure, the processing mode with the lowest processing time cost can be dynamically selected according to various characteristics of the data, i.e., the size of the data segment, the compression level, and the deduplication rate. By applying the second mode (chain processing) in the case of low deduplication rate or the first mode (independent processing) in the case of high deduplication rate, the data processing amount (throughput) per unit time can be significantly improved. The scheme can maximize the performance of the storage space to improve the data processing speed, thereby enhancing the user's experience of using the storage system applying the scheme.
[0061] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. For example, the electronic device 500 can be used to implement the computing device 105 shown in FIG. 1. Figure 1 As shown, the device 500 includes a central processing unit (CPU) 501 that can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 502 or computer program instructions loaded into a random access memory (RAM) 503 from a storage unit 508. Various programs and data required for operations of the device 500 can also be stored in the RAM 503. The CPU 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0062] Various components in the device 500 are connected to the I / O interface 505, including an input unit 506, e.g., a keyboard, a mouse, etc., an output unit 507, e.g., various types of displays, speakers, etc., the storage unit 508, e.g., a magnetic disk, an optical disk, etc., and a communication unit 509, e.g., a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0063] The processing unit 501 performs various methods and processes described above, such as any of the processes 300-400. For example, in some embodiments, any of the processes 300-400 can be implemented as a computer software program or computer program product tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, portions or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded onto the RAM 503 and executed by the CPU 501, one or more steps of any of the processes 300-400 described above can be performed. Alternatively, in other embodiments, the CPU 501 can be configured to perform any of the processes 300-400 by other means, such as by way of firmware.
[0064] The present disclosure can be a method, apparatus, system, and / or computer program product. The computer program product can include a computer-readable storage medium (or media) having computer readable program instructions stored therein.
[0065] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, a non-transitory storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch cards or punched tape, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.
[0066] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0067] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0068] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0069] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data storage cycles that change state. The instructions can be executed by one or more processors of a computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions which execute via the one or more processors of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0070] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data storage cycles that change state. The instructions can be executed by one or more processors of a computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions which execute via the one or more processors of the computer or other programmable data processing apparatus create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0071] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0072] The above-described implementations of the disclosure are merely meant to be illustrative and not exhaustive, and are not limited to the implementations disclosed. Many modifications and variations of the described implementations are possible, without departing from the scope and spirit of the described implementations. The selection of terms herein is intended to best explain the principles of the implementations, the practical application, or technical improvement over the technology found in the market, or to enable others skilled in the art to understand the implementations disclosed herein.
Claims
1. A method for data processing in a storage system, comprising: determining, by a data segment size monitor of the storage system, sizes of a plurality of data segments included in to-be-processed data, wherein the to-be-processed data is to-be-backed-up data in the storage system; determining, based on the sizes, a first time required for performing a matching operation for each data segment, the matching operation being used to determine non-duplicate data segments; determining, based on the size of each data segment and a compression level for the to-be-processed data, a second time required for performing a compression operation for each data segment, wherein different compression levels correspond to different compression rates, different storage spaces required for compressed data, and different time requirements for processing; determining, based on the first time, the second time, and a deduplication rate for the to-be-processed data, a target mode for processing the plurality of data segments from a first mode and a second mode, wherein in the first mode, the compression operation is performed only on the non-duplicate data segments of the plurality of data segments based on a time required for processing the data using the first mode, the first time, the second time, and the deduplication rate for the to-be-processed data, and wherein in the second mode, the compression operation is performed on each data segment of the plurality of data segments; and performing the matching operation and the compression operation based on the target mode by a fast auxiliary technology card to back up the to-be-processed data; wherein in response to the target mode being the first mode, performing the matching operation comprises: performing a fingerprint matching operation on the plurality of data segments to obtain a plurality of data fingerprints of the plurality of data segments; and comparing the plurality of data fingerprints with existing data in a storage of the storage system to determine the non-duplicate data segments. 2.The method of claim 1, wherein determining the target mode comprises: determining the target mode to be the first mode if it is determined that the deduplication rate is greater than a ratio of the first time to the second time; and determining the target mode to be the second mode if it is determined that the deduplication rate is less than the ratio of the first time to the second time. 3.The method of claim 1, wherein the target mode is the first mode, the method further comprising: performing the compression operation on the non-duplicate data segments according to the compression level. 4.The method of claim 1, wherein the target mode is the second mode, the method further comprising: performing the matching operation and the compression operation on each data segment of the plurality of data segments according to the compression level. 5.The method of claim 1, wherein determining the first time comprises: determining an average value of sizes of data segments processed in a historical time period; and determining the first time based on the average value. 6.The method of claim 1, wherein determining the second time comprises: obtaining a compression level mapping table, the compression level mapping table comprising a plurality of compression operation times corresponding to a plurality of sizes of data segments and a plurality of compression levels; and determining the second time based on the compression level mapping table. determining the second time from the plurality of compression operation times based on the compression level mapping table, the size of the data segment, and the compression level.
7. A storage system comprising: a data segment size monitor; a fast auxiliary technology card; a processor; and a memory coupled with the processor, the memory having stored therein instructions which, when executed by the processor, cause the storage system to perform acts comprising: causing the data segment size monitor to determine sizes of a plurality of data segments included in data to be processed, wherein the data to be processed is data to be backed up in the storage system; determining a first time required to perform a matching operation for each data segment based on the sizes, the matching operation being used to determine non-duplicate data segments; determining a second time required to perform a compression operation for each data segment based on the size of each data segment and a compression level for the data to be processed, wherein different compression levels correspond to different compression rates, different storage space required for compressed data, and different time requirements for processing; determining a target mode for processing the plurality of data segments from a first mode and a second mode based on the first time, the second time, and a deduplication rate for the data to be processed, wherein in the first mode, the compression operation is performed only for the non-duplicate data segments of the plurality of data segments based on a time required to process the data using the first mode, the first time, the second time, and the deduplication rate for the data to be processed, and wherein in the second mode, the compression operation is performed for each data segment of the plurality of data segments; and performing the matching operation and the compression operation based on the target mode via the fast auxiliary technology card to back up the data to be processed; wherein in response to the target mode being the first mode, performing the matching operation comprises: performing a fingerprint matching operation on the plurality of data segments to obtain a plurality of data fingerprints for the plurality of data segments; and comparing the plurality of data fingerprints with existing data in a memory of the storage system to determine the non-duplicate data segments.
8. The storage system of claim 7, wherein determining the target mode comprises: determining the target mode to be the first mode if it is determined that the deduplication rate is greater than a ratio of the first time to the second time; and determining the target mode to be the second mode if it is determined that the deduplication rate is less than the ratio of the first time to the second time.
9. The storage system of claim 7, wherein the target mode is the first mode, the acts further comprising: performing the compression operation on the non-duplicate data segments according to the compression level.
10. The storage system of claim 7, wherein the target mode is the second mode, the acts further comprising: performing the matching operation and the compression operation on each data segment of the plurality of data segments according to the compression level.
11. The storage system of claim 7, wherein determining the first time comprises: determining an average of sizes of data segments processed in a historical time period; and determining the first time based on the average.
12. The storage system of claim 7, wherein determining the second time comprises: obtaining a compression level mapping table comprising a plurality of compression operation times corresponding to a plurality of sizes of data segments and a plurality of compression levels; and determining the second time from the plurality of compression operation times based on the compression level mapping table, the size of the data segment, and the compression level.
13. A computer program product tangibly stored on a computer readable medium and comprising machine executable instructions that, when executed, cause a machine to perform the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Intelligent configuration storage backup method suitable for cloud storage
CN102200936A
Adaptive caching replacement manager with dynamic updating granulates and partitions for shared flash-based storage system
US20180067961A1