Data deduplication method and device, equipment, storage medium and program product

By obtaining the priority of deleting target data and the busyness of the system, dynamically adjusting the data deleting strategy, the problem of low data deleting efficiency in the existing technology is solved, and more efficient data management and storage optimization are achieved.

CN120215803APending Publication Date: 2025-06-27SUGON INFORMATION IND +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311810215.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The data processing efficiency of existing data deletion methods is not high, and it is impossible to effectively utilize the characteristics of the target data and the operation of the system to optimize the deletion strategy.

Method used

By obtaining the priority of the redeletion of the target data to be stored and the busyness of the current system, the redeletion strategy of the target data is determined to optimize the data redeletion processing.

Benefits of technology

It improves the efficiency of data de-deletion, and can dynamically adjust the de-deletion strategy according to the characteristics of the target data and the operation of the system to avoid resource waste and system performance impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215803A_ABST
    Figure CN120215803A_ABST
Patent Text Reader

Abstract

The invention relates to a data deduplication method and device, equipment, a storage medium and a program product. The method comprises the following steps: acquiring deduplication information, wherein the deduplication information comprises a deduplication priority of to-be-stored target data and / or a busy degree of a current system; then, according to the deduplication information, determining a deduplication strategy of the target data; by adopting the method, the deduplication strategy is determined according to the characteristics of the target data and / or the running condition of the current system, and the data deduplication efficiency is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data deduplication, and in particular, to a data deduplication method, device, equipment, storage medium, and program product. Background Art

[0002] Data deduplication is a data management technology that can reduce the occupancy of storage space and improve the processing efficiency of data by identifying or deleting duplicate data blocks or files and replacing the duplicate data with references to existing data, thereby saving the space occupied by duplicate data. With the increasing growth of data volume and storage requirements, data deduplication has become increasingly important.

[0003] In related technologies, data deduplication is usually performed by online deduplication or post-processing deduplication. Among them, online deduplication performs duplicate detection and deduplication on data before writing the data to the disk, and post-processing deduplication reads the data out after writing the data to the disk for duplicate detection and deduplication, and then writes the data to the disk after deduplication.

[0004] However, the current data deduplication methods have low data processing efficiency. Summary of the Invention

[0005] Based on this, in view of the above technical problems, it is necessary to provide a data deduplication method, device, equipment, storage medium, and program product that can improve the data deduplication efficiency.

[0006] In a first aspect, the present application provides a data deduplication method, which includes:

[0007] Obtain deduplication information, where the deduplication information includes the deduplication priority of the target data to be stored and / or the busy degree of the current system;

[0008] Determine the deduplication strategy of the target data according to the deduplication information.

[0009] In this embodiment, by obtaining deduplication information, where the deduplication information includes the deduplication priority of the target data to be stored and / or the busy degree of the current system; then, according to the deduplication information, determine the deduplication strategy of the target data. In this way, in the embodiment of the present application, the deduplication priority of the target data can reflect the characteristics of the target data itself, and the busy degree of the current system can reflect the running situation of the current system. By obtaining the deduplication priority of the target data and / or the busy degree of the current system and considering the characteristics of the target data itself and / or the running situation of the current system in the process of determining the deduplication strategy, compared with the common method of performing data deduplication only by one of online deduplication or post-processing deduplication, and determining the deduplication strategy according to the characteristics of the target data itself and / or the running situation of the current system, the data deduplication efficiency is higher.

[0010] In one embodiment, the obtaining of the deduplication information includes:

[0011] Determine a real-time evaluation result of the target data according to the target data, and determine a data complexity evaluation result of the target data according to the target data, where the real-time evaluation result is used to characterize the real-time requirement of the target data;

[0012] Determine the deduplication priority according to the real-time evaluation result and the data complexity evaluation result.

[0013] In this embodiment, the deduplication priority of the target data is determined according to the real-time evaluation result and the data complexity evaluation result of the target data, comprehensively considering the real-time requirement and data characteristics of the target data, realizing the deduplication grading of the target data, so that the deduplication strategy of the target data can be determined according to the deduplication priority, in order to achieve the best data management and storage optimization effect.

[0014] In one embodiment, the determining of the real-time evaluation result of the target data according to the target data includes:

[0015] Determine a first parameter evaluation result of the real-time parameter of the target data according to the target data, where the real-time parameter includes at least one of the update frequency, data usage scenario, and data transmission delay of the target data;

[0016] Determine the real-time evaluation result according to the first parameter evaluation results of the real-time parameters.

[0017] In this embodiment, the real-time evaluation result of the target data is determined according to the update frequency, data usage scenario, and data transmission delay of the target data, so that the deduplication priority can be determined according to the real-time evaluation result, comprehensively considering different characteristics of the target data, and the calculation method has a wide application range.

[0018] In one embodiment, the determining of the data complexity evaluation result of the target data according to the target data includes:

[0019] Determine a second parameter evaluation result of the data complexity parameter of the target data according to the target data, where the data complexity parameter includes at least one of the data structure, data type, data quality, data volume, and data correlation of the target data;

[0020] Determine the data complexity evaluation result according to the second parameter evaluation results of the data complexity parameters.

[0021] In this embodiment, the evaluation result of the data complexity of the target data is determined according to the data structure, data type, data quality, data volume, and data correlation of the target data, so that the deduplication priority can be determined according to the evaluation result of the data complexity. Different characteristics of the target data are comprehensively considered, and the calculation method has a wide range of applications.

[0022] In one embodiment, the obtaining of the deduplication information includes:

[0023] Obtain the task status of the current system;

[0024] Determine the busy degree of the current system according to the task status of the current system.

[0025] In this embodiment, the busy degree of the current system is determined through the task status of the current system, and then the deduplication strategy of the target data can be determined according to the busy degree of the current system, so as to ensure that the determined deduplication strategy does not affect the operation efficiency of the system.

[0026] In one embodiment, the deduplication information includes the deduplication priority. The determining of the deduplication strategy of the target data according to the deduplication information includes:

[0027] Divide the first deduplication level for the target data;

[0028] Determine the deduplication strategy of the target data according to the first deduplication level and the deduplication priority of the target data.

[0029] In this implementation, the deduplication strategy of the target data is determined according to the first deduplication level and the deduplication priority of the target data. When a large number of target data need to be processed, data deduplication processing of the target data is realized according to the characteristics of the target data, and online deduplication processing is preferentially performed on the target data with a higher deduplication priority, so as to avoid the resources required by the data with a lower deduplication priority occupying the data with a higher deduplication priority.

[0030] In one embodiment, the deduplication information includes the busy degree of the current system. The determining of the deduplication strategy of the target data according to the deduplication information includes:

[0031] Divide the busy degree levels for the current system;

[0032] Determine the deduplication strategy of the target data according to the busy degree levels and the busy degree of the current system.

[0033] In this embodiment, the deduplication strategy of the target data is determined through different busy degree levels and the busy degree of the current system, and the data volume of different deduplication methods can be balanced, so as not to affect the system performance during data deduplication processing.

[0034] In one embodiment, the deduplication information includes the deduplication priority and the busy degree of the current system. Determining the deduplication policy for the target data according to the deduplication information includes:

[0035] Dividing a second deduplication level according to the target data and the busy degree of the current system;

[0036] Determining the deduplication policy for the target data according to the second deduplication level and the deduplication priority of the target data.

[0037] In this embodiment, by obtaining the busy degree of the current system and jointly determining the deduplication policy for the target data as online deduplication or post-processing deduplication according to the busy degree of the current system and the deduplication priority, while ensuring the system operation efficiency, the deduplication efficiency of the data is improved, and the performance requirements of the computer system and the deduplication efficiency requirements of the data can be balanced. Then, by comparing the target deduplication priority threshold corresponding to the busy degree of the current system with the deduplication priorities of the target data, the deduplication policies of the target data are determined, different deduplication methods are selected for different target data according to their respective deduplication priorities, and the target data is shunted according to the characteristics of the target data itself, and the data deduplication method is more flexible.

[0038] In one embodiment, determining the deduplication policy for the target data according to the deduplication information includes:

[0039] Selecting one of the online deduplication and post-processing deduplication policies as the deduplication policy for the target data according to the deduplication information.

[0040] In this embodiment, the target data is shunted for deduplication through online deduplication or post-processing deduplication.

[0041] In a second aspect, the present application further provides a data storage method, and the method includes:

[0042] Using the data deduplication method according to any one of the first aspects above to determine the deduplication policy for the target data;

[0043] Performing deduplication processing on the target data according to the deduplication policy, and storing the target data after the deduplication processing in a storage system.

[0044] In this embodiment, the target data is first deduplicated and then stored. By deleting duplicate data, the storage space can be saved and the data storage efficiency can be improved.

[0045] In a third aspect, the present application further provides a data compression method, and the method includes:

[0046] Using the data deduplication method described in any item of the above first aspect, determine the deduplication strategy for the target data;

[0047] Perform deduplication processing on the target data according to the deduplication strategy, and perform data compression processing on the target data after deduplication processing.

[0048] In this embodiment, before performing data compression on the target data, perform deduplication processing on the target data. By deleting duplicate data, the amount of data to be compressed can be reduced, thereby improving the efficiency of data compression.

[0049] Fourth aspect, the present application also provides a data transmission method, and the method includes:

[0050] Using the data deduplication method described in any item of the above first aspect, determine the deduplication strategy for the target data;

[0051] Perform deduplication processing on the target data according to the deduplication strategy, and send the target data after deduplication processing to the receiving end.

[0052] In this embodiment, before data transmission, the data sending end performs deduplication processing on the target data, and then sends the target data after deduplication processing to the receiving end, which can reduce the amount of data during the transmission of the target data, reduce the data transmission burden, and improve the efficiency of data transmission.

[0053] Fifth aspect, the present application also provides a data deduplication device, and the device includes:

[0054] The first acquisition module is used to acquire deduplication information, and the deduplication information includes the deduplication priority of the target data to be stored and / or the busy degree of the current system;

[0055] The first determination module is used to determine the deduplication strategy of the target data according to the deduplication information.

[0056] Sixth aspect, the present application also provides a data storage device, and the device includes:

[0057] The second determination module is used to use the data deduplication method described in any item of the above first aspect to determine the deduplication strategy of the target data;

[0058] The storage module is used to perform deduplication processing on the target data according to the deduplication strategy, and send the target data after deduplication processing to the receiving end.

[0059] Seventh aspect, the present application also provides a data compression device, and the device includes:

[0060] The third determination module is used to use the data deduplication method described in any item of the above first aspect to determine the deduplication strategy of the target data;

[0061] A compression module, configured to perform deduplication processing on the target data according to the deduplication policy, and store the target data after the deduplication processing in a storage system.

[0062] In a fifth aspect, the present application further provides a data transmission device, which includes:

[0063] A fourth determination module, configured to determine a deduplication policy for target data by using the data deduplication method according to any one of the first aspects above;

[0064] A sending module, configured to perform deduplication processing on the target data according to the deduplication policy, and send the target data after the deduplication processing to a receiving end.

[0065] In a sixth aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method according to the first aspect above are implemented.

[0066] In a seventh aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to the first aspect above are implemented.

[0067] In an eighth aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method according to the first aspect above are implemented.

[0068] For the above data deduplication method, device, equipment, storage medium and program product, by obtaining deduplication information, the deduplication information includes the deduplication priority of the target data to be stored and / or the busy degree of the current system; then, according to the deduplication information, a deduplication policy for the target data is determined. In this way, in the embodiments of the present application, the deduplication priority of the target data can reflect the characteristics of the target data itself, and the busy degree of the current system can reflect the running situation of the current system. By obtaining the deduplication priority of the target data and / or the busy degree of the current system, the characteristics of the target data itself and / or the running situation of the current system are considered in the process of determining the deduplication policy. Compared with the common method of performing data deduplication only through one of online deduplication or post-processing deduplication, determining the deduplication policy according to the characteristics of the target data itself and / or the running situation of the current system, the efficiency of data deduplication is higher. Description of the Drawings

[0069] To more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the accompanying drawings required for the description of the embodiments or related technologies. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0070] Figure 1 It is an application environment diagram of the data deduplication method in an embodiment;

[0071] Figure 2 It is a schematic flowchart of the data deduplication method in an embodiment;

[0072] Figure 3 It is a schematic flowchart of the data deduplication method in another embodiment;

[0073] Figure 4 It is a schematic flowchart of the data deduplication method in another embodiment;

[0074] Figure 5 It is a schematic diagram for calculating the real-time evaluation result in another embodiment;

[0075] Figure 6 It is a schematic flowchart of the data deduplication method in another embodiment;

[0076] Figure 7 It is a schematic flowchart of the data deduplication method in another embodiment;

[0077] Figure 8 It is a schematic flowchart of the data deduplication method in another embodiment;

[0078] Figure 9 It is a schematic flowchart of the data deduplication method in another embodiment;

[0079] Figure 10 It is a schematic flowchart of the data deduplication method in another embodiment;

[0080] Figure 11 It is a schematic flowchart of the data deduplication method in another embodiment;

[0081] Figure 12 It is a schematic flowchart of the data deduplication method in another embodiment;

[0082] Figure 13 It is an overall block diagram of the data deduplication system in another embodiment;

[0083] Figure 14 It is a schematic flowchart of the data deduplication method in another embodiment;

[0084] Figure 15 It is a structural block diagram of a data deduplication device in an embodiment. Specific implementation manners

[0085] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0086] Currently, data deduplication technology is widely used in application scenarios such as data backup, archiving systems, cloud storage, and virtualization environments. Data deduplication technology can significantly reduce the storage space requirements, lower the storage cost, and at the same time improve the data transmission and processing efficiency. Data deduplication technology divides data into blocks and generates unique identifiers, such as hash values or fingerprints, by using various algorithms, such as hash functions, fingerprint algorithms, sliding windows, etc., and then determines whether the data already exists. Based on the determination result, it decides whether to delete or retain the data. If the data does not exist, the data is retained; if the data already exists, a reference is used to point the existing data to the existing copy to save storage space and avoid duplicate storage.

[0087] In a storage system, data deduplication processing is usually performed in an online deduplication manner or a post-processing deduplication manner. Among them, online deduplication performs duplicate detection and deduplication processing on data before the data is written to the disk. By detecting and deduplicating data in real time, it can timely reduce data redundancy and improve the data reduction ratio. However, in the case of a large amount of data to be processed or high-concurrency data writing, it may increase the burden on the system, extend the system response time, and affect the overall performance of the system. Post-processing deduplication is to read the data out after the data is written to the disk for duplicate detection and deduplication processing, and then write it back to the disk after the deduplication processing. Compared with online deduplication, post-processing deduplication does not have an obvious impact on the system performance and can maintain a high system response speed. However, since the data is first written to the disk and then read out for deduplication processing, this method may cause some duplicate data to be deduplicated after being stored for a period of time. Therefore, the storage space cannot be immediately released during the data writing process, and the real-time performance of the deduplication effect will be correspondingly reduced. Common data deduplication methods all choose one of online deduplication or post-processing deduplication. It can be seen that the current data deduplication methods have low data deduplication efficiency.

[0088] In view of this, the data deduplication method provided by the embodiments of the present application obtains deduplication information, where the deduplication information includes the deduplication priority of the target data to be stored and / or the busy degree of the current system; then, according to the deduplication information, a deduplication policy for the target data is determined. In this way, in the embodiments of the present application, the deduplication priority of the target data can reflect the characteristics of the target data itself, and the busy degree of the current system can reflect the operating conditions of the current system. By obtaining the deduplication priority of the target data and / or the busy degree of the current system, the characteristics of the target data itself and / or the operating conditions of the current system are considered in the process of determining the deduplication policy. Compared with the common method of performing data deduplication only through one of online deduplication or post-processing deduplication, determining the deduplication policy according to the characteristics of the target data itself and / or the operating conditions of the current system, the efficiency of data deduplication is higher.

[0089] The data deduplication method provided in this embodiment can be applied to a computer device, which can be a terminal or a server. Taking the terminal as an example, its internal structure diagram can be as Figure 1 shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit, and an input device. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface, the display unit, and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements a data deduplication method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the shell of the computer device, or an external keyboard, a touchpad, or a mouse, etc.

[0090] Those skilled in the art can understand that Figure 1 the structure shown in merely represents a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0091] In an exemplary embodiment, as Figure 2 shown, a data deduplication method is provided. Taking the computer device in Figure 1 as an example, the method includes the following steps 201 to step 202. Among them:

[0092] Step 201, obtain deduplication information.

[0093] Among them, the deduplication information includes the deduplication priority of the target data to be stored and / or the busy degree of the current system. The target data is the data that needs to be stored. There may be data duplication in the target data. For example, data duplication caused by copying data in different applications, data repeated upload, etc. If the target data is directly stored in the storage system of the computer device, it may cause a large amount of duplicate data in the storage system, which will not only occupy the storage space of the storage system, but also may increase the burden in the data transmission, backup and processing processes. Therefore, when the target data is stored in the storage system, the target data needs to be deduplicated through data deduplication technology. Optionally, the deduplication priority of the target data is related to at least one of the real-time requirement and data complexity of the target data. If the real-time requirement of the target data is high, that is, the target data needs to be processed in time, it can be understood that the priority requirement of the corresponding deduplication priority of the target data is higher. If the data complexity of the target data is high, that is, the target data is relatively complex and it takes a long time to process the target data, the priority requirement of the corresponding deduplication priority of the target data is lower. The deduplication priority is the priority for the target data to perform data deduplication processing. Optionally, the deduplication priority can be represented by different levels, such as the first level, the second level, and the third level, etc., or can be represented by specific numerical values, such as positive integers.

[0094] Among them, the busy degree of the current system refers to the ability of the computer system to process tasks and respond to tasks in the current time period. Optionally, the busy degree of the current system can be determined by obtaining the number of tasks running in the current system.

[0095] Step 202, determine the deduplication strategy of the target data according to the deduplication information.

[0096] After obtaining the deduplication priority of the target data and / or the busy degree of the current system, according to the deduplication priority of the target data and / or the busy degree of the current system, determine the deduplication strategy of the target data, that is, whether the target data adopts online deduplication or post-processing deduplication. Optionally, the target data can be multiple pieces of data. By determining the deduplication strategies of the respective target data, data diversion of the respective target data is achieved.

[0097] In the above embodiments, by obtaining deduplication information, the deduplication information includes the deduplication priority of the target data to be stored and / or the busy degree of the current system; then, according to the deduplication information, a deduplication strategy for the target data is determined. In this way, in the embodiments of the present application, the deduplication priority of the target data can reflect the characteristics of the target data itself, and the busy degree of the current system can reflect the operating conditions of the current system. By obtaining the deduplication priority of the target data and / or the busy degree of the current system, the characteristics of the target data itself and / or the operating conditions of the current system are considered in the process of determining the deduplication strategy. Compared with the common method of only performing data deduplication through one of online deduplication or post-processing deduplication, determining the deduplication strategy according to the characteristics of the target data itself and / or the operating conditions of the current system, the efficiency of data deduplication is higher.

[0098] In one embodiment, as Figure 3 shown in the embodiment, step 201 may include step 301 and step 302:

[0099] Step 301, determine a real-time evaluation result of the target data according to the target data, and determine a data complexity evaluation result of the target data according to the target data.

[0100] Among them, the real-time evaluation result is used to characterize the real-time requirement of the target data. Optionally, the real-time requirement of the target data may be related to the importance degree, update frequency, data usage scenario, and data transmission delay of the target data, etc. By respectively scoring the importance degree, update frequency, data usage scenario, and data transmission delay of the target data, etc., the real-time evaluation result of the target data can be determined. The data complexity evaluation result of the target data may be related to the characteristics of the target data, such as the data structure, data type, data quality, data volume, and data correlation of the target data, etc. By respectively scoring the data structure, data type, data quality, data volume, and data correlation of the target data, etc., the data complexity evaluation result of the target data can be determined.

[0101] Step 302, determine the deduplication priority of the target data according to the real-time evaluation result and the data complexity evaluation result.

[0102] Optionally, the level ranges corresponding to the real-time evaluation result and the data complexity evaluation result can be represented by 0-N respectively, where level 0 represents the highest and level N represents the lowest. Among them, the priority of the real-time evaluation result and the deduplication priority is positively correlated, that is, the higher the level of the real-time evaluation result, the higher the priority of the deduplication priority. The priority of the data complexity evaluation result and the deduplication priority is negatively correlated, that is, the higher the level of the data complexity evaluation result, due to the more complex data processing, the lower the priority of the deduplication priority. Therefore, by calculating the real-time evaluation result minus the data complexity evaluation result and adding N, the final deduplication priority is obtained, where the deduplication priority range is 0-2N.

[0103] In this embodiment, according to the real-time evaluation result and the data complexity evaluation result of the target data, the deduplication priority of the target data is determined, comprehensively considering the real-time requirements and data characteristics of the target data, realizing the deduplication grading of the target data, so that the deduplication strategy of the target data can be determined according to the deduplication priority, in order to achieve the best data management and storage optimization effect.

[0104] In one embodiment, based on Figure 3 the embodiment shown, the process of determining the real-time evaluation result of the target data according to the target data in step 301 is as Figure 4 shown, and may include step 401 and step 402:

[0105] Step 401, according to the target data, determine the first parameter evaluation result of the real-time parameter of the target data.

[0106] Among them, the real-time parameters include at least one of the update frequency of the target data, the data usage scenario, and the data transmission delay. According to the update frequency of the target data, the first parameter evaluation result corresponding to the update frequency can be determined. Optionally, if the update frequency of the target data is high, it is determined that the real-time requirement of the target data is high. According to the data usage scenario of the target data, the first parameter evaluation result corresponding to the data usage scenario can be determined. Optionally, if the data usage scenario of the target data has a high requirement for real-time, the real-time requirement of the target data is high. According to the data transmission delay of the target data, the first parameter evaluation result corresponding to the data transmission delay can be determined. Optionally, the three real-time parameters of the update frequency, data usage scenario, and data transmission delay of the target data are usually interrelated. For example, when the update frequency of the target data is high, it usually corresponds to a data usage scenario with a high real-time requirement for the target data, and at the same time, the data transmission delay of the target data is also low.

[0107] Step 402, according to the first parameter evaluation results of the respective real-time parameters, determine the real-time evaluation result.

[0108] Determine the first-parameter evaluation results corresponding to each real-time parameter according to the parameters included in the real-time parameters, and then determine the real-time evaluation result. Optionally, the average of the first-parameter evaluation results corresponding to each real-time parameter can be calculated to determine the real-time evaluation result of the target data, or the first-parameter evaluation results of each real-time parameter can be summed to determine the real-time evaluation result of the target data.

[0109] Optionally, perform statistical processing on the first-parameter evaluation results of each real-time parameter to obtain the real-time evaluation result.

[0110] Optionally, the statistical processing can be weighted summation. When there are multiple real-time parameters, perform weighted summation on the first-parameter evaluation results of the multiple real-time parameters to obtain the real-time evaluation result. For an example of the above calculation process, please refer to Figure 5 . When the real-time parameters include the update frequency of the target data, the data usage scenario, and the data transmission delay at the same time, where the weights of the update frequency, the data usage scenario, and the data transmission delay can be m1, m2, and m3, and m1 + m2 + m3 = 1. Taking a target data as an example, the score value of the first-parameter evaluation result of the update frequency corresponding to the target data is a, the score value of the first-parameter evaluation result of the data usage scenario is b, and the score value of the first-parameter evaluation result of the data transmission delay is c, where the value ranges of a, b, and c are all from 0 to N. Perform weighted summation processing on the first-parameter evaluation results of each real-time parameter to obtain the real-time evaluation result as am1 + bm2 + cm3, and round it to obtain the real-time evaluation result of this target data.

[0111] In this embodiment, the real-time evaluation result of the target data is determined according to the update frequency, the data usage scenario, and the data transmission delay of the target data, so that the deduplication priority can be determined according to the real-time evaluation result, comprehensively considering different characteristics of the target data, and the calculation method has a wide range of applications.

[0112] In one embodiment, based on Figure 3 the embodiment shown, the process in which the computer device determines the evaluation result of the data complexity of the target data in step 301 is as Figure 6 shown, and it may include step 601 and step 602:

[0113] Step 601, determine the second-parameter evaluation result of the data complexity parameter of the target data according to the target data.

[0114] Among them, the data complexity parameter includes at least one of the data structure, data type, data quality, data volume, and data correlation of the target data.

[0115] Optionally, determine the second parameter evaluation result corresponding to the data structure according to the data structure of the target data. The data structure is the basis of algorithm efficiency, and different data structures will affect the time complexity and space complexity of the algorithm. By analyzing the relationships, hierarchies, and organizational methods among the target data, the second parameter evaluation result of the data structure can be determined. For example, in a storage system, for the two data structures of linked lists and arrays, since arrays can use contiguous memory space, they have an advantage in storage space and lower complexity, while linked lists require additional storage space to save pointers to the next node, so the complexity of linked lists is higher. If the data structure contains complex structures such as nesting, the corresponding complexity will be even higher. Therefore, based on the complexity of the data structure of the target data, the second parameter evaluation result of the data structure can be determined. The scoring value of the second parameter evaluation result can be d, and its value range is from 0 to N.

[0116] Optionally, determine the second parameter evaluation result corresponding to the data type according to the data type of the target data. Among them, the data type of the target data refers to the data type stored in memory, which can include basic data types and complex data types. Among them, basic data types can include integers, boolean types, and characters, etc., while complex data types can include arrays, objects, and strings, etc. Since different data types have different characteristics in terms of storage, calculation, and operation, and these characteristics will also affect the complexity of data processing. For example, when storing, integer data types usually require more storage space than character data types. Different data types have different characteristics. Based on the data type of the target data, the second parameter evaluation result of the data type of the target data can be determined. The scoring value of the second parameter evaluation result can be e, and its value range is from 0 to N.

[0117] Optionally, determine the second-parameter evaluation result corresponding to the data quality according to the data quality of the target data. Among them, the data quality has a certain impact on the data memory complexity. The data quality can be evaluated according to the integrity, consistency, and timeliness of the target data. Among them, the integrity of the data can include whether aspects such as the quantity, type, and range of the data are complete and correct, and the integrity of the data can be evaluated by comparing the recorded values of the data statistics with the current data. The consistency of the data can be evaluated by the recorded values of the data statistics, and can include verifying whether the date format is correct and whether the data type conforms to the specification, etc. The timeliness of the data can be evaluated by checking the update time of the data. If the data is not updated in a timely manner, that is, it has not been updated for a long time, the timeliness of the data is low. When the quality of the data is low, data cleaning and preprocessing operations may be required during the data processing process, thereby increasing the difficulty of data integration during the data processing process and increasing the data memory complexity. Based on the data quality of the target data, determine the second-parameter evaluation result of the data quality of the target data. The score value of the second-parameter evaluation result can be f, and its value range is from 0 to N.

[0118] Optionally, determine the second-parameter evaluation result corresponding to the data volume according to the data volume of the target data. Among them, the data volume has a greater impact on the data complexity. As the data volume of the target data increases, the complexity of storing and processing the target data in memory will increase. At the same time, the increase in the data volume will occupy a larger memory space and increase the memory complexity. Based on the data volume of the target data, determine the second-parameter evaluation result of the data volume of the target data. The score value of the second-parameter evaluation result can be g, and its value range is from 0 to N.

[0119] Optionally, determine the second-parameter evaluation result corresponding to the data correlation according to the data correlation of the target data. Among them, the data correlation has a greater impact on the data complexity. When there are complex correlation relationships between the target data to be processed, such as tree structures or network structures, etc., the complex correlation relationships will lead to an increase in the memory complexity. For example, if the target data contains multiple tables and there are complex correlation relationships and hierarchical relationships between the tables, the system requires more memory space when storing and processing the target data, and at the same time increases the complexity of processing. Based on the data correlation relationship of the target data, determine the second-parameter evaluation result of the data correlation relationship of the target data. The score value of the second-parameter evaluation result can be h, and its value range is from 0 to N.

[0120] Step 602, the computer device determines the data complexity evaluation result according to the second-parameter evaluation results of each data complexity parameter.

[0121] Determine the data complexity evaluation result of the target data according to the second parameter evaluation results of each data complexity parameter. Optionally, the average of the second parameter evaluation results corresponding to each data complexity parameter can be calculated to determine the data complexity evaluation result of the target data, or the sum of the second parameter evaluation results corresponding to each data complexity parameter can be calculated to determine the data complexity result of the target data.

[0122] Optionally, perform statistical processing on the second parameter evaluation results of each data complexity parameter to obtain the data complexity evaluation result.

[0123] Optionally, the statistical processing can be weighted summation. When there are multiple data complexity parameters, perform weighted summation on the second parameter evaluation results of the multiple data complexity parameters to obtain the complexity evaluation result. For an example of the above calculation process, please refer to Figure 7 . When the data complexity parameters include the data structure, data type, data quality, data volume, and data correlation of the target data, where the weights of the target data structure, data type, data quality, data volume, and data correlation are set to n1, n2, n3, n4, and n5 respectively, and n1 + n2 + n3 + n4 + n5 = 1. Taking a target data as an example, the score value of the second parameter evaluation result of the data structure corresponding to the target data is d, the score value of the second parameter evaluation result of the data type is e, the score value of the second parameter evaluation result of the data quality is f, the score value of the second parameter evaluation result of the data volume is g, and the score value of the second parameter evaluation result of the data correlation is h. Perform weighted summation on the second parameter evaluation results to obtain the data complexity evaluation result as dn1 + en2 + fn3 + gn4 + hn5, and round it to obtain the data complexity evaluation result of the target data.

[0124] In this embodiment, the data complexity evaluation result of the target data is determined according to the data structure, data type, data quality, data volume, and data correlation of the target data, so that the deduplication priority can be determined according to the data complexity evaluation result, comprehensively considering different characteristics of the target data, and the calculation method has a wide range of applications.

[0125] In one embodiment, as Figure 8 shown, step 201 may include step 801 and step 802:

[0126] Step 801, obtain the task status of the current system.

[0127] Obtain the task status processed by the current computer system within a specific time, as well as the mutual association and mutual influence between tasks.

[0128] Step 802: Determine the busyness level of the current system according to the task status of the current system.

[0129] Optionally, according to the capabilities of the computer system to process tasks and respond to tasks, the busyness level of the system can be divided into a first load state, a second load state, and a third load state. Among them, the first load state is a low load state, in which the computer system does not run tasks or only needs to process a very small number of tasks for most of the time, and this state is usually called the idle state. The second load state is a medium load state, in which the computer system needs to process a certain number of tasks, but there is no obvious mutual association or influence between the tasks, and this state is usually called the non-blocking state. The third load state is a high load state, in which the computer system needs to process a large number of tasks, and there may be mutual association or influence between these tasks, and this state is usually called the blocking state.

[0130] In this embodiment, the busyness level of the current system is determined through the task status of the current system, and then the deduplication policy for the target data can be determined according to the busyness level of the current system, so as to ensure that the determined deduplication policy does not affect the operating efficiency of the system.

[0131] In one embodiment, step 202 of determining the deduplication policy for the target data according to the deduplication information may include the following steps A1:

[0132] Step A1: Select one of the online deduplication and post-processing deduplication policies as the deduplication policy for the target data according to the deduplication information.

[0133] Optionally, when determining the deduplication policy, due to the characteristics of online deduplication and post-processing deduplication, it is necessary to make a choice according to the specific situation. When the system and the data to be processed have high requirements for real-time performance and can accept a certain attenuation of system performance, online deduplication can be selected. Online deduplication can delete duplicate data in a timely manner and reduce system redundancy. If the system and the data to be processed have high requirements for system performance and can accept a certain time delay in data processing, post-processing deduplication can be selected. Post-processing deduplication will not have an obvious impact on system performance and can ensure the response speed of the system, but there will be a certain delay in data processing and the storage space cannot be released immediately. Therefore, in order to comprehensively consider system performance and data deduplication efficiency, obtain the deduplication information of the target data, and select one of the online deduplication and post-processing deduplication policies for data deduplication processing.

[0134] In one embodiment, when the deduplication information includes the deduplication priority, step 202 as Figure 9 shown may include step 901 and step 902:

[0135] Step 901: Divide the first deduplication level for the target data.

[0136] When there are multiple target data that need to be deduplicated, divide the multiple target data into the first deduplication level. Optionally, the first deduplication level can include multiple levels, and different levels can correspond to different deduplication strategies.

[0137] Step 902, determine the deduplication strategy of the target data according to the first deduplication level and the deduplication priority of the target data.

[0138] After dividing the first deduplication level, the deduplication strategy of each target data can be determined according to the deduplication priority and the first deduplication level of each target data. Optionally, the calculation of the deduplication priority of the target data is as shown in the above embodiments. For example, the deduplication priority can be represented by 0-2N, where 0 represents the highest and 2N represents the lowest. The deduplication priority corresponding to the first deduplication level can be K (0 < K < 2N). When the deduplication priority of the target data is less than K, it means that the deduplication priority level of the target data is relatively high. At this time, it can be determined that the deduplication strategy of the target data is the online deduplication method, and the data deduplication process of the target data can be preferentially performed. When the deduplication priority of the target data is greater than K, it means that the deduplication priority level of the target data is relatively low. At this time, it can be determined that the deduplication strategy of the target data is the post-processing deduplication method, that is, in order not to affect the system performance, the target data is first written to the disk, and the data deduplication process is performed after a period of time.

[0139] In this embodiment, the deduplication strategy of the target data is determined according to the first deduplication level and the deduplication priority of the target data. When a large number of target data need to be processed, the data deduplication process of the target data is realized according to the characteristics of the target data, and the online deduplication process of the target data with a relatively high deduplication priority is preferentially performed, so as to avoid the data with a relatively low deduplication priority occupying the resources required by the data with a relatively high deduplication priority.

[0140] In the embodiments of the present application, the deduplication information includes the busy degree of the current system, such as Figure 10 As shown, according to the deduplication information, the steps of determining the deduplication strategy of the target data include step 1001 and step 1002:

[0141] Step 1001, divide the busy degree level of the current system.

[0142] Optionally, as described in the above embodiments, the system can be divided into different busy degree levels according to the task state of the system operation, such as high load state, heavy load state, and low load state, etc.

[0143] Step 1002, determine the deduplication strategy of the target data according to the busy degree level and the busy degree of the current system.

[0144] Obtain the busy degree of the current system, and determine the deduplication strategy for the target data according to different busy degree levels and the busy degree of the current system. Optionally, different busy degree levels can correspond to different deduplication strategies. For example, when the busy degree level is in a high-load state, it indicates that there are more tasks running in the system at this time. Therefore, in order not to affect the performance of the system, the corresponding deduplication strategy can be post-processing deduplication. When the busy degree level is in a low-load state, it indicates that there are fewer tasks running in the system at this time. Therefore, in order to improve the deduplication efficiency, the corresponding deduplication strategy can be online deduplication. When the busy degree level is in a medium-load state, a certain proportion of the target data can be deduplicated online, and another proportion of the target data can be deduplicated by post-processing.

[0145] In this embodiment, by determining the deduplication strategy for the target data based on different busy degree levels and the busy degree of the current system, it is possible to balance the data volumes of different deduplication methods, so as not to affect the system performance during the data deduplication process.

[0146] In one embodiment, the deduplication information includes the deduplication priority and the busy degree of the current system. As Figure 11 shown, the steps of determining the deduplication strategy for the target data according to the deduplication information include step 1101 and step 1102:

[0147] Step 1101, divide the second deduplication level according to the target data and the busy degree of the current system.

[0148] Optionally, the second deduplication level can be set with different deduplication priority thresholds according to the busy degree of different systems.

[0149] Step 1102, determine the deduplication strategy for the target data according to the second deduplication level and the deduplication priority of the target data.

[0150] Since there are certain defects in both single online deduplication and post-processing deduplication, the above is adjusted through the busy degree of the system, and the data ratio of different deduplication methods is controlled according to the busy degree of the system. However, although this method can balance the data volumes of different deduplication methods, it does not consider the characteristics of the data itself, which may cause data with low real-time requirements, that is, non-urgent data, to occupy the resources required for data processing with high real-time requirements, that is, urgent data. That is, data with low real-time requirements is deduplicated by post-processing, while data with high real-time requirements is deduplicated online. Therefore, some data processing requirements may not be met in time.

[0151] In view of this, in order to achieve a balance between the performance of the computer system and the efficiency of data deduplication processing, the deduplication strategy for the target data can be jointly determined according to the current system busyness level and the deduplication priority of the target data. Optionally, when the current system is in a low-load state, more target data can be allowed to have an online deduplication strategy, that is, increase the proportion of target data for online deduplication, thereby improving the efficiency of data deduplication. The deduplication strategy for each target data can be determined to be online deduplication or post-processing deduplication according to the deduplication priority of each target data. When the current system is in a high-load state, in order not to affect the system performance, more target data needs to have a post-processing deduplication strategy, that is, increase the proportion of target data for post-processing deduplication, and determine whether each target data is subject to online deduplication or post-processing deduplication according to the deduplication priority of each target data.

[0152] Optionally, obtain the target deduplication priority threshold corresponding to the current system busyness level, and compare the deduplication priority with the target deduplication priority threshold to obtain a comparison result.

[0153] Among them, the target deduplication priority threshold is the deduplication priority threshold corresponding to the current system busyness level. Optionally, when the deduplication priority is represented by a level, the deduplication priority threshold is also a level. When the deduplication priority is represented by a specific value, the deduplication priority threshold is also in numerical form.

[0154] Optionally, the priority level of the deduplication priority is negatively correlated with the magnitude of the deduplication priority, and the system busyness level is negatively correlated with the magnitude of the corresponding deduplication priority threshold. For example, when the deduplication priority is represented by a value, the larger the deduplication priority, the lower the priority level of the deduplication priority, and the smaller the deduplication priority, the higher the priority level of the deduplication priority. For example, the deduplication priority can be represented by 0 - 2N, where 0 represents the highest and 2N represents the lowest. The higher the system busyness level, that is, the busier it is, the smaller the corresponding deduplication priority threshold, and the lower the system busyness level, that is, the more idle it is, the larger the corresponding deduplication priority threshold.

[0155] Optionally, the priority level of the deduplication priority can also be positively correlated with the magnitude of the deduplication priority, and the system busyness level can also be positively correlated with the magnitude of the corresponding deduplication priority threshold, that is, contrary to the above, the embodiments of the present application do not limit this.

[0156] Then, according to the comparison result, determine the deduplication strategy to be online deduplication or post-processing deduplication. When there are multiple target data, compare the deduplication priorities of each target data with the target deduplication priority threshold, and according to each comparison result, determine the deduplication strategy of each target data to be online deduplication or post-processing deduplication, so as to achieve the diversion of each target data.

[0157] In this embodiment, by obtaining the busy degree of the current system, and jointly determining the deduplication strategy of the target data as online deduplication or post-processing deduplication according to the busy degree of the current system and the deduplication priority, the target data is deduplicated and split according to the busy degree of the current system, which can improve the deduplication efficiency of the data while ensuring the system operation efficiency, and can balance the performance requirements of the computer system and the deduplication efficiency requirements of the data. Then, by comparing the target deduplication priority threshold corresponding to the busy degree of the current system with the deduplication priorities of the target data, the deduplication strategies of the target data are determined, and different deduplication methods are selected for different target data according to their respective deduplication priorities, and the target data is split according to the characteristics of the target data itself, making the data deduplication method more flexible.

[0158] Optionally, if the comparison result is that the deduplication priority is less than or equal to the target deduplication priority threshold, the computer device determines that the deduplication strategy is online deduplication.

[0159] Taking the case where the priority degree of the deduplication priority in the above embodiment is negatively correlated with the size of the deduplication priority, and the busy degree of the system is negatively correlated with the size of the corresponding target deduplication priority threshold as an example, if the deduplication priority of the target data is less than or equal to the target deduplication priority threshold, that is, the priority degree of the target data deduplication priority is higher, so to improve the data deduplication efficiency, the target data deduplication strategy is determined to be online deduplication.

[0160] If the comparison result is that the deduplication priority is greater than the target deduplication priority threshold, the computer device determines that the deduplication strategy is post-processing deduplication.

[0161] If the deduplication priority of the target data is greater than the target deduplication priority threshold, that is, the priority degree of the target data deduplication priority is lower, so as not to affect the system performance, the target data deduplication strategy is determined to be post-processing deduplication.

[0162] In this embodiment, according to the comparison result of the deduplication priority and the target deduplication priority threshold, the deduplication strategy of the target data is determined, which can improve the data deduplication efficiency without affecting the system performance.

[0163] In the embodiment of the present application, please refer to Figure 12 , which shows the flowchart of the data deduplication method provided by the embodiment of the present application. The data deduplication method includes the following steps:

[0164] Step 1201, obtain deduplication information.

[0165] Step 1202, determine the real-time evaluation result of the target data according to the target data, and determine the data complexity evaluation result of the target data according to the target data.

[0166] Step 1203: Determine the deduplication priority of the target data according to the real-time evaluation result and the data complexity evaluation result.

[0167] Step 1204: Divide the second deduplication level according to the target data and the current system's busyness.

[0168] Step 1205: Determine the deduplication strategy of the target data according to the second deduplication level and the deduplication priority of the target data.

[0169] In the embodiments of the present application, the deduplication method provided by the embodiments of the present application can be applied in a deduplication system. The overall block diagram of the deduplication system is as Figure 13 shown. The data input module is used to obtain multiple target data to be stored. Then, the pattern recognition module and the data classification module are used to analyze each target data to determine the deduplication priority of each target data. On the other hand, the bandwidth acquisition module and the reference sampling module are used to obtain the performance metrics of the current system, such as the current running tasks. The load classification module is used to classify the busyness of the current system according to the current system's tasks. Finally, the data shunting module determines the deduplication strategy of each target data as online deduplication or post-processing deduplication according to the busyness of the current system and the deduplication priority, so as to achieve data shunting for each target data.

[0170] In one embodiment, the process of determining the deduplication strategy of each target data as online deduplication or post-processing deduplication according to the busyness of the current system and the deduplication priority, so as to achieve data shunting for each target data is as Figure 14 shown. Different levels of the busyness of the current system correspond to different target deduplication priority thresholds. Optionally, when the busyness of the current system is in a high-load state, the target deduplication priority threshold is A, that is, at this time, only the deduplication strategy of the target data with levels 0 - A is online deduplication, and the deduplication strategies of other levels of target data are post-processing deduplication. When the busyness of the current system is in a medium-load state, the target deduplication priority threshold is B, that is, at this time, only the deduplication strategy of the target data with levels 0 - B is online deduplication, and the deduplication strategies of other levels of target data are post-processing deduplication. When the busyness of the current system is in a low-load state, the target deduplication priority threshold is C, that is, at this time, only the deduplication strategy of the target data with levels 0 - C is online deduplication, and the deduplication strategies of other levels of target data are post-processing deduplication. Among them, (A < B < C < 2N).

[0171] In this way, according to the different degrees of busyness of the current system, the deduplication strategy of the target data is adjusted according to the deduplication priority. On the premise of meeting the system performance, the efficiency of data deduplication is improved, data redundancy is minimized to the greatest extent, the storage efficiency is increased, and it has high practicability. Further, due to dynamic deduplication diversion according to different deduplication priorities and the degree of busyness of the system, data with a high priority is deduplicated online, and data with a low priority is post-processed for deduplication, thus ensuring the deduplication efficiency of data with a high priority and the overall efficiency of the system.

[0172] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.

[0173] In an embodiment of the present application, a data storage method is further provided, and the method includes:

[0174] Using the data deduplication method described in any one of the above method embodiments, determine the deduplication strategy of the target data.

[0175] Perform deduplication processing on the target data according to the deduplication strategy, and store the target data after deduplication processing in the storage system.

[0176] In this embodiment, the target data is first deduplicated and then stored. By deleting duplicate data, the storage space can be saved and the data storage efficiency can be improved.

[0177] In an embodiment of the present application, a data compression method is further provided, and the method includes:

[0178] Using the data deduplication method described in any one of the above method embodiments, determine the deduplication strategy of the target data.

[0179] Perform deduplication processing on the target data according to the deduplication strategy, and perform data compression processing on the target data after deduplication processing.

[0180] In this embodiment, the target data is first deduplicated before data compression. By deleting duplicate data, the amount of data to be compressed can be reduced, thereby improving the data compression efficiency.

[0181] In an embodiment of the present application, a data transmission method is further provided, and the method includes:

[0182] Using the data deduplication method in any one of the above method embodiments, determine the deduplication strategy for the target data.

[0183] Perform deduplication processing on the target data according to the deduplication strategy, and send the target data after deduplication processing to the receiving end.

[0184] In this embodiment, before data transmission, the data sending end performs deduplication processing on the target data, and then sends the target data after deduplication processing to the receiving end, which can reduce the amount of data during the transmission of the target data, reduce the data transmission burden, and improve the efficiency of data transmission.

[0185] Based on the same inventive concept, an embodiment of the present application further provides a data deduplication device for implementing the data deduplication method involved above. The solution provided by the device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the data deduplication device provided below can refer to the limitations on the data deduplication method in the above text, and will not be repeated here.

[0186] In an exemplary embodiment, as Figure 15 shown, a data deduplication device is provided, including:

[0187] A first acquisition module 1501, configured to acquire deduplication information, where the deduplication information includes the deduplication priority of the target data to be stored and / or the busyness of the current system;

[0188] A first determination module 1502, configured to determine the deduplication strategy for the target data according to the deduplication information.

[0189] In one of the embodiments, the first determination module 1502 is specifically configured to determine a real-time evaluation result of the target data according to the target data, and determine a data complexity evaluation result of the target data according to the target data, where the real-time evaluation result is used to characterize the real-time requirement of the target data; determine the deduplication priority of the target data according to the real-time evaluation result and the data complexity evaluation result.

[0190] In one of the embodiments, the first determination module 1502 is specifically configured to determine a first parameter evaluation result of the real-time parameter of the target data according to the target data, where the real-time parameter includes at least one of the update frequency of the target data, the data usage scenario, and the data transmission delay; determine the real-time evaluation result according to the first parameter evaluation results of the respective real-time parameters.

[0191] In one embodiment, the first determination module 1502 is specifically configured to determine a second parameter evaluation result of a data complexity parameter of the target data according to the target data, where the data complexity parameter includes at least one of a data structure, a data type, a data quality, a data volume, and a data correlation of the target data; and determine the data complexity evaluation result according to the second parameter evaluation results of the data complexity parameters.

[0192] In one embodiment, the first acquisition module 1501 is specifically configured to acquire a task status of the current system; and determine a busy degree of the current system according to the task status of the current system.

[0193] In one embodiment, the deduplication information includes the deduplication priority. The first determination module 1502 is specifically configured to divide a first deduplication level for the target data; and determine a deduplication strategy for the target data according to the first deduplication level and the deduplication priority of the target data.

[0194] In one embodiment, the deduplication information includes the busy degree of the current system. The first determination module 1502 is specifically configured to divide a busy degree level for the current system; and determine a deduplication strategy for the target data according to the busy degree level and the busy degree of the current system.

[0195] In one embodiment, the deduplication information includes the deduplication priority and the busy degree of the current system. The first determination module 1502 is specifically configured to divide a second deduplication level according to the target data and the busy degree of the current system; and determine a deduplication strategy for the target data according to the second deduplication level and the deduplication priority of the target data.

[0196] In one embodiment, the first determination module 1502 is specifically configured to select one of an online deduplication and a post - processing deduplication strategy as the deduplication strategy for the target data according to the deduplication information.

[0197] In an exemplary embodiment, a data storage device is provided, including:

[0198] A second determination module, configured to determine a deduplication strategy for target data by using the data deduplication method according to any one of the above - mentioned first aspects;

[0199] A storage module, configured to perform deduplication processing on the target data according to the deduplication strategy, and store the target data after the deduplication processing into a storage system.

[0200] In an exemplary embodiment, a data compression device is provided, and the device includes:

[0201] A third determination module, configured to determine a deduplication policy for target data by using the data deduplication method described in any one of the above first aspects;

[0202] A compression module, configured to perform deduplication processing on the target data according to the deduplication policy, and perform data compression processing on the target data after the deduplication processing.

[0203] In an exemplary embodiment, a data transmission device is provided, and the device includes:

[0204] A fourth determination module, configured to determine a deduplication policy for target data by using the data deduplication method described in any one of the above first aspects;

[0205] A sending module, configured to perform deduplication processing on the target data according to the deduplication policy, and send the target data after the deduplication processing to a receiving end.

[0206] Each module in the above data deduplication device can be implemented in whole or in part by software, hardware, and their combination. Each of the above modules can be embedded in a processor in a computer device in a hardware form or be independent of the processor, or can be stored in a memory in the computer device in a software form, so that the processor can call and execute operations corresponding to each of the above modules.

[0207] In an exemplary embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented: obtaining deduplication information, where the deduplication information includes the deduplication priority of target data to be stored and / or the busy degree of the current system; determining a deduplication policy for the target data according to the deduplication information.

[0208] In one of the embodiments, when the processor executes the computer program, the following steps are further implemented: determining a real-time evaluation result of the target data according to the target data, and determining a data complexity evaluation result of the target data according to the target data, where the real-time evaluation result is used to characterize the real-time requirement of the target data; determining the deduplication priority of the target data according to the real-time evaluation result and the data complexity evaluation result.

[0209] In one of the embodiments, when the processor executes the computer program, the following steps are further implemented: determining a first parameter evaluation result of real-time parameters of the target data according to the target data, where the real-time parameters include at least one of an update frequency of the target data, a data usage scenario, and a data transmission delay; determining the real-time evaluation result according to the first parameter evaluation results of the respective real-time parameters.

[0210] In one embodiment, when the processor executes the computer program, the following steps are further implemented: determining a second parameter evaluation result of the data complexity parameter of the target data according to the target data, where the data complexity parameter includes at least one of the data structure, data type, data quality, data volume, and data correlation of the target data; determining the data complexity evaluation result according to the second parameter evaluation results of the data complexity parameters.

[0211] In one embodiment, when the processor executes the computer program, the following steps are further implemented: obtaining the task status of the current system; determining the busy degree of the current system according to the task status of the current system.

[0212] In one embodiment, the deduplication information includes the deduplication priority. When the processor executes the computer program, the following steps are further implemented: dividing a first deduplication level for the target data; determining a deduplication strategy for the target data according to the first deduplication level and the deduplication priority of the target data.

[0213] In one embodiment, the deduplication information includes the busy degree of the current system. When the processor executes the computer program, the following steps are further implemented: dividing a busy degree level for the current system; determining a deduplication strategy for the target data according to the busy degree level and the busy degree of the current system.

[0214] In one embodiment, the deduplication information includes the deduplication priority and the busy degree of the current system. When the processor executes the computer program, the following steps are further implemented: dividing a second deduplication level according to the target data and the busy degree of the current system; determining a deduplication strategy for the target data according to the second deduplication level and the deduplication priority of the target data.

[0215] In one embodiment, when the processor executes the computer program, the following steps are further implemented: selecting one of the online deduplication and post - processing deduplication strategies as the deduplication strategy for the target data according to the deduplication information.

[0216] In one embodiment, when the processor executes the computer program, the following steps are further implemented: determining a deduplication strategy for the target data by using the data deduplication method described in any one of the above - mentioned method embodiments; performing deduplication processing on the target data according to the deduplication strategy, and storing the target data after deduplication processing in the storage system.

[0217] In one embodiment, when the processor executes a computer program, the following steps are further implemented: using the data deduplication method described in any one of the above method embodiments, determining a deduplication policy for the target data; performing deduplication processing on the target data according to the deduplication policy, and performing data compression processing on the target data after the deduplication processing.

[0218] In one embodiment, when the processor executes a computer program, the following steps are further implemented: using the data deduplication method described in any one of the above method embodiments, determining a deduplication policy for the target data; performing deduplication processing on the target data according to the deduplication policy, and sending the target data after the deduplication processing to the receiving end.

[0219] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining deduplication information, where the deduplication information includes the deduplication priority of the target data to be stored and / or the busy degree of the current system; determining a deduplication policy for the target data according to the deduplication information.

[0220] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: determining a real-time evaluation result of the target data according to the target data, and determining a data complexity evaluation result of the target data according to the target data, where the real-time evaluation result is used to characterize the real-time requirement of the target data; determining the deduplication priority of the target data according to the real-time evaluation result and the data complexity evaluation result.

[0221] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: according to the target data, determining a first parameter evaluation result of the real-time parameter of the target data, where the real-time parameter includes at least one of the update frequency, data usage scenario, and data transmission delay of the target data; determining the real-time evaluation result according to the first parameter evaluation results of the respective real-time parameters.

[0222] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: according to the target data, determining a second parameter evaluation result of the data complexity parameter of the target data, where the data complexity parameter includes at least one of the data structure, data type, data quality, data volume, and data correlation of the target data; determining the data complexity evaluation result according to the second parameter evaluation results of the respective data complexity parameters.

[0223] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: obtaining the task state of the current system; determining the busy degree of the current system according to the task state of the current system.

[0224] In one embodiment, the deduplication information includes the deduplication priority. When the computer program is executed by a processor, the following steps are implemented: dividing a first deduplication level for the target data; determining a deduplication policy for the target data according to the first deduplication level and the deduplication priority of the target data.

[0225] In one embodiment, the deduplication information includes the busy degree of the current system. When the computer program is executed by a processor, the following steps are implemented: dividing a busy degree level for the current system; determining a deduplication policy for the target data according to the busy degree level and the busy degree of the current system.

[0226] In one embodiment, the deduplication information includes the deduplication priority and the busy degree of the current system. When the computer program is executed by a processor, the following steps are implemented: dividing a second deduplication level according to the target data and the busy degree of the current system; determining a deduplication policy for the target data according to the second deduplication level and the deduplication priority of the target data.

[0227] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: selecting one of the online deduplication and post - processing deduplication policies as the deduplication policy for the target data according to the deduplication information.

[0228] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: using the data deduplication method described in any one of the above - mentioned method embodiments to determine a deduplication policy for the target data; performing deduplication processing on the target data according to the deduplication policy, and storing the target data after deduplication processing in a storage system.

[0229] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: using the data deduplication method described in any one of the above - mentioned method embodiments to determine a deduplication policy for the target data; performing deduplication processing on the target data according to the deduplication policy, and performing data compression processing on the target data after deduplication processing.

[0230] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: using the data deduplication method described in any one of the above - mentioned method embodiments to determine a deduplication policy for the target data; performing deduplication processing on the target data according to the deduplication policy, and sending the target data after deduplication processing to a receiving end.

[0231] In one embodiment, a computer program product is provided, including a computer program which, when executed by a processor, implements the following steps: obtaining deduplication information, where the deduplication information includes the deduplication priority of target data to be stored and / or the busyness level of the current system; determining a deduplication policy for the target data according to the deduplication information.

[0232] In one of the embodiments, when the computer program is executed by a processor, the following steps are implemented: determining a real-time evaluation result of the target data according to the target data, and determining a data complexity evaluation result of the target data according to the target data, where the real-time evaluation result is used to characterize the real-time requirement of the target data; determining the deduplication priority of the target data according to the real-time evaluation result and the data complexity evaluation result.

[0233] In one of the embodiments, when the computer program is executed by a processor, the following steps are implemented: determining a first parameter evaluation result of a real-time parameter of the target data according to the target data, where the real-time parameter includes at least one of the update frequency, data usage scenario, and data transmission delay of the target data; determining the real-time evaluation result according to the first parameter evaluation results of the respective real-time parameters.

[0234] In one of the embodiments, when the computer program is executed by a processor, the following steps are implemented: determining a second parameter evaluation result of a data complexity parameter of the target data according to the target data, where the data complexity parameter includes at least one of the data structure, data type, data quality, data volume, and data correlation of the target data; determining the data complexity evaluation result according to the second parameter evaluation results of the respective data complexity parameters.

[0235] In one of the embodiments, when the computer program is executed by a processor, the following steps are implemented: obtaining the task status of the current system; determining the busyness level of the current system according to the task status of the current system.

[0236] In one of the embodiments, the deduplication information includes the deduplication priority. When the computer program is executed by a processor, the following steps are implemented: dividing a first deduplication level for the target data; determining a deduplication policy for the target data according to the first deduplication level and the deduplication priority of the target data.

[0237] In one of the embodiments, the deduplication information includes the busyness level of the current system. When the computer program is executed by a processor, the following steps are implemented: dividing a busyness level for the current system; determining a deduplication policy for the target data according to the busyness level and the busyness level of the current system.

[0238] In one embodiment, the deduplication information includes the deduplication priority and the busy degree of the current system. When the computer program is executed by a processor, the following steps are implemented: dividing a second deduplication level according to the target data and the busy degree of the current system; determining a deduplication policy for the target data according to the second deduplication level and the deduplication priority of the target data.

[0239] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: selecting one of online deduplication and post-processing deduplication policies as the deduplication policy for the target data according to the deduplication information.

[0240] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: determining a deduplication policy for the target data by using the data deduplication method described in any one of the above method embodiments; performing deduplication processing on the target data according to the deduplication policy, and storing the target data after deduplication processing in a storage system.

[0241] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: determining a deduplication policy for the target data by using the data deduplication method described in any one of the above method embodiments; performing deduplication processing on the target data according to the deduplication policy, and performing data compression processing on the target data after deduplication processing.

[0242] In one embodiment, when the computer program is executed by a processor, the following steps are implemented: determining a deduplication policy for the target data by using the data deduplication method described in any one of the above method embodiments; performing deduplication processing on the target data according to the deduplication policy, and sending the target data after deduplication processing to a receiving end.

[0243] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0244] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0245] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0246] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A data deduplication method, characterized in that, The method includes: Obtaining deduplication information, where the deduplication information includes the deduplication priority of the target data to be stored and / or the busyness degree of the current system; Determining a deduplication strategy for the target data according to the deduplication information.

2. The method according to claim 1, wherein The obtaining of the deduplication information includes: Determining a real-time evaluation result of the target data according to the target data, and determining an evaluation result of the data complexity degree of the target data according to the target data, where the real-time evaluation result is used to characterize the real-time requirement of the target data; Determining the deduplication priority of the target data according to the real-time evaluation result and the evaluation result of the data complexity degree.

3. The method according to claim 2, wherein The determining of the real-time evaluation result of the target data according to the target data includes: Determining a first parameter evaluation result of the real-time parameters of the target data according to the target data, where the real-time parameters include at least one of the update frequency, data usage scenario, and data transmission delay of the target data; Determining the real-time evaluation result according to the first parameter evaluation results of the real-time parameters.

4. The method according to claim 2, wherein The determining of the evaluation result of the data complexity degree of the target data according to the target data includes: Determining a second parameter evaluation result of the data complexity parameters of the target data according to the target data, where the data complexity parameters include at least one of the data structure, data type, data quality, data volume, and data correlation of the target data; Determining the evaluation result of the data complexity degree according to the second parameter evaluation results of the data complexity parameters.

5. The method according to claim 1, characterized in that, The obtaining of the deduplication information includes: Obtaining the task status of the current system; Determining the busyness degree of the current system according to the task status of the current system.

6. The method according to claim 1, wherein The deduplication information includes the deduplication priority. The determining of the deduplication strategy for the target data according to the deduplication information includes: Dividing a first deduplication level for the target data; Determining the deduplication strategy for the target data according to the first deduplication level and the deduplication priority of the target data.

7. The method according to claim 1, wherein The deduplication information includes the busyness degree of the current system. The determining of the deduplication strategy for the target data according to the deduplication information includes: Dividing a busyness degree level for the current system; Determining the deduplication strategy for the target data according to the busyness degree level and the busyness degree of the current system.

8. The method according to claim 1, characterized in that The deduplication information includes the deduplication priority and the busyness degree of the current system. The determining of the deduplication strategy for the target data according to the deduplication information includes: Dividing a second deduplication level according to the target data and the busyness degree of the current system; Determining the deduplication strategy for the target data according to the second deduplication level and the deduplication priority of the target data.

9. The method according to claim 1, characterized in that The determining of the deduplication strategy for the target data according to the deduplication information includes: Selecting one of the online deduplication and post-processing deduplication strategies as the deduplication strategy for the target data according to the deduplication information.

10. A data storage method, characterized in that, The method includes: Determining a deduplication strategy for target data by using the data deduplication method according to any one of claims 1-9; Deduplicate the target data according to the deduplication policy, and store the deduplicated target data in the storage system.

11. A data compression method, characterized in that, The method includes: Determine the deduplication policy of the target data by using the data deduplication method according to any one of claims 1-9; Deduplicate the target data according to the deduplication policy, and perform data compression processing on the deduplicated target data.

12. A data transmission method, characterized in that, The method includes: Determine the deduplication policy of the target data by using the data deduplication method according to any one of claims 1-9; Deduplicate the target data according to the deduplication policy, and send the deduplicated target data to the receiving end.

13. A data deduplication device, characterized in that, The device includes: A first acquisition module, configured to acquire deduplication information, where the deduplication information includes the deduplication priority of the target data to be stored and / or the busy degree of the current system; A first determination module, configured to determine the deduplication policy of the target data according to the deduplication information.

14. A data storage device, characterized in that, The device includes: A second determination module, configured to determine the deduplication policy of the target data by using the data deduplication method according to any one of claims 1-9; A storage module, configured to deduplicate the target data according to the deduplication policy, and store the deduplicated target data in the storage system.

15. A data compression device, characterized in that, The device includes: A third determination module, configured to determine the deduplication policy of the target data by using the data deduplication method according to any one of claims 1-9; A compression module, configured to deduplicate the target data according to the deduplication policy, and perform data compression processing on the deduplicated target data.

16. A data transmission device, characterized in that, The device includes: A fourth determination module, configured to determine the deduplication policy of the target data by using the data deduplication method according to any one of claims 1-9; A sending module, configured to deduplicate the target data according to the deduplication policy, and send the deduplicated target data to the receiving end.

17. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.

18. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.

19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.