Data processing method and device

By performing key generation calculations on the sub-block data of plaintext blocks, it ensures that highly similar plaintext blocks are encrypted with the same key, which solves the problem of redundant storage in encrypted data deduplication technology, and achieves the reduction of storage costs and the improvement of transmission performance.

CN120296751APending Publication Date: 2025-07-11HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410045760.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing encrypted data deduplication technology cannot effectively eliminate the redundancy between ciphertext blocks corresponding to similar plaintext blocks, resulting in higher storage costs.

Method used

Key generation calculations are performed on the sub-block data of plaintext blocks to ensure that highly similar plaintext blocks are encrypted with the same key, thereby retaining the similarity of blocks so as to subsequently remove redundant data between similar ciphertext blocks.

Benefits of technology

By preserving block similarity, reduce storage space, reduce storage costs, and improve data transfer performance and encryption speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296751A_ABST
    Figure CN120296751A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, and relates to the field of computers. The method is applied to first equipment and comprises the steps that a first secret key corresponding to a first plaintext block is determined, the first secret key is obtained after secret key generation calculation is conducted on N sub-block data of the first plaintext block, the first secret key is the same as a second secret key, the second secret key is the secret key corresponding to a second plaintext block similar to the first plaintext block, and N is a positive integer; the second key is obtained by performing key generation calculation on W sub-block data of the second plaintext block, and both N and W are positive integers greater than 1; and based on the first key, encrypting the first plaintext block to obtain a first ciphertext block. The method and the device are used for reserving the similarity of the blocks, so that redundant data between similar ciphertext blocks can be removed subsequently, the storage space is reduced, and the storage cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and in particular, to a data processing method and apparatus. Background Art

[0002] In the era of big data, data has exploded in growth, and the demand for data storage is also increasing continuously. To reduce storage overhead, currently commonly used Outsourced Storage Systems (OSS) adopt encrypted data deduplication technology to perform data deduplication on encrypted data. However, since the key of the encrypted data deduplication technology is obtained by performing a hash calculation on the data content of the plaintext block, this means that for duplicate plaintext blocks, the keys used during encryption are the same, and the resulting ciphertext blocks are also duplicate; while for similar plaintext blocks, the keys used during encryption are completely different, and the resulting ciphertext blocks will also become completely different.

[0003] Since the similarity of similar plaintext blocks has been severely damaged after encryption, the traditional encrypted data deduplication technology can only eliminate duplicate ciphertext blocks, and cannot eliminate the redundancy between ciphertext blocks corresponding to similar plaintext blocks. When the number of similar plaintext blocks occupies a large proportion, it will occupy a large storage space, thereby resulting in a high storage cost. Summary of the Invention

[0004] This application provides a data processing method and apparatus, which are used to retain the similarity of blocks, so as to facilitate subsequent removal of redundant data between similar ciphertext blocks, reduce the storage space, and lower the storage cost.

[0005] To achieve the above object, this application adopts the following technical solutions.

[0006] In a first aspect, an embodiment of this application provides a data processing method, which is applied to a first device and includes: determining a first key corresponding to a first plaintext block, where the first key is obtained by performing key generation calculations on N sub-block data of the first plaintext block respectively, the first key is the same as a second key, the second key is a key corresponding to a second plaintext block having similarity with the first plaintext block, and the second key is obtained by performing key generation calculations on W sub-block data of the second plaintext block respectively, and both N and W are positive integers greater than 1; encrypting the first plaintext block based on the first key to obtain a first ciphertext block.

[0007] In the above method, since the first key is obtained by performing key generation calculations on the N sub-block data of the first plaintext block respectively, and the second key is obtained by performing key generation calculations on the W sub-block data of the second plaintext block respectively, this means that the embodiments of the present application propose a new key generation method, that is, the key is not directly obtained by performing key generation calculations on the plaintext block itself, but is obtained by performing key generation calculations on multiple sub-block data included in the plaintext block itself. Among them, the second plaintext block in the embodiments of the present application refers to a plaintext block similar to the first plaintext block, that is, a plaintext block with a similarity to the first plaintext block reaching a similarity threshold (for example, 80%), where the similarity here can be the similarity determined based on character comparison, or the similarity determined based on super eigenvalue comparison, or the similarity determined in other forms, and it will not be limited here. In other words, there are a certain number of common factors between the first plaintext block and the second plaintext block, and the first key is the same as the second key, which means that this new key generation method can effectively ensure that for highly similar plaintext blocks, the finally determined keys are the same. Then, subsequently, when using the same key to encrypt similar plaintext blocks, the similarity existing in the plaintext blocks can be retained, that is, the obtained ciphertext blocks are also similar. Thus, it can be seen that the data processing method proposed in the embodiments of the present application can retain the similarity of blocks, so as to facilitate subsequent removal of redundant data between similar ciphertext blocks, thereby reducing storage space and storage costs.

[0008] In one implementation, the first plaintext block is any plaintext block obtained by dividing the original data into blocks, and the first plaintext block includes N sub-block data; determining the first key corresponding to the first plaintext block includes: determining the hash values corresponding to the N sub-block data respectively; and determining the target hash value among the N hash values as the first key corresponding to the first plaintext block.

[0009] In the above implementation, the key generation calculation can take hash calculation as an example, that is, the first key is the target hash value selected from the hash values corresponding to the N sub-block data respectively, rather than the hash value obtained by performing hash calculation on the data content of the first plaintext block. This key generation method makes the keys obtained from similar plaintext blocks the same, so as to effectively retain the similarity of blocks.

[0010] In one implementation, each of the N sub-block data is obtained by sliding the first sliding window on the first plaintext block, and the first sliding window and the second sliding window are the same sliding window, and the second sliding window is used to divide the original data into blocks.

[0011] In the above implementation, since the first sliding window is the same as the second sliding window, this means that the generation of the first key does not require a new sliding window to calculate the first plaintext block again, but directly selects from the N hash values calculated during the chunking process, which can greatly reduce the computational overhead and improve the key generation speed.

[0012] In one implementation, the first device is provided with a buffer for storing target hash values during the chunking process; determining the first key corresponding to the first plaintext block includes: traversing the original data using the first sliding window; determining the data covered by the first sliding window at position j as sub-block data j, and determining the hash value H of sub-block data j j , where j is used to indicate the distance between the current window position and the first boundary point position, and j is a positive integer less than or equal to N; updating the target hash value based on the hash value H j ; until it is determined that the position j reaches the second boundary point, determining the data block between the second boundary point and the first boundary point as the first plaintext block, and determining the target hash value in the updated buffer as the first key corresponding to the first plaintext block.

[0013] In the above implementation, the first device records the target hash values during the chunking process by setting a buffer, which means that the first device can couple the key generation process with the data chunking process, that is, when the chunking of the first plaintext block ends, the first key of the first plaintext block is also determined from the buffer, which can eliminate the additional computational overhead generated during the key generation process and improve the key generation speed.

[0014] In one implementation, the target hash value is the maximum hash value or the minimum hash value.

[0015] In the above implementation, a new key generation method is implemented by using the maximum hash value or the minimum hash value as the key for the similarity of the reserved blocks, so as to facilitate the subsequent implementation of differential coding of similar ciphertext blocks, thereby reducing the storage cost.

[0016] In one implementation, the first key is the same as the third key corresponding to the third plaintext block, the third plaintext block is the plaintext block detected in the first device and having similarity with the first plaintext block, and the third key is obtained by performing key generation calculations on P sub-block data in the third plaintext block respectively, where P is a positive integer greater than 1; encrypting the first plaintext block based on the first key to obtain the first ciphertext block includes: determining the first difference data between the first plaintext block and the third plaintext block; encrypting the first difference data based on the first key, and determining the encrypted first difference data as the first ciphertext block.

[0017] In the above implementation manner, before encryption, the first device may first detect whether there are similar blocks locally on the first device. If so, it can directly encrypt the differential data between the similar blocks. In this way, when the data needs to be transmitted to other devices (for example, the second device) subsequently, the data transmission volume can be effectively reduced and the transmission performance can be improved.

[0018] In one implementation manner, the first plaintext block is any plaintext block obtained after the original data is segmented. The original data includes the first data and the second data. The first plaintext block is the last segmented data in the first data. The method further includes: segmenting the second data until a fourth plaintext block is obtained. The fourth plaintext block is the next plaintext block of the first plaintext block in the original data; determining a fourth key corresponding to the fourth plaintext block, where the fourth key is obtained by respectively performing key generation calculations on Q sub-block data in the fourth plaintext block, and Q is a positive integer greater than 1; encrypting the fourth plaintext block based on the fourth key.

[0019] In the above implementation manner, the first device does not need to wait to obtain all the plaintext blocks in the original data before performing the encryption process. Instead, after each plaintext block is segmented, the corresponding key is also generated at the same time. In this way, it can be encrypted immediately, thereby reducing the I / O overhead and further accelerating the encryption process.

[0020] In one implementation manner, the method further includes: obtaining R local ciphertext blocks, where the R local ciphertext blocks include the first ciphertext block, and R is a positive integer greater than 1; performing a similarity detection on the R local ciphertext blocks to obtain a similarity result; if the similarity result indicates that there is a second ciphertext block in the R local ciphertext blocks that is similar to the first ciphertext block, then performing differential encoding on the first ciphertext block and the second ciphertext block.

[0021] In the above implementation manner, the first device can perform differential compression on the similar ciphertext blocks locally, so that the redundant data in the similar ciphertext blocks can be effectively eliminated locally on the first device, so as to further reduce the storage space occupied by the encrypted data locally on the first device and reduce the storage cost of the first device.

[0022] In a second aspect, an embodiment of the present application provides a data processing method, which is applied to a second device. The method includes: obtaining M ciphertext blocks, where the M ciphertext blocks include the first ciphertext block, and M is a positive integer greater than 1; performing a similarity detection on the M ciphertext blocks, and determining the ciphertext block that is similar to the first ciphertext block as the second ciphertext block. The first key corresponding to the first ciphertext block is the same as the second key corresponding to the second ciphertext block; performing differential encoding on the first ciphertext block and the second ciphertext block.

[0023] In the above implementation manner, since the first key is the same as the second key, this means that the similarity of the block can be retained between the first ciphertext block and the second ciphertext block. Therefore, compared with storing M ciphertext blocks, the second device can greatly reduce the space occupied during data storage and lower the storage cost of the second device by removing the redundant data of the similar ciphertext blocks.

[0024] In one implementation manner, the M ciphertext blocks are the ciphertext blocks counted by the second device during the compression period, and the M ciphertext blocks are sent by X first devices, where X is a positive integer.

[0025] In the above implementation manner, the second device can not only perform differential compression on the ciphertext blocks sent by a certain first device, but also perform differential compression on the ciphertext blocks sent by multiple first devices. When the number of ciphertext blocks is large, more second ciphertext blocks can be detected, which can greatly reduce the redundant data of the similar ciphertext blocks and lower the storage cost.

[0026] In one implementation manner, if the second device is a network edge device, then all X first devices are located in the area corresponding to the network edge device.

[0027] In the above implementation manner, differential compression of similar ciphertext blocks is realized in the field of network deduplication, and the storage cost of the network edge device in the field of network deduplication is reduced.

[0028] In a third aspect, an embodiment of the present application provides a data processing device for implementing the method in the foregoing first aspect or any implementation manner in the first aspect. Specifically, the device includes: a determination unit configured to determine a first key corresponding to a first plaintext block, where the first key is obtained by respectively performing key generation calculations on N sub-block data of the first plaintext block, the first key is the same as a second key, the second key is a key corresponding to a second plaintext block having similarity to the first plaintext block, and the second key is obtained by respectively performing key generation calculations on W sub-block data of the second plaintext block, and both N and W are positive integers greater than 1; an encryption unit configured to encrypt the first plaintext block based on the first key to obtain a first ciphertext block.

[0029] In one implementation manner, the first plaintext block is any plaintext block obtained by partitioning the original data, and the first plaintext block includes N sub-block data; the determination unit configured to determine the first key corresponding to the first plaintext block includes: a first determination subunit configured to determine the hash values respectively corresponding to the N sub-block data; the first determination subunit is further configured to determine the target hash value among the N hash values as the first key corresponding to the first plaintext block.

[0030] In one implementation, each of the N sub-block data is obtained by sliding a first sliding window over a first plaintext block. The first sliding window and the second sliding window are the same sliding window, and the second sliding window is used to block the original data.

[0031] In one implementation, a first device is provided with a buffer for storing a target hash value during the blocking process, and a determination unit for determining a first key corresponding to the first plaintext block, including: a traversal subunit for traversing the original data using the first sliding window; a second determination subunit for determining the data covered by the first sliding window at position j as sub-block data j, and determining the hash value H of sub-block data j j , where j is used to indicate the distance between the current window position and the position of the first boundary point, and j is a positive integer less than or equal to N; an update subunit for updating the target hash value based on the hash value H j . The second determination subunit is further configured to, when it is determined that the position j reaches the second boundary point, determine the data block between the second boundary point and the first boundary point as the first plaintext block, and determine the target hash value in the updated buffer as the first key corresponding to the first plaintext block.

[0032] In one implementation, the target hash value is the maximum hash value or the minimum hash value.

[0033] In one implementation, the first key is the same as the third key corresponding to the third plaintext block. The third plaintext block is a plaintext block detected in the first device and having similarity with the first plaintext block. The third key is obtained by respectively performing key generation calculations on P sub-block data in the third plaintext block, where P is a positive integer greater than 1. The encryption unit is configured to encrypt the first plaintext block based on the first key to obtain a first ciphertext block, including: a third determination subunit for determining the first difference data between the first plaintext block and the third plaintext block; an encryption subunit for encrypting the first difference data based on the first key and determining the encrypted first difference data as the first ciphertext block.

[0034] In one implementation, the first plaintext block is any plaintext block obtained by blocking the original data. The original data includes first data and second data. The first plaintext block is the last segmented data in the first data. The apparatus further includes: a blocking unit for blocking the second data until a fourth plaintext block is obtained. The fourth plaintext block is the next plaintext block of the first plaintext block in the original data. The determination unit is further configured to determine a fourth key corresponding to the fourth plaintext block. The fourth key is obtained by respectively performing key generation calculations on Q sub-block data in the fourth plaintext block, where Q is a positive integer greater than 1. The encryption unit is further configured to encrypt the fourth plaintext block based on the fourth key.

[0035] In one implementation, the device further includes: an obtaining unit, configured to obtain R local ciphertext blocks, where the R local ciphertext blocks include a first ciphertext block, and R is a positive integer greater than 1; a detecting unit, configured to perform similarity detection on the R local ciphertext blocks to obtain a similarity result; and an encoding unit, configured to perform differential encoding on the first ciphertext block and a second ciphertext block if the similarity result indicates that there is a second ciphertext block similar to the first ciphertext block among the R local ciphertext blocks.

[0036] In a fourth aspect, an embodiment of the present application provides a data processing device for implementing the method according to the second aspect or any implementation manner of the second aspect. Specifically, the device includes: an obtaining unit, configured to obtain M ciphertext blocks, where the M ciphertext blocks include a first ciphertext block, and M is a positive integer greater than 1; a detecting unit, configured to perform similarity detection on the M ciphertext blocks and determine a ciphertext block similar to the first ciphertext block as a second ciphertext block, where a first key corresponding to the first ciphertext block is the same as a second key corresponding to the second ciphertext block; and an encoding unit, configured to perform differential encoding on the first ciphertext block and the second ciphertext block.

[0037] In one implementation, the M ciphertext blocks are the ciphertext blocks statistically obtained by the second device during a compression period, and the M ciphertext blocks are sent by X first devices, where X is a positive integer.

[0038] In one implementation, if the second device is a network edge device, then all the X first devices are located in the area corresponding to the network edge device.

[0039] In a fifth aspect, an embodiment of the present application provides a data processing system, which includes a first device and a second device; the first device is configured to execute the first aspect or any possible implementation manner of the first aspect, and the second device is configured to execute each module of the data processing method according to the second aspect or any possible implementation manner of the second aspect.

[0040] Among them, the data processing system has the functions of implementing the methods in the above first aspect and second aspect. The functions can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions. In a possible design, the data processing system includes: a first device determines a first key corresponding to a first plaintext block, and the first key is obtained by performing key generation calculations on N sub-block data of the first plaintext block respectively; the first device encrypts the first plaintext block based on the first key to obtain a first ciphertext block; the first device sends the first ciphertext block to a second device; the second device obtains M ciphertext blocks including the first ciphertext block; the second device performs similarity detection on the M ciphertext blocks, and determines the ciphertext block having similarity with the first ciphertext block as a second ciphertext block, and the second key corresponding to the second ciphertext block is the same as the first key, and the second key is obtained by performing key generation calculations on W sub-block data of a second plaintext block respectively, and N, M, and W are all positive integers greater than 1; the second device performs differential encoding on the first ciphertext block and the second ciphertext block.

[0041] In a sixth aspect, an embodiment of the present application provides a data processing device, including a memory and a processor. The memory is used to store computer instructions, and the processor is used to call and run the computer instructions from the memory to implement the method in the first aspect or any implementation manner in the first aspect, or to implement the method in the second aspect or any implementation manner in the second aspect.

[0042] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a processor, the method in the first aspect or any implementation manner in the first aspect is implemented, or the method in the second aspect or any implementation manner in the second aspect is implemented.

[0043] In an eighth aspect, an embodiment of the present application provides a computer program product, and the computer program product includes instructions. When the instructions run on a computer, the computer is caused to execute the method in the first aspect or any implementation manner in the first aspect, or to execute the method in the second aspect or any implementation manner in the second aspect.

[0044] The technical effects generated by the above second aspect to eighth aspect and any implementation manner in each aspect can refer to the first aspect and the corresponding implementation manner in the first aspect, and the repeated parts will not be elaborated here. Description of the Drawings

[0045] Figure 1 It is a schematic structural diagram of a system architecture provided by an embodiment of the present application;

[0046] Figure 2It is a schematic diagram of a scenario provided by an embodiment of the present application for encrypting similar plaintext blocks using the same key;

[0047] Figure 3 It is a method interaction diagram for data processing provided by an embodiment of the present application;

[0048] Figure 4 It is a schematic diagram of a scenario provided by an embodiment of the present application for encrypting plaintext blocks based on a sliding window;

[0049] Figure 5 It is a schematic diagram of a scenario provided by an embodiment of the present application for coupling the data chunking process and the encryption process;

[0050] Figure 6 It is a method schematic for data processing provided by an embodiment of the present application Figure 1 ;

[0051] Figure 7 It is a schematic diagram of a scenario provided by an embodiment of the present application for encrypting plaintext blocks;

[0052] Figure 8 It is a schematic diagram of a scenario provided by an embodiment of the present application for pipelining the chunking process and the encryption process;

[0053] Figure 9 It is a method schematic for data processing provided by an embodiment of the present application Figure 2 ;

[0054] Figure 10 It is a method schematic for data processing provided by an embodiment of the present application Figure 3 ;

[0055] Figure 11 It is a structural schematic of a data processing device provided by an embodiment of the present application Figure 1 ;

[0056] Figure 12 It is a structural schematic of a data processing device provided by an embodiment of the present application Figure 2 ;

[0057] Figure 13 It is a structural schematic of a data processing device provided by an embodiment of the present application Figure 3 ;

[0058] Figure 14 It is a structural schematic of a data processing device provided by an embodiment of the present application Figure 4 。 Detailed implementation manners

[0059] The following will describe the technical solutions in the embodiments of the present application in combination with the accompanying drawings in the embodiments of the present application. Among them, in order to clearly describe the technical solutions in the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that the terms such as "first" and "second" do not limit the quantity and execution order, and the terms such as "first" and "second" do not necessarily limit to be different. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "exemplary" or "for example" aims to present relevant concepts in a specific way for easy understanding.

[0060] To facilitate the understanding of the technical solutions provided in the embodiments of the present application, first, relevant terms related to data processing in the embodiments of the present application are introduced:

[0061] 1. Encrypted Deduplication (hereinafter referred to as encrypted deduplication) technology

[0062] The encrypted deduplication technology can usually organize the original data into plaintext chunks (hereinafter referred to as plaintext blocks), and then encrypt the plaintext blocks according to the keys of the plaintext blocks to obtain ciphertext chunks (hereinafter referred to as ciphertext blocks), and then eliminate the duplicate ciphertext blocks.

[0063] 2. Delta Compression technology

[0064] The delta compression technology first detects the similarity of the blocks, and then eliminates data redundancy by storing the differences between the similar blocks.

[0065] 3. Encrypted Delta Compression (EDC)

[0066] Encrypted delta compression means realizing delta compression on the basis of the encrypted deduplication technology.

[0067] The following briefly introduces the related technologies of the embodiments of the present application.

[0068] In the current technical solution, differential compression is another popular data compression technology in storage systems and is widely used in backup / archiving applications. It can eliminate redundancy between non-repetitive but highly similar blocks. For example, a storage system can use differential compression technology to detect the similarity of multiple plaintext blocks to be stored. If it is detected that plaintext block P1 and plaintext block P2 are similar, then plaintext block P1 can be used as the base-chunk. After differential compression, the storage system only stores plaintext block P1, the difference between plaintext block P1 and plaintext block P2, and the mapping information between plaintext block P1 and plaintext block P2. This can eliminate duplicate data between similar blocks and reduce storage costs.

[0069] However, in special scenarios with high data security requirements, the data to be differentially compressed is often already encrypted plaintext blocks (i.e., ciphertext blocks). Since the key is obtained by performing a hash calculation on the data content of the plaintext block based on an encryption hash function, but this encryption hash function often exhibits the avalanche effect, which means that even if there are minor changes between two plaintext blocks, it will result in completely different finally generated hash values, that is, the keys generated by similar plaintext blocks are completely different. And using such completely different keys to encrypt similar plaintext blocks will destroy the similarity between the similar plaintext blocks, that is, two completely different ciphertext blocks are obtained. That is, the ciphertext blocks obtained from similar plaintext blocks are not similar, which means that differential compression cannot be applied to eliminate redundant data between ciphertext blocks and can only be applied to eliminate redundant data between plaintext blocks, resulting in higher storage costs.

[0070] Thus, in scenarios with high data security requirements, how to preserve the similarity of blocks for subsequent encrypted differential compression is the key to further reducing storage costs.

[0071] To solve the above problems, the embodiments of the present application provide a data processing method. In this method, a first device can determine a first key corresponding to a first plaintext block. Here, the first key is obtained by performing key generation calculations on N sub-block data of the first plaintext block respectively. The first key is the same as the second key, and the second key is the key corresponding to a second plaintext block that is similar to the first plaintext block. The second key is obtained by performing key generation calculations on W sub-block data of the second plaintext block respectively. Both N and W are positive integers greater than 1. Further, the first device can encrypt the first plaintext block based on the first key to obtain a first ciphertext block.

[0072] In the embodiments of the present application, since the first key is obtained by respectively performing key generation calculations on N sub-block data of the first plaintext block, and the second key is obtained by respectively performing key generation calculations on W sub-block data of the second plaintext block, this means that the embodiments of the present application propose a new key generation method, that is, the key is not directly obtained by performing key generation calculation on the plaintext block itself, but is obtained by respectively performing key generation calculations on multiple sub-block data included in the plaintext block itself. Among them, the first key is the same as the second key, which means that this new key generation method can effectively ensure that for highly similar plaintext blocks, the finally determined keys are the same. Then, subsequently, when using the same key to encrypt similar plaintext blocks, the similarity existing in the plaintext blocks can be retained, that is, the obtained ciphertext blocks are also similar. Thus, it can be seen that the data processing method proposed in the embodiments of the present application can retain the similarity of blocks, so as to facilitate subsequent removal of redundant data between similar ciphertext blocks, thereby reducing the storage space and storage cost.

[0073] It can be understood that the data processing method provided in the embodiments of the present application can be applied to other fields such as data storage scenarios or network deduplication scenarios, etc., and will not be exemplified one by one here.

[0074] In the data storage scenario, the embodiments of the present application can be deployed as an encryption module in the client of the storage system, or can be written into the entire storage system as a complete encrypted differential compression module, which will not be limited here. When the storage system includes a first device and a second device, the first device here can be a client for encrypting plaintext blocks, and the second device can be a server providing storage services. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0075] With the development of cloud computing, outsourcing data storage has become an inevitable trend and fashion. For example, outsourcing storage to the cloud is a common method for enterprises to save self-management overhead, which can relieve the pressure on storage capacity. In other words, the data processing method provided in the embodiments of the present application can be an encrypted data differential compression method for cloud storage. Among them, cloud storage is a new concept extended and developed on the basis of the cloud computing concept. A distributed cloud storage system refers to a storage system that combines a large number of different types of storage devices (storage devices are also called storage nodes) in the network through functions such as cluster applications, grid technologies, and distributed file systems, and collaborates through application software or application interfaces to jointly provide data storage and business access functions to the outside world.

[0076] In the network deduplication scenario, the second device here can be a network edge device (e.g., a switch) for governing a certain area, and the first device (i.e., the data sender) can be any device (e.g., a user terminal) in the area governed by the network edge device. To ensure the security of data transmission, the first device can use the above new encryption method to encrypt the plaintext block to obtain a ciphertext block that can retain similarity. After receiving ciphertext blocks sent by multiple first devices, the second device can perform differential compression on the received ciphertext blocks and send the data obtained after differential compression to the third device (i.e., the data receiver) to reduce the network transmission load.

[0077] The following introduces the network architecture applying the data processing method provided by the embodiments of the present application:

[0078] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of a system architecture provided by the embodiments of the present application. As Figure 1 shown, this system architecture can be the architecture corresponding to the storage system. Among them, this system architecture provided by the embodiments of the present application mainly includes four modules, specifically including the Chunking module, the Encrypting module, the Deduplication module, and the Delta Compression module.

[0079] It should be noted that in a special case, the Chunking module, the Encrypting module, the Deduplication module, and the Delta Compression module here can all be deployed in one device. For example, the above four modules can all be deployed in a server; or when the resource computing power of the client is strong enough, the above four modules can all be deployed in the client. Generally, due to the limited resource computing power of the client, to save the management overhead of the corresponding objects of the client (e.g., enterprises, schools, etc.), the deployment method shown in Figure 1 is usually adopted, that is, the Chunking module and the Encrypting module are deployed in the client, and the Deduplication module and the Delta Compression module are deployed in the server.

[0080] As Figure 1 shown, the storage system can include a server 100 and a client cluster. Among them, the client cluster here can include one or more clients, and the number of clients will not be limited here. As Figure 1As shown, the client cluster can be network-connected to the above-mentioned server 100 respectively, so that each client can perform data interaction with the server 100 through this network connection. Herein, the network connection is not limited to the connection method, and can be directly or indirectly connected through a wired communication method, or can be directly or indirectly connected through a wireless communication method, or can also be connected through other methods, which are not limited in this application.

[0081] Among them, each client in the client cluster can run on an intelligent terminal with data processing capabilities, and the intelligent terminal can include: smart phones, tablet computers, laptop computers, desktop computers, smart speakers, smart watches, vehicle-mounted terminals, smart TVs, etc. It should be understood that the client here can include social clients, multimedia clients (such as video clients), entertainment clients (such as game clients), information flow clients, education clients, live broadcast clients, etc. Among them, the client can be an independent client or an embedded sub-client integrated in a certain client (such as a social client, an education client, and a multimedia client, etc.), which is not limited herein.

[0082] As Figure 1 shown, the server 100 in the embodiment of this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Among them, the number of servers will not be limited in the embodiment of this application.

[0083] For the convenience of description, the embodiment of this application takes the Figure 1 deployment method shown as an example to explain the above four modules respectively:

[0084] Chunking Module: For any client, the chunking module may specifically include a data chunking process and a key generation process. Among them, in the data chunking process, the chunking algorithm used to implement data chunking can be a sliding window-based hash algorithm (for example, Rabin-based Chunking, Rabin chunking algorithm). The Rabin chunking algorithm can take the size of the target average block (for example, 8KB) and the sliding window size (for example, 48 bytes) as inputs. By sliding the sliding window over the original data, the sub-block data at each position is determined, and by calculating the hash value of the sub-block data, the block boundaries are identified, thereby realizing the chunking of the original data. Among them, in the key generation process, the key here refers to a similarity-preserving encryption key (Similarity-Preserving Key, SP-Key), which is obtained by performing key generation calculations on the multiple sub-block data included in a certain plaintext block respectively. The embodiments of the present application can be described by taking any plaintext block (for example, plaintext block P1) in the original data as an example, that is, the key of the plaintext block P1 can be the target hash value (for example, the maximum hash value or the minimum hash value) directly selected from the hash values corresponding to the multiple sub-block data after the data chunking process is completed. Optionally, in order to effectively improve the client throughput rate and improve the system performance, the embodiments of the present application can also set a buffer to realize the generation of the key. The buffer here can be used to update the target hash value of the current block in the chunking process.

[0085] Encryption Module: The encryption module can use the SP-Key to encrypt the plaintext block, and realizes similarity-preserving encryption on the client. The embodiments of the present application can implement bitwise XOR encryption operations based on the Convergent All-or-nothing Transform (CAONT) method used in the encryption deduplication method. It can be understood that in order to reduce the additional I / O overhead of the system and accelerate the process of generating the encryption key, the embodiments of the present application can use multi-threading to implement a chunking encryption pipeline, so that the chunking module and the encryption module can run synchronously.

[0086] Data Deduplication Module: The data deduplication module aims to eliminate the redundancy of duplicate ciphertext blocks. In this data deduplication module, the embodiment of the present application implements a fingerprint index, maps the fingerprint of each ciphertext block to the physical address storing the ciphertext block, and uses a key-value pair-based storage database (e.g., LevelDB) to store the fingerprint index of the ciphertext blocks. Among them, the key is used to store the fingerprint of the ciphertext block, and the value is used to store the physical address of the ciphertext block. For storage on the disk, containers can be used to reduce the disk I / O overhead. The containers pack fine-grained unique blocks (with an average unit of KB) into a coarse-grained container data structure (with an average unit of MB) according to the positions of these ciphertext blocks.

[0087] Differential Compression Module: The differential compression module aims to reduce the redundancy of similar ciphertext blocks. This differential compression module specifically includes three stages: Similarity Detection stage, Basic Block Reading stage, and Delta Encoding stage. The similarity detection method in the embodiment of the present application can be a detection method based on Super-features (SF), or other detection methods, which will not be limited here. In the delta encoding stage, based on LevelDB, the basic blocks generated after delta encoding, their corresponding delta data, and the index data corresponding to the delta data can be stored.

[0088] It can be understood that the embodiment of the present application proposes and proves two theoretical supports for similarity-preserving encryption. Among them, the first theorem shows that after similar plaintext blocks are encrypted using the same hash key (e.g., bitwise XOR encryption), the generated ciphertext blocks still remain similar. The second theorem shows that for the set of hash values generated by multiple highly similar plaintext blocks, the target hash value (the maximum hash value or the minimum hash value) has a very high probability of being the same.

[0089] Based on these two theorems, the embodiment of the present application further designs a practical scheme for similarity-preserving encryption. This scheme uses a sliding window algorithm to calculate the set of hash values of the blocks, and according to the proposed theorem, selects the target hash value as the encryption key for similarity preservation to preserve the similarity of the blocks.

[0090] Furthermore, for the convenience of understanding the proof process of the first theorem, please refer to Figure 2 , Figure 2 which is a schematic diagram of a scenario where the same key is used to encrypt similar plaintext blocks provided by the embodiment of the present application. As shown in Figure 2As shown, the embodiments of the present application may take two plaintext blocks as an example, specifically including plaintext block P1 (for example, "100101001001") and plaintext block P2 (for example, "100101001010"). According to observation, both plaintext block P1 and plaintext block P2 are 12 bits, and the first 10 bits have the same value, which means that plaintext block P1 and plaintext block P2 belong to two similar plaintext blocks.

[0091] Among them, the encryption method adopted in the embodiments of the present application is bitwise exclusive OR, that is, when the values at the same digit are the same, the exclusive OR result is 0; when the values at the same digit are different, the exclusive OR result is 1. It should be understood that when using Figure 2 the shown key K (for example, "001100110011") to encrypt plaintext block P1, the ciphertext block C1 corresponding to plaintext block P1 (for example, "101001111010") can be obtained. After using the same key K to encrypt plaintext block P2, the ciphertext block C2 corresponding to plaintext block P2 (for example, "101001111001") can be obtained.

[0092] Based on the above observation, the first 10 bits of ciphertext block C1 and ciphertext block C2 still remain the same, which means that when using the same key to encrypt two similar plaintext blocks, with the help of bitwise exclusive OR, the same parts of the similar plaintext blocks will remain the same in the ciphertext blocks, that is, ciphertext block C1 and ciphertext block C2 are still similar. Obviously, for multiple similar plaintext blocks, this encryption method using the same key is still applicable, that is, multiple similar plaintext blocks are encrypted with the same key, and multiple similar ciphertext blocks can be generated.

[0093] The embodiments of the present application can summarize the first theorem, that is, n similar plaintext blocks (for example, plaintext block P1, plaintext block P2,..., plaintext block P n ), after being encrypted by bitwise exclusive OR with the same key, the n generated ciphertext blocks (for example, ciphertext block C1, ciphertext block C2,..., ciphertext block C n ) are also similar.

[0094] Since each first device (for example, the client) does not know whether its plaintext block is similar to the data of other first devices, and at the same time needs to ensure that similar blocks can be differentially compressed on the second device (for example, the cloud server), it is necessary to generate the same hash value for possible similar plaintext blocks for encryption on the client. For differential compression, similarity detection for identifying similar blocks is a key step. Similarity detection methods generally involve calculating and comparing the super feature values of each data block, and these super features are composed of multiple features calculated based on Rabin fingerprints. Specifically, the feature calculation method based on the Rabin sliding window algorithm can be seen in the following formula (1):

[0095]

[0096] Among them, L is used to represent the length of the data block, m i and a i are both predefined random values. Rabin j is used to represent the hash value corresponding to position j, which is obtained by calculating the content (i.e., sub-block data) of the sliding window (the window size can be 48 bytes) located at position j using the hash algorithm.

[0097] It can be understood that the Broder theorem proves that for the two hash value sets H(A) and H(B), the probability that H(A) and H(B) have the same maximum (or minimum) hash element is equal to the similarity between set A and set B. Among them, H(A) is the set of hash values obtained by calculating each element inside set A, and H(B) is the set of hash values obtained by calculating each element inside set B. Further, traditional similarity detection methods can be based on the Broder theorem (also known as the Min-Hash theorem, Min-Hash) and use the SF of the blocks to detect highly similar blocks.

[0098] According to the above formula (1), it can be seen that the Broder theorem actually implies that the similarity degree of two blocks is highly correlated with their maximum hash values, but the Broder theorem is not applicable to the case of n blocks. Therefore, the embodiments of the present application can generalize the Broder theorem from 2 similar blocks to the case of n similar blocks, so as to prove whether the similarity of n (n>2) similar blocks is also highly correlated with their maximum hash values. Based on this, the embodiments of the present application propose and prove the second theorem, which can be specifically referred to the following formula (2):

[0099]

[0100] Among them, H(A) is the set of corresponding hash values generated by all elements in set A through the hash function H; H(B) is the set of corresponding hash values generated by all elements in set B through the hash function H; and so on, H(N) is the set of corresponding hash values generated by all elements in set N through the hash function H; max(S) can be used to represent the maximum element in set S.

[0101] Among them, the proof process of the second theorem can be referred to the following formulas (3)-(6):

[0102] Proof: According to the definition of H, if then it can be obtained that:

[0103] H(α) ∈ H(A) ∪ H(B) ∪... ∪ H(N) (3)

[0104] According to formula (3) and the definition of H, we have:

[0105]

[0106] Let H(α max ) be the maximum H(α), that is such that:

[0107] H(α max ) = max{H(A) ∪... ∪ H(N)} (5)

[0108] When α = α max , according to the above formula (4), we have:

[0109]

[0110] Therefore, according to the above formula (5) and formula (6), we have:

[0111]

[0112] The second theorem states that the probability that n blocks have the same maximum hash value is related to how many elements they have in common (i.e., their degree of similarity). Therefore, n similar plaintext blocks may have the same maximum hash value. Similarly, n similar plaintext blocks may also have the same minimum hash value. Therefore, the embodiments of the present application can use their same target hash value (e.g., the maximum hash value or the minimum hash value) as their same key.

[0113] Based on the above first theorem and second theorem, the embodiments of the present application can solve the problem that similarity is destroyed after encryption. That is to say, after using the target hash value as the key to retain similarity and performing bitwise XOR encryption on the plaintext blocks, the similarity existing in the plaintext blocks can be retained. Among them, the specific implementation manner of removing redundant data of similar ciphertext blocks by the embodiments of the present application through a new encryption difference compression technology can be seen in the following Figures 3 - 10 corresponding embodiment method.

[0114] Furthermore, please refer to Figure 3 , Figure 3 which is an interaction diagram of a method for data processing provided by the embodiments of the present application. As Figure 3 shown, this method can be jointly executed by the first device and the second device. Among them, in the outsourcing storage scenario, the first device here can be any client in the client cluster shown above Figure 1 , and the second device here can be the above server 100. This method can at least include S101 - S106:

[0115] S101, the first device determines the first key corresponding to the first plaintext block.

[0116] Among them, the first plaintext block belongs to any one of the plaintext blocks in the original data. Here, the first key is obtained by performing key generation calculations on N sub-block data in the first plaintext block respectively. N is a positive integer greater than 1. Among them, the sub-block data here can be obtained by sliding a sliding window on the first plaintext block, or can be obtained by dividing the first plaintext block. Here, it will not be limited.

[0117] In an optional manner, the first device can divide the original data into blocks to obtain the first plaintext block in the original data. Among them, the first plaintext block here can include N sub-block data; further, the first device can determine the hash values corresponding to the N sub-block data respectively, and determine the target hash value among the N hash values as the first key corresponding to the first plaintext block. Among them, the target hash value can be the maximum hash value among the N sub-block data, or can be the minimum hash value among the N sub-block data. Here, it will not be limited.

[0118] For example, after obtaining the first plaintext block, the first device can use a sliding window to traverse the first plaintext block to obtain N sub-block data, and then can use a key generation algorithm (for example, a hash algorithm) to perform hash calculations on the N sub-block data respectively to obtain the hash values corresponding to the N sub-block data respectively. Further, the first device can determine the maximum hash value (or minimum hash value) among the N hash values as the key of the first plaintext block.

[0119] It can be understood that since the data block division process identifies the block boundaries by performing hash calculations on the data where the sliding window slides on the original data. For example, the first device can use the Rabin sliding window algorithm to calculate the hash fingerprint (Fingerprint, fp) of each block to find the boundary point. If the hash fingerprint meets the block division condition, it is considered that the sliding window has reached the boundary point of the current block, and then the original data can be cut into plaintext blocks at the boundary point. The block division condition can refer to the following formula (7):

[0120] fp mod D = r (7)

[0121] Among them, D is used for the expected average block size (for example, 8KB), and r is used to represent a predefined value.

[0122] The keys for each plaintext block in the embodiments of the present application also need to perform a hashing calculation on each sub-block data of the plaintext block. Calculating the hash values of each sub-block data is a rather time-consuming process, which easily leads to additional computational overhead in the process of generating keys. For example, when using a sliding window of 48 bytes to calculate a plaintext block of 8KB, the number of calculations will exceed 100 times, and the calculation of hash values itself also has a high time complexity. Therefore, in order to reduce the amount of calculation and improve the key generation speed, during the encryption process, the first device can directly determine the key by using the same sliding window as the chunking process and the hash values calculated by the chunking process. For the sake of distinction, in the embodiments of the present application, the sliding window used in the key generation process can be called the first sliding window, and the sliding window used in the data chunking process can be called the second sliding window.

[0123] In other words, each of the above N sub-block data is obtained by the first sliding window sliding on the first plaintext block. The first sliding window and the second sliding window are the same sliding window, and the second sliding window is used to chunk the original data.

[0124] For ease of understanding, further, please refer to Figure 4 , Figure 4 is a schematic diagram of a scenario for encrypting a plaintext block based on a sliding window provided by the embodiments of the present application. As Figure 4 shown, the encryption process and the chunking process in the embodiments of the present application are independent of each other, that is, the key generation process needs to wait until the chunking process is completed before starting to run.

[0125] As Figure 4 shown, during the chunking process, the first device can use the sliding window W1 (i.e., the second sliding window) to split the original data D. For example, the first device slides the sliding window W1 on the original data D to determine the sub-block data j corresponding to different positions, and determines whether the hash fingerprint of the current block satisfies the above formula (7) by calculating the hash values of each sub-block data j. Where j is used to indicate the distance between the current window position of the sliding window W1 and the position of the first boundary point, and j is a positive integer.

[0126] If the position of the first boundary point is the position B1 shown in Figure 4 , this means that the first device has split out the plaintext block P1 from the original data D. If the sliding window W1 slides from the position B1 to the position B2, and the first device determines that the modulus of the current hash fingerprint is equal to r, at this time, the first device can determine that the position B2 belongs to the new boundary point (i.e., the second boundary point), and then can split the original data D again at the new boundary point to obtain the next plaintext block (for example, the plaintext block P2).

[0127] After the chunking process is completed, all plaintext chunks are stored in the system waiting for the encryption process to start. After the encryption process starts, it is necessary to use the sliding window W2 to generate corresponding keys for all plaintext chunks. It can be understood that when the sliding window W1 slides from position B1 to position B2, the first device has already calculated the hash values corresponding to the N sub-chunk data in the plaintext chunk P2. Therefore, in order to reduce the computational overhead, the first device does not need to use a new sliding window to calculate the plaintext chunk P2 again, but can directly determine the hash values corresponding to the N sub-chunk data in the chunking process as the hash value set associated with the plaintext chunk P2. This hash value set can be {H 21 , H 22 , …, H 2n}. In other words, the sliding window W1 here and the sliding window W2 are the same sliding window.

[0128] Among them, H 21 can be used to represent the hash value obtained after performing a hash calculation on the data covered by the sliding window W1 at the first position (i.e., the head) of the plaintext chunk P2; H 2n can be used to represent the hash value obtained after performing a hash calculation on the data covered by the sliding window W1 at the nth position (i.e., the tail) of the plaintext chunk P2. Then, the first device can use the target hash value (i.e., the maximum hash value or the minimum hash value) in the above hash value set as the key corresponding to the plaintext chunk P2. Among them, the key corresponding to the plaintext chunk P2 is used to encrypt the plaintext chunk P2 to obtain the ciphertext chunk corresponding to the plaintext chunk P2 (for example, Figure 4 the ciphertext chunk C2 shown).

[0129] It can be seen that in the encryption process, in the embodiments of the present application, there is no need to use a new sliding window to calculate a certain plaintext chunk again, but directly select from the N hash values calculated in the chunking process, which can greatly reduce the computational overhead, improve the key generation speed, and thus improve the encryption speed.

[0130] In another alternative, since the data process needs to calculate the hash fingerprint based on the sliding window to determine the boundary points of the original data to split the original data into plaintext chunks, and the key generation process also needs to calculate the hash value through a sliding window to generate the key, therefore, in order to effectively improve the client throughput and system performance, the embodiments of the present application can propose a key generation scheme that couples the key generation process with the data chunking process, thereby eliminating the additional computational overhead generated by the key generation process and improving the encryption speed. That is, the embodiments of the present application can set a buffer in the first device, and this buffer is used to store the target hash value (the maximum hash value or the minimum hash value) during the chunking process.

[0131] Exemplarily, the first device may use a first sliding window to traverse the original data, and then determine the data covered by the first sliding window at position j as sub-block data j, and determine the hash value H of sub-block data j. j Where j is used to indicate the distance between the current window position and the first boundary point position (i.e., the starting position of the current block), and j is a positive integer less than or equal to N. Further, the first device may update the target hash value based on the hash value H j . Until it is determined that the position j reaches the second boundary point (i.e., the termination position of the current block), the data block between the second boundary point and the first boundary point is determined as the first plaintext block, and then the target hash value in the updated buffer can be determined as the first key corresponding to the first plaintext block.

[0132] For ease of understanding, further, please refer to Figure 5 , Figure 5 which is a schematic diagram of a scenario that couples the data block process and the encryption process provided by an embodiment of the present application. As Figure 5 shown, the original data D is the data stream to be encrypted by the first device. In the embodiment of the present application, the key of the current block can be generated while the block process of each plaintext block is completed, thereby eliminating the computational overhead caused by key generation. For ease of description, the first plaintext block in the embodiment of the present application takes the plaintext block P2 as an example, and the target hash value (for example, the maximum hash value or the minimum hash value) here can be stored in the buffer of the first device.

[0133] It should be understood that the first device may use Figure 5 the first sliding window shown to traverse the original data D. If the current window position of the first sliding window is Figure 5 the position B1 shown, it can be considered that the first device has sliced out the first plaintext block (for example, the plaintext block P1) from the original data D. At this time, the boundary point corresponding to the position B1 can be called the first boundary point.

[0134] Further, the first device can continue to slide on the original data D to slice the second plaintext block. It can be understood that the first device can determine the data covered by the first sliding window at position 1 as sub-block data 1, and perform a hash calculation on the sub-block data 1 to obtain the hash value H corresponding to the sub-block data 1 21 . At this time, the first device can directly store the hash value H 21 into the buffer. Then, the first device needs to determine the hash fingerprint of the current block based on the hash value H 21 to determine whether the position 1 has reached the second boundary point (i.e., the next boundary point of the first boundary point).

[0135] If position 1 reaches the second boundary point, the first device can determine the data block between the second boundary point and the first boundary point as the plaintext block P2, and then can use the hash value H 21 as the key corresponding to the plaintext block P2.

[0136] If position 1 does not reach the second boundary point, it is necessary to continue sliding the first sliding window on the original data D. Among them, the first device can determine the data covered by the first sliding window at position 2 as the sub-block data 2, and perform a hash calculation on the sub-block data 2 to obtain the hash value H 22 corresponding to the sub-block data 2. At this time, the first device can update the hash value in the buffer based on the hash value H 21 . For example, the first device needs to update the hash value H 22 by comparing it with the hash value H 21 stored in the buffer, and store the maximum hash value of these two hash values (for example, the hash value H 22 ) directly into the buffer. Then, the first device needs to determine the hash fingerprint of the current block based on the hash value H 22 to determine whether position 2 has reached the second boundary point (i.e., the next boundary point of the first boundary point), and so on.

[0137] Until the first device determines that Figure 5 the position n (i.e., position B2) shown reaches the second boundary point, the first device can determine the data block between the second boundary point and the first boundary point as the plaintext block P2. Here, the plaintext block P2 refers to the data block between position B2 and position B1.

[0138] At this time, the first device can determine the target hash value in the updated buffer as the key corresponding to the plaintext block P2, so as to encrypt the plaintext block P2 based on the key corresponding to the plaintext block P2 to obtain the ciphertext block C2.

[0139] In other words, the embodiments of the present application can obtain the SP-Key of the plaintext block in three steps: (1) At the beginning of the block division of the current block, record the target hash value in the current block; (2) During the sliding process from the head to the tail of the block, continuously update the target hash value of the current block; (3) When the block division is completed (i.e., the first sliding window reaches the boundary point), use the target hash value of the current block stored in the buffer as its SP-Key.

[0140] It can be seen that by coupling the two stages of block - based fingerprint calculation and encryption key generation in the embodiments of the present application, encryption can be directly performed when the block division of a plaintext block is completed, accelerating the process of generating the encryption key and reducing the additional computational overhead. In addition, by designing an encryption pipeline in the embodiments of the present application to eliminate the additional overhead, in - place calculation is realized, and the calculation rate is increased. This acceleration strategy can effectively improve the throughput of the client and enhance the system performance.

[0141] S102. The first device encrypts the first plaintext block based on the first key to obtain the first ciphertext block.

[0142] It should be understood that after determining the first key, the first device can directly encrypt the first plaintext block itself. For example, the first device can refer to the encryption method shown above Figure 2 (for example, using the bit - by - bit exclusive - or encryption method) to encrypt the first plaintext block itself.

[0143] Optionally, in order to reduce the data transmission volume between the first device and the second device and improve the transmission performance, the first device can detect whether there is a plaintext block similar to the first plaintext block before encryption. If not, it encrypts the first plaintext block itself based on the first key. If so, instead of encrypting the first plaintext block itself, it encrypts the differential data between the first plaintext block and its similar block.

[0144] Among them, in the embodiments of the present application, the plaintext block detected in the first device and similar to the first plaintext block can be called the third plaintext block. The third key corresponding to the third plaintext block is obtained by performing key generation calculations on P sub - block data in the third plaintext block respectively, where P is a positive integer greater than 1. According to the above - mentioned second theorem, for the set of hash values generated by multiple highly similar blocks, the maximum or minimum hash value in it is very likely to be the same. Therefore, the third key is the same as the first key.

[0145] It can be understood that the first device can determine the differential data (i.e., the first differential data) between the first plaintext block and the third plaintext block, and then encrypt the first differential data based on the first key, and determine the encrypted first differential data as the first ciphertext block.

[0146] S103. The first device sends the first ciphertext block to the second device.

[0147] S104. The second device obtains M ciphertext blocks including the first ciphertext block.

[0148] Among them, M is a positive integer greater than 1. These M ciphertext blocks are the ciphertext blocks counted by the second device within a compression period (for example, one day). The M ciphertext blocks are sent by X first devices, and X is a positive integer. For example, in a data storage scenario, the second device can be a cloud server deployed in the cloud, and each of the X first devices can be a client having a network connection relationship with the cloud server; in a network deduplication scenario, the second device here can be a network edge device (for example, a switch), and the X first devices are all located in the area corresponding to the network edge device.

[0149] S105, the second device performs similarity detection on the M ciphertext blocks, and determines the ciphertext block having similarity with the first ciphertext block as the second ciphertext block.

[0150] Among them, the second key corresponding to the second ciphertext block here is obtained by performing key generation calculations on the W sub-block data of the second plaintext block respectively. W is a positive integer greater than 1. According to the above second theorem, for a set of hash values generated by multiple highly similar blocks, the maximum hash value or the minimum hash value among them has a very high probability of being the same. Therefore, the second key here is the same as the first key, which means that the embodiments of the present application can generate the same key for similar plaintext blocks to facilitate encryption while preserving similarity.

[0151] S106, the second device performs differential encoding on the first ciphertext block and the second ciphertext block.

[0152] Specifically, when regarding the first ciphertext block as the basic block, the second device can first determine the differential data (i.e., the second differential data) between the second ciphertext block and the first ciphertext block, and then obtain the index data (i.e., the first index data) associated with the second differential data. Among them, the first index data here is used to point to the first ciphertext block. Further, the first device can store the first ciphertext block, the second differential data, and the first index data.

[0153] Taking the number of ciphertext blocks as 5 as an example, it can specifically include ciphertext block C1, ciphertext block C2, ciphertext block C3, ciphertext block C4, and ciphertext block C5. Then, the second device can first call the data deduplication module to delete the duplicate ciphertext blocks of the first ciphertext block, and then call the differential compression module to perform differential compression on the similar ciphertext blocks of the first ciphertext block.

[0154] For example, the first ciphertext block here can be taken as ciphertext block C1. When the second device invokes the data deduplication module and determines that the fingerprint indexes of ciphertext block C1 and ciphertext block C2 are the same, then ciphertext block C2 can be determined as the same ciphertext block as ciphertext block C1 (i.e., duplicate ciphertext block). At this time, the second device can delete ciphertext block C2. When the second device invokes the differential compression module and determines that both ciphertext block C3 and ciphertext block C5 belong to ciphertext blocks similar to ciphertext block C1 (i.e., the second ciphertext block), for ciphertext block C3, the second device does not need to store ciphertext block C3 itself, but through differential encoding, stores differential data 1 (the differential data between ciphertext block C3 and ciphertext block C1) and an index data 1 for pointing to ciphertext block C1. Similarly, for ciphertext block C5, the second device can store differential data 2 (the differential data between ciphertext block C5 and ciphertext block C1) and an index data 2 for pointing to ciphertext block C1 through differential encoding.

[0155] Based on this, the data finally stored by the second device are ciphertext block C1, ciphertext block C4, differential data 1, index data 1, differential data 2, and index data 2 respectively. Compared with the above 5 ciphertext blocks, the space occupied during data storage is greatly reduced, and the storage cost is lowered.

[0156] It can be seen that based on the support of two theorems, the embodiment of the present application encrypts by using the target hash value (calculating the maximum hash value or minimum hash value of similar blocks through a sliding window) that can retain block similarity as the key of the plaintext block. Thus, on the basis of encrypted deduplication, encrypted differential compression is achieved, so as to effectively eliminate redundant data between similar ciphertext blocks subsequently, reduce the storage space, and lower the storage cost.

[0157] Next, in conjunction with the accompanying drawings, from the perspective of a single network device, the method provided by the embodiment of the present application will be introduced.

[0158] Please refer to Figure 6 , Figure 6 which is a schematic diagram of a method for data processing provided by the embodiment of the present application. Figure 1 . As Figure 6 shown, this method can be executed by a first device, and the first device can be any client in the client cluster shown above Figure 1 , and no limitation will be imposed here. This method can at least include S201 - S202:

[0159] S201, determine the first key corresponding to the first plaintext block.

[0160] Among them, the first key here is obtained by performing key generation calculations on N sub-block data of the first plaintext block respectively. The first key is the same as the second key. The second key is the key corresponding to the second plaintext block that is similar to the first plaintext block. The second key is obtained by performing key generation calculations on W sub-block data of the second plaintext block respectively. Both N and W are positive integers greater than 1.

[0161] Among them, the key in the embodiment of the present application can be the target hash value (i.e., the hash value used to preserve similarity) selected from multiple hash values associated with the plaintext block itself. The target hash value can be the maximum hash value or the minimum hash value. For example, when the target hash value is the maximum hash value, the first key here can be the maximum hash value selected from the hash values corresponding to N sub-block data respectively, and the second key can be the maximum hash value selected from the hash values corresponding to M sub-block data respectively; when the target hash value is the minimum hash value, the first key here can be the minimum hash value selected from the hash values corresponding to N sub-block data respectively, and the second key can be the minimum hash value selected from the hash values corresponding to M sub-block data respectively.

[0162] It can be understood that if the first device is client A and the first plaintext block is any plaintext block in a certain original data (for example, original data D1) in client A, then the second plaintext block may be another plaintext block in the original data D1, or may be any plaintext block in another original data (for example, original data D2) in client A, or may also be any plaintext block in the original data (for example, original data D3) in another first device (for example, client B different from client A). In other words, the first plaintext block and the second plaintext block here can belong to the same original data or different original data, and no limitation will be imposed here.

[0163] S202. Encrypt the first plaintext block based on the first key to obtain the first ciphertext block.

[0164] Among them, the encryption method here can be Figure 2 the bitwise exclusive OR shown.

[0165] Among them, the specific implementation manners of S201 - S202 can refer to the descriptions of S101 - S102 in the corresponding embodiments above, and will not be elaborated here. Figure 3 For ease of understanding, further, please refer to

[0166] For ease of understanding, further, please refer to Figure 7 , Figure 7 which is a schematic diagram of a scenario for encrypting a plaintext block proposed in the embodiment of the present application. As Figure 7As shown, the plaintext blocks included in the original data here can be taken as 3 examples, specifically including plaintext block P1, plaintext block P2, and plaintext block P3.

[0167] As Figure 7 shown, the encryption process of the first device can be divided into the following 3 steps. Step 1: Calculate a set of hash values for each plaintext block; Step 2: Use the set of hash values to generate a key for each plaintext block; Step 3: Encrypt each plaintext block with its own key respectively.

[0168] For plaintext block P1, the first device can use Figure 7 the sliding window shown (for example, 48Bytes) to slide on plaintext block P1 to calculate the hash values corresponding to each sub-block data in plaintext block P1. When the sliding window slides from the head of plaintext block P1 to the tail of plaintext block P1, the first device can obtain a set of hash values associated with plaintext block P1 (for example, set 1), and this set 1 can be {H 11, H 12 , …, H 1n}. Then, the first device can select a target hash value (for example, Figure 7 the hash value H 1i ) shown as the key of plaintext block P1. Here, the hash value H 1i is the maximum or minimum hash value in set 1. Further, the first device can use the key of plaintext block P1 to encrypt plaintext block P1 through the Figure 2 encryption method shown to obtain Figure 7 the ciphertext block C1 shown.

[0169] For plaintext block P2, the first device can use Figure 7 the sliding window shown (for example, 48Bytes) to slide on plaintext block P2 to calculate the hash values corresponding to each sub-block data in plaintext block P2. When the sliding window slides from the head of plaintext block P2 to the tail of plaintext block P2, the first device can obtain a set of hash values associated with plaintext block P2 (for example, set 2), and this set 2 can be {H 21, H 22 , …, H 2n}. Then, the first device can select a target hash value (for example, Figure 7 the hash value H 2j ) shown as the key of plaintext block P2. Here, the hash value H 2j is the maximum or minimum hash value in set 2. Further, the first device can use the key of plaintext block P2 to encrypt plaintext block P2 through the Figure 2The encryption method shown uses the key of the plaintext block P2 to encrypt the plaintext block P2 to obtain Figure 7 the ciphertext block C2 shown.

[0170] For the plaintext block P3, the first device may use Figure 7 the sliding window shown (e.g., 48 Bytes) to slide on the plaintext block P3 to calculate the hash values respectively corresponding to each sub-block data in the plaintext block P3. When the sliding window slides from the head of the plaintext block P3 to the tail of the plaintext block P3, the first device may obtain a hash value set associated with the plaintext block P3 (e.g., set 3), and this set 3 may be {H 31, H 32 , …, H 3n}. Then, the first device may select a target hash value (e.g., Figure 7 the hash value H shown 3m ) from this set 3 as the key of the plaintext block P3, where the hash value H 3m is the maximum hash value or the minimum hash value in the set 3. Further, the first device may use the encryption method shown above Figure 2 to encrypt the plaintext block P3 using the key of the plaintext block P3 to obtain Figure 7 the ciphertext block C3 shown.

[0171] It can be understood that if both the plaintext block P2 and the plaintext block P3 are plaintext blocks similar to the plaintext block P1, then according to the above second theorem, the keys determined by these three plaintext blocks will probably be the same, that is, the key generation method proposed in the embodiments of the present application makes the keys obtained from similar plaintext blocks the same, thus effectively retaining the similarity of the blocks.

[0172] However, the existing encryption and deduplication scheme needs to first perform block processing on the original data. After obtaining all the plaintext blocks, the plaintext blocks are stored from memory to disk. After waiting for the encryption process to start, the system then retrieves the plaintext blocks from the disk to memory and encrypts them with their respective keys, so it will cause a large amount of I / O overhead. In the embodiments of the present application, after each plaintext block is itself block-divided, the corresponding key is also generated at the same time. To solve the above problem, the embodiments of the present application may encrypt each plaintext block immediately after it is block-divided, thereby reducing the I / O overhead and further accelerating the encryption process.

[0173] Among them, after the first device divides the original data into chunks to obtain the first plaintext chunk, the plaintext chunks that have been sliced off can be referred to as the first data, and the data that has not been sliced can be referred to as the second data (i.e., the remaining data). That is, the original data can include the first data and the second data, and here the first plaintext chunk is the last sliced data in the first data. Based on this, the first device can continue to divide the second data into chunks until the fourth plaintext chunk is obtained. Among them, the fourth plaintext chunk here is the next plaintext chunk of the first plaintext chunk in the original data. At this time, the first device can determine the fourth key corresponding to the fourth plaintext chunk, and then, based on the fourth key, encrypt the fourth plaintext chunk. Among them, the fourth key here is obtained by performing key generation calculations on Q sub-block data in the fourth plaintext chunk respectively, and Q is a positive integer greater than 1. The specific implementation manner for the first device to encrypt the fourth plaintext chunk can refer to the specific implementation manner for encrypting the first plaintext chunk above, and will not be elaborated here.

[0174] For ease of understanding, further, please refer to Figure 8 , Figure 8 which is a schematic diagram of a scenario provided by an embodiment of the present application for pipelining the chunking process and the encryption process. As Figure 8 shown, the first device in the embodiment of the present application can be a client. Among them, the data processing workflow of the client is divided into two stages: the chunking stage and the encryption stage.

[0175] In the chunking stage, the client can continuously divide the original data into a series of plaintext chunks, and at the same time refer to the Figure 5 key generation method shown above to generate the key (SP-Key) corresponding to each plaintext chunk one by one. For example, Figure 8 the key K1 shown is the key corresponding to the plaintext chunk P1 determined by the client when dividing to obtain the plaintext chunk P1; the key K2 is the key corresponding to the plaintext chunk P2 determined by the client when dividing to obtain the plaintext chunk P2; the key K3 is the key corresponding to the plaintext chunk P3 determined by the client when dividing to obtain the plaintext chunk P3; and so on.

[0176] In the encryption stage, the client does not need to wait for the original data division to be completed, but immediately encrypts the currently obtained plaintext chunk using the key of the plaintext chunk. For example, when dividing the original data to obtain the plaintext chunk P1, while continuing to divide the original data, encrypt the plaintext chunk P1 based on the key K1 to obtain the ciphertext chunk corresponding to the plaintext chunk P1 (for example, the ciphertext chunk C1).

[0177] Since the above two stages are executed block by block, they can also be encrypted in a pipelined manner block by block. As Figure 8As shown, the chunking stage of the plaintext block P2 (i.e., the fourth plaintext block) can run concurrently with the encryption stage of its previous plaintext block (i.e., the first plaintext block, e.g., plaintext block P1). Similarly, the chunking stage of the plaintext block P3 can run concurrently with the encryption stage of its previous plaintext block (plaintext block P2).

[0178] This means that for any plaintext block, when its chunking process is completed, its encryption process can immediately run in memory, as Figure 8 shown. In the embodiments of the present application, not only can the chunking process be coupled with the key generation process, but also the chunking process can be pipelined with the encryption process. This encryption method can eliminate the additional I / O overhead caused by extracting from disk to memory, thereby accelerating the encryption process of EDC.

[0179] Please refer to Figure 9 , Figure 9 which is a schematic diagram of a method for data processing provided by the embodiments of the present application. Figure 2 . As Figure 9 shown, this method can be executed by a first device, which can be a computer device with strong resource computing power. The first device is deployed with a chunking module, an encryption module, a data deduplication module, and a differential compression module. Here, the first device can be a client or a server, and will not be limited herein. This method can at least include S301 - S305:

[0180] S301, determine the first key corresponding to the first plaintext block.

[0181] Among them, the first plaintext block is any plaintext block obtained by the first device calling the chunking module to chunk the original data. Here, the first key belongs to the first hash value set, and the first hash value set refers to the hash value set associated with the first plaintext block. The first hash value set includes the hash values obtained by respectively performing hash calculations on N sub-block data of the first plaintext block, and N is a positive integer greater than 1. Among them, the first key can be the target hash value selected from the first hash value set, and here the target hash value can be the maximum hash value or the minimum hash value.

[0182] S302, encrypt the first plaintext block based on the first key to obtain the first ciphertext block.

[0183] Among them, the first ciphertext block can be the ciphertext block obtained by the first device calling the encryption module to encrypt the first plaintext block. The encryption method adopted by the encryption module can be the bitwise XOR encryption method shown above Figure 2 .

[0184] It can be understood that if the number of bits of the first key is the same as that of the first plaintext block, the first device can directly use the bitwise XOR encryption method to encrypt the first plaintext block to obtain the first ciphertext block. If the number of bits of the first key is less than that of the first plaintext block, the first device needs to pad the first key and then encrypt the first plaintext block based on the padded first key to obtain the first ciphertext block. Among them, the number of bits of the padded first key is the same as that of the first plaintext block.

[0185] For example, if the first key is "1001" and the first plaintext block is "110110011100", then the first device can pad the first key to "100110011001" and then encrypt the first plaintext block based on the padded first key to obtain the first ciphertext block (for example, "010000000101").

[0186] Among them, the specific implementation manners of S301 - S302 can refer to the descriptions of S101 - S102 in the corresponding embodiments above, which will not be elaborated here. Figure 3 The descriptions of S101 - S102 in the corresponding embodiments above, which will not be elaborated here.

[0187] S303, obtain R local ciphertext blocks.

[0188] Among them, these R local ciphertext blocks include the first ciphertext block, and R is a positive integer greater than 1.

[0189] It can be understood that if there is a ciphertext block identical to the first ciphertext block among these R local ciphertext blocks, the first device can call the data deduplication module to retain the first ciphertext block and delete the ciphertext blocks identical to the first ciphertext block.

[0190] S304, perform similarity detection on the R local ciphertext blocks to obtain a similarity result.

[0191] Among them, this similarity result is the result obtained by the first device calling the differential compression module to perform similarity detection on the R local ciphertext blocks.

[0192] S305, if the similarity result is that there is a second ciphertext block similar to the first ciphertext block among the R local ciphertext blocks, perform differential encoding on the first ciphertext block and the second ciphertext block.

[0193] Among them, the first key here is the same as the second key corresponding to the second ciphertext block. The second key belongs to a second hash value set associated with the second plaintext block. This second hash value set includes hash values obtained by respectively performing hash calculations on W sub-block data of the second plaintext block, where W is a positive integer greater than 1. Specifically, the first device can use the first ciphertext block as a basic block, and then determine the difference data (i.e., the third difference data) between the second ciphertext block and the first ciphertext block, and obtain index data (i.e., the second index data) associated with the third difference data. Here, the second index data is used to point to the first ciphertext block. Then, the first device can store the first ciphertext block, the third difference data, and the second index data.

[0194] It can be seen that when the resource computing power of the first device is strong enough, the first device can not only couple the chunking process and the encryption process, but also perform a pipelined design on the chunking process and the encryption process, thereby eliminating the additional computational overhead of key generation and the I / O overhead of traditional encryption schemes. Therefore, the encryption process of EDC can be accelerated. In addition, in order to reduce the storage cost of the first device, the first device can combine the above first theorem and second theorem to further reduce the redundancy of encrypted data deduplication, that is, locally implement encrypted difference compression to effectively eliminate redundant data in similar ciphertext blocks locally on the first device.

[0195] Further, please refer to Figure 10 , Figure 10 which is a schematic diagram of a method for data processing provided by an embodiment of this application. Figure 3 As Figure 10 shown, this method can be executed by a second device, and the second device can be the server 100 shown above. Figure 1 This method can at least include S401 - S403:

[0196] S401, obtain M ciphertext blocks.

[0197] Among them, the M ciphertext blocks include the first ciphertext block, where M is a positive integer greater than 1. These M ciphertext blocks are ciphertext blocks counted by the second device within a compression period (for example, one day), and the M ciphertext blocks are sent by X first devices, where X is a positive integer. For example, in a data storage scenario, the second device can be a cloud server deployed in the cloud, and each of the X first devices can be a client having a network connection relationship with the cloud server; in a network deduplication scenario, the second device here can be a network edge device (for example, a switch), and all of the X first devices are located in the area corresponding to the network edge device.

[0198] S402, perform similarity detection on the M ciphertext blocks, and determine the ciphertext blocks having similarity with the first ciphertext block as the second ciphertext blocks.

[0199] Among them, the second key corresponding to the second ciphertext block here is obtained by performing key generation calculations on the W sub-block data of the second plaintext block respectively. W is a positive integer greater than 1. According to the above second theorem, the second key is the same as the first key.

[0200] S403. Perform differential encoding on the first ciphertext block and the second ciphertext block.

[0201] Specifically, when regarding the first ciphertext block as the basic block, the second device can first determine the second differential data between the second ciphertext block and the first ciphertext block, and then obtain the index data associated with the second differential data (i.e., the first index data). Among them, the first index data here is used to point to the first ciphertext block. Further, the first device can store the first ciphertext block, the second differential data, and the first index data.

[0202] Among them, the specific implementation manners of S401 - S403 can refer to the descriptions of S104 - S106 in the corresponding embodiments above. Figure 3 Details will not be described here.

[0203] It can be seen that compared with storing M ciphertext blocks, the data processing method provided by the embodiments of the present application can remove redundant data of similar ciphertext blocks by combining the first theorem and the second theorem, which can greatly reduce the space occupied during data storage and reduce the storage cost.

[0204] Further, please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a data processing device provided by the embodiments of the present application. Figure 1 . As Figure 11 shown, the data processing device 1 may include at least one of a determination unit 1101 and an encryption unit 1102. These units can perform the response functions of each device in the above method embodiments.

[0205] In a possible implementation manner, the data processing device 1 may be used to implement Figure 6 the function of the first device in

[0206] Specifically, the determination unit 1101 is used to determine the first key corresponding to the first plaintext block. The first key is obtained by performing key generation calculations on the N sub-block data of the first plaintext block respectively. The first key is the same as the second key. The second key is the key corresponding to the second plaintext block that is similar to the first plaintext block. The second key is obtained by performing key generation calculations on the W sub-block data of the second plaintext block respectively. Both N and W are positive integers greater than 1; the encryption unit 1102 is used to encrypt the first plaintext block based on the first key to obtain the first ciphertext block.

[0207] Among them, for the specific implementation manners of the determining unit 1101 and the encrypting unit 1102, reference can be made to the descriptions of steps S201 - S202 in the corresponding embodiments above, and details will not be elaborated here. In addition, the beneficial effects of adopting the same method will not be elaborated either. Figure 6

[0208] Further, please refer to Figure 12 , Figure 12 which is a schematic structural diagram of a data processing device provided by an embodiment of the present application. Figure 2 As Figure 12 shown, the data processing device 2 may include at least one of a determining unit 1201, an encrypting unit 1202, a blocking unit 1203, an obtaining unit 1204, a detecting unit 1205, and a coding unit 1206. These units may perform the response functions of the respective devices in the above method embodiments.

[0209] In a possible implementation manner, the data processing device 2 may be used to implement Figure 6 or Figure 9 the functions of the first device in

[0210] Specifically, the determining unit 1201 is configured to determine a first key corresponding to a first plaintext block. The first key is obtained by respectively performing key generation calculations on N sub-block data of the first plaintext block. The first key is the same as a second key, and the second key is a key corresponding to a second plaintext block having similarity to the first plaintext block. The second key is obtained by respectively performing key generation calculations on W sub-block data of the second plaintext block. Both N and W are positive integers greater than 1. The encrypting unit 1202 is configured to encrypt the first plaintext block based on the first key to obtain a first ciphertext block.

[0211] In one implementation manner, the first plaintext block is any plaintext block obtained by blocking the original data. The first plaintext block includes N sub-block data. The determining unit 1201 is configured to determine the first key corresponding to the first plaintext block, including: a first determining subunit 12011 configured to determine the hash values respectively corresponding to the N sub-block data; the first determining subunit 12011 is further configured to determine the target hash value among the N hash values as the first key corresponding to the first plaintext block.

[0212] In one implementation manner, each of the N sub-block data is obtained by sliding a first sliding window on the first plaintext block. The first sliding window and a second sliding window are the same sliding window, and the second sliding window is used to block the original data.

[0213] ​In one implementation, the first device is provided with a buffer, which is used to store the target hash value during the chunking process; a determination unit 1201, which is used to determine the first key corresponding to the first plaintext block, including: a traversal subunit 12012, a second determination subunit 12013, and an update subunit 12014.

[0214] The traversal subunit 12012 is used to traverse the original data by using a first sliding window; the second determination subunit 12013 is used to determine the data covered by the first sliding window at position j as sub-block data j, and determine the hash value H of the sub-block data j j , where j is used to indicate the distance between the current window position and the position of the first boundary point, and j is a positive integer less than or equal to N; the update subunit 12014 is used to update the target hash value based on the hash value H j . The second determination subunit 12013 is further used to, until it is determined that the position j reaches the second boundary point, determine the data block between the second boundary point and the first boundary point as the first plaintext block, and determine the target hash value in the updated buffer as the first key corresponding to the first plaintext block.

[0215] In one implementation, the target hash value is the maximum hash value or the minimum hash value.

[0216] In one implementation, the first key is the same as the third key corresponding to the third plaintext block. The third plaintext block is the plaintext block detected in the first device and having similarity with the first plaintext block. The third key is obtained by respectively performing key generation calculations on P sub-block data in the third plaintext block, where P is a positive integer greater than 1; an encryption unit 1202, which is used to encrypt the first plaintext block based on the first key to obtain a first ciphertext block, including: a third determination subunit 12021 and an encryption subunit 12022.

[0217] The third determination subunit 12021 is used to determine the first difference data between the first plaintext block and the third plaintext block; the encryption subunit 12022 is used to encrypt the first difference data based on the first key, and determine the encrypted first difference data as the first ciphertext block.

[0218] In one implementation, the first plaintext block is any plaintext block obtained after dividing the original data into blocks. The original data includes first data and second data. The first plaintext block is the last segmented data in the first data. The apparatus further includes: a block division unit 1203 configured to divide the second data into blocks until a fourth plaintext block is obtained. The fourth plaintext block is the next plaintext block of the first plaintext block in the original data; a determination unit 1201 further configured to determine a fourth key corresponding to the fourth plaintext block. The fourth key is obtained by respectively performing key generation calculations on Q sub-block data in the fourth plaintext block, where Q is a positive integer greater than 1; an encryption unit 1202 further configured to encrypt the fourth plaintext block based on the fourth key.

[0219] In one implementation, the apparatus further includes: an acquisition unit 1204 configured to acquire R local ciphertext blocks. The R local ciphertext blocks include a first ciphertext block, where R is a positive integer greater than 1; a detection unit 1205 configured to perform similarity detection on the R local ciphertext blocks to obtain a similarity result; an encoding unit 1206 configured to, if the similarity result indicates that there is a second ciphertext block similar to the first ciphertext block among the R local ciphertext blocks, perform differential encoding on the first ciphertext block and the second ciphertext block.

[0220] Wherein, the specific implementation manners of the determination unit 1201, the encryption unit 1202, the block division unit 1203, the acquisition unit 1204, the detection unit 1205, and the encoding unit 1206 may refer to the descriptions of steps S201 - S202 in the corresponding embodiments above or the descriptions of steps S301 - S305 in the corresponding embodiments above Figure 6 and will not be elaborated herein. In addition, the description of the beneficial effects of adopting the same method will not be elaborated either. Figure 9

[0221] Further, please refer to Figure 13 , Figure 13 which is a schematic structural diagram of a data processing apparatus provided by an embodiment of the present application. Figure 3 As Figure 13 shown, the data processing apparatus 3 may include at least one of an acquisition unit 1301, a detection unit 1302, and an encoding unit 1303. These units may perform the response functions of each device in the above method embodiments.

[0222] In a possible implementation, the data processing apparatus 3 may be used to implement the function of the second device in Figure 10 .

[0223] ​Specifically, an obtaining unit 1301 is configured to obtain M ciphertext blocks, where the M ciphertext blocks include a first ciphertext block, and M is a positive integer greater than 1; a detecting unit 1302 is configured to perform similarity detection on the M ciphertext blocks, and determine a ciphertext block having similarity with the first ciphertext block as a second ciphertext block, where a first key corresponding to the first ciphertext block is the same as a second key corresponding to the second ciphertext block; an encoding unit 1303 is configured to perform differential encoding on the first ciphertext block and the second ciphertext block.

[0224] In one implementation, the M ciphertext blocks are ciphertext blocks statistically counted by a second device during a compression period, and the M ciphertext blocks are sent by X first devices, where X is a positive integer.

[0225] In one implementation, if the second device is a network edge device, then the X first devices are all located in a region corresponding to the network edge device.

[0226] Among them, for the specific implementation manners of the obtaining unit 1301, the detecting unit 1302, and the encoding unit 1303, reference may be made to the descriptions of steps S401 - S403 in the corresponding embodiments above, and details will not be elaborated here. In addition, the beneficial effects of adopting the same method will not be elaborated either. Figure 10

[0227] Further, please refer to Figure 14 , Figure 14 which is a schematic structural diagram of a data processing device provided by an embodiment of the present application. Figure 4 As Figure 14 shown, specifically, the data processing device 4 may be used to implement Figure 6 or Figure 9 the functions of the first device in Figure 10 , or the data processing device 4 may be used to implement

[0228] the functions of the second device in Figure 14 Figure 14 Refer to Figure 14

[0229] ​​​The processor 1401 may be a central processor unit (CPU), a network processor (NP), or a combination of a CPU and an NP. The processor 1401 may also include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above-mentioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0230] The communication interface 1402 is used to receive and send data. Specifically, the communication interface 1402 may include a receiving interface and a sending interface. Among them, the receiving interface may be used to receive data, and the sending interface may be used to send data. The number of communication interfaces 1402 may be one or more.

[0231] The memory 1403 may include volatile memory, such as random-access memory (RAM); the memory 3003 may also include non-volatile memory, such as flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); the memory 1403 may also include a combination of the above types of memory.

[0232] Optionally, the memory 1403 stores an operating system and programs, executable modules, or data structures, or subsets thereof, or extended sets thereof. Among them, the programs may include various operation instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks. The processor 1401 may read the programs in the memory 1403 to implement the method provided in the embodiments of the present application.

[0233] Among them, the memory 1403 may be a storage device in the data processing device 4 or a storage device independent of the data processing device 4.

[0234] The bus system 1404 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus system 3004 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 14 it is represented by only one thick line in Figure 14 , but this does not mean that there is only one bus or one type of bus.

[0235] In some possible embodiments, the above data processing device can be implemented as a virtualized device. For example, the virtualized device can be a virtual machine (VM) running a program for data processing functions, and the virtual machine is deployed on a hardware device (for example, a physical server). A virtual machine refers to a complete computer system with complete hardware system functions simulated by software and running in a completely isolated environment. The virtual machine can be configured as a data processing device. For example, the function of the data processing device can be implemented based on a general physical server combined with network functions virtualization (NFV) technology. Those skilled in the art can virtualize a data processing device with the above functions on a general physical server by reading this application, which will not be elaborated here.

[0236] It should be noted that the data processing device mentioned in the embodiments of this application can be the first device, the second device, or a chip for implementing the method of this application, and the embodiments of this application do not make specific limitations. When the data processing device is a chip, the interface circuit in the chip can be used to perform reception or transmission operations, and the processor in the chip can be used to perform processing operations.

[0237] In a specific implementation, the embodiments of this application also provide a chip, including a processor and an interface circuit. The interface circuit is used to receive instructions and transmit them to the processor; the processor can be used to perform the operations of each data processing device in the above data processing method. Among them, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the chip implements the method in any of the above method embodiments.

[0238] Optionally, there can be one or more processors in the chip. The processor can be implemented by hardware and / or software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc. When implemented by software, the processor can be a general-purpose processor that implements by reading the software code stored in the memory.

[0239] Optionally, the memory in the chip may also be one or more. The memory may be integrated with the processor or may be separately provided from the processor, which is not limited in this application. Exemplarily, the memory may be a non-transitory processor, such as a read-only memory (ROM), which may be integrated with the processor on the same chip or may be separately provided on different chips. This application does not specifically limit the type of the memory and the setting manner of the memory and the processor.

[0240] Exemplarily, the chip may be a field programmable gate array (FPGA), may be an application-specific integrated circuit (ASIC), may also be a system on chip (SoC), may also be a central processor unit (CPU), may also be a network processor (NP), may also be a digital signal processing circuit (DSP), may also be a micro controller unit (MCU), may also be a programmable logic device (PLD) or other integrated chips.

[0241] An embodiment of this application also provides a computer-readable storage medium, including instructions or a computer program, which, when running on a processor, enables the processor to execute the data processing method provided in the foregoing embodiment.

[0242] An embodiment of this application also provides a computer program product including instructions or a computer program, which, when running on a processor, enables a data processing device to execute the data processing method provided in the foregoing embodiment.

[0243] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and drawings of this application are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way may be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.

[0244] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0245] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical service division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0246] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of the embodiments of the present application.

[0247] Those skilled in the art should be able to realize that in the above one or more examples, the services described in the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these services can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transmission of a computer program from one place to another. The storage media can be any available medium accessible by a general-purpose or special-purpose computer.

[0248] The above specific implementation manners further elaborate the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above is only the specific implementation manner of the present application.

[0249] The above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A data processing method, characterized in that, Applied to a first device, the method includes: Determine a first key corresponding to a first plaintext block, where the first key is obtained by performing key generation calculations on N sub-block data of the first plaintext block respectively. The first key is the same as a second key, and the second key is a key corresponding to a second plaintext block similar to the first plaintext block. The second key is obtained by performing key generation calculations on W sub-block data of the second plaintext block respectively. Both N and W are positive integers greater than 1; Based on the first key, encrypt the first plaintext block to obtain a first ciphertext block.

2. The method according to claim 1, wherein The first plaintext block is any one of the plaintext blocks obtained by dividing the original data, and the first plaintext block includes N sub-block data; The determining the first key corresponding to the first plaintext block includes: Determine the hash values corresponding to the N sub-block data respectively; Determine the target hash value among the N hash values as the first key corresponding to the first plaintext block.

3. The method according to claim 2, characterized in that Each sub-block data in the N sub-block data is obtained by sliding a first sliding window on the first plaintext block. The first sliding window and a second sliding window are the same sliding window, and the second sliding window is used to divide the original data into blocks.

4. The method according to claim 1, wherein The first device is provided with a buffer, and the buffer is used to store the target hash value during the block division process; The determining the first key corresponding to the first plaintext block includes: Traverse the original data using the first sliding window; Determine the data covered by the first sliding window at position j as sub-block data j, and determine the hash value H of the sub-block data j j , where j is used to indicate the distance between the current window position and the position of the first boundary point, and j is a positive integer less than or equal to N; Based on the hash value H j , update the target hash value; Until it is determined that the position j reaches the second boundary point, determine the data block between the second boundary point and the first boundary point as the first plaintext block, and determine the target hash value in the updated buffer as the first key corresponding to the first plaintext block.

5. The method according to any one of claims 2 to 4, characterized in that The target hash value is the maximum hash value or the minimum hash value.

6. The method according to any one of claims 1-5, characterized in that, The first key is the same as a third key corresponding to a third plaintext block. The third plaintext block is a plaintext block detected in the first device and similar to the first plaintext block. The third key is obtained by performing key generation calculations on P sub-block data of the third plaintext block respectively. P is a positive integer greater than 1; The encrypting the first plaintext block based on the first key to obtain a first ciphertext block includes: Determine the first difference data between the first plaintext block and the third plaintext block; Based on the first key, encrypt the first difference data, and determine the encrypted first difference data as the first ciphertext block.

7. The method according to any one of claims 1-6, characterized in that, The first plaintext block is any one of the plaintext blocks obtained by dividing the original data. The original data includes first data and second data, and the first plaintext block is the last segmented data in the first data. The method further includes: Divide the second data into blocks until a fourth plaintext block is obtained. The fourth plaintext block is the next plaintext block of the first plaintext block in the original data; Determine a fourth key corresponding to the fourth plaintext block. The fourth key is obtained by performing key generation calculations on Q sub-block data of the fourth plaintext block respectively. Q is a positive integer greater than 1; Encrypt the fourth plaintext block based on the fourth key.

8. The method according to any one of claims 1 to 7, characterized in that, The method further includes: Obtain R local ciphertext blocks, where the R local ciphertext blocks include the first ciphertext block, and R is a positive integer greater than 1; Perform similarity detection on the R local ciphertext blocks to obtain a similarity result; If the similarity result indicates that there is a second ciphertext block in the R local ciphertext blocks that is similar to the first ciphertext block, then perform differential coding on the first ciphertext block and the second ciphertext block.

9. A data processing method, characterized in that, When applied to a second device, the method includes: Obtain M ciphertext blocks, where the M ciphertext blocks include a first ciphertext block, and M is a positive integer greater than 1; Perform similarity detection on the M ciphertext blocks, and determine the ciphertext block that is similar to the first ciphertext block as the second ciphertext block, where the first key corresponding to the first ciphertext block and the second key corresponding to the second ciphertext block are the same; Perform differential coding on the first ciphertext block and the second ciphertext block.

10. The method according to claim 9, wherein The M ciphertext blocks are the ciphertext blocks counted by the second device during a compression period, and the M ciphertext blocks are sent by X first devices, where X is a positive integer.

11. The method according to claim 10, characterized in that, If the second device is a network edge device, then all of the X first devices are located in the area corresponding to the network edge device.

12. A data processing system, characterized in that, It includes a first device and a second device; wherein, the first device is configured to execute the method according to any one of claims 1-8, and the second device is configured to execute the method according to any one of claims 9-11.

13. A data processing device, characterized in that, It includes a memory and a processor, the memory is configured to store computer instructions, and the processor is configured to call and run the computer instructions from the memory to implement the method according to any one of claims 1-8, or to implement the method according to any one of claims 9-11.

14. A computer program product, characterized in that, The computer program product includes instructions, when the instructions run on a computer, causing the computer to execute the method according to any one of claims 1-8, or to execute the method according to any one of claims 9-11.