Group-oriented privacy protection block-level duplicate removal and multi-copy dynamic auditing method

The group private key and tag update factor are generated through the private key generation center, combined with data block division and multi-set hash function, solving the problem of high overhead in the multi-replica storage strategy, realizing multi-replica data integrity audit and data deduplication, supporting dynamic management of group members, providing data privacy protection and lightweight key updates.

CN120474692APending Publication Date: 2025-08-12SUQIAN COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510505656.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art has problems such as high audit overhead and high computing resource utilization in multi-replica storage strategies, and it is impossible to effectively realize multi-replica storage audit of group shared files.

Method used

The private key generation center is used to generate the group private key and tag update factor, and the block-level deduplication credentials are generated through data block division, encryption and generating block-level deduplication credentials, combined with the multi-set hash function to realize data deduplication and multi-replica dynamic auditing, and a flexible label generation mechanism is designed to perform multi-replica block-level dynamic operations.

Benefits of technology

It realizes multi-replica data integrity audit protection, reduces storage and update overhead, supports dynamic management of group members, provides data privacy protection, and reduces cloud overhead through explicit text hybrid audit.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474692A_ABST
    Figure CN120474692A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data cloud storage, and particularly relates to a group-oriented privacy protection block-level deduplication and multi-copy dynamic auditing method, which comprises the following steps of: S1, system initialization: generating system public parameters and a main private key; s2, generating a group key; s3, data processing: performing data processing on the group sharing document and then storing a data set; s4, auditing interaction: judging the integrity of the data stored on the cloud server by verifying the legality of the auditing evidence; s5, data dynamic operation: the cloud server performs dynamic operation on the specified multi-copy block after identity verification; s6, dynamic group members: the group users apply for joining or quitting the group; and S7, data de-duplication: performing data de-duplication based on the block-level de-duplication voucher and the file label. According to the method, multi-copy data integrity auditing protection is realized, a plurality of copies are flexibly stored, and the expenditure related to storage and updating is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data cloud storage, and in particular relates to a group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method. Background Art

[0002] With the surge in data volumes, cloud storage has become the preferred data backup method for users. However, data corruption and loss are becoming increasingly prominent. For users who do not keep local copies of their data, ensuring the integrity of cloud-stored data is crucial. To this end, data integrity auditing schemes have emerged, allowing users to verify data integrity through random challenges. A variety of auditing strategies already exist in the existing technology, which can meet the diverse needs of single-copy storage environments, including data updates, data privacy protection, group data sharing, and key updates. Group data sharing is a hot research topic under single-copy auditing. The core issue is how to sign shared documents for groups with multiple group users. Currently, research on group document auditing focuses on privacy protection and dynamic member management. Privacy protection mainly involves encrypting the entire file through various encryption technologies. Although existing research can successfully support member joining and leaving, the resulting label update overhead is a significant burden for users.

[0003] At the same time, single-copy storage still has potential risks, which may cause data files to be permanently damaged or lost, so multi-copy strategies have emerged. Multi-copy strategies significantly improve data availability and recoverability by storing multiple copies of data files on cloud storage servers. However, if the existing group data sharing single-copy audit scheme applied to the single-copy strategy is directly applied to the multi-copy scenario, the audit overhead will increase linearly. For example, due to the limitations of the private key generation and label construction methods of the group data sharing single-copy audit scheme, when applied to the multi-copy strategy, if the group user members increase or decrease, it is necessary to generate and update multiple secret keys for each user member, and to construct and update the authenticatable labels one by one. These behaviors will result in high label update overhead and cannot adapt to the needs of variable copy storage and mixed plaintext and ciphertext auditing.

[0004] Furthermore, research indicates that data stored in cloud servers contains a significant amount of redundant data. Improving cloud storage efficiency and reducing cloud storage space have become hot topics. Therefore, data deduplication is an effective means of improving cloud storage performance. A commonly adopted solution is for data owners to employ convergent encryption, using the file's hash value as the encryption key to create encrypted data blocks. This approach ensures that users with the same file generate identical encrypted data blocks. Once uploaded to the cloud, cloud servers can use these encrypted data blocks to deduplicate data. To support data deduplication and enable auditing of cloud storage data, some research has attempted to build data integrity schemes based on ciphertext deduplication within single-copy storage, primarily with the goal of improving deduplication efficiency. Furthermore, some research has attempted to integrate blockchain technology to enable audit solutions that share audit results by uploading audit challenges and evidence to the blockchain. However, while some single-copy audit schemes that support deduplication exist, as well as multi-copy audit schemes that meet specific scenarios, none of them can achieve multi-copy storage auditing for group-shared files while reducing cloud overhead. Summary of the Invention

[0005] The purpose of the present invention is to provide a group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method to solve the technical problems that the existing technology has obvious high overhead and occupies more computing resources in the application of copy auditing and data deduplication in multi-copy storage strategies.

[0006] The group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method is applied to a system model, which includes:

[0007] Private Key Generation Center: PKG stands for Private Key Generation Center, which is responsible for generating group private keys and label update factors;

[0008] Group users: The group ID of the group is GID, and the number of users is Θ, U θ represents the θth user, U1 represents the group manager, who is responsible for the initial processing of group shared files and the dynamic joining and exit of group members. When θ∈(1,Θ], U θ Indicates that ordinary group users can access group shared documents stored in the cloud;

[0009] Cloud Server: CS stands for Cloud Server, which provides the ultimate multi-copy storage service for the group;

[0010] The method comprises the following steps:

[0011] S1. System initialization: PKG initializes the system and generates system public parameters and master private key;

[0012] S2, Group key generation: U1 generates an aggregate identity and interacts with PKG to obtain the group private key;

[0013] S3, Data Processing: U1 processes the group shared document and stores the data set; data processing includes: ① data partitioning to generate data blocks, ② encryption and copy generation, the encryption process generates block-level deduplication certificates corresponding to the data blocks, ③ calculation of block labels and file labels, and ④ duplication checking;

[0014] S4, Audit Interaction: U1 generates a random challenge and sends it to CS. CS generates the corresponding audit evidence and returns it to U1. U1 verifies the legitimacy of the audit evidence to determine the integrity of the data stored on CS.

[0015] S5, Data dynamics: U1 sends a data update request to CS. After identity verification, CS performs dynamic operations on the specified multi-copy block;

[0016] S6. Group member dynamics: Group users apply to join or leave the group. PKG assists U1 in updating the corresponding group private key and calculating the update factor. The update factor is then sent to CS to update the block label.

[0017] S7. Data deduplication: Deduplication is performed based on block-level deduplication credentials and whether there are duplicates in file tags.

[0018] Preferably, in step S1, the system public parameters are G1 and G2 are two multiplicative cyclic groups of order q, where q is a large prime number, and there exists a computable bilinear pairing e:G1×G1→G2; g represents a generator in G1, and h, H, H1, H2 are four hash functions: ①h: ②H: ③H1: l represents the bit length, ④H2:{0,1} * →G1, where u1,u2,...,u s ∈G1, is a randomly selected s parameter, s is the number of regions into which the data block is divided; is a secret parameter, is a public parameter, public value Y = g γ ∈G1, master private key

[0019] Steps S4-S7 can be executed separately in any order.

[0020] Preferably, step S2 includes the following sub-steps:

[0021] S2.1, U1 identifies all users in the group and obtains Θ IDsθ ∈{0,1} l , and then generate the aggregate identity: Then send it to PKG;

[0022] S2.2. PKG selects secret parameters And Θ public parameters g1,g2,...,g Θ ∈G1, and calculate the group private key sk and public value R: R = g r ∈G1; PKG sends the group private key sk to U1 through a private channel;

[0023] S2.3, U1 by calculation Verify the validity of the group private key sk. If it passes, U1 accepts the group private key sk and keeps it secret. If it fails, U1 will resend the request to PKG.

[0024] Preferably, F represents a group shared document, and step S3 includes:

[0025] S3.1, U1 first divides F into several data blocks, and then divides the data blocks into regions. The operation on F uses the block level as the basic unit of data processing;

[0026] S3.2, block-level deduplication credentials are the encryption keys corresponding to the data blocks. U1 chooses to store a different number of copies for each data block;

[0027] S3.3, U1 partially encrypts F and creates multiple copies simultaneously, where the copies include multiple replica blocks corresponding to the original data block;

[0028] S3.4, U1 calculates file tags and prevents files from being added to the cloud repeatedly;

[0029] S3.5, U1 calculates the block label;

[0030] S3.6. Perform a consistency check on the received data based on the file tag and block tag, and store the corresponding data set if the consistency check passes.

[0031] Preferably, in step S3.1, U1 first divides F into data blocks, and then further divides the data blocks into regions: F = {m i} 1≤i≤n ={m ij} 1≤i≤n,1≤j≤s , where m i Represents the i-th data block. m ij represents the data in the jth zone belonging to the i-th data block. In step S3.2, F is divided into two sets: F = F1∪F2, where F1 is the data set in F that can be made public, and F2 is the data set in F that contains private data and needs to be encrypted; the i-th data block mi With ν i replica blocks; in step S3.3, the encryption and replica block generation strategies are integrated to achieve one-time processing, C represents the processed replica block set, and the corresponding expression is:

[0032]

[0033] Where E is the encryption algorithm, Refers to the i-th data block m in F as the original file i The ζ-th replica block, m i The corresponding replica block is

[0034] Preferably, in step S3.4, the z-th replica block of the i-th data block is divided into s zones, expressed as: Λ i Indicates the ν corresponding to the i-th data block i The concatenation of all regions of a replica block is expressed as:

[0035]

[0036] U1 calculates the file label τ C For file-level deduplication, the calculation formula is:

[0037]

[0038] Among them, mh is the multi-set hash function corresponding to the hash function h; then, U1 combines the group identifier GID with the file label τ C Upload to CS together, CS uses file tag τ C Check whether the file already exists in the cloud and whether the group ID exists in the check record table, which records the correspondence between the cloud storage file and the owner's identity.

[0039] Preferably, step S3.5 specifically includes: U1 randomly selects a secret parameter And calculate the public value N = g η , then calculate the block label, block label σ i The calculation formula is:

[0040]

[0041] Where B=H2(GID||s); Finally, the complete block label set Φ={σ i} 1≤i≤n ; Step S3.6 specifically includes: U1 takes the data set {C, Φ, τ C ,GID} and the set of replica numbers {ν1,ν2,...,ν n} and upload them to CS together, and label the file τ C Stored locally; when CS receives the uploaded information, it first uses the replica block set C to calculate the verification tag The calculation formula is: The verification label τ C With file label τ C Perform consistency comparison to verify file tags, and then use the expression: Verify the consistency of the block tag and the uploaded block. If both are verified successfully, CS stores the corresponding data set. Otherwise, CS refuses to store and notifies U1.

[0042] Preferably, step S4 includes the following sub-steps:

[0043] S4.1. U1 generates a random challenge Q, which is expressed as where i∈I c , I c is a set of c integers randomly selected from the set of integers [1, n], It is from U1 sends the random challenge Q to CS.

[0044] S4.2. CS generates verification evidence P, which is expressed as: P = {{μ j} 1≤j≤s ,σ}, where CS sends evidence P to U1;

[0045] S4.3, U1 checks the integrity of the stored data using the following expression:

[0046]

[0047] If the equation is true, it proves that the data integrity on CS has passed the test and outputs "TRUE"; otherwise, it means that the test has not passed and outputs "FALSE".

[0048] Preferably, in step S5, U1 processes the new data block and sends a dynamic update request to CS. After verification, CS performs dynamic operations and calculates the file tag, and returns the newly calculated file tag to U1. U1 checks the newly calculated file tag and the file tag previously stored locally. If the verification passes, the received file tag replaces the file tag previously stored locally. If it fails, CS is required to respond again.

[0049] Preferably, step S6 includes: user U α When exiting or joining a group, U1 will α Sent to PKG; PKG updates the group private key, the calculation formula is: Update factor when exiting a group Update factor when joining a group And the new private key Sent to U1 through a private channel; U1 has a new private key The validity of is verified, the expression is: If it fails, U1 resends the request to PKG; if it passes, U1 accepts the group private key And keep it secret, and then Sent to CS, CS updates the block label of the relevant group shared files, the calculation formula is: For group-shared files, only n group multiplication operations are required to update the labels of each block.

[0050] The present invention has the following advantages:

[0051] (1) Multi-copy data integrity audit protection: By launching integrity challenges to the cloud server during the audit, group users can detect whether the cloud server actually holds a copy of the data without holding a local copy of the data block, thereby achieving multi-copy data integrity protection.

[0052] (2) Flexible multi-copy storage: Different replica blocks are stored according to the value of the file data block. Compared with the fixed number of replica storage solutions, it has better performance. For data blocks with lower importance, the number of backup replica blocks can be reduced, reducing the related storage and update overhead.

[0053] (3) Dynamic update of multi-replica block level: Through the designed label generation mechanism and multi-set hash function, the consistency of multi-replica block level dynamic operations (including block update, deletion and insertion) between the cloud and the local end is achieved. In addition, the computational overhead of the local end and the cloud is constant and has nothing to do with the number of data blocks and replicas.

[0054] (4) Group member dynamics: By designing the group private key to bind to the member identity, a new group private key and update factor are generated with the help of PKG. When members join or leave, the group key and the label corresponding to the cloud group file can be updated in real time.

[0055] (5) Data privacy protection: Message locking encryption technology is used to encrypt data blocks containing privacy. Without the original data blocks, entities other than group users cannot infer the plaintext information.

[0056] (6) Hybrid audit of plaintext and ciphertext: Different from the traditional method of completely encrypting the entire file, the scheme in this paper allows encryption of some data blocks containing private information, and combines encryption technology with multi-copy block generation strategy to realize hybrid audit of plaintext and ciphertext for multi-copy storage in cloud.

[0057] (7) Local block-level and file-level deduplication: The encryption key generated by the message-locked encryption technology is used as the local block-level deduplication tag, and the file tag generated by the multi-set hash function is used as the cloud and local file deduplication tag to avoid group users uploading duplicate data one after another, reducing the communication between users and cloud servers and cloud storage overhead.

[0058] (8) Lightweight key update and label conversion overhead: When the group private key is updated or a member joins or leaves, the group private key can be quickly updated through a single multiplication operation using the update factor (a constant) sent by the group manager. At the same time, the cloud only needs n group multiplication operations to quickly implement label conversion, that is, converting the block copy signature signed by the previous group private key to the corresponding block copy signature under the new private key. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a basic flow chart of the group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method of the present invention.

[0060] Figure 2 It is a schematic diagram of the interaction mode of the system model applying the present invention. DETAILED DESCRIPTION

[0061] The following is a further detailed description of the specific implementation methods of the present invention through the description of the embodiments with reference to the accompanying drawings, so as to help those skilled in the art to have a more complete, accurate and in-depth understanding of the inventive concept and technical solution of the present invention.

[0062] like Figure 1-Figure 2 As shown, the present invention provides a group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method. The system model in which this method is applied includes:

[0063] (1) Private Key Generator (PKG): PKG stands for Private Key Generation Center and is responsible for generating the system public parameters and the group private key sk. During the system initialization phase, it generates the group private key sk based on the aggregated identity of the given group users. It is also responsible for generating new group private keys sk and label update factors when the group changes dynamically.

[0064] (2) Group users (User, U): For group users, it is necessary to outsource shared document data to the cloud for multi-copy block-level storage. For a group, the group identifier is GID, the number of users is Θ, U θ Represents the θth user. All group users are divided into two categories: ① Ordinary group users, when θ∈(1,Θ], U θU1 represents an ordinary group user who can access group shared documents stored in the cloud, send integrity challenges Q to the CS (cloud server) on behalf of the group administrator, and can automatically leave the group as needed; ②U1 represents the group administrator, who is responsible for the initial processing of group shared files and the dynamic joining and leaving of group members. He is also responsible for sending integrity challenges Q to the CS (cloud server) during the audit period and sending data update requests req during data updates.

[0065] (3) Cloud Server (CS): CS stands for Cloud Server. Cloud servers with ample storage space provide the ultimate multi-copy storage service for the group. During the audit, they need to respond to integrity challenges from group users and perform block-level dynamic operations on the cloud multi-copy data based on data update requests sent by the group manager.

[0066] Specifically, the group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method includes the following steps.

[0067] S1. System Initialization: PKG initializes the system and generates system public parameters and master private key. This step includes the following sub-steps.

[0068] S1.1. Choose two multiplicative cyclic groups G1 and G2 of order q, and select a computable bilinear pairing e:G1×G1→G2, where q is a large prime number.

[0069] S1.2. Select a generator of G1, denoted by g, and choose four hash functions: ①h: ②H: ③H1: ④H2:{0,1} * →G1.

[0070] S1.3. PKG selects secret parameters And public parameters And calculate the public value Y = g γ ∈G1 and the master private key in

[0071] S1.4, randomly select s parameters u1,u2,...,u s ∈G1, and output the system public parameters as The master private key msk is held secretly by PKG.

[0072] S2. Group key generation: U1 generates an aggregated identity and interacts with PKG to obtain the group private key. This step includes the following sub-steps.

[0073] S2.1, U1 identifies all users in the group and obtains Θ IDs θ ∈{0,1} l , where l represents the bit length, and then generates the aggregated identity: Then send it to PKG to obtain the group private key.

[0074] S2.2, when receiving Θ ID θ Afterwards, PKG selects secret parameters And Θ public parameters g1,g2,...,g Θ ∈G1 and calculate the group private key sk and public value R. The calculation formulas of the group private key sk and public value R are: R = g r ∈G1; for convenience, the letter A is used instead in the subsequent steps Afterwards, PKG sends the group private key sk to U1 through a private channel.

[0075] S2.3, U1 by calculation Verify the validity of the group private key sk, where e is the bilinear pairing selected previously; if passed, U1 accepts the group private key sk and keeps it secret; if not, U1 will resend the request to PKG.

[0076] S3. Data Processing: U1 processes the group-shared document and stores the data set. F represents the group-shared document. Data processing includes: ① data partitioning to generate data blocks, ② encryption and copy generation, ③ calculation of block and file labels, and ④ duplicate checking. This step includes the following substeps.

[0077] S3.1, U1 first divides F into several data blocks, and then divides the data blocks into zones. The operations on F use the block level as the basic unit of data processing.

[0078] This step specifically includes: U1 first divides F into data blocks: F={m i} 1≤i≤n , where m i Represents the i-th data block, the number of data blocks is n, and the data blocks are further divided into regions: F = {m i} 1≤i≤n ={m ij} 1≤i≤n,1≤j≤s ,in represents the data in the jth zone of the i-th data block, where s is the number of zones. To reduce computational and communication overhead, this step further divides the data block into zones. Subsequent operations use the block level as the basic unit of data processing. For example, n block tags are ultimately generated for each data block in the file, rather than n × s zone tags for each zone. This process reduces the number of operations on the divided data.

[0079] S3.2. Generate block-level deduplication credentials. The block-level deduplication credentials are the encryption keys corresponding to the data blocks. U1 chooses to store a different number of copies for each data block.

[0080] This step specifically includes: using ν i Represents the i-th data block m i The number of copies that need to be stored, there is a copy number set ν i ∈[1,2,...,ν max ], ν max Select the maximum number of copies to be stored for U1; divide F into two sets: F = F1∪F2, where F1 is the set of data that can be made public in F, and F2 is the set of data in F that needs to be encrypted, that is, the set of data containing private data. To protect data privacy and achieve block-level deduplication, U1 generates the corresponding block-level deduplication certificate / encryption key k i , calculated as k i =H(m i1 ||m i2 ||...||m is ), H is the hash function selected previously. Block-level deduplication certificate / encryption key is represented by k i It is both a block-level deduplication certificate and an encryption key. U1 stores k locally. i Used to implement local block-level deduplication and subsequent decryption of private data.

[0081] S3.3, U1 partially encrypts F and creates multiple copies at the same time. The copies include multiple replica blocks corresponding to the data blocks. The encryption and replica block generation strategies are integrated to achieve one-time processing.

[0082] This step specifically involves: To further reduce local overhead, U1 chooses to partially encrypt F. This saves overhead compared to encrypting the entire file. It also creates different replicas, enabling multiple copies of data for improved data recoverability and availability. Replicas are composed of replica blocks, and the number of replica blocks corresponding to each data block is determined based on step S3.2. This step integrates encryption and replica block generation strategies to achieve a single-step process, simultaneously protecting data privacy and achieving replica generation.

[0083] C represents the set of processed replica blocks, and the corresponding expression is:

[0084]

[0085] Where E is the efficient symmetric encryption algorithm AES-256, ζ∈[1,ν i ], Refers to the i-th data block m in F as the original file i The ζ-th replica block, that is, m iA total of ν i replica blocks, m i The corresponding replica block set is When any group of users U θ Download any replica block When the user only needs to XOR it with the copy number to get the corresponding data; if the corresponding data is meaningful, it can be read directly; if it is meaningless, it can be decrypted using the decryption algorithm E. For non-group users, even if they get the copy block that needs to be protected, (i.e. the replica block is encrypted), due to the lack of the encryption key k calculated from the original plaintext i , so the original plaintext cannot be obtained, thus protecting data privacy. At the same time, at any time, it is only necessary to ensure that any replica block of the corresponding data block is correct, and other damaged replica blocks corresponding to the data block can be restored. U1 can select different replica numbers for each data block. i , to meet flexible storage needs. For example, for a file with n data blocks, if a traditional multi-copy fixed storage solution is used, assuming that one of the most valuable data blocks requires s copies, then the user needs to pay for the storage space occupied by n×s copies of the blocks. If this method is used, different replica numbers can be flexibly selected for each data block. i , assuming that the i-th data block is s i replica blocks, then 1≤s i ≤s, for all n data blocks, the storage space required for all replica blocks is ∑s i , obviously ∑s i The more data blocks there are, the more obvious the advantage of the solution of the present invention in terms of overhead.

[0086] S3.4, U1 calculates file tags and prevents files from being added to the cloud repeatedly.

[0087] This step specifically includes: the z-th replica block of the i-th data block is divided into s zones, which is expressed as:

[0088]

[0089] Λ i Indicates the ν corresponding to the i-th data block i The concatenation of all regions of a replica block is expressed as:

[0090]

[0091] U1 calculates the file label τ C For file-level deduplication, the calculation formula is:

[0092]

[0093] Here, h is the hash function selected previously, and mh is the multiset hash function corresponding to h. A multiset hash function is a function that maps a multiset (a set that allows elements to appear repeatedly) to a fixed-length hash value.

[0094] Afterwards, U1 combines the group identifier GID with the file label τ C Upload to CS together, CS uses file tag τ C Check whether the file already exists in the cloud. If not, return the message "does not exist" to U1, and U1 will perform subsequent file upload processing; if it exists, check the τ in the record table CT. C Check whether the corresponding row contains a GID. The CT records the file tag and its corresponding relationship with the identity (group ID) that can access the file. After checking the CT, if the GID does not exist in the CT, the GID is added to the record corresponding to the group shared file. If the GID exists, it is ignored and a "yes" message is returned to U1. U1 does not process the file further after receiving the message. This prevents duplicate files from being added to the cloud.

[0095] S3.5, U1 calculates the block label.

[0096] This step specifically includes: U1 randomly selects a secret parameter And calculate the public value N = g η , then calculate the block label, block label σ i The calculation formula is:

[0097]

[0098] in B=H2(GID||s), H1 and H2 are the hash functions selected previously. Finally, the complete block label set Φ={σ i} 1≤i≤n , that is, for an original file with n data blocks, a set of n labels is finally generated.

[0099] S3.6. Perform a consistency check on the received data based on the file tag and block tag, and store the corresponding data set if the consistency check passes.

[0100] This step specifically includes: U1 takes the data set {C, Φ, τ C ,GID} and the set of replica numbers {ν1,ν2,...,ν n} together to CS, and then label the file τ C Stored locally for subsequent data update judgment and local file-level deduplication.

[0101] When CS receives the upload information, it first uses the replica block set C to calculate the verification tag The calculation formula is: Then verify the label With file label τ C Compare them. If they are different, CS refuses to store them and notifies U1. If they are the same, CS further verifies the consistency of the block tag and the uploaded block. The expression is as follows:

[0102]

[0103] in, B = H2(GID||s). If the equation holds, CS stores the corresponding data set; otherwise, CS refuses to store the data and notifies U1. To save computational overhead for subsequent audit challenges, CS can store the corresponding A and B values for the corresponding group, eliminating the need for repeated calculations.

[0104] S4. Audit Interaction: U1 generates a random challenge and sends it to CS. CS generates the corresponding audit evidence and returns it to U1. U1 verifies the legitimacy of the audit evidence to determine the integrity of the data stored on CS. This step includes the following sub-steps.

[0105] S4.1. U1 generates a random challenge Q, which is expressed as where i∈I c , I c is a set of c integers randomly selected from the set of integers [1, n], It is from U1 sends the random challenge Q to CS.

[0106] S4.2. CS generates verification evidence P, which is expressed as: P = {{μ j} 1≤j≤s ,σ}, where CS sends evidence P to U1.

[0107] S4.3, U1 checks the integrity of the stored data using the following expression:

[0108]

[0109] If the equation is true, it proves that the data integrity on CS has passed the test and outputs "TRUE"; otherwise, it means that the test has not passed and outputs "FALSE".

[0110] This step enables the solution of the present invention to support public auditing of data integrity. Here, for convenience, the group administrator is directly used to perform the challenge. Of course, ordinary group members can also be used to perform the challenge or outsource the audit to a third-party entity as needed.

[0111] S5. Data dynamics: U1 sends a data update request to CS. After identity verification, CS performs dynamic operations on the specified multi-copy block.

[0112] In this step, U1 processes the new data block and sends a dynamic update request to the CS. After verification, the CS performs dynamic operations and calculates the file tag. It then returns the newly calculated file tag to U1. U1 checks the newly calculated file tag against the previously stored file tag. If verification passes, the received file tag replaces the previously stored file tag. If verification fails, the CS is requested to respond again. This step includes the following substeps.

[0113] S5.1. Data update: Determine whether to update through double verification of block tags and file tags, and update multiple copies of the corresponding data block stored in the cloud after the double verification is passed.

[0114] This step specifically includes: U1 needs to convert the a-th data block m in F a Update to update data block First update the data block Further divided into zones, the expression is: Then generate the corresponding block-level deduplication certificate / encryption key based on this content The calculation formula is:

[0115] When updating a data block, there are two cases where a replica block is generated: 1) If Belong to the data set that can be made public, then calculate ν a replica blocks, the calculation formula is: 2) If The data set belongs to privacy, calculate ν a replica blocks, the calculation formula is:

[0116] The obtained replica blocks are further divided into zones, and the expression is: Regenerate block tags The calculation formula is:

[0117]

[0118] Next, U1 uploads the dynamic update request req to the cloud storage server CS. The expression of the updated req is:

[0119] When CS receives the upload information, it first uses the record table CT to determine the legitimacy of the user identity. If it is not legal, the execution is rejected. If the user identity is legal, the consistency between the received file tag and the locally stored file tag is checked. If they are inconsistent, the execution is rejected. Otherwise, the consistency between the updated data block and the block tag is verified: If the verification fails, the execution is rejected; otherwise, the as well as Calculate new file labels The calculation formula is as follows:

[0120]

[0121] in,

[0122]

[0123] CS use and Replace the previous σ a With τ C , and and h(Λ a ||ν a ) is sent to U1 for update operation verification.

[0124] U1 uses the locally stored τ C as well as With the received and h(Λ a ||ν a ) for verification, the expression is: This checks the correctness of the CS operation. If the verification passes, U1 considers the CS operation to be correct and saves the new file label. and the new encryption key Otherwise, U1 asks CS to respond again.

[0125] S5.2. Data deletion: Determine whether to delete the data by verifying the file tag, and delete the corresponding group shared files after the verification is passed.

[0126] This step specifically includes: U1 needs to delete the bth data block U1 uploads the dynamic update request req to CS. The expression of the deleted req is:

[0127] When CS receives the uploaded information, it first uses the registration table CT to determine the legitimacy of the user's identity. If it is illegal, it refuses to execute. Otherwise, it checks the consistency between the received file label and the stored file label. If they are different, it refuses to execute. If they are the same, it uses the file label τ C as well as Calculate new file labels The calculation formula is: in,

[0128]

[0129] Afterwards, CS is deleted from storage and and use Instead of τ C , will also and h(Λ b ||ν b ) is sent to U1 for deletion operation verification.

[0130] U1 uses the locally stored τ C With the received and h(Λ b ||ν b ) for verification, the expression is: In order to check the correctness of the operation, if the verification is passed, U1 believes that the CS operation is correct and uses Instead of τ C ; Otherwise, U1 requires CS to respond again.

[0131] S5.3. Data insertion: Determine whether to insert data by verifying the consistency between the updated data block and the block label, and insert data into the corresponding group shared file after the verification is passed.

[0132] This step specifically includes: U1 needs to insert an update data block The data block Further dividing the area, the expression is: Then generate block-level deduplication credentials / encryption keys based on this content The calculation formula is:

[0133] When inserting a data block, there are two cases where a copy block is generated: 1) If Belong to the publicly available data set, calculation replica blocks, the calculation formula is: 2) If Data collection belonging to privacy, calculation replica blocks, the calculation formula is:

[0134] The obtained replica blocks are then further divided into zones, expressed as: Then generate the block tags The calculation formula is:

[0135]

[0136] U1 uploads the dynamic update request req to the cloud storage server CS. The expression of req inserted is: Among them, d means that the data block will be updated Insert at block position d.

[0137] When CS receives the uploaded information, it first uses the registration table CT to determine the legitimacy of the user identity. If it is illegal, it refuses to execute; otherwise, it verifies the consistency of the updated data block and the block label. The expression is:

[0138]

[0139] If the verification fails, the execution is rejected; otherwise, the previously stored file label τ is used. C as well as Calculate new file labels The calculation formula is: in,

[0140]

[0141] CS Storage and and use Instead of τ C , will also Sent to U1 for insertion operation verification.

[0142] U1 uses the locally stored τ C as well as With the received To verify, the expression is: in Depend on If the verification is successful, U1 considers that the CS operation is correct and uses Instead of τ C ; Otherwise, U1 requires CS to respond again.

[0143] S6. Group Member Dynamics: When a group user applies to join or leave the group, PKG assists U1 in updating the corresponding group private key and calculating the update factor, which is then sent to CS to update the block label. This step includes the following sub-steps.

[0144] S6.1. Group member exit: User Uα When exiting the group, U1 will α Sent to PKG. PKG updates the group private key, the calculation formula is: Update Factor PKG will be the new private key Sent to U1 through a private channel, U1 has a new private key The validity of is verified, the expression is: If it fails, U1 resends the request to PKG; if it passes, U1 accepts the group private key And keep it secret, and then Sent to CS, CS updates the block label of the relevant group shared files, the calculation formula is: For group-shared files, only n group multiplication operations are required to update the labels of each block.

[0145] S6.2. Group member joining: User U β When joining a group, U1 will use the corresponding ID β Sent to PKG. PKG updates the group private key, the calculation formula is: Update Factor PKG will be the new private key Sent to U1 through a private channel, U1 has a new private key The validity of is verified, the expression is: If it fails, U1 resends the request to PKG; if it passes, U1 accepts the group private key and keep it secret, and Sent to CS, CS updates the block label of the relevant group shared files, the calculation formula is: For group-shared files, only n group multiplication operations are required to update the labels of each block.

[0146] S7, Data Deduplication: Deduplication is performed based on the block-level deduplication credentials and whether the file tags are duplicated. This step includes the following sub-steps.

[0147] S7.1. Block-level deduplication: U1 is based on the block-level deduplication certificate k calculated in step S3. i , compare the block-level deduplication certificates stored locally. If there is no duplication in the block-level deduplication certificates, proceed with subsequent data block processing; if there is duplication, subsequent operations such as division, encryption to generate copies, block labels and file labels only need to be performed on non-duplicate data blocks.

[0148] S7.2. File-level deduplication: U1 calculates the file label τ based on the multi-set hash function C , file label τ C Stored on the local user terminal and CS at the same time, by comparing the file label τC Perform file-level deduplication on both the client and CS. Before uploading files to the cloud, the client first detects file duplication based on file tags to save cloud storage overhead.

[0149] The above steps S1-S3 need to be executed in sequence, while steps S4-S7 can be executed separately in any order.

[0150] The present invention is described above by way of example in conjunction with the accompanying drawings. It is obvious that the specific implementation of the present invention is not limited to the above-mentioned method. As long as various non-substantial improvements are made using the inventive concept and technical solution of the present invention, or the inventive concept and technical solution are directly applied to other occasions without improvement, they are all within the scope of protection of the present invention.

Claims

1. A privacy-preserving block-level deduplication and multi-copy dynamic auditing method for groups, characterized by: Applied to a system model, the system model includes: Private key generation center: PKG represents the private key generation center, which is responsible for generating the group private key and label update factor; Group users: The group identifier of the group is GID, and the number of users is Θ, U θ represents the θth user, U1 represents the group manager, who is responsible for the initial processing of group shared files and the dynamic joining and exit of group members. When θ∈(1,Θ], U θ Indicates that ordinary group users can access group shared documents stored in the cloud; Cloud Server: CS stands for Cloud Server, which provides the ultimate multi-copy storage service for the group; The method comprises the following steps: S1. System initialization: PKG initializes the system and generates system public parameters and master private key; S2, Group key generation: U1 generates an aggregate identity and interacts with PKG to obtain the group private key; S3, Data Processing: U1 processes the group shared document and stores the data set; data processing includes: ① data partitioning to generate data blocks, ② encryption and copy generation, the encryption process generates block-level deduplication certificates corresponding to the data blocks, ③ calculation of block labels and file labels, and ④ duplication checking; S4, Audit Interaction: U1 generates a random challenge and sends it to CS. CS generates the corresponding audit evidence and returns it to U1. U1 verifies the legitimacy of the audit evidence to determine the integrity of the data stored on CS. S5, Data dynamics: U1 sends a data update request to CS. After identity verification, CS performs dynamic operations on the specified multi-copy block; S6. Group member dynamics: Group users apply to join or leave the group. PKG assists U1 in updating the corresponding group private key and calculating the update factor. The update factor is then sent to CS to update the block label. S7. Data deduplication: Deduplication is performed based on block-level deduplication credentials and file tags to determine if there are duplicates. Steps S4-S7 can be executed separately in any order.

2. The privacy-preserving block-level deduplication and multi-copy dynamic auditing method for groups according to claim 1 is characterized by: In step S1, the system public parameters are G1 and G2 are two multiplicative cyclic groups of order q, where q is a large prime number, and there exists a computable bilinear pairing e:G1×G1→G2; g represents a generator in G1, and h, H, H1, H2 are four hash functions: ① ② ③ l represents the bit length, ④H2:{0,1} * →G1, where u1,u2,...,u s ∈G1, is a randomly selected s parameter, s is the number of regions into which the data block is divided; is a secret parameter, is a public parameter, public value Y = g γ ∈G1, master private key 3. The privacy-preserving block-level deduplication and multi-copy dynamic auditing method for groups according to claim 2 is characterized by: Step S2 includes the following sub-steps: S2.1, U1 identifies all users in the group and obtains Θ IDs θ ∈{0,1} l , and then generate the aggregate identity: Then send it to PKG; S2.

2. PKG selects secret parameters And Θ public parameters g1,g2,...,g Θ ∈G1, and calculate the group private key sk and public value R: R = g r ∈G1; PKG sends the group private key sk to U1 through a private channel; S2.3, U1 passed the inspection Verify the validity of the group private key sk. If it passes, U1 accepts the group private key sk and keeps it secret. If it fails, U1 will resend the request to PKG.

4. The group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method according to claim 3 is characterized by: F represents a group shared document, and step S3 includes: S3.1, U1 first divides F into several data blocks, and then divides the data blocks into regions. The operation on F uses the block level as the basic unit of data processing; S3.2, block-level deduplication credentials are the encryption keys corresponding to the data blocks. U1 chooses to store a different number of copies for each data block; S3.3, U1 partially encrypts F and creates multiple copies simultaneously, where the copies include multiple replica blocks corresponding to the original data block; S3.4, U1 calculates file tags and prevents files from being added to the cloud repeatedly; S3.5, U1 calculates the block label; S3.

6. Perform a consistency check on the received data based on the file tag and block tag, and store the corresponding data set if the consistency check passes.

5. The group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method according to claim 4 is characterized by: In step S3.1, U1 first divides F into data blocks, and then further divides the data blocks into regions: F = {m i } 1≤i≤n ={m ij } 1≤i≤n,1≤j≤s , where m i represents the i-th data block; m ij represents the data in the jth zone belonging to the i-th data block. In step S3.2, F is divided into two sets: F = F1∪F2, where F1 is the data set in F that can be made public, and F2 is the data set in F that contains private data and needs to be encrypted; the i-th data block m i With ν i replica blocks; in step S3.3, the encryption and replica block generation strategies are integrated to achieve one-time processing, C represents the processed replica block set, and the corresponding expression is: Where E is the encryption algorithm, Refers to the i-th data block m in F as the original file i The ζ-th replica block, m i The corresponding replica block set is 6. The group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method according to claim 5 is characterized by: In step S3.4, the z-th replica block of the i-th data block is divided into s zones, expressed as: Λ i Indicates the ν corresponding to the i-th data block i The concatenation of all regions of a replica block is expressed as: U1 calculates the file label τ C For file-level deduplication, the calculation formula is: Among them, mh is the multi-set hash function corresponding to the hash function h; then, U1 combines the group identifier GID with the file label τ C Upload to CS together, CS uses file tag τ C Check whether the file already exists in the cloud and whether the group ID exists in the check record table, which records the correspondence between the cloud storage file and the owner's identity.

7. The group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method according to claim 6 is characterized by: Step S3.5 specifically includes: U1 randomly selects a secret parameter And calculate the public value N = g η , then calculate the block label, block label σ i The calculation formula is: Where B=H2(GID||s); finally, we get the complete block label set Φ={σ i } 1≤i≤n ; Step S3.6 specifically includes: U1 takes the data set {C, Φ, τ C ,GID} and the set of replica numbers {ν1,ν2,...,ν n } and upload them to CS together, and label the file τ C Stored locally; when CS receives the uploaded information, it first uses the replica block set C to calculate the verification tag The calculation formula is: Verify the label With file label τ C Perform a consistency comparison to verify the file tag, and then use the expression: Verify the consistency of the block tag and the uploaded block. If both are verified successfully, CS stores the corresponding data set. Otherwise, CS refuses to store and notifies U1.

8. The group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method according to claim 7 is characterized by: Step S4 includes the following sub-steps: S4.

1. U1 generates a random challenge Q, which is expressed as where i∈I c , I c is a set of c integers randomly selected from the set of integers [1, n], It is from U1 sends the random challenge Q to CS. S4.

2. CS generates verification evidence P, which is expressed as: P = {{μ j } 1≤j≤s ,σ}, where CS sends evidence P to U1; S4.3, U1 checks the integrity of the stored data using the following expression: If the equation holds true, it proves that the data integrity on CS has passed the test and outputs "TRUE"; otherwise, it fails the test and outputs "FALSE".

9. The group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method according to claim 8, characterized in that: In step S5, U1 processes the new data block and sends a dynamic update request to CS. After verification, CS performs dynamic operations and calculates the file tag, and returns the newly calculated file tag to U1. U1 checks the newly calculated file tag and the file tag previously stored locally. If the verification passes, the received file tag replaces the file tag previously stored locally. If it fails, CS is required to respond again.

10. The group-oriented privacy-preserving block-level deduplication and multi-copy dynamic auditing method according to claim 9, characterized in that: Step S6 includes: User U α When exiting or joining a group, U1 will α Sent to PKG; PKG updates the group private key, the calculation formula is: Update factor when exiting a group Update factor when joining a group And the new private key Sent to U1 through a private channel; U1 has a new private key The validity of is verified, the expression is: If it fails, U1 resends the request to PKG; if it passes, U1 accepts the group private key And keep it secret, and then Sent to CS, CS updates the block label of the relevant group shared files, the calculation formula is: For group-shared files, only n group multiplication operations are required to update the block label.