A Hybrid Spatial Subdivision Encryption Method and Device

Through the hybrid space subdivision encryption method, combined with strict and non-strict space subdivision algorithms, the problems of data authenticity and availability in anonymous data release in the big data era are solved, and the security, privacy and availability of data are balanced, and the company's compliance and customer trust are enhanced.

CN118337404BActive Publication Date: 2025-07-18INSPUR ZHUOSHU BIG DATA IND DEV CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410182722.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-07-18
Estimated Expiration
2044-02-19

AI Technical Summary

Technical Problem

The existing anonymous data release technology ensures data security while maintaining the authenticity and availability of data, especially in the era of big data, where there is insufficient protection of personal privacy.

Method used

The hybrid space subdivision encryption method is adopted, combined with strict and non-strict space subdivision algorithms, and the initial division is performed through a partial optimal greedy algorithm, which is strictly divided until it cannot continue, and then non-strict division is performed until the k-anonymity requirement is met, and the data availability is evaluated using discernible metrics, classified metrics and standardized deterministic penalties.

Benefits of technology

Improve the availability of anonymous data, ensure the security and privacy of data, while enhancing the company's legality, compliance and customer trust, and enhancing brand image and customer loyalty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118337404B_ABST
    Figure CN118337404B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data security and privacy protection, and specifically provides a hybrid spatial sub-division encryption method and device, which has the following steps: S1. Perform a binary partition operation on the overall data, and adopt a partial optimal greedy algorithm to ensure that the data information loss degree is always maintained in the two sub-spaces generated after each partition; S2. Perform strict sub-space partitioning on each respective sub-space in turn until the sub-spaces in the last layer can no longer meet the requirements of strict sub-space partitioning, and stop the strict partitioning operation; S3. At this point, several sub-anonymous data groups that have undergone strict sub-space partitioning are generated, and then viewed in turn. If there are any that meet the requirements of non-strict sub-space partitioning, perform non-strict sub-space partitioning operations again until there are no more sub-anonymous group data groups that meet the requirements of non-strict sub-space partitioning. Compared with the prior art, the present invention combines a strict spatial sub-division algorithm with a non-strict spatial sub-division algorithm, making the processed anonymous data have higher usability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data security and privacy protection, and particularly provides a hybrid space sub-division encryption method and device. Background Art

[0002] In the big data era, many companies or organizations need to regularly release data externally. For example, universities need to release scientific research results externally, and so on. In the past decade, with the rapid development of science and technology, especially computer technology, people have entered the big data era, and the collection and analysis of data have become increasingly convenient. However, this has also posed a major threat to people's personal privacy security. For example, by mining personal medical insurance information, it can be learned whether a certain patient has a certain major disease, and this process has violated the patient's personal privacy. Some studies have pointed out that by performing a linking operation on a registration form and a medical insurance information form with individual identifiers hidden through a quasi-identifier, the personal identity information of more than 80% of American citizens can be identified. Therefore, how to protect data during the data release process has become an urgent problem to be solved.

[0003] The privacy protection technology for data release has always been a research hotspot. Technologies such as data perturbation, data encryption, and data anonymization have emerged successively. Data perturbation is to perform random perturbation by adding noise to the original data; while data encryption protects privacy by hiding sensitive data, but it is not widely used due to its high cost; and the data anonymization release technology based on the K-anonymity model has become a research hotspot in the academic community while ensuring the security of the data and ensuring the authenticity of the data itself. Summary of the Invention

[0004] The present invention aims at the above-mentioned deficiencies of the prior art and provides a hybrid space sub-division encryption method with strong practicability.

[0005] A further technical task of the present invention is to provide a hybrid space sub-division encryption device with reasonable design, safety and applicability.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0007] A hybrid space sub-division encryption method has the following steps:

[0008] S1. Perform a binary partition operation on the overall data. When performing the binary partition on the data, use a partial optimal greedy algorithm to ensure that the data information loss degree is always maintained in the two sub-spaces generated after each partition;

[0009] S2. Successively perform strict sub-space partitions on the respective sub-spaces until the sub-spaces in the last layer can no longer meet the requirements of strict sub-space partitions, and then stop the strict partition operation;

[0010] S3. So far, several anonymous sub-groups are generated after strict subspace division, and then they are checked in turn. If there are any that meet the requirements of non-strict subspace division, the non-strict subspace division operation is performed again until there are no more anonymous sub-groups that meet the requirements of non-strict subspace division.

[0011] Furthermore, in step S1, a space [C 1,x1 ,C 1,y1 ]×[C 2,x2 ,C 2,y2 ]×……×[C d,xd ,C d,yd The level of ] is

[0012] According to the definition, if [C 1,0 ,C 1,t1 ]×[C 2,0 ,C 2,t2 ]×……×[C d,0 ,C d,td ]'s spatial level l satisfies

[0013] After defining the calculation order, the following data structure is used to store relevant information for each space S:

[0014] (num,bestcut,Opt,level)

[0015] Where num represents the number of points contained in the subspace, and is a sign of whether the subspace S satisfies the k-anonymity requirement; bestcut refers to the point that satisfies Opt(S) = Opt(S<bestcut)+Opt(S> Opt is the information loss caused by the optimal partition of subspace S; level is the level of subspace S.

[0016] Furthermore, when a strict partitioning technique is used to perform a two-partition operation on the overall data, a table T with d quasi-identifiers and k values are input, and all tuples are mapped to the space Ω(T,ψ); a k-anonymized data table is output.

[0017] Furthermore, in step S2, input: a table T with d quasi-identifiers, a privacy and security protection mechanism P, an information loss metric χ, a space Ω(T,Ψ), and output: a data table obtained after anonymization operation.

[0018] Furthermore, in step S3, the data is divided into a mixed space and encrypted, with the input being: data table T, k value, information loss quality standard χ, space Ω(T,Ψ), and the output being: k-anonymized published data table.

[0019] Further, a Discernibility Measure (DM), a Classification Measure (CM), and a Normalized Certainty Penalty (NCP) are used to determine the usability of the encrypted data.

[0020] Further, for a given data table T, assuming the final released data after processing is T’(P1, P2, …, P m ), where P1, P2, …, P m are m anonymous groups, according to the DM metric function, the data table is penalized with the following value:

[0021]

[0022] Further, for the Classification Measure (CM), the first step of the classification index is to divide the entire data into two groups using a specific criterion. For a given data table T, assuming the final released data after processing is T’(P1, P2, …, P m ), where P1, P2, …, P m are m anonymous groups, and each record in P i , 1 ≤ i ≤ m, also belongs to two sets respectively. The one with the larger number is denoted as majority(P i ), and the one with the smaller number is denoted as minority(P i ). For the CM metric function, then this table will be penalized with the following value:

[0023]

[0024] Further, for the given data table T, assuming the final released data after anonymization is T’(t1, t2, …, t n ), where t1, t2, …, t n are n records. Assume that T’ has d quasi-identifier (QI) attributes. If the j-th attribute A j is a numeric attribute, and assume the value of the j-th attribute of t i is [y ij , z ij , then t i is penalized on attribute A j as follows:

[0025]

[0026] where |A j | represents the size of the attribute domain of attribute A j ;

[0027] If the j-th attribute A j is a non-numeric attribute, assume ti The value of the j-th attribute is X ij , then t i receives the following penalty on attribute A j :

[0028]

[0029] where |A j | represents the size of the attribute domain of attribute A j , and |X ij | represents the number of leaf nodes included by X ij in the semantic tree. Thus, the penalty received by the i-th record t i is as follows:

[0030]

[0031] And the penalty received by the entire table after publication is as follows:

[0032]

[0033] A hybrid spatial sub-division encryption device, comprising: at least one memory and at least one processor;

[0034] The at least one memory is used for storing machine-readable programs;

[0035] The at least one processor is used for calling the machine-readable programs to execute a hybrid spatial sub-division encryption method.

[0036] Compared with the prior art, a website verification code, a method and a device for visual interpretation of pictures according to the present invention have the following outstanding beneficial effects:

[0037] By combining a strict spatial sub-division algorithm with a non-strict spatial sub-division algorithm, the processed anonymous data has higher availability. The present invention is based on the k-anonymity model and the hybrid division technology, and organically combines the strict spatial sub-division algorithm with the non-strict spatial sub-division algorithm to form an encryption algorithm based on hybrid spatial sub-division. This algorithm is expected to provide the company with stronger data protection capabilities, ensuring the security, privacy and availability of the company's data.

[0038] The present invention can increase customers' trust in the company's products and data, and ensure the company's legal compliance in terms of law and ethics, which helps to improve the company's reputation and brand image, and obtain customer stickiness and loyalty. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.

[0040] Appendix Figure 1 is a schematic flow chart of a hybrid spatial sub - division encryption method. Specific implementation manners

[0041] To enable those skilled in the art of this technology to better understand the solution of the present invention, the following will further elaborate on the present invention in combination with specific implementation manners. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0042] The following gives an optimal embodiment:

[0043] As Figure 1 shown, a hybrid spatial sub - division encryption method in this embodiment has the following steps:

[0044] S1. Perform a binary partition operation on the overall data. When performing the binary partition on the data, use the partial optimal greedy algorithm to ensure that the two sub - spaces generated after each partition always maintain the data information loss degree.

[0045] S2. Sequentially perform strict sub - space partitions on their respective sub - spaces until the sub - spaces of the last layer can no longer meet the requirements of strict sub - space partition, and then stop the strict partition operation.

[0046] S3. At this point, several sub - anonymized data groups after strict sub - space partition are generated. Then, check them sequentially. If there are any that meet the requirements of non - strict sub - space partition, perform the non - strict sub - space partition operation again until there are no more sub - anonymized data groups that meet the non - strict sub - space partition.

[0047] Define the level of the space [C 1,x1 , C 1,y1 ×[C 2,x2 , C 2,y2 ×……×[C d,xd , C d,yd as

[0048] According to the definition, it can be known that such as [C 1,0 , C 1,t1 ×[C 2,0 , C2,t2 × …… × [C d,0 , C d,td The spatial hierarchy l of ] satisfies

[0049] After defining the calculation order, for each space S, the following data structure is used to store relevant information:

[0050] (num, bestcut, Opt, level).

[0051] Among them, num represents the number of points contained in the subspace, which is a flag indicating whether the subspace S meets the k-anonymity requirement; bestcut refers to the dividing line that satisfies Opt(S) = Opt(S<bestcut)+Opt(S>bestcut); Opt is the information loss caused by the optimal partition of the subspace S; level is the level of the subspace S.

[0052] For the overall data, a strict partitioning technique is used for the bipartition operation:

[0053] Input: A table T with d quasi-identifiers, a k value, mapping all tuples to the space Ω(T, ψ).

[0054] Output: A k-anonymized data table.

[0055] l ← d - 1;

[0056]

[0057] l ← l + 1;

[0058] For all Space S such that S Ω(T, ψ) ∧ S.level = l, Calculate S.number, S.bestcut ← none, Opt(S) ← X(S);

[0059] if S.number ≥ 2k then;

[0060] For S1, S2 that satisfy the relaxed disjoint condition;

[0061] / * Both S1 and S2 satisfy the relaxed disjoint condition * /

[0062] if S1 Y S2 = S ∧ S1.number ≥ k ∧ S2.number ≥ k then;

[0063] if Opt(S1)+Opt(S2)<Opt(S) then;

[0064] / * Update the optimal value and record the optimal solution * /

[0065] Opt(S) ← Opt(S1)+Opt(S2);

[0066] S.bestcut ← special point ρ of S1 and S2;

[0067] end if;

[0068] end if;

[0069] end for;

[0070] end if;

[0071] end for;

[0072] end while;

[0073] Build k - anonymization recursively through bestcut;

[0074] / * Build the k - anonymity table using bestcut. * /

[0075] Perform a non - strict partition on each space that cannot be further subdivided:

[0076] Input: Table T with d quasi - identifiers, privacy protection mechanism P, information loss metric standard χ, space Ω(T, Ψ);

[0077] Output: The data table obtained after anonymization.

[0078] l ← d - 1;

[0079]

[0080] l ← l + 1;

[0081] for all Space S: [C1,x1,C1,y1]*[C2,x2,C2,y2]*……*[Cd,xd,Cd,yd] s.t. S Ω(T, Ψ) ∧ S.level = l;

[0082] if S satisfies P then;

[0083] S.best.cut ← none, Opt(S) ← χ(S);

[0084] for l ≤ i ≤ d do;

[0085] for x i ≤ j ≤ y i do;

[0086] if S < c(i,j) satisfies P and S > c(i,j) satisfies P then;

[0087] / * Update the optimal value and record the optimal solution * /

[0088] Opt(S) ← Opt(S < c(i,j)) + Opt(S > c(i,j));

[0089] S.bestcut ← c(i,j);

[0090] end if;

[0091] end if;

[0092] end for;

[0093] end for;

[0094] end if;

[0095] end for;

[0096] end while;

[0097] Build k - anonymization recursively through bestcut;

[0098] / * Build the k - anonymized table using bestcut. * /

[0099] Perform mixed - space partitioning and encryption on the data:

[0100] Input: data table T, k value, information loss quality standard χ, space Ω(T, Ψ);

[0101] Output: the published data table after k - anonymization.

[0102] Temp ← Ω(T, Ψ);

[0103] if Temp can be strictly partitioned into two subspaces S1, S2 such that S1 ≥ k ∧ S2 ≥ k then;

[0104] / *Strictly partition the subspace* /

[0105] Find a partition to minimize|χ(S1)-χ(S2)|;

[0106] T1←all elements in S1;

[0107] T2←all elements in S2;

[0108] return Hybird Recoding(T1)YHybird Recoding(T2);

[0109] / *Find the strict partition solution for each subspace* /

[0110] end if;

[0111] else if|Temp|≥2k then;

[0112] Partition Temp non-strictly according to some criterion;

[0113] end else。

[0114] Use a metric function to judge the usability of the data after being encrypted.

[0115] (1) Discernibility Metric (DM);

[0116] For a given data table T, assume the final released data after processing is T’(P1, P2, ……, P m ), where P1, P2, ……, P m are its m anonymous groups. According to the DM metric function, the data table is penalized with the following value:

[0117]

[0118] (2) Classification Metric (CM);

[0119] The first step of the classification index is to use a specific criterion to partition the entire data into two groups. For a given data table T, assume the final released data after processing is T’(P1, P2, ……, P m ), where P1, P2, ……, P m are its m anonymous groups, and the records in each P i (1≤i≤m) also belong to two sets respectively. The one with the larger number is denoted as majority(Pi ) and the smaller quantity is denoted as minority(P i ). For the CM metric function, then this table will be penalized by the following values [7] :

[0120]

[0121] (3) Normalized Certainty Penalty (NCP);

[0122] For the given data table T, assume that the final released data after anonymization is T’(t1,t2,……,t n ), where t1,t2,……,t n are its n records. Assume that T’ has d QI attributes. If the j-th attribute A j is a numeric attribute, assume that the value of the j-th attribute of t i is [y ij ,z ij , then t i is penalized on the attribute A j as follows:

[0123]

[0124] where |A j | represents the size of the attribute domain of the attribute A j . If the j-th attribute A j is a non-numeric attribute, assume that the value of the j-th attribute of t i is X ij , then t i is penalized on the attribute A j as follows:

[0125]

[0126] where |A j | represents the size of the attribute domain of the attribute A j , and |X ij | represents the number of leaf nodes included in X ij in the semantic tree.

[0127] Thus, the penalty received by the i-th record t i is as follows:

[0128]

[0129] And the penalty received by the entire table after release is as follows:

[0130]

[0131] Based on the above method, a hybrid space sub-division encryption device in this embodiment includes: at least one memory and at least one processor;

[0132] The at least one memory is used for storing machine-readable programs;

[0133] The at least one processor is used for calling the machine-readable program and executing a hybrid space sub-division encryption method.

[0134] The above specific embodiments are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific embodiments. Any appropriate changes or substitutions made by those of ordinary skill in the art that comply with the claims of a method and device for website verification codes and visual interpretation of pictures of the present invention shall fall within the patent protection scope of the present invention.

[0135] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A hybrid spatial sub-division encryption method, characterized in that, The steps are as follows: S1. Perform a binary partitioning operation on the overall data. When performing the binary partitioning on the data, use the partial optimal greedy algorithm to ensure that the two subspaces generated after each partition always maintain the data information loss degree; Define the space [C 1,x1, C 1,y1 × [C 2,x2 , C 2,y2 × …… × [C d,xd , C d,yd has a level of ; According to the definition, such as [C 1,0 ,C 1,t1 ×[C 2,0 ,C 2,t2 ×……×[C d,0 ,C d,td has a spatial hierarchy l that satisfies d ≤ l ≤ ; After defining the calculation order, use the following data structure to store relevant information for each space S: (num, bestcut, Opt, level) Among them, num represents the number of points contained in the subspace and is a flag indicating whether the subspace S meets the k-anonymity requirement; bestcut refers to the dividing line that satisfies Opt(S)=Opt(S<bestcut)+Opt(S>bestcut); Opt is the information loss caused by the optimal partition of the subspace S; level is the level of the subspace S; When performing the binary partitioning operation on the overall data using the strict partitioning technique, input the table T with d quasi-identifiers and the k value, and map all tuples to the space Ω(T,ψ); output the k-anonymized data table; S2. Sequentially perform strict subspace partitioning on each respective subspace until the subspaces in the last layer can no longer meet the requirements of strict subspace partitioning, and then stop the strict partitioning operation; Input: Table T with d quasi-identifiers, privacy and security protection mechanism P, information loss measurement criteria , space , Output: The data table obtained after anonymization processing S3. At this point, several sub-anonymous data groups after strict subspace partitioning are generated. Then, check them sequentially. If there are any that meet the requirements of non-strict subspace partitioning, perform the non-strict subspace partitioning operation again until there are no more sub-anonymous group data groups that meet the non-strict subspace partitioning; Perform mixed spatial partitioning and encryption processing on the data. Input: data table T, k value, information loss quality standard , space , Output: the published data table after k-anonymization.

2. The hybrid spatial sub-division encryption method according to claim 1, wherein Use the discriminability metric DM, the classification metric CM, and the normalized certainty penalty NCP to judge the usability of the data after encryption processing.

3. A hybrid spatial sub-division encryption method according to claim 2, characterized in that , for a given data table T, assume that the final released data after processing is T’(P1, P2, ……, P m ), where P1, P2, ……, P m are m anonymous groups. According to the DM metric function, the data table will receive a penalty with the following value: 。 4. A hybrid spatial sub-division encryption method according to claim 3, characterized in that , the classification metric CM, the first step of the classification index is to use specific criteria to divide the entire data into two groups. For a given data table T, assume that the final published data after processing is T’ (P1, P2, ……, P m ), where P1, P2, ……, P m are m anonymous groups, and each record in P i , , also belongs to two sets respectively. The one with a larger quantity is denoted as majority(P i ), and the one with a smaller quantity is denoted as minority(P i ). For the CM metric function, then this table will be penalized by the following values: 。 5. A hybrid spatial sub-division encryption method according to claim 4, characterized in that The standardized deterministic penalty NCP for the given data table T, assuming that the final published data after anonymization is T’(t1, t2, ……, t n ), where t1, t2, ……, t n are n records. Assume that T’ has d QI attributes. If the j-th attribute A j is a numerical attribute, assume that the value of the j-th attribute of t i is [y ij , z ij , then t i suffers the following penalty on attribute A j : ; Among them, |A j | represents the size of the attribute domain of attribute A j ; If the j-th attribute A j is a non-numeric attribute, assuming that i the value of the j-th attribute of t ij is X i , then t j receives the following penalty on attribute A ; Among them, |A j | represents the size of the attribute domain of attribute A j , and |X ij | represents the number of leaf nodes included in X ij in the semantic tree. Thus, the penalty received by the i-th record t i is as follows: ; And the penalty received by the entire table after publication is as follows: 。 6. A hybrid spatial sub-division encryption device, characterized in that, Including: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program and execute the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Secondary k-anonymity privacy protection algorithm for differentiating quasi-identifier attributes

    CN106021541A

  • K-anonymous privacy protection method adopting density-based partition

    CN107292195A