Personalized Protection Method Based on Generalization of Multi-Sensitive Associated Data

By building a generalized hierarchical tree and bucketing operations, combining Aprior algorithm and pseudo-data, the privacy leakage problem in the associated data structure is solved, and efficient personalized privacy protection is achieved.

CN114611142BActive Publication Date: 2025-07-11WUHAN HAOZE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210235772.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-10
Publication Date
2025-07-11
Estimated Expiration
2042-03-10

AI Technical Summary

Technical Problem

In the associated data structure, the prior art is prone to the problem of privacy leakage.

Method used

By building a generalization hierarchy tree, using the Aprior algorithm to establish association rules, combining bucketing and pseudo-data to increase diversity, perform association segmentation and lossless connection of sensitive data, and protect the association relationship of sensitive attributes.

Benefits of technology

It effectively reduces the probability of user privacy being disclosed, improves data security, and meets personalized privacy protection needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114611142B_ABST
    Figure CN114611142B_ABST
Patent Text Reader

Abstract

The present invention discloses a method belonging to the field of data technology, specifically a personalized protection method based on generalization of multi-sensitive associated data, including the following specific steps: S1, according to the sensitive data set constructed in the knowledge graph, extract a sequence A of sensitive feature attributes of a certain column of users; S2, construct a generalization hierarchy tree based on the sensitivity of the data in the attribute sequence A; S3, at the same time, users can, according to their own personalized needs, eliminate the information loss caused by over-generalization by setting the sensitivity level through introducing the sensitive value weight for sensitive nodes in the tree, establish the association rules between attributes by using the Aprior algorithm, and measure the association degree of sensitive nodes by using the amount of information, so as to protect the association relationship of sensitive attributes. The present invention has the advantages of eliminating the association relationship of user sensitive features, reducing the probability of privacy disclosure, and improving security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data technology, and specifically to a personalized protection method based on generalization of multi-sensitive associated data. Background Art

[0002] A data structure is a way for a computer to store and organize data. A data structure refers to a collection of data elements that have one or more specific relationships with each other. There are generally two interpretations of an associated data structure: 1. Representing associated data with a data structure, such as an associative array; 2. Connecting relevant data structures through a method.

[0003] However, when associating data, the problem of privacy leakage is likely to occur.

[0004] Therefore, we propose a personalized protection method based on generalization of multi-sensitive associated data. Summary of the Invention

[0005] In view of the above and / or problems existing in the existing personalized protection method based on generalization of multi-sensitive associated data, the present invention is proposed.

[0006] Therefore, the object of the present invention is to provide a personalized protection method based on generalization of multi-sensitive associated data, which can solve the above-mentioned existing problems.

[0007] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solution:

[0008] A personalized protection method based on generalization of multi-sensitive associated data, which includes the following specific steps:

[0009] S1. According to the sensitive data set constructed in the knowledge graph, extract a certain column of sensitive feature attribute sequence A of the user;

[0010] S2. Based on the sensitivity of the data in the attribute sequence A, construct a generalization hierarchy tree;

[0011] S3. At the same time, the user can set the sensitivity level by introducing the sensitive value weight for the sensitive nodes in the tree according to their own personalized needs to eliminate the information loss caused by over-generalization, establish the association rules between attributes by using the Aprior algorithm, and measure the association degree of the sensitive nodes by using the amount of information, so as to protect the association relationship of the sensitive attributes;

[0012] S4. Take the composite sensitive data composed of the user's multi-dimensional sensitive data as a high-dimensional vector, and map the values of each dimension of the composite sensitive data vector to different buckets according to the random decision method for the records in the relationship table. Through bucketing, the sensitive data of the user is associated and segmented. Secondly, perform a grouping operation on the sensitive data table of each bucket, add pseudo-data to make each group meet L-diversity;

[0013] S5. Finally, perform a lossless join through the group ID of the sensitive data table to complete the one-to-many relationship between the user and the sensitive data, and reduce the probability of a certain user's privacy being disclosed to 1 / L.

[0014] As a preferred solution of the personalized protection method based on multi-sensitive associated data generalization according to the present invention, wherein: in S1, a certain column of sensitive feature attribute sequence A of the user is extracted by using a crawler method.

[0015] As a preferred solution of the personalized protection method based on multi-sensitive associated data generalization according to the present invention, wherein: in S2, the sensitivity of the data in the attribute sequence A is set by the user himself.

[0016] As a preferred solution of the personalized protection method based on multi-sensitive associated data generalization according to the present invention, wherein: in S2, the generalization hierarchy tree can be defined as a quadruple ITAX(A) = {r A ,L A ,IN A ,R A};

[0017] Among them, r A represents the root node of the hierarchy tree, L A represents the set of leaf nodes of the hierarchy tree, IN A represents the set of intermediate nodes of the hierarchy tree. The elements in the set represent the generalization nodes of various values of the sensitive attribute A. The higher the generalization level where the node is located, the lower the sensitivity, the larger the boundary of the sensitive data, and the more difficult it is for the attacker to attack. R A represents the association relationship between the nodes in the hierarchy tree.

[0018] As a preferred solution of the personalized protection method based on multi-sensitive associated data generalization according to the present invention, wherein: in S2, the insensitive ones in the generalization hierarchy tree are used as the root nodes, and the sensitive ones are used as the leaf nodes.

[0019] As a preferred solution of the personalized protection method based on multi-sensitive associated data generalization according to the present invention, wherein: in S4, the calculation formula for bucketing is:

[0020]

[0021] Among them, represents the sum of the elements representing rows and columns.

[0022] Compared with the prior art:

[0023] 1. According to the sensitive data set constructed in the knowledge graph, extract a certain column of sensitive feature attribute sequence A of the user. Based on the sensitivity of the data in the attribute sequence A, construct a generalization hierarchy tree. At the same time, the user can, according to their own personalized needs, set the sensitivity level by introducing sensitive value weights for sensitive nodes in the tree to eliminate the information loss caused by over-generalization. Use the Aprior algorithm to establish the association rules between attributes, and use the amount of information to measure the association degree of sensitive nodes, so as to protect the association relationship of sensitive attributes, which has the effect of hiding information;

[0024] 2. Take the composite sensitive data composed of the user's multi-dimensional sensitive data as a high-dimensional vector. Map the values of each dimension of the composite sensitive data vector to different buckets according to the random decision method for the records in the relationship table. Through bucketing, the sensitive data of the user is associated and segmented. Secondly, perform a grouping operation on the sensitive data table of each bucket, add pseudo-data, so that each group satisfies L-diversity. Finally, perform a lossless join through the group ID of the sensitive data table to complete the one-to-many relationship between the user and the sensitive data, reducing the probability of a certain user's privacy being disclosed to 1 / L, which has the effect of eliminating the association relationship of the user's sensitive features, reducing the probability of privacy being disclosed, and improving security. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is the sensitive data generalization tree diagram of the present invention;

[0026] Figure 2 It is the concealment process diagram of the sensitive nodes of the present invention;

[0027] Figure 3 It is the multi-sensitive attribute bucketing and dimensionality reduction privacy protection diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0028] To make the objectives, technical solutions and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0029] The present invention provides a personalized protection method based on the generalization of multi-sensitive associated data. Please refer to Figures 1 - 3 , and the specific steps are as follows:

[0030] S1. According to the sensitive data set (user privacy data set, interest preference) constructed in the knowledge graph, extract a certain column of sensitive feature attribute sequence A of the user (using the crawler method);

[0031] S2. Build a generalization hierarchy tree based on the sensitivity of the data in the attribute sequence A (set by the user himself / herself) (the insensitive ones are used as the root nodes, and the sensitive ones are used as the leaf nodes).

[0032] As Figure 1 shown: The generalization hierarchy tree can be defined as a quadruple ITAX(A) = {r A , L A , IN A , R A};

[0033] Among them, r A represents the root node of the hierarchy tree, L A represents the set of leaf nodes of the hierarchy tree, IN A represents the set of intermediate nodes of the hierarchy tree. The elements in the set represent the generalization nodes of various values of the sensitive attribute A. The higher the generalization level of the node (the higher the level in the tree, set by the user himself / herself), the lower the sensitivity, the larger the boundary of the sensitive data, and the more difficult it is for the attacker to attack. R A represents the association relationship between the nodes in the hierarchy tree;

[0034] S3. At the same time, the user can, according to his / her own personalized needs, set the sensitivity level for the sensitive nodes in the tree by introducing the sensitive value weight (set by the user himself / herself. For example, if I want to protect the hotel where I stayed during my travel as privacy and don't want others to know what hotel I stayed in, I will set the privacy level of the hotel to high) to eliminate the information loss caused by over-generalization, use the Aprior algorithm to establish the association rules between attributes, and use the information quantity to measure the association degree of the sensitive nodes (which can be calculated using the entropy calculation formula), so as to protect the association relationship of the sensitive attributes (as Figure 2 shown);

[0035] As Figure 3 shown: Suppose the sensitive node G in the second layer is the highest-risk node, then the user can set the sensitivity of this node to the highest weight and hide all the nodes of this subtree, so as to achieve personalized protection on the basis of generalization;

[0036] S4. The composite sensitive data composed of the user's multi-dimensional (assuming the location is one-dimensional, with the location of home, hotel, and tourist location in this dimension, and the preference is one-dimensional, with shopping preference, diet preference, and tourist preference) sensitive data is used as a high-dimensional vector. The records in the relationship table are mapped to different buckets respectively according to the random decision method for each dimension value of the composite sensitive data vector (randomly assigned and distributed in row-column format). The sensitive data of the user is associated and segmented through bucketing. Secondly, grouping operations are performed on the sensitive data table of each bucket (since it is a vector and there are many buckets, we take the data of one bucket as a group), and pseudo-data is added (randomly filled data for interference) to make each group satisfy L-diversity (as shown in Figure 3 shown);

[0037] The calculation formula for bucketing is:

[0038]

[0039] where, represents the sum of the elements representing rows and columns;

[0040] S5. Finally, a lossless join is performed through the group ID of the sensitive data table to complete the one-to-many relationship between the user and the sensitive data, reducing the probability of a certain user's privacy being disclosed to 1 / L.

[0041] Although the present invention has been described above with reference to the embodiments, various improvements can be made to it and components therein can be replaced with equivalents without departing from the scope of the present invention. In particular, as long as there is no structural conflict, the various features in the disclosed embodiments of the present invention can be combined with each other in any way. The exhaustive description of these combinations is not given in this specification only for the sake of saving space and resources. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.

Claims

1. A personalized protection method based on generalization of multi-sensitive associated data, characterized in that, The specific steps are as follows: S1. Extract a sequence A of sensitive feature attributes of a certain column of the user according to the sensitive data set constructed in the knowledge graph; S2. Construct a generalization hierarchy tree based on the sensitivity of the data in the attribute sequence A; S3. At the same time, the user can, according to their own personalized needs, set the sensitivity level for the sensitive nodes in the tree by introducing sensitive value weights to eliminate the information loss caused by over-generalization, establish association rules between attributes using the Aprior algorithm, and use the amount of information to measure the association degree of sensitive nodes, so as to protect the association relationship of sensitive attributes; S4. Take the composite sensitive data composed of the user's multi-dimensional sensitive data as a high-dimensional vector, map the values of each dimension of the composite sensitive data vector to different buckets according to the random decision method for the records in the relationship table, and perform association segmentation on the user's sensitive data through bucketing. Secondly, perform grouping operations on the sensitive data table in each bucket, add pseudo-data, so that each group satisfies L-diversity; S5. Finally, perform a lossless join through the group ID of the sensitive data table to complete the one-to-many relationship between the user and the sensitive data, and reduce the probability of a certain user's privacy being disclosed to 1 / L.

2. The personalized protection method based on multi-sensitive associated data generalization according to claim 1, characterized in that In S1, a web crawler method is used to extract a sequence A of sensitive feature attributes of a certain column of the user.

3. The personalized protection method based on multi-sensitive associated data generalization according to claim 1, characterized in that In S2, the sensitivity based on the data in the attribute sequence A is set by the user himself.

4. The personalized protection method based on generalization of multi-sensitive associated data according to claim 1, characterized in that, In S2, the generalization hierarchy tree can be defined as a quadruple ITAX(A) = {r A , L A , IN A , R A}; Among them, r A represents the root node of the hierarchical tree, L A represents the set of leaf nodes of the hierarchical tree, IN A represents the set of intermediate nodes of the hierarchical tree. The elements in the set represent the generalization nodes of various values of the sensitive attribute A. The higher the generalization level of the node, the lower the sensitivity, the larger the boundary of the sensitive data, and the more difficult it is for the attacker to attack. R A represents the association relationship between the nodes in the hierarchical tree.

5. The personalized protection method based on multi-sensitive associated data generalization according to claim 1, characterized in that, In S2, the insensitive ones in the generalization hierarchy tree are used as the root nodes, and the sensitive ones are used as the leaf nodes.

6. The personalized protection method based on generalization of multi-sensitive associated data according to claim 1, characterized in that In S4, the calculation formula for bucketing is: Among them, represents the sum of the elements representing rows and columns.

Citation Information

Patent Citations

  • Generalization method for weighted social network

    CN104317904A

  • Data recorder

    CN106384404A