Data desensitization method and device, electronic equipment and storage medium
By generating data access strategies and risk prediction models, and combining random encryption and anonymization, the problem of de-identified datasets being easily cracked was solved, enabling secure data flow and efficient utilization across departments.
Patent Information
- Application Number
- CN202511162704.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-18
AI Technical Summary
In existing technologies, de-identified datasets are easily cracked, leading to a sharp contradiction between customer privacy protection and efficient data utilization. This is especially true in highly sensitive scenarios such as cross-border payments and supply chain finance, where data security is insufficient across departments and institutions.
Data access policies are generated by obtaining the attribute information of users with access rights, the dataset is encrypted using a randomly generated master key, and the risk probability of a field being illegally decrypted is determined by a pre-built risk prediction model. Based on the risk probability threshold, fields to be anonymized are selected and anonymized to generate a desensitized dataset.
It improves the security of de-identified datasets, prevents them from being easily cracked, and meets the needs for secure data flow across departments and organizations.
Smart Images

Figure CN120979737A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence and financial technology, and in particular to a data anonymization method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the diversification of banking business scenarios and the normalization of cross-border data flow, the contradiction between customer privacy protection and efficient data utilization has become increasingly acute. This contradiction is particularly prominent in highly sensitive scenarios such as cross-border payments and supply chain finance. Against this backdrop, how to achieve the secure flow of data across departments and institutions while ensuring compliance has become a core issue that urgently needs to be addressed in the fintech field.
[0003] In existing technologies, to protect data security, encrypted datasets are typically anonymized directly to obtain de-identified datasets. However, due to the increasing complexity of privacy breach attacks, attackers can easily bypass the anonymization protection of de-identified datasets by linking multiple data sources, such as publicly available social network information with some de-identified transaction records. Summary of the Invention
[0004] This invention provides a data anonymization method, apparatus, electronic device, and storage medium, which improves the security of anonymized datasets.
[0005] In a first aspect, embodiments of the present invention provide a data anonymization method, comprising:
[0006] Obtain each authorized user corresponding to the dataset to be encrypted, and generate a data access policy corresponding to the dataset to be encrypted based on the user attribute information of each authorized user;
[0007] Based on the data access policy and the randomly generated master key, the dataset to be encrypted is encrypted to obtain the encrypted dataset corresponding to the dataset to be encrypted.
[0008] By using a pre-built risk prediction model, the probability of each field in the encrypted dataset being illegally decrypted is determined;
[0009] Based on the comparison results between each risk probability and the preset risk probability threshold, the fields to be anonymized are determined in the encrypted dataset;
[0010] Anonymize the fields to be anonymized in the encrypted dataset to obtain the de-anonymized dataset.
[0011] Secondly, embodiments of the present invention also provide a data desensitization device, comprising:
[0012] The access policy generation module is used to obtain each authorized user corresponding to the dataset to be encrypted, and generate a data access policy corresponding to the dataset to be encrypted based on the user attribute information of each authorized user.
[0013] The first encryption module is used to encrypt the dataset to be encrypted according to the data access policy and the randomly generated master key, so as to obtain the encrypted dataset corresponding to the dataset to be encrypted.
[0014] The risk probability determination module is used to determine the probability of each field in the encrypted dataset being illegally decrypted through a pre-built risk prediction model;
[0015] The field to be anonymized module is used to determine the field to be anonymized in the encrypted dataset based on the comparison result between each risk probability and a preset risk probability threshold.
[0016] The second encryption module is used to anonymize the fields to be anonymized in the encrypted dataset, resulting in a de-anonymized dataset.
[0017] Thirdly, embodiments of the present invention also provide an electronic device, the electronic device comprising:
[0018] At least one processor; and
[0019] A memory that is communicatively connected to at least one processor; wherein,
[0020] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the data desensitization method provided in any embodiment of the present invention.
[0021] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions that are used to cause a processor to execute and implement the data desensitization method provided in any embodiment of the present invention.
[0022] The technical solution of this invention encrypts the dataset to be encrypted using a data access strategy and a randomly generated master key to obtain an encrypted dataset corresponding to the dataset to be encrypted. Then, a pre-built risk prediction model is used to determine the risk probability of each field in the encrypted dataset being illegally decrypted. Based on each risk probability, fields to be anonymized are determined in the encrypted dataset, and these fields are anonymized to obtain a de-identified dataset. This avoids the situation in existing technologies where the de-identified dataset is often easily cracked due to direct anonymization of the encrypted dataset, thus improving the security of the de-identified dataset.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a flowchart of a data desensitization method provided in Embodiment 1 of the present invention;
[0026] Figure 2 This is a flowchart of another data desensitization method provided in Embodiment 2 of the present invention;
[0027] Figure 3 This is a schematic diagram of a data desensitization device according to Embodiment 3 of the present invention;
[0028] Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. Detailed Implementation
[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0031] Example 1
[0032] Figure 1 This is a flowchart of a data desensitization method according to Embodiment 1 of the present invention. This embodiment is applicable to situations where data desensitization processing is performed. The method can be executed by a data desensitization device, which can be implemented in hardware and / or software and can be configured in an electronic device such as a computer.
[0033] like Figure 1 As shown, this embodiment discloses a data anonymization method, including:
[0034] S110. Obtain each authorized user corresponding to the dataset to be encrypted, and generate a data access policy corresponding to the dataset to be encrypted based on the user attribute information of each authorized user.
[0035] In this embodiment, the dataset to be encrypted can be understood as a set of fields that need to be encrypted, such as customer information datasets, transaction log datasets, and credit business datasets. Authorized users can be understood as users who are authorized to access the dataset to be encrypted. User attribute information may include the user's department, position, and job level. Data access policies can be used to reflect the conditions that users must meet to access the dataset to be encrypted.
[0036] In this step, specifically, target attribute information can be obtained from the user attribute information of each authorized user, and a data access policy corresponding to the dataset to be encrypted can be generated based on the target attribute information of each authorized user.
[0037] For example, suppose that the users authorized to access the dataset to be encrypted include a first user and a second user, and that the user attribute information of both the first user and the second user includes the user's department, position, and job level. In this case, the user's department and position can be used as the target attribute information. Thus, if the first user and the second user both belong to the business department, and the first user and the second user's positions are business manager and salesperson, respectively, the following data access policy can be generated for the dataset to be encrypted: (Department = Business Department) AND (Position = Business Manager OR Salesperson).
[0038] S120. Based on the data access policy and the randomly generated master key, encrypt the dataset to be encrypted to obtain the encrypted dataset corresponding to the dataset to be encrypted.
[0039] In this embodiment, the encrypted dataset may include ciphertext obtained by encrypting the dataset to be encrypted using a master key, and ciphertext components corresponding to each user attribute in the data access policy.
[0040] In this step, specifically, attribute-based encryption technology can be used to encrypt the dataset to be encrypted according to the data access policy and a randomly generated master key, thereby obtaining an encrypted dataset corresponding to the dataset to be encrypted.
[0041] S130. Determine the probability of each field in the encrypted dataset being illegally decrypted using a pre-built risk prediction model.
[0042] Specifically, in this step, an adjacency matrix corresponding to the encrypted dataset is generated based on the correlation between fields in the encrypted dataset, and a field feature matrix corresponding to the encrypted dataset is generated based on the field features of each field in the encrypted dataset. Then, a pre-built risk prediction model can be used to determine the probability of each field in the encrypted dataset being illegally decrypted based on the adjacency matrix and the field feature matrix. The risk prediction model can be a graph sampling and aggregation (GraphSAGE) model.
[0043] S140. Based on the comparison results between each risk probability and the preset risk probability threshold, determine the field to be anonymized in the encrypted dataset.
[0044] In this embodiment, the field to be anonymized can be understood as the field in the encrypted dataset that needs to be anonymized.
[0045] In this step, specifically, fields in the encrypted dataset with a risk probability greater than a preset risk probability threshold can be identified as fields to be anonymized. The preset risk probability threshold can be set based on historical experience; for example, it can be set to 0.7.
[0046] S150. Anonymize the fields to be anonymized in the encrypted dataset to obtain the de-anonymized dataset.
[0047] In this step, specifically, the fields to be anonymized in the encrypted dataset can be anonymized by adding Laplace noise to the fields to be anonymized, replacing the fields to be anonymized with other fields, or performing generalization processing on the fields to be anonymized, thus obtaining a de-anonymized dataset.
[0048] Furthermore, when anonymizing fields in an encrypted dataset by adding Laplace noise to the fields to be anonymized, resulting in a de-anonymized dataset, the Laplace noise can be added to the fields to be anonymized using the following specific calculation formula:
[0049]
[0050] Among them, f v ' represents the field to be anonymized after adding Laplace noise, f vFor the field to be anonymized, Δf = max(f v )-min(f v ), max(f v ) represents the maximum value of the feature field v, min(f) v ) represents the minimum value of the feature field v, and ∈ represents a predefined privacy budget value. It is worth noting that the constraints of the above calculation formula need to ensure that the anonymized data satisfies k-anonymity. The k in the above k-anonymity can be determined according to user needs and historical experience. For example, k in k-anonymity can be set to any positive integer greater than or equal to 5.
[0051] Optionally, before encrypting the dataset to be encrypted according to the data access policy and the randomly generated master key to obtain the encrypted dataset corresponding to the dataset to be encrypted, the method further includes: selecting random integers from a predefined prime order group that correspond to each user attribute of the authorized user; generating a private key corresponding to the dataset to be encrypted according to the randomly generated master key and the random integers, so that the authorized user can decrypt the encrypted dataset to be encrypted according to the private key.
[0052] Specifically, we can choose bilinear groups G1 and G2, with order p and generators g ∈ G1, and base them on the formula... An element is randomly selected from a predefined group of prime numbers as the master key. Based on the formula... Randomly select an integer from a predefined prime group as the private key to generate an integer, based on the formula... From a predefined group of prime numbers, select random integers corresponding to each user attribute for each user with access rights. Here, α is the master key, r is the integer generated from the private key, and β... i It is a random integer. The group is a predefined prime group. Then, integers and random integers can be generated using the generators in the bilinear group, the master key, and the private key to generate the private key corresponding to the dataset to be encrypted.
[0053] For example, the private key corresponding to the dataset to be encrypted can be generated using the following specific calculation formula:
[0054]
[0055] Among them, SK S α is the private key, g is a generator in the bilinear group, α is the master key, r is the integer generated by the private key, and β is the master key. i S is a random integer, and S represents the set of user attributes corresponding to the user with access rights.
[0056] The advantage of this setup is that it generates a private key corresponding to the dataset to be encrypted by using random integers corresponding to each user attribute of the authorized user and a randomly generated master key. This facilitates access to the encrypted dataset by authorized users while preventing internal misuse of the data.
[0057] The technical solution of this embodiment obtains the authorized users corresponding to the dataset to be encrypted, and generates a data access policy corresponding to the dataset to be encrypted based on the user attribute information of the authorized users; encrypts the dataset to be encrypted according to the data access policy and a randomly generated master key to obtain an encrypted dataset; determines the risk probability of each field in the encrypted dataset being illegally decrypted through a pre-built risk prediction model; determines the fields to be anonymized in the encrypted dataset based on the comparison results of each risk probability and a preset risk probability threshold; and anonymizes the fields to be anonymized in the encrypted dataset to obtain a desensitized dataset. This technical means solves the problem in the prior art of directly anonymizing the dataset to be encrypted to obtain a desensitized dataset, which often makes the desensitized dataset easy to crack, thus improving the security of the desensitized dataset.
[0058] Example 2
[0059] Figure 2 This is a flowchart of another data desensitization method provided in Embodiment 2 of the present invention. This embodiment is a further optimization and extension based on the above embodiments and can be combined with various optional technical solutions in the above embodiments.
[0060] like Figure 2 As shown, this embodiment discloses a data anonymization method, including:
[0061] S210. Obtain each authorized user corresponding to the dataset to be encrypted, and generate a data access policy corresponding to the dataset to be encrypted based on the user attribute information of each authorized user.
[0062] S220. Obtain the version number of the data access strategy, and obtain the version random number corresponding to the version number of the data access strategy from the predefined prime order group.
[0063] In this step, specifically, it can be done through a formula. Retrieve a random number corresponding to the version number of the data access policy. Where s is the random number for the version. It is a predefined prime group.
[0064] S230. Based on the version random number, the randomly generated master key, and the attribute public key corresponding to each user attribute information in the data access policy, encrypt the dataset to be encrypted to obtain the encrypted dataset corresponding to the dataset to be encrypted.
[0065] Specifically, in this step, ciphertext corresponding to the dataset to be encrypted is generated based on the generator of the bilinear group, the version random number, and the randomly generated master key. Group elements related to the version random number are generated based on the generator of the bilinear group and the version random number. Ciphertext components corresponding to each user attribute in the data access policy are then generated based on the attribute public key corresponding to each user attribute information in the data access policy. Finally, an encrypted dataset corresponding to the dataset to be encrypted is generated based on the ciphertext corresponding to the dataset to be encrypted, the group elements related to the version random number, and the ciphertext components.
[0066] For example, the encrypted dataset corresponding to the dataset to be encrypted can be generated using the following specific calculation formula:
[0067]
[0068] Where C is the encrypted dataset, C0 is the ciphertext corresponding to the dataset to be encrypted, C1 is the group element related to the version random number, {C i} represents the ciphertext component corresponding to each user attribute in the data access policy. m is the dataset to be encrypted, and e(g,g) αs This is a symmetric key combining the master key and a random version number, where g is a pre-selected generator in a bilinear group, α is the master key, and s is the random version number. This represents the attribute public key corresponding to each user attribute information in the data access policy.
[0069] By setting up the above, various business scenarios such as loan approval and anti-money laundering monitoring can be fully integrated to formulate data access policies. These policies can then be used to implement fine-grained access control over data, thereby preventing unauthorized internal access and improving data security.
[0070] S240. Based on the degree of correlation between fields in the encrypted dataset, generate an adjacency matrix corresponding to the encrypted dataset, and based on the field features of each field in the encrypted dataset, generate a field feature matrix corresponding to the encrypted dataset.
[0071] In this step, specifically, the adjacency element values corresponding to each pair of fields in the encrypted dataset are determined by comparing the correlation between the fields in the encrypted dataset with a preset correlation threshold. The preset correlation threshold can be determined based on historical experience; for example, it can be set to 0.5. Adjacency element values can be 0 or 1. Then, an adjacency matrix corresponding to the encrypted dataset is generated based on each adjacency element value, and a field feature matrix corresponding to the encrypted dataset is generated based on the field values of each field in the encrypted dataset. The adjacency matrix can be represented as A∈{0,1}. n×n .
[0072] S250. Take each field in the encrypted dataset as the center field in turn, and determine all the neighbor fields corresponding to each center field according to the adjacency matrix.
[0073] S260. Based on the field feature matrix, determine the central field feature of each central field, and the neighbor field features of all neighbor fields corresponding to each central field.
[0074] S270. Using a pre-built risk prediction model, determine the probability of each field in the encrypted dataset being illegally decrypted based on the central field characteristics of each central field and the neighbor field characteristics of all neighbor fields corresponding to each central field.
[0075] In this step, specifically, a pre-built risk prediction model can be used to determine the aggregate features of each central field based on the central field features of each central field and the neighbor field features of all neighbor fields corresponding to each central field. Based on the aggregate features of each central field, the probability of each central field being illegally decrypted can be determined.
[0076] By using the above settings, we can aggregate the features of the neighbor fields, thereby improving the accuracy of determining the probability of risk.
[0077] Optionally, a pre-built risk prediction model is used to determine the probability of each field in the encrypted dataset being illegally decrypted based on the central field features of each central field and the neighbor field features of all neighbor fields corresponding to each central field. This includes: determining the current aggregation feature of each central field at the current aggregation layer based on the central field features of each central field and the neighbor field features of all neighbor fields corresponding to each central field; determining whether the current aggregation layer is a predefined loop termination layer; if not, after using the current aggregation feature of the central field as the central field feature of the central field and updating the current aggregation layer, the operation of determining the current aggregation feature of each central field at the current aggregation layer based on the pre-built risk prediction model, based on the central field features of each central field and the neighbor field features of all neighbor fields corresponding to each central field, is returned until the current aggregation layer is a predefined loop termination layer. Then, based on the current aggregation feature of each central field, the probability of each field in the encrypted dataset being illegally decrypted is determined.
[0078] Specifically, the risk propagation formula in the risk prediction model can be used to determine the current aggregation feature of each central field at the current aggregation layer, based on the central field features of each central field and the neighbor field features of all neighbor fields corresponding to each central field. After determining the current aggregation feature of each central field at the current aggregation layer, if the current aggregation layer is not a predefined loop termination layer, the above feature aggregation operation is repeated. If the current aggregation layer is a predefined loop termination layer, the risk probability of each field in the encrypted dataset being illegally decrypted is determined using the risk calculation formula in the risk prediction model, based on the current aggregation feature of each central field. The numerical range of the risk probability is [0,1].
[0079] For example, the current aggregation feature of each central field at the current aggregation level can be determined using the following risk propagation formula:
[0080]
[0081] in, W is the embedding vector of field v at layer k, representing the current aggregated feature of field v after propagation through k layers, which integrates the information of field v itself and the structural information of its neighbors. (k)Let be the trainable weight matrix of the k-th layer. Its purpose is to map the concatenated features (information about field v itself and information about its neighbors' structures) to a new embedding space, learning the non-linear pattern of risk propagation. The CONCAT function represents a vector concatenation operation, used to combine the features of field v itself with the aggregated features of its neighbors, preserving independent information. The MEAN function is used to average the embeddings of neighbor fields, thereby smoothing neighbor information, reducing noise, and capturing local correlation patterns. The ReLU function, as an activation function, is mainly used to enhance the model's expressive power and filter out negative correlations. For the old feature embedding of field v at the (k-1)th layer, The old feature embedding of field v's neighbor field u at the (k-1)th layer. It is the set of neighboring fields of field v, which defines the local scope of risk propagation, allowing only highly correlated fields in the encrypted dataset to participate in the calculation.
[0082] When the current aggregation layer is the loop termination layer, the probability of each field in the encrypted dataset being illegally decrypted can be determined using the following risk calculation formula:
[0083]
[0084] Among them, R v For the probability of risk, W out Let be a trainable weight matrix of dimension 1*d, and σ be the activation function. The current aggregated feature of the center field at the loop termination level.
[0085] The advantage of this setup is that by aggregating field features multiple times, we can make full use of the direct and indirect neighbor information of the field to determine the probability of the field being illegally decrypted, thereby improving the accuracy of determining the probability of the field being illegally decrypted.
[0086] S280. Based on the comparison results between each risk probability and the preset risk probability threshold, determine the field to be anonymized in the encrypted dataset.
[0087] S290. Anonymize the fields to be anonymized in the encrypted dataset to obtain the de-anonymized dataset.
[0088] Optionally, after anonymizing the fields to be anonymized in the encrypted dataset to obtain the de-identified dataset, the method further includes: obtaining all operations performed during the de-identification process of the encrypted dataset to obtain the de-identified dataset, as well as the timestamp corresponding to each operation; constructing a de-identification operation log corresponding to the de-identified dataset based on all operations, the timestamp corresponding to each operation, and the user attribute information of the user with access rights, and generating a hash value corresponding to the de-identification operation log so that the user can read the de-identification operation log.
[0089] Specifically, algorithms such as secure hashing or cyclic redundancy check can be used to generate a hash value corresponding to the de-identified operation log. For example, the hash value corresponding to the de-identified operation log can be generated using the following specific calculation formula:
[0090] H = SHA3-256(L)
[0091] Where H represents the hash value, SHA3 represents the secure hash algorithm, and L represents the de-identification operation log.
[0092] The advantage of this setup is that by recording the data anonymization operation logs during the data anonymization process and generating corresponding hash values, the data anonymization process can be made tamper-proof and traceable. This not only meets the regulatory requirements of relevant departments but also facilitates the verification of the data anonymization process by relevant personnel.
[0093] Optionally, after generating the hash value corresponding to the data masking operation log, in response to a user's verification request for the data masking process, the data masking operation log can be retrieved, and the operation and user attribute information recorded in the log can be verified to conform to predefined data masking specifications. Then, if the operation and user attribute information recorded in the log conform to the predefined data masking specifications, access requests from authorized users can continue to be responded to. If the operation and user attribute information recorded in the log do not conform to the predefined data masking specifications, the data access policy should be updated. The new data access policy must override the old data access policy, i.e. After updating the data access policy, the new policy version number corresponding to the new policy and the old policy version number corresponding to the old policy can be obtained. A re-encryption key is then generated based on the master key, the old policy version number, and the new policy version number. Finally, the old encrypted dataset can be re-encrypted using the re-encryption key to obtain a re-encrypted dataset corresponding to the dataset to be encrypted, thus preventing unauthorized access to the dataset during the re-encryption process.
[0094] For example, the re-encryption key can be generated using the following specific calculation formula:
[0095]
[0096] Where rk is the re-encryption key, α is the master key, and s old s represents the old strategy version number. new This is the version number of the new strategy.
[0097] Then, the old encrypted dataset can be re-encrypted using the following specific calculation formula to obtain the re-encrypted dataset:
[0098]
[0099] Among them, C new For re-encrypting datasets, C old Given the old encrypted dataset, g is a pre-selected generator in a bilinear group, rk is the re-encryption key, and S... new For the new strategy version number, s old This is the old strategy version number. T represents the encrypted component corresponding to each user attribute in the new data access policy. new For the new data access strategy.
[0100] The above settings allow for the separate updating of data access policies for specific data segments, and the corresponding data can be re-encrypted according to the new data access policy without re-encrypting the entire dataset. This avoids situations where high latency in policy switching makes it difficult to meet real-time business needs under the high-pressure environment of tens of millions of transactions per day in banks. It can support real-time business scenarios as well as high-frequency transaction scenarios, ensuring that access permission adjustments are seamless.
[0101] Optionally, after anonymizing the fields to be anonymized in the encrypted dataset to obtain the de-identified dataset, the method further includes: in response to access operations of different authorized users on the de-identified dataset, generating de-identified data views corresponding to different authorized users, and displaying the corresponding de-identified data views to different authorized users.
[0102] Specifically, in response to access operations of de-identified datasets by different authorized users, the system can obtain data presentation methods corresponding to different authorized users, and generate de-identified data views corresponding to different authorized users based on these presentation methods. The data presentation formats can be various, such as lists and graphs.
[0103] By setting up the above, different anonymized data views can be provided to different authorized users, thereby improving the user experience for all authorized users.
[0104] Optionally, after anonymizing the fields to be anonymized in the encrypted dataset to obtain the de-identified dataset, the method further includes: in response to an access operation of an authorized user on the de-identified dataset, generating a first de-identified data view corresponding to the authorized user and a second de-identified data view corresponding to the business scenario in which the authorized user is located, and displaying the first de-identified data view and the second de-identified data view to the authorized user.
[0105] The advantage of this setup is that it generates different de-identified data views for authorized users based on roles and scenarios, making it easier for users to intuitively understand the purpose of the de-identified dataset and quickly apply it in practice.
[0106] Optionally, after anonymizing the fields to be anonymized in the encrypted dataset to obtain a de-identified dataset, the method further includes: training an anti-fraud model in conjunction with a third-party credit reporting agency to solve the data silo problem, while ensuring that customer privacy data is not disclosed to the third-party credit reporting agency.
[0107] The technical solution of this embodiment obtains the authorized users corresponding to the dataset to be encrypted, and generates a data access policy corresponding to the dataset based on the user attribute information of the authorized users; obtains the version number of the data access policy, and obtains a version random number corresponding to the version number of the data access policy from a predefined prime group; encrypts the dataset to be encrypted based on the version random number, a randomly generated master key, and the attribute public key corresponding to each user attribute information in the data access policy, resulting in an encrypted dataset corresponding to the dataset to be encrypted. This can be fully integrated with various business scenarios such as loan approval and anti-money laundering monitoring to formulate data access policies, and then use the formulated data access policies to perform fine-grained access control on the data, thereby avoiding unauthorized internal access and improving data security. Secondly, through a pre-built risk prediction model, based on the central field features of each central field and the neighbor field features of all neighbor fields corresponding to each central field, the risk probability of each field in the encrypted dataset being illegally decrypted is determined. The features of neighbor fields can be aggregated and learned, thereby improving the accuracy of determining the risk probability.
[0108] The requirements state that the customer-related information collected by this invention is information and data authorized by the customer or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, take necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation entry points for indebted companies to choose to authorize or refuse.
[0109] Example 3
[0110] Figure 3 This is a schematic diagram of a data desensitization device according to Embodiment 3 of the present invention. This embodiment is applicable to the situation of data desensitization processing. The data desensitization device can be implemented in hardware and / or software and can be configured in electronic devices such as computers.
[0111] like Figure 3 As shown, the data anonymization device disclosed in this embodiment includes: an access policy generation module 31, a first encryption module 32, a risk probability determination module 33, a field to be anonymized determination module 34, and a second encryption module 35, wherein:
[0112] The access policy generation module 31 is used to obtain each authorized user corresponding to the dataset to be encrypted, and generate a data access policy corresponding to the dataset to be encrypted based on the user attribute information of each authorized user.
[0113] The first encryption module 32 is used to encrypt the dataset to be encrypted according to the data access strategy and the randomly generated master key, so as to obtain the encrypted dataset corresponding to the dataset to be encrypted.
[0114] Risk probability determination module 33 is used to determine the risk probability of each field in the encrypted dataset being illegally decrypted through a pre-built risk prediction model;
[0115] The field to be anonymized module 34 is used to determine the field to be anonymized in the encrypted dataset based on the comparison result between each risk probability and a preset risk probability threshold.
[0116] The second encryption module 35 is used to anonymize the fields to be anonymized in the encrypted dataset to obtain the de-anonymized dataset.
[0117] The technical solution in this embodiment, through the cooperation of the access policy generation module 31, the first encryption module 32, the risk probability determination module 33, the field to be anonymized determination module 34, and the second encryption module 35, solves the problem that the existing technology directly anonymizes the dataset to be encrypted to obtain the de-identified dataset, which often makes the de-identified dataset easy to crack, and improves the security of the de-identified dataset.
[0118] Optionally, the first encryption module 32 is specifically used to: obtain the version number of the data access policy, and obtain a version random number corresponding to the version number of the data access policy from a predefined prime number group; encrypt the dataset to be encrypted according to the version random number, the randomly generated master key, and the attribute public key corresponding to each user attribute information in the data access policy, to obtain an encrypted dataset corresponding to the dataset to be encrypted.
[0119] Optionally, the risk probability determination module 33 includes:
[0120] The matrix generation unit is used to generate an adjacency matrix corresponding to the encrypted dataset based on the degree of correlation between the fields in the encrypted dataset, and to generate a field feature matrix corresponding to the encrypted dataset based on the field features of the fields in the encrypted dataset.
[0121] The neighbor field determination unit is used to sequentially take each field in the encrypted dataset as the center field and determine all neighbor fields corresponding to each center field according to the adjacency matrix.
[0122] The field feature acquisition unit is used to determine the central field feature of each central field and the neighbor field features of all neighbor fields corresponding to each central field based on the field feature matrix.
[0123] The risk probability determination unit is used to determine the risk probability of each field in the encrypted dataset being illegally decrypted by using a pre-built risk prediction model, based on the central field characteristics of each central field and the neighbor field characteristics of all neighbor fields corresponding to each central field.
[0124] Optionally, the risk probability determination unit is specifically used to: determine the current aggregation feature of each central field at the current aggregation layer by using a pre-built risk prediction model, based on the central field feature of each central field and the neighbor field features of all neighbor fields corresponding to each central field; determine whether the current aggregation layer is a predefined loop termination layer; if not, after using the current aggregation feature of the central field as the central field feature of the central field and updating the current aggregation layer, return to execute the operation of determining the current aggregation feature of each central field at the current aggregation layer by using the pre-built risk prediction model, based on the central field feature of each central field and the neighbor field features of all neighbor fields corresponding to each central field, until the current aggregation layer is a predefined loop termination layer, and determine the risk probability of each field in the encrypted dataset being illegally decrypted based on the current aggregation feature of each central field.
[0125] Optionally, the device also includes a de-identification log generation module, which is used to: obtain all operations performed during the process of de-identifying the dataset to be encrypted, as well as the timestamp corresponding to each operation; construct a de-identification operation log corresponding to the de-identified dataset based on all operations, the timestamp corresponding to each operation, and the user attribute information of the user with access rights, and generate a hash value corresponding to the de-identification operation log so that the user can read the de-identification operation log.
[0126] Optionally, the device also includes a de-identified data display module, which is used to: generate de-identified data views corresponding to different authorized users in response to their access operations on the de-identified dataset, and display the corresponding de-identified data views to the different authorized users.
[0127] Optionally, the device also includes a private key generation module, which is used to: select random integers corresponding to each user attribute of the authorized user from a predefined prime order group; and generate a private key corresponding to the dataset to be encrypted based on the randomly generated master key and the random integers, so that the authorized user can decrypt the encrypted dataset based on the private key.
[0128] The data desensitization apparatus provided in this embodiment of the invention can execute the data desensitization method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution. Content not described in detail in this embodiment can be referred to the description in any method embodiment of this application.
[0129] Example 4
[0130] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement embodiments of the present invention is shown. For example... Figure 4 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0131] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0132] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as data desensitization methods.
[0133] In some embodiments, the data anonymization method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data anonymization method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data anonymization method by any other suitable means (e.g., by means of firmware).
[0134] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0135] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0136] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0137] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0138] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0139] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0140] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0141] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A data anonymization method, characterized in that, The method includes: Obtain each authorized user corresponding to the dataset to be encrypted, and generate a data access policy corresponding to the dataset to be encrypted based on the user attribute information of each authorized user; Based on the data access strategy and the randomly generated master key, the dataset to be encrypted is encrypted to obtain an encrypted dataset corresponding to the dataset to be encrypted. By using a pre-built risk prediction model, the probability of each field in the encrypted dataset being illegally decrypted is determined; Based on the comparison results between each risk probability and a preset risk probability threshold, the fields to be anonymized are determined in the encrypted dataset; The fields to be anonymized in the encrypted dataset are anonymized to obtain a de-anonymized dataset.
2. The method according to claim 1, characterized in that, Based on the data access policy and the randomly generated master key, the dataset to be encrypted is encrypted to obtain an encrypted dataset corresponding to the dataset to be encrypted, including: Obtain the version number of the data access strategy, and obtain a version random number corresponding to the version number of the data access strategy from a predefined prime number group; The dataset to be encrypted is encrypted using the version random number, the randomly generated master key, and the attribute public key corresponding to each user attribute information in the data access policy, to obtain the encrypted dataset corresponding to the dataset to be encrypted.
3. The method according to claim 1, characterized in that, Using a pre-built risk prediction model, the probability of each field in the encrypted dataset being illegally decrypted is determined, including: Based on the degree of correlation between the fields in the encrypted dataset, an adjacency matrix corresponding to the encrypted dataset is generated, and based on the field characteristics of the fields in the encrypted dataset, a field feature matrix corresponding to the encrypted dataset is generated. Each field in the encrypted dataset is taken as a center field in turn, and all neighbor fields corresponding to each center field are determined according to the adjacency matrix. Based on the field feature matrix, determine the central field feature of each central field, and the neighbor field features of all neighbor fields corresponding to each central field; By using a pre-built risk prediction model, the probability of each field in the encrypted dataset being illegally decrypted is determined based on the central field characteristics of each central field and the neighbor field characteristics of all neighbor fields corresponding to each central field.
4. The method according to claim 3, characterized in that, By using a pre-built risk prediction model, based on the central field characteristics of each central field and the neighbor field characteristics of all neighbor fields corresponding to each central field, the probability of each field in the encrypted dataset being illegally decrypted is determined, including: By using a pre-built risk prediction model, the current aggregation feature of each central field at the current aggregation level is determined based on the central field features of each central field and the neighbor field features of all neighbor fields corresponding to each central field. Determine if the current aggregation level is the predefined loop termination level; If not, after using the current aggregation feature of the central field as the central field feature of the central field and updating the current aggregation layer, the process returns to execute the operation of determining the current aggregation feature of each central field at the current aggregation layer based on the central field feature of each central field and the neighbor field features of all neighbor fields corresponding to each central field, until the current aggregation layer reaches the predefined loop termination layer. Based on the current aggregation characteristics of each central field, determine the probability of each field in the encrypted dataset being illegally decrypted.
5. The method according to any one of claims 1-4, characterized in that, After anonymizing the fields to be anonymized in the encrypted dataset to obtain the de-anonymized dataset, the process further includes: Obtain all operations performed during the process of de-identifying the dataset to be encrypted to obtain the de-identified dataset, as well as the timestamp corresponding to each operation; Based on all the operations, the timestamp corresponding to each operation, and the user attribute information of the user with access rights, a de-identified operation log corresponding to the de-identified dataset is constructed, and a hash value corresponding to the de-identified operation log is generated so that the user can read the de-identified operation log.
6. The method according to any one of claims 1-4, characterized in that, After anonymizing the fields to be anonymized in the encrypted dataset to obtain the de-anonymized dataset, the process further includes: In response to access operations of different authorized users on the de-identified dataset, a de-identified data view corresponding to each authorized user is generated and displayed to each authorized user.
7. The method according to any one of claims 1-4, characterized in that, Before encrypting the dataset to be encrypted according to the data access policy and the randomly generated master key to obtain the encrypted dataset corresponding to the dataset to be encrypted, the method further includes: From a predefined group of prime numbers, select random integers corresponding to each user attribute of the user with access rights; Based on the randomly generated master key and the random integer, a private key corresponding to the dataset to be encrypted is generated, so that the authorized user can decrypt the encrypted dataset using the private key.
8. A data anonymization device, characterized in that, The device includes: The access policy generation module is used to obtain each authorized user corresponding to the dataset to be encrypted, and generate a data access policy corresponding to the dataset to be encrypted based on the user attribute information of each authorized user. The first encryption module is used to encrypt the dataset to be encrypted according to the data access strategy and the randomly generated master key to obtain an encrypted dataset corresponding to the dataset to be encrypted. The risk probability determination module is used to determine the risk probability of each field in the encrypted dataset being illegally decrypted through a pre-built risk prediction model; The field to be anonymized module is used to determine the field to be anonymized in the encrypted dataset based on the comparison result between each risk probability and a preset risk probability threshold; The second encryption module is used to anonymize the fields to be anonymized in the encrypted dataset to obtain a de-anonymized dataset.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data desensitization method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the data desensitization method according to any one of claims 1-7.