Enterprise data linkage desensitization method and device and computer equipment
By constructing a sparse risk hypergraph and performing field desensitization processing, the problem of low desensitization efficiency of multi-field high-order correlation risks is solved, and high-precision and high-efficiency data desensitization is achieved, ensuring that the data retains maximum availability under given risks.
Patent Information
- Application Number
- CN202511215143.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-08-28
AI Technical Summary
Existing technologies are unable to effectively avoid multi-field and multi-dimensional high-order correlation risks in data desensitization, resulting in high computational complexity, insufficient accuracy, and low desensitization efficiency.
By constructing a sparse risk hypergraph, identifying the data type and desensitization impact category of sensitive fields, using the weight value of the sparse risk hypergraph to desensitize the fields, and combining the weighted vertex cover algorithm to generate the minimum desensitized field set, we ensure that data availability is retained to the maximum extent within the established residual risk upper limit.
Under limited computational complexity and storage space, it provides a high-precision and high-efficiency multi-field correlation desensitization solution, improving desensitization efficiency and data availability.
Smart Images

Figure CN120744983A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of big data security and privacy protection, and in particular to a method, apparatus and computer equipment for the linked desensitization of enterprise data. Background Art
[0002] With the widespread adoption of cloud computing and big data analytics, the risk of data leakage has significantly increased in industries such as finance, healthcare, and government when engaging in data sharing, model training, and outsourcing services. The data services industry is increasingly demanding data desensitization technologies. Traditional data desensitization methods, which rely on two-dimensional relationship processing techniques such as differential privacy, K-anonymity, and L-diversity, can only statically mask single fields or pairs of fields, making it difficult to capture the high-order correlation risks associated with field combinations. Therefore, effectively mitigating these high-order correlation risks is a current research priority.
[0003] Existing technology uses hypergraph-based methods to capture high-order correlation risks. However, in current data desensitization solutions, the huge computational complexity overhead brought about by multi-field and multi-dimensional high-order desensitization and the requirement for quantitative calculation accuracy of correlation risks in designing desensitization solutions are contradictory, resulting in low data desensitization efficiency. Summary of the Invention
[0004] Based on this, it is necessary to provide a method, device, computer equipment, computer-readable storage medium and computer program product for the linkage desensitization of enterprise data to address the above technical issues.
[0005] In a first aspect, the present application provides a method for collaborative desensitization of enterprise data, including:
[0006] Obtaining enterprise data that needs to be desensitized, and encrypting the enterprise data to obtain an enterprise encryption map;
[0007] Based on the enterprise encrypted graph, a sparse risk hypergraph is constructed, and weight values of each hyperedge set of the sparse risk hypergraph are identified;
[0008] Based on the sparse risk hypergraph and the weight values of each hyperedge set of the sparse risk hypergraph, the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field are identified, and based on the data type of each sensitive field and the desensitization impact category of each sensitive field, field desensitization processing is performed on each sensitive field separately to obtain desensitized enterprise data.
[0009] Optionally, encrypting the enterprise data to obtain an enterprise encryption map includes:
[0010] In the enterprise data, the field data of each field after deduplication is screened, and based on the field data of each field, the field name of each field and the field description of each field are identified;
[0011] Perform string concatenation on the field name and field description of each field to obtain new field data corresponding to each field, and build an enterprise data map based on the new field data corresponding to each field;
[0012] The enterprise data graph is encrypted through multiple encryption strategies to obtain an enterprise encrypted graph.
[0013] Optionally, constructing a sparse risk hypergraph based on the enterprise encrypted graph includes:
[0014] Based on the enterprise encryption map, a sampling matrix is screened, and uniqueness calculation is performed on the sampling matrix to obtain a uniqueness ratio corresponding to each field combination of the sampling matrix;
[0015] Based on the uniqueness ratio corresponding to each of the field combinations, each target field combination is screened by a combined screening strategy, and each of the target fields is used as each hyperedge set;
[0016] Based on the hyperedge sets, a sparse risk hypergraph is constructed.
[0017] Optionally, identifying the weight values of each hyperedge set of the sparse risk hypergraph includes:
[0018] Based on each of the hyperedge sets, calculating the normalized mutual information of each of the hyperedge sets by a normalization algorithm;
[0019] Based on the normalized mutual information of each of the hyperedge sets, a weight value of each of the hyperedge sets is calculated using a hyperedge weight estimation algorithm.
[0020] Optionally, identifying the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field based on the sparse risk hypergraph and the weight values of each hyperedge set of the sparse risk hypergraph includes:
[0021] Based on the weight value of each hyperedge set, performing risk inference processing on each hyperedge set to obtain risk distribution information of the sparse risk hypergraph;
[0022] Based on the risk distribution information of the sparse risk hypergraph, identifying sensitive fields in each field of each superset, and identifying the data type of each sensitive field of each superset;
[0023] Based on the data type of each sensitive field, a scoring value corresponding to each sensitive field and a desensitization impact category of each sensitive field are determined.
[0024] Optionally, based on the data type of each sensitive field and the desensitization impact category of each sensitive field, performing field desensitization processing on each sensitive field to obtain desensitized enterprise data includes:
[0025] Based on the desensitization impact category of each sensitive field and the data type of each sensitive field, calculating the field desensitization cost value corresponding to each sensitive field using a field desensitization cost algorithm;
[0026] Based on the field desensitization cost value corresponding to each of the sensitive fields, identifying, in the enterprise data, each minimum desensitization coverage subset corresponding to the enterprise data by a desensitization coverage algorithm;
[0027] Based on the data type of each sensitive field in each desensitizing coverage subset, the field of each minimum desensitizing coverage subset is desensitized using the desensitization strategy corresponding to each data type to obtain desensitized enterprise data.
[0028] In a second aspect, the present application also provides a linkage desensitization device for enterprise data, including:
[0029] An acquisition module is used to acquire enterprise data that needs to be desensitized and encrypt the enterprise data to obtain an enterprise encryption map;
[0030] An identification module, configured to construct a sparse risk hypergraph based on the enterprise encrypted graph and identify weight values of each hyperedge set of the sparse risk hypergraph;
[0031] The desensitization module is used to identify the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field based on the sparse risk hypergraph and the weight values of each hyperedge set of the sparse risk hypergraph, and perform field desensitization processing on each sensitive field based on the data type of each sensitive field and the desensitization impact category of each sensitive field to obtain desensitized enterprise data.
[0032] Optionally, the acquisition module is specifically configured to:
[0033] In the enterprise data, the field data of each field after deduplication is screened, and based on the field data of each field, the field name of each field and the field description of each field are identified;
[0034] Perform string concatenation on the field name and field description of each field to obtain new field data corresponding to each field, and build an enterprise data map based on the new field data corresponding to each field;
[0035] The enterprise data graph is encrypted through multiple encryption strategies to obtain an enterprise encrypted graph.
[0036] Optionally, the identification module is specifically configured to:
[0037] Based on the enterprise encryption map, a sampling matrix is screened, and uniqueness calculation is performed on the sampling matrix to obtain a uniqueness ratio corresponding to each field combination of the sampling matrix;
[0038] Based on the uniqueness ratio corresponding to each of the field combinations, each target field combination is screened by a combined screening strategy, and each of the target fields is used as each hyperedge set;
[0039] Based on the hyperedge sets, a sparse risk hypergraph is constructed.
[0040] Optionally, the identification module is specifically configured to:
[0041] Based on each of the hyperedge sets, calculating the normalized mutual information of each of the hyperedge sets by a normalization algorithm;
[0042] Based on the normalized mutual information of each of the hyperedge sets, a weight value of each of the hyperedge sets is calculated using a hyperedge weight estimation algorithm.
[0043] Optionally, the desensitization module is specifically used to:
[0044] Based on the weight value of each hyperedge set, performing risk inference processing on each hyperedge set to obtain risk distribution information of the sparse risk hypergraph;
[0045] Based on the risk distribution information of the sparse risk hypergraph, identifying sensitive fields in each field of each superset, and identifying the data type of each sensitive field of each superset;
[0046] Based on the data type of each sensitive field, a scoring value corresponding to each sensitive field and a desensitization impact category of each sensitive field are determined.
[0047] Optionally, the desensitization module is specifically used to:
[0048] Based on the desensitization impact category of each sensitive field and the data type of each sensitive field, calculating the field desensitization cost value corresponding to each sensitive field using a field desensitization cost algorithm;
[0049] Based on the field desensitization cost value corresponding to each of the sensitive fields, identifying, in the enterprise data, each minimum desensitization coverage subset corresponding to the enterprise data by a desensitization coverage algorithm;
[0050] Based on the data type of each sensitive field in each desensitizing coverage subset, the field of each minimum desensitizing coverage subset is desensitized using the desensitization strategy corresponding to each data type to obtain desensitized enterprise data.
[0051] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the methods described in the first aspect when executing the computer program.
[0052] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of any one of the methods in the first aspect.
[0053] In a fifth aspect, the present application provides a computer program product, wherein the computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of any one of the methods in the first aspect are implemented.
[0054] The above-mentioned method, apparatus, and computer device for linked desensitization of enterprise data obtain enterprise data requiring desensitization and encrypt the enterprise data to obtain an encrypted enterprise graph. Based on the encrypted enterprise graph, a sparse risk hypergraph is constructed, and the weight values of each hyperedge set of the sparse risk hypergraph are identified. Based on the sparse risk hypergraph and the weight values of each hyperedge set of the sparse risk hypergraph, the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field are identified. Based on the data type and desensitization impact category of each sensitive field, field desensitization is performed on each sensitive field to obtain desensitized enterprise data. This solution abstracts a large number of sensitive fields and their combinations into a hypergraph structure, and uses rapidly calculated uniqueness as a screening criterion to make the hypergraph structure sufficiently sparse. This allows the potential risks brought about by the linkage of multiple fields to be fully demonstrated while maintaining limited computational complexity and storage space, breaking through the traditional limitation of focusing only on single fields or field pairs. Then, based on the risk weights predicted by the hypergraph and combined with the field data costs, this solution intelligently generates the minimum desensitized field set through a weighted vertex cover algorithm to ensure that data availability is retained to the maximum extent within the established residual risk upper limit, thereby providing a multi-field sparse data desensitization solution with high precision, high efficiency and high availability implementation potential, and comprehensively improving the desensitization efficiency of desensitized fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0056] Figure 1 1. A flowchart of a method for linked desensitization of enterprise data in one embodiment;
[0057] Figure 2 This is a flowchart of an example of linked desensitization of enterprise data in one embodiment;
[0058] Figure 3 This is a structural block diagram of a linkage desensitization device for enterprise data in one embodiment;
[0059] Figure 4 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0060] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0061] The linkage desensitization method for enterprise data provided in the embodiment of the present application can be applied to terminals, servers, and systems including terminals and servers, and is implemented through the interaction between terminals and servers. Among them, the terminal can be, but is not limited to, various personal computers, laptops, mid-range computers, etc. Among them, the terminal abstracts a large number of sensitive fields and their combinations into a hypergraph structure, and uses the uniqueness of fast calculation as a screening benchmark to make the hypergraph structure sparse enough, so that it can fully demonstrate the potential risks brought by the linkage of multiple fields under limited computational complexity and storage space, breaking through the traditional limitation of focusing only on single fields or field pairs. Then, based on the risk weight predicted by the hypergraph and combined with the field data cost, this solution intelligently generates the minimum desensitized field set through a weighted vertex cover algorithm to ensure that the data availability is retained to the maximum extent within the established residual risk upper limit, thereby providing a multi-field sparse data desensitization solution with high precision, high efficiency and high availability landing potential, and comprehensively improving the desensitization efficiency of desensitized fields.
[0062] In an exemplary embodiment, Figure 1 As shown, a method for linkage desensitization of enterprise data is provided, which is described by taking the application of the method to a terminal as an example, and includes the following steps S101 to S103. Among them:
[0063] Step S101: obtain enterprise data that needs to be desensitized, and encrypt the enterprise data to obtain an enterprise encryption map.
[0064] In this embodiment, the terminal, in response to the system's information upload operation, obtains enterprise data requiring data desensitization. This enterprise data may include, but is not limited to, charts, text, and other data. Before obtaining this enterprise data, a user must upload the associated relational tables or knowledge graphs. The terminal then encrypts the enterprise data to produce an encrypted enterprise graph. This encryption method may include, but is not limited to, homomorphic encryption or multi-party encryption. The specific encryption process will be described in detail later.
[0065] Step S102: construct a sparse risk hypergraph based on the enterprise encrypted graph, and identify the weight value of each hyperedge set of the sparse risk hypergraph.
[0066] In this embodiment, the terminal constructs a sparse risk hypergraph based on the enterprise encrypted graph and identifies the weights of each hyperedge set in the sparse risk hypergraph. This sparse risk hypergraph is constructed by constructing hyperedge sets based on the data arrangement of high-risk enterprise data, thereby generating a sparse risk hypergraph. The terminal then calculates the weights of each hyperedge set in the sparse risk hypergraph using normalized information calculation. The specific calculation process will be described in detail later.
[0067] Step S103: Based on the sparse risk hypergraph and the weight values of each hyperedge set of the sparse risk hypergraph, identify the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field, and based on the data type of each sensitive field and the desensitization impact category of each sensitive field, perform field desensitization processing on each sensitive field separately to obtain desensitized enterprise data.
[0068] In this embodiment, the terminal identifies the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field based on the sparse risk hypergraph and the weight values of each hyperedge set of the sparse risk hypergraph, and performs field desensitization processing on each sensitive field based on the data type of each sensitive field and the desensitization impact category of each sensitive field to obtain desensitized enterprise data. Among them, the field desensitization method depends on the data type corresponding to different sensitive fields, that is, the desensitization schemes corresponding to sensitive fields of different data types are different. The desensitization schemes include but are not limited to pseudonymization, data masking or local suppression schemes. Among them, as shown in Table 1, the relationship table corresponding to the data type of each sensitive field and the desensitization impact category is:
[0069] Table 1: Correspondence table of field types
[0070]
[0071] Based on the above solution, by abstracting massive sensitive fields and their combinations into a hypergraph structure and using the rapidly calculated uniqueness as a screening benchmark to make the hypergraph structure sufficiently sparse, we can fully demonstrate the potential risks brought about by the linkage of multiple fields within limited computational complexity and storage space, breaking through the traditional limitation of focusing only on single fields or field pairs. This solution then uses a weighted vertex cover algorithm to intelligently generate a minimum set of desensitized fields based on the risk weights predicted by the hypergraph and the field data cost, ensuring that data availability is retained to the maximum extent possible within the established residual risk limit. This provides a multi-field sparse data desensitization solution with high precision, high efficiency, and high availability, and comprehensively improves the desensitization efficiency of desensitized fields.
[0072] Optionally, the enterprise data is encrypted to obtain an enterprise encryption map, including: filtering the field data of each field after deduplication in the enterprise data, and identifying the field name of each field and the field description of each field based on the field data of each field; performing string concatenation on the field name of each field and the field description of each field to obtain new field data corresponding to each field, and constructing an enterprise data map based on the new field data corresponding to each field; performing data encryption on the enterprise data map through multiple encryption strategies to obtain an enterprise encryption map.
[0073] In this embodiment, the terminal screens the field data of each field after deduplication in the enterprise data, and identifies the field name of each field and the field description of each field based on the field data of each field. Specifically, all fields are first extracted from multiple relational tables or existing knowledge graphs, and fields related to enterprise privacy and enterprise secrets or fields that may be combined for desensitization are pre-screened through a local large model, and the locally deployed DeepSeek-R1-14B large model is used to perform semantic deduplication on the above fields from multiple tables to obtain a deduplicated field set. ,in , M is the total number of fields after deduplication. In field screening, the information density of each field is counted, and fields with a cardinality of 1 or a missing rate exceeding 95% are statically eliminated to ensure that each All of them have statistical value.
[0074] Then, the terminal performs string concatenation on the field name of each field and the field description of each field to obtain the new field data corresponding to each field, and builds the enterprise data map based on the new field data corresponding to each field. Specifically, in order to ensure the data security of hypergraph construction and network training, the terminal performs string concatenation on the field name and field description of each field, and uses the pre-trained language model deployed locally offline to embed the semantic information into the vector map. .
[0075] Finally, the terminal encrypts the target enterprise data map through multiple encryption strategies to obtain the enterprise encryption map. Perform homomorphic encryption or multi-party encryption operations to make It can no longer be inversely solved by the embedded model. Among them, V is stored in the metadata table as a row vector, and each row contains the field embedding vector , data type , cardinality , missing rate ,get .
[0076] Based on the above solution, by preprocessing and encrypting enterprise data, we can avoid the risk of data leaving the domain and prevent data leakage in subsequent processes that would cause field names and field descriptions to be reverse-crackered by external models, thereby achieving differential privacy protection for private domain data.
[0077] Optionally, a sparse risk hypergraph is constructed based on the enterprise encrypted graph, including: based on the enterprise encrypted graph, screening the sampling matrix, and performing uniqueness calculation on the sampling matrix to obtain the uniqueness ratio corresponding to each field combination of the sampling matrix; based on the uniqueness ratio corresponding to each field combination, screening each target field combination through a combined screening strategy, and using each target field as each hyperedge set; based on each hyperedge set, a sparse risk hypergraph is constructed.
[0078] In this embodiment, the terminal screens a sampling matrix based on the enterprise encryption graph and calculates the uniqueness of the sampling matrix to obtain the uniqueness percentage corresponding to each field combination in the sampling matrix. Then, based on the uniqueness percentage corresponding to each field combination, the terminal uses a combined screening strategy to select target field combinations and uses each target field as a hyperedge set. Finally, the terminal constructs a sparse risk hypergraph based on each hyperedge set.
[0079] Specifically, in order to construct a sparse risk hypergraph, the terminal extracts M fields based on the fields of each enterprise data. Row, get the sampling matrix .
[0080] Then, based on the above sampling matrix, the terminal system can further calculate the uniqueness of the secondary and tertiary field combinations. The expression is:
[0081]
[0082] Among them, the field combination Is a certain order The field combination, ,and A tuple representing the values of the field combination E. The meaning is: for the field combination E, there is OK( An enterprise) can be uniquely identified by the value of the field combination E.
[0083] Among them, the uniqueness ratio It will be used for initial screening. The terminal will calculate the risk of each field combination before building the risk hypermap. , and only keep it in the hypergraph The field combination of is used as a hyperedge. , which makes the hyperedges on the hypergraph sparser and greatly reduces the computational complexity. The hyperedge set of the hypergraph is obtained, which can be written as . Then, the weight of the hyperedge E represents the re-identification risk of the hyperedge.
[0084] Based on the above solution, by constructing a sparse risk hypergraph, massive sensitive fields and their combinations are abstracted into a hypergraph structure. At the same time, the fast-calculated uniqueness is used as a screening benchmark to make the hypergraph structure sufficiently sparse. In this way, the potential risks brought about by the linkage of multiple fields can be fully demonstrated under limited computational complexity and storage space, breaking through the traditional limitation of focusing only on single fields or field pairs.
[0085] Optionally, identifying the weight values of each hyperedge set of the sparse risk hypergraph includes: calculating the normalized mutual information of each hyperedge set based on each hyperedge set through a normalization algorithm; and calculating the weight value of each hyperedge set through a hyperedge weight estimation algorithm based on the normalized mutual information of each hyperedge set.
[0086] In this embodiment, the terminal calculates the normalized mutual information of each hyperedge set based on each hyperedge set through a normalization algorithm. Specifically, the normalized mutual information is a reliable weight evaluation benchmark, and its expression is:
[0087]
[0088] In order to reduce the amount of calculation, the terminal uses t=1 to approximate the normalized mutual information. While initializing the weights of all hyperedges in the hypergraph, we calculate the exact normalized mutual information of 5% of the randomly sampled hyperedges using formula (2): As labeled datasets, these data constitute the training set and test set of HyperGNN.
[0089] Then, the terminal calculates the weight value of each hyperedge set based on the normalized mutual information of each hyperedge set through the hyperedge weight estimation algorithm. The characteristics of Given, and do not set the output head on the node, but only set the output head on the hyperedge E. Use the two-layer HyperSAGE as the backbone network to calculate the feature representation of each hyperedge:
[0090]
[0091] And use an MLP as the output head of HyperGNN to get the weight estimate of the hyperedge:
[0092]
[0093] Based on the above scheme, by combining the normalization algorithm, the normalized mutual information of each hyperedge set is calculated, so as to obtain the weight value of each hyperedge set, thereby improving the accuracy of risk analysis of each hyperedge set.
[0094] Optionally, based on the sparse risk hypergraph and the weight values of each hyperedge set of the sparse risk hypergraph, the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field are identified, including: based on the weight values of each hyperedge set, risk inference processing is performed on each hyperedge set to obtain risk distribution information of the sparse risk hypergraph; based on the risk distribution information of the sparse risk hypergraph, the sensitive fields in each field of each superset are identified, and the data type of each sensitive field of each superset is identified; based on the data type of each sensitive field, the scoring value corresponding to each sensitive field and the desensitization impact category of each sensitive field are determined.
[0095] In this embodiment, the terminal performs risk inference processing on each hyperedge set based on the weight value of each hyperedge set to obtain the risk distribution information of the sparse risk hypergraph; specifically, the terminal considers the loss function , for HyperGNN and MLP head Perform back propagation update. After training is completed, perform a forward inference on all hyperedges to obtain the final risk distribution .
[0096] Then, based on the risk distribution information of the sparse risk hypergraph, the terminal identifies the sensitive fields within each field of each superset and the data type of each sensitive field in each superset. Based on the data type of each sensitive field, the terminal determines the score corresponding to each sensitive field and the desensitization impact category of each sensitive field. This risk distribution information includes the risk value corresponding to each field. The terminal then selects fields with risk values below a risk threshold preset in the terminal as sensitive fields. Then, using Table 1, it performs type adaptation to obtain the score corresponding to each sensitive field and the desensitization impact category of each sensitive field.
[0097] Based on the above scheme, a specially designed sparse risk hypergraph is applied. The network can be trained with 5% labeled data calculated based on mutual information, and can automatically capture the complex interaction patterns of field combinations in the hypergraph, and generate an interpretable risk score for each high-order risk association, thereby improving the accuracy and flexibility of desensitization decisions.
[0098] Optionally, based on the data type of each sensitive field and the desensitization impact category of each sensitive field, field desensitization processing is performed on each sensitive field separately to obtain desensitized enterprise data, including: based on the desensitization impact category of each sensitive field and the data type of each sensitive field, the field desensitization cost value corresponding to each sensitive field is calculated through a field desensitization cost algorithm; based on the field desensitization cost value corresponding to each sensitive field, the minimum desensitization coverage subsets corresponding to the enterprise data are identified in the enterprise data through a desensitization coverage algorithm; based on the data type of each sensitive field in each desensitization coverage subset, the desensitization strategy corresponding to each data type is used to perform field desensitization processing on each minimum desensitization coverage subset to obtain desensitized enterprise data.
[0099] In this embodiment, the terminal calculates the field desensitization cost value corresponding to each sensitive field based on the desensitization impact category of each sensitive field and the data type of each sensitive field through a field desensitization cost algorithm. Specifically, the field desensitization cost algorithm is:
[0100]
[0101] in, For desensitized fields, is the field desensitization cost value of the desensitized field. is the rating value, Is the data type.
[0102] Then, based on the field desensitization cost value corresponding to each sensitive field, the terminal uses the desensitization coverage algorithm to identify the minimum desensitization coverage subsets corresponding to the enterprise data. The screening method for each minimum desensitization coverage subset is:
[0103]
[0104] in, is the minimum desensitization coverage subset, A collection of fields in enterprise data.
[0105] Finally, the terminal desensitizes the fields in each minimum desensitization coverage subset based on the data type of each sensitive field in each desensitization coverage subset and applies the desensitization policy corresponding to each data type, thereby obtaining desensitized enterprise data. The terminal presets the correspondence between different data types and desensitization policies. Then, using this correspondence, the terminal adapts the desensitization policy corresponding to each data type.
[0106] Based on the above solution, by combining different data types and adapting various desensitization strategies, the accuracy and effect of data desensitization are improved. Secondly, this solution comprehensively estimates the desensitization cost based on the field usage value and risk intensity, and automatically selects desensitization targets by solving weighted vertex cover with a greedy algorithm, maximizing the availability of business data while ensuring that the overall desensitization risk is controllable.
[0107] This application also provides an example of linkage desensitization of enterprise data, such as Figure 2 As shown, the specific processing process includes the following steps:
[0108] Step S201: Acquire enterprise data that needs to be desensitized.
[0109] Step S202 : Filter the field data of each field after deduplication in the enterprise data, and identify the field name of each field and the field description of each field based on the field data of each field.
[0110] In step S203, string concatenation is performed on the field name of each field and the field description of each field to obtain new field data corresponding to each field, and an enterprise data map is constructed based on the new field data corresponding to each field.
[0111] Step S204: encrypt the enterprise data graph using multiple encryption strategies to obtain an encrypted enterprise graph.
[0112] Step S205 , based on the enterprise encryption graph, screen the sampling matrix, and calculate the uniqueness of the sampling matrix to obtain the uniqueness ratio corresponding to each field combination of the sampling matrix.
[0113] In step S206 , based on the uniqueness ratio corresponding to each field combination, each target field combination is screened by a combined screening strategy, and each target field is used as each hyperedge set.
[0114] Step S207: construct a sparse risk hypergraph based on each hyperedge set.
[0115] Step S208 : Based on each hyperedge set, calculate the normalized mutual information of each hyperedge set by a normalization algorithm.
[0116] Step S209 : Calculate the weight value of each hyperedge set based on the normalized mutual information of each hyperedge set by using a hyperedge weight estimation algorithm.
[0117] Step S210 : performing risk inference processing on each hyperedge set based on the weight value of each hyperedge set to obtain risk distribution information of a sparse risk hypergraph.
[0118] Step S211 : Based on the risk distribution information of the sparse risk hypergraph, the sensitive fields in each field of each superset are identified, and the data type of each sensitive field of each superset is identified.
[0119] Step S212: Based on the data type of each sensitive field, determine the score value corresponding to each sensitive field and the desensitization impact category of each sensitive field.
[0120] Step S213 : Based on the desensitization impact category of each sensitive field and the data type of each sensitive field, a field desensitization cost value corresponding to each sensitive field is calculated using a field desensitization cost algorithm.
[0121] Step S214: Based on the field desensitization cost value corresponding to each sensitive field, a desensitization coverage algorithm is used to identify the minimum desensitization coverage subsets corresponding to the enterprise data in the enterprise data.
[0122] Step S215 , based on the data type of each sensitive field in each desensitizing coverage subset, and using the desensitizing strategy corresponding to each data type, perform field desensitization processing on each minimum desensitizing coverage subset to obtain desensitized enterprise data.
[0123] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0124] Based on the same inventive concept, the embodiment of the present application also provides a linkage desensitization device for enterprise data for implementing the linkage desensitization method for enterprise data involved above. The implementation solution provided by the device is similar to the implementation solution described in the above method, so the specific limitations in the embodiments of the linkage desensitization device for one or more enterprise data provided below can be found in the limitations of the linkage desensitization method for enterprise data above, and will not be repeated here.
[0125] In an exemplary embodiment, Figure 3 As shown, a linkage desensitization device for enterprise data is provided, including: an acquisition module 310, an identification module 320 and a desensitization module 330, wherein:
[0126] An acquisition module 310 is used to acquire enterprise data that needs to be desensitized and encrypt the enterprise data to obtain an enterprise encryption map;
[0127] An identification module 320 is configured to construct a sparse risk hypergraph based on the enterprise encrypted graph and identify weight values of each hyperedge set of the sparse risk hypergraph;
[0128] The desensitization module 330 is used to identify the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field based on the sparse risk hypergraph and the weight values of each hyperedge set of the sparse risk hypergraph, and perform field desensitization processing on each sensitive field based on the data type of each sensitive field and the desensitization impact category of each sensitive field to obtain desensitized enterprise data.
[0129] Optionally, the acquisition module 310 is specifically configured to:
[0130] In the enterprise data, the field data of each field after deduplication is screened, and based on the field data of each field, the field name of each field and the field description of each field are identified;
[0131] Perform string concatenation on the field name and field description of each field to obtain new field data corresponding to each field, and build an enterprise data map based on the new field data corresponding to each field;
[0132] The enterprise data graph is encrypted through multiple encryption strategies to obtain an enterprise encrypted graph.
[0133] Optionally, the identification module 320 is specifically configured to:
[0134] Based on the enterprise encryption map, a sampling matrix is screened, and uniqueness calculation is performed on the sampling matrix to obtain a uniqueness ratio corresponding to each field combination of the sampling matrix;
[0135] Based on the uniqueness ratio corresponding to each of the field combinations, each target field combination is screened by a combined screening strategy, and each of the target fields is used as each hyperedge set;
[0136] Based on the hyperedge sets, a sparse risk hypergraph is constructed.
[0137] Optionally, the identification module 320 is specifically configured to:
[0138] Based on each of the hyperedge sets, calculating the normalized mutual information of each of the hyperedge sets by a normalization algorithm;
[0139] Based on the normalized mutual information of each of the hyperedge sets, a weight value of each of the hyperedge sets is calculated using a hyperedge weight estimation algorithm.
[0140] Optionally, the desensitization module 330 is specifically configured to:
[0141] Based on the weight value of each hyperedge set, performing risk inference processing on each hyperedge set to obtain risk distribution information of the sparse risk hypergraph;
[0142] Based on the risk distribution information of the sparse risk hypergraph, identifying sensitive fields in each field of each superset, and identifying the data type of each sensitive field of each superset;
[0143] Based on the data type of each sensitive field, a scoring value corresponding to each sensitive field and a desensitization impact category of each sensitive field are determined.
[0144] Optionally, the desensitization module 330 is specifically configured to:
[0145] Based on the desensitization impact category of each sensitive field and the data type of each sensitive field, calculating the field desensitization cost value corresponding to each sensitive field using a field desensitization cost algorithm;
[0146] Based on the field desensitization cost value corresponding to each of the sensitive fields, identifying, in the enterprise data, each minimum desensitization coverage subset corresponding to the enterprise data by a desensitization coverage algorithm;
[0147] Based on the data type of each sensitive field in each desensitizing coverage subset, the field of each minimum desensitizing coverage subset is desensitized using the desensitization strategy corresponding to each data type to obtain desensitized enterprise data.
[0148] Each module in the aforementioned enterprise data linkage desensitization device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0149] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 4As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be achieved via Wi-Fi, mobile cellular networks, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a method for detecting surface defects of ultra-precision optical components. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0150] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0151] In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor performs the steps of any one of the methods according to the first aspect when executing the computer program.
[0152] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the methods in the first aspect are described.
[0153] In one embodiment, a computer program product is provided, comprising a computer program, wherein the computer program is executed by a processor to perform the steps of any one of the methods according to the first aspect.
[0154] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0155] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processors (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.
[0156] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0157] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for linked desensitization of enterprise data, characterized in that: The method comprises: Obtaining enterprise data that needs to be desensitized, and encrypting the enterprise data to obtain an enterprise encryption map; Based on the enterprise encrypted graph, a sparse risk hypergraph is constructed, and weight values of each hyperedge set of the sparse risk hypergraph are identified; Based on the sparse risk hypergraph and the weight values of each hyperedge set of the sparse risk hypergraph, the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field are identified, and based on the data type of each sensitive field and the desensitization impact category of each sensitive field, field desensitization processing is performed on each sensitive field separately to obtain desensitized enterprise data.
2. The method according to claim 1, characterized in that The step of encrypting the enterprise data to obtain an enterprise encryption graph includes: In the enterprise data, the field data of each field after deduplication is screened, and based on the field data of each field, the field name of each field and the field description of each field are identified; Perform string concatenation on the field name and field description of each field to obtain new field data corresponding to each field, and build an enterprise data map based on the new field data corresponding to each field; The enterprise data graph is encrypted through multiple encryption strategies to obtain an enterprise encrypted graph.
3. The method according to claim 2, characterized in that The step of constructing a sparse risk hypergraph based on the enterprise encrypted graph includes: Based on the enterprise encryption map, a sampling matrix is screened, and uniqueness calculation is performed on the sampling matrix to obtain a uniqueness ratio corresponding to each field combination of the sampling matrix; Based on the uniqueness ratio corresponding to each of the field combinations, each target field combination is screened by a combined screening strategy, and each of the target fields is used as each hyperedge set; Based on the hyperedge sets, a sparse risk hypergraph is constructed.
4. The method according to claim 3, characterized in that The identifying weight values of each hyperedge set of the sparse risk hypergraph includes: Based on each of the hyperedge sets, calculating the normalized mutual information of each of the hyperedge sets by a normalization algorithm; Based on the normalized mutual information of each of the hyperedge sets, a weight value of each of the hyperedge sets is calculated using a hyperedge weight estimation algorithm.
5. The method according to claim 1, wherein The identifying, based on the sparse risk hypergraph and the weight values of the hyperedge sets of the sparse risk hypergraph, the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field includes: Based on the weight value of each hyperedge set, performing risk inference processing on each hyperedge set to obtain risk distribution information of the sparse risk hypergraph; Based on the risk distribution information of the sparse risk hypergraph, identifying sensitive fields in each field of each superset, and identifying the data type of each sensitive field of each superset; Based on the data type of each sensitive field, a scoring value corresponding to each sensitive field and a desensitization impact category of each sensitive field are determined.
6. The method according to claim 1, characterized in that Based on the data type of each sensitive field and the desensitization impact category of each sensitive field, field desensitization processing is performed on each sensitive field to obtain desensitized enterprise data, including: Based on the desensitization impact category of each sensitive field and the data type of each sensitive field, calculating the field desensitization cost value corresponding to each sensitive field using a field desensitization cost algorithm; Based on the field desensitization cost value corresponding to each of the sensitive fields, identifying, in the enterprise data, each minimum desensitization coverage subset corresponding to the enterprise data by a desensitization coverage algorithm; Based on the data type of each sensitive field in each desensitizing coverage subset, the field of each minimum desensitizing coverage subset is desensitized using the desensitization strategy corresponding to each data type to obtain desensitized enterprise data.
7. A linkage desensitization device for enterprise data, characterized in that: The device comprises: An acquisition module is used to acquire enterprise data that needs to be desensitized and encrypt the enterprise data to obtain an enterprise encryption map; An identification module, configured to construct a sparse risk hypergraph based on the enterprise encrypted graph and identify weight values of each hyperedge set of the sparse risk hypergraph; The desensitization module is used to identify the data type of each sensitive field in the enterprise data and the desensitization impact category of each sensitive field based on the sparse risk hypergraph and the weight values of each hyperedge set of the sparse risk hypergraph, and perform field desensitization processing on each sensitive field based on the data type of each sensitive field and the desensitization impact category of each sensitive field to obtain desensitized enterprise data.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Combined desensitization checking method for medical data
CN117494193A
Establishment method of data leakage prevention system
CN119377995A
Dynamic data desensitization system for employee performance assessment
CN120372691A
User information storage control method, electronic device, and non-volatile computer-readable storage medium
US20250111085A1
Data access method and device, and storage medium and electronic device
WO2022012669A1
Cited By
Sensitive data detection method and device based on hierarchical information enhancement and graph convolutional network
CN121479840A