Dynamic data desensitization system for employee performance appraisal

By identifying the risk field combination in employee performance appraisal data, calculating the sensitivity and permissions of the query party, and generating a joint desensitization strategy, the problem of insufficient risk prevention and control of field combinations in the existing technology is solved, and efficient and safe data management is achieved.

CN120372691BActive Publication Date: 2025-08-19HANGZHOU JINYUAN BIAOJU TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510885154.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-19
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

The existing technology lacks an effective prevention and control mechanism for field combination risks in employee performance appraisal data management, resulting in high risk of data leakage and increased system complexity and cost.

Method used

The field analysis module is used to identify the risk field combination, and the sensitivity and permissions of each query party are calculated through the query analysis module, a joint desensitization strategy is generated, and the desensitization operation is performed through the desensitization execution module, optimizing the desensitization strategy to meet the needs of different query parties.

Benefits of technology

It realizes accurate desensitization of employee performance appraisal data, reduces data leakage risks, reduces storage costs and system complexity, and improves the flexibility and security of data use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372691B_ABST
    Figure CN120372691B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of data processing, and discloses a dynamic data desensitization system for employee performance appraisal, including a field analysis module, a query analysis module, a desensitization strategy module, and a desensitization execution module; wherein: the field analysis module is used to identify risk field combinations between fields of object data; the query analysis module is used to calculate the sensitivity of each query party to each field in the risk field combination; the desensitization strategy module synchronously generates a joint desensitization strategy for each query party for the risk field combination based on the sensitivity of each query party to each field in the risk field combination and the desensitization constraint conditions; the desensitization execution module is used to execute the joint desensitization strategy, and perform joint desensitization of the risk field combination for each query party. The present application optimizes the desensitization strategy and improves the efficiency and security of employee performance appraisal data management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and specifically to a dynamic data desensitization system for employee performance appraisal. Background Art

[0002] The current industry still has many shortcomings in employee performance appraisal data management. Existing technologies are significantly inadequate in preventing the risk of data correlation leakage. As enterprises' data analysis needs deepen, the correlations between multiple fields can lead to potential information leakage. Most solutions are still limited to protecting single fields and fail to establish effective prevention and control mechanisms for risks associated with field combinations. Existing data masking methods are often not precise or dynamic enough. Some enterprises use simple static masking methods, such as uniformly masking all object data or deleting sensitive fields. While this method can protect data privacy to a certain extent, it reduces data availability and cannot meet the diverse data needs of different querying parties. For example, when reviewing employee performance data for their department, a department manager may need to understand the general distribution of employee performance. However, statically masked data may not provide sufficient information, affecting the department manager's management decision-making efficiency. The existing data masking strategy development process lacks global optimization considerations. Most companies independently calculate the most advantageous desensitization strategy for each querying party without considering the needs of other querying parties. As a result, the same risk field combination requires generating multiple versions of desensitization strategies for different querying parties. This not only increases storage costs, but also requires maintaining multiple sets of desensitization processes, increasing system complexity and operating costs.

[0003] For example, the Chinese patent application with publication number CN116502265A discloses a data security system based on big data, which involves the field of data security and includes a monitoring center, wherein the monitoring center is connected to an account management module, a virtual system module, a risk assessment module, and a data security module; the account management module helps employees register accounts, establish a mirror virtual machine, and employees log in, verify the IP address of the employee's computer, and map the company's website data to the employee's computer. The risk assessment module assesses the risk level of the employee's copy operation behavior, and takes corresponding measures on the employee's computer according to the risk level. If the risk level is level three, the core data is desensitized and a key is set. This technical solution can prevent the data within the company from being leaked. At the same time, if the desensitized core data is leaked, it cannot be viewed, which ensures the security of the core data. However, there are still problems raised in the background technology of this application:

[0004] The above patents all have the problem raised by this background technology: they fail to establish an effective prevention and control mechanism for field combination risks.

[0005] The information disclosed in this background technology section is only intended to enhance the understanding of the overall background of the application and should not be regarded as an admission or any form of suggestion that the information constitutes the prior art already known to ordinary technicians in this field. Summary of the Invention

[0006] The technical problem to be solved by this application is to overcome the defects of the existing technology, provide a dynamic data desensitization system for employee performance appraisal, optimize the desensitization strategy, and improve the efficiency and security of employee performance appraisal data management.

[0007] To solve the above technical problems, this application provides the following technical solutions:

[0008] A dynamic data desensitization system for employee performance appraisal includes a field analysis module, a query analysis module, a desensitization strategy module, and a desensitization execution module; wherein:

[0009] The field analysis module is used to identify risky field combinations between fields in object data;

[0010] The query analysis module is used to calculate the sensitivity of each querying party to each field in the risk field combination; the query analysis module is also used to obtain the query constraint authority of each querying party and map the query constraint authority to the desensitization constraint condition of the querying party for each field;

[0011] The desensitization strategy module generates a joint desensitization strategy for each query party based on the sensitivity of each field in the risk field combination and the desensitization constraints of each query party.

[0012] The desensitization execution module is used to execute the joint desensitization strategy and perform joint desensitization of risk field combinations for each querying party.

[0013] As a preferred solution of the dynamic data desensitization system for employee performance appraisal described in this application, wherein: the field analysis module includes an association combination unit;

[0014] The association combination unit is configured with an association identification strategy; the association identification strategy is used to identify the association field combination between the fields of the object data; specifically includes:

[0015] Obtain an object dataset; calculate each marginal probability of each field and each joint marginal probability of any m fields based on the object dataset; m is an integer greater than 1;

[0016] Based on each marginal probability of each field and each joint marginal probability of any m fields, the mutual information value of any m fields is calculated as the correlation value of the corresponding m fields;

[0017] The association combination unit is further configured with an association threshold; the association identification strategy further includes: if the association value of any m fields is greater than the association threshold, the corresponding m fields are an association field combination.

[0018] As a preferred solution of the dynamic data desensitization system for employee performance appraisal described in this application, wherein: the field analysis module also includes a risk identification unit;

[0019] The risk identification unit is configured with a risk identification strategy; the risk identification strategy is used to identify risk field combinations in the associated field combinations, specifically including:

[0020] For any combination of associated fields, count the number of people in all sub-combinations; among all sub-combinations with non-zero number of people, take the number of people in the sub-combination with the least number of people as the anonymity of the corresponding associated field combination;

[0021] The risk identification unit is further configured with an anonymity threshold; the risk identification strategy further comprises: if the anonymity of any associated field combination is less than the anonymity threshold, marking the corresponding associated field combination as a risky field combination.

[0022] As a preferred solution of the dynamic data desensitization system for employee performance appraisal described in this application, wherein: the query analysis module includes a sensitivity unit;

[0023] The sensitivity unit is used to calculate the sensitivity of each querying party to each field in the risk field combination; the querying party includes HR, department managers, and senior executives;

[0024] The sensitivity unit calculates the sensitivity of the querying party to each field in the risk field combination in the following manner: obtaining the desensitization operation records of the querying party for each field in each risk field combination, and counting the proportion of desensitization operation records with a desensitization intensity level of 3 in the desensitization operation records as the sensitivity of the corresponding field; the desensitization intensity levels include level 1, level 2, and level 3, and the larger the value of the desensitization intensity level, the higher the desensitization intensity.

[0025] As a preferred solution of the dynamic data desensitization system for employee performance appraisal described in this application, wherein: the query analysis module further includes a desensitization constraint unit;

[0026] The desensitizing constraint unit is used to obtain the query constraint authority of each query party; the query constraint authority is the authority range of the corresponding query party to view the object data;

[0027] The desensitization constraint unit is further configured to map the query constraint authority to a desensitization constraint condition for each field of the query party; the desensitization constraint condition includes a constraint condition for the desensitization intensity level of the corresponding field;

[0028] The desensitizing constraint unit is configured with a mapping rule; based on the mapping rule, the desensitizing constraint unit maps the query constraint authority to the desensitizing constraint condition of the query party for each field.

[0029] As a preferred solution of the dynamic data desensitization system for employee performance appraisal described in this application, wherein: the desensitization strategy module includes a joint optimization unit and a calculation unit; wherein the calculation unit is configured with a benefit calculation algorithm; the joint optimization unit is used to construct a joint desensitization strategy set; the calculation unit calculates the benefit value of each query party in each joint desensitization strategy based on the benefit calculation algorithm; the joint optimization unit generates a Pareto frontier of the joint desensitization strategy based on the benefit value of each query party, and selects a joint desensitization strategy based on the Pareto frontier.

[0030] The joint desensitization strategy set includes all joint desensitization strategies; any joint desensitization strategy includes specifying the desensitization risk level of each field in the risk field combination for each querying party;

[0031] The joint optimization unit obtains the desensitization constraint conditions of each field set by the querying party, and prunes the joint desensitization strategy set to eliminate the joint desensitization strategies that do not meet the desensitization constraint conditions.

[0032] As a preferred solution of the dynamic data desensitization system for employee performance appraisal described in this application, wherein: the calculation unit calculates the profit value of each querying party based on the profit calculation algorithm, specifically including:

[0033] Calculate the data availability of each querying party after executing the joint desensitization strategy;

[0034] Calculate the privacy risk value of each querying party after executing the joint desensitization strategy;

[0035] The data availability and privacy risk value of each querying party are weighted and summed to obtain the benefit value of each querying party, wherein the weight coefficient of data availability is positive and the weight coefficient of privacy risk value is negative.

[0036] As a preferred solution of the dynamic data desensitization system for employee performance appraisal described in this application, the profit calculation algorithm also includes calculating the data availability of any querying party, specifically including:

[0037] The data availability of each field is calculated and summed to obtain the data availability of the querying party; wherein, the method of calculating the data availability of any field includes: calculating the statistical error after desensitization of the field and normalizing it; the calculation unit is also configured with an error threshold for each field of each querying party; the normalized value of the statistical error of each field is divided by the error threshold of the corresponding field to obtain the error rate; the data availability of the field is 1 minus the error rate.

[0038] The profit calculation algorithm also includes calculating the privacy risk value of any querying party, specifically including:

[0039] Calculate the privacy risk value of each field after desensitization and sum them up to obtain the privacy risk value of the querying party. The privacy risk value of any field is the ratio of the querying party's sensitivity to the field to the desensitization intensity level of the field.

[0040] As a preferred solution of the dynamic data desensitization system for employee performance appraisal described in this application, the joint optimization unit generates a Pareto frontier of the joint desensitization strategy based on the benefit value of each query party, specifically including:

[0041] Generate a benefit value sequence for each joint desensitization strategy; the benefit value sequence includes the benefit value of each query party in the joint desensitization strategy;

[0042] For any two payoff value sequences P and Q, if every payoff value in P is not less than the corresponding item in Q, and at least one payoff value in P is greater than the corresponding item in Q, then P is the dominated solution of Q, and P is a non-dominated solution; if there is no dominated solution in P, then P is a non-dominated solution; the joint desensitization strategies corresponding to all non-dominated solutions constitute the Pareto frontier of the joint desensitization strategy;

[0043] The joint optimization unit is further configured to obtain query records of each querying party; based on the query records, count the frequency of each querying party simultaneously querying all fields in the risk field combination, and select the querying party with the highest frequency as the leading querying party of the risk field combination;

[0044] Based on the benefit value sequence of each joint desensitization strategy in the Pareto front, the joint desensitization strategy with the largest benefit value for the dominant query party is located as the joint desensitization strategy selected by the joint optimization unit.

[0045] As a preferred solution of the dynamic data desensitization system for employee performance appraisal described in this application, the joint optimization unit generates a Pareto frontier of the joint desensitization strategy based on the benefit value of each query party, specifically including:

[0046] Generate a benefit value sequence for each joint desensitization strategy; the benefit value sequence includes the benefit value of each query party in the joint desensitization strategy;

[0047] For any two payoff value sequences P and Q, if every payoff value in P is not less than the corresponding item in Q, and at least one payoff value in P is greater than the corresponding item in Q, then P is the dominated solution of Q, and P is a non-dominated solution; if there is no dominated solution in P, then P is a non-dominated solution; the joint desensitization strategies corresponding to all non-dominated solutions constitute the Pareto frontier of the joint desensitization strategy;

[0048] The joint optimization unit is further configured to obtain query records of each querying party; based on the query records, count the frequency of each querying party simultaneously querying all fields in the risk field combination, and select the querying party with the highest frequency as the leading querying party of the risk field combination;

[0049] Based on the benefit value sequence of each joint desensitization strategy in the Pareto front, the joint desensitization strategy with the largest benefit value for the dominant query party is located as the joint desensitization strategy selected by the joint optimization unit.

[0050] Compared with the prior art, the beneficial effects achieved by this application are as follows:

[0051] This application determines the associated field combinations by calculating the mutual information values between fields, and then calculates the anonymity to filter out risky field combinations. It can accurately locate field combinations that are frequently queried and prone to leaking employee information, providing a basis for subsequent targeted desensitization.

[0052] The field sensitivity is calculated for different query parties, and the query constraint permissions are mapped to desensitization constraints. This allows the system to customize the desensitization strength of each field based on the needs and permissions of each query party, thereby meeting the query party's reasonable use of data, preventing the leakage of sensitive information, and improving the flexibility and security of data use.

[0053] Compared with the traditional independent calculation of desensitization strategies, the multi-query joint optimization method of this application avoids the situation where multiple versions of desensitization strategies are generated for the same risk field combination, reduces storage costs, reduces system complexity, and improves operational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be derived from these drawings without inventive effort. Among them:

[0055] Figure 1 A schematic diagram of the structure of the dynamic data desensitization system for employee performance appraisal provided for this application;

[0056] Figure 2 Workflow diagram of the dynamic data desensitization system for employee performance appraisal provided for this application. DETAILED DESCRIPTION

[0057] The technical solution of the present application is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Unless there is a conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.

[0058] This embodiment introduces a dynamic data desensitization system for employee performance appraisal. Figure 1 The system includes a field analysis module, a query analysis module, a desensitization strategy module, and a desensitization execution module; Figure 2 ,The work of each module is detailed as follows.

[0059] The field analysis module is used to identify risky field combinations between fields in object data;

[0060] The field analysis module includes an association combination unit and a risk identification unit;

[0061] The association combination unit is configured with an association identification strategy; the association identification strategy is used to identify the association field combination between the fields of the object data; specifically includes:

[0062] Obtain an object dataset; calculate each marginal probability of each field and each joint marginal probability of any m fields based on the object dataset; m is an integer greater than 1;

[0063] The object dataset includes all fields of all assessment objects, the values of each field, and historical query records. The marginal probability of any field is the probability of that field's corresponding value appearing in historical query records. For example, if one value of the field "Department" is Sales Department, the corresponding marginal probability is the ratio of the number of times Sales Department appears in historical query records to the number of times the field "Department" appears in historical query records. If one value of the field "Performance Level" is Grade A, the joint marginal probability of the fields "Department" and "Performance Level" is the ratio of the number of times Sales Department and Grade A appear together in a historical query record to the number of times the fields "Department" and "Performance Level" appear together in a historical query record.

[0064] Based on each edge probability of each field and each joint edge probability of any m fields, the mutual information value of any m fields is calculated as the association value of the corresponding m fields; the mutual information value is calculated based on the calculation formula of mutual information in information theory.

[0065] The association combination unit is also configured with an association threshold. The association identification strategy further includes: if the association value of any m fields is greater than the association threshold, then the corresponding m fields are considered an association field combination. For example, if the association field combination consists of the fields "department" + "performance level", then m = 2.

[0066] The preferred fields of partial object data in this embodiment include the name, work number, department, performance level (A / B / C / D), years of service, salary, project team number, performance score (0-100 points), and reward and punishment records of the assessment object.

[0067] The risk identification unit is configured with a risk identification strategy; the risk identification strategy is used to identify risk field combinations in the associated field combinations, specifically including:

[0068] For any combination of associated fields, the number of people in all subcombinations is counted. The number of people in the subcombination with the fewest people among all subcombinations with non-zero numbers is used as the anonymity level for that combination. For example, for the combination "Department" + "Performance Level," if Department's values include Sales and Technology, and Performance Level's values include A and B, then there are four possible combinations of any Department value with any Performance Level value, meaning there are four subcombinations for this associated field. If, in this subcombination, the number of people associated with Sales Department + Level A is 8 (i.e., among all employees, 8 have both a Department and an A Performance Level), the number of people associated with Technology Department + Level A is 5, and the number of people in all other subcombinations is 0, then the anonymity level for the combination "Department" + "Performance Level" is 5. Anonymity is a direct measure of privacy protection strength. A higher anonymity level indicates better privacy protection, ensuring that the associated field combination cannot be accurately attributed to individual employees.

[0069] The risk identification unit is further configured with an anonymity threshold; the risk identification strategy further comprises: if the anonymity of any associated field combination is less than the anonymity threshold, marking the corresponding associated field combination as a risky field combination.

[0070] This example uses two rounds of screening: first, by comparing the correlation values of multiple fields, we identify frequently queried field combinations. Then, we calculate the anonymity level to identify field combinations that are likely to leak employee information. Highly queried and potentially leaking employee information require joint desensitization.

[0071] The query analysis module is used to calculate the sensitivity of each querying party to each field in the risk field combination; the query analysis module is also used to obtain the query constraint authority of each querying party and map the query constraint authority to the desensitization constraint condition of the querying party for each field;

[0072] The query analysis module includes a sensitivity unit and a desensitization constraint unit;

[0073] The sensitivity unit is used to calculate the sensitivity of each querying party to each field in the risk field combination; the querying party includes HR, department managers, and senior executives;

[0074] The sensitivity unit calculates the sensitivity of the querying party to each field in the risk field combination in the following manner: obtaining the desensitization operation records of the querying party for each field in each risk field combination, and counting the proportion of desensitization operation records with a desensitization intensity level of 3 in the desensitization operation records as the sensitivity of the corresponding field; the desensitization intensity levels include level 1, level 2, and level 3, and the larger the value of the desensitization intensity level, the higher the desensitization intensity.

[0075] Sensitivity quantifies the confidentiality level of each field. If a field is frequently masked at level 3, it indicates that the field is highly sensitive and therefore has a high confidentiality level for the querying party.

[0076] The desensitizing constraint unit is used to obtain the query constraint authority of each query party; the query constraint authority is the authority scope of the corresponding query party to view the object data; some of the preferred query constraint authorities in this embodiment are as follows: for HR, it is allowed to view department-level statistical values (such as average performance); it is prohibited to view personal detailed data; for department managers, it is allowed to view the desensitized performance range of the top 50% of employees in the department; cross-departmental data association is prohibited.

[0077] The desensitization constraint unit is further configured to map the query constraint authority to a desensitization constraint condition for each field of the query party; the desensitization constraint condition includes a constraint condition for the desensitization intensity level of the corresponding field;

[0078] The desensitization constraint unit is configured with a mapping rule; the desensitization constraint unit maps the query constraint authority to the desensitization constraint condition of each field of the query party based on the mapping rule. Some of the preferred mapping rules of this embodiment are as follows: the query constraint authority is to prohibit the viewing of personal detailed data, and the mapped desensitization constraint condition is that the desensitization intensity level of the fields "name" and "employee number" is 3, that is, they must be completely desensitized; the query constraint authority is to allow the viewing of the desensitized performance range of the top 50% of employees in the department, and the mapped desensitization constraint condition is that the desensitization level of the field "performance score" is not greater than 2, that is, the performance score is allowed to be desensitized to a certain extent, but the ranking is retained.

[0079] The desensitization strategy module generates a joint desensitization strategy for each query party based on the sensitivity of each field in the risk field combination and the desensitization constraints of each query party.

[0080] The desensitization strategy module includes a joint optimization unit and a calculation unit; wherein the calculation unit is configured with a benefit calculation algorithm; the joint optimization unit is used to construct a joint desensitization strategy set; the calculation unit calculates the benefit value of each query party in each joint desensitization strategy based on the benefit calculation algorithm; the joint optimization unit generates a Pareto frontier of the joint desensitization strategy based on the benefit value of each query party, and selects a joint desensitization strategy based on the Pareto frontier;

[0081] The joint desensitization policy set includes all joint desensitization policies; any joint desensitization policy includes specifying the desensitization risk level of each field in the risk field combination for each querying party; for example, a joint desensitization policy for "department + performance score" includes: HR's desensitization intensity level for departments is level 2, and the desensitization intensity level for performance scores is level 3; department managers' desensitization intensity level for departments is level 2, and the desensitization intensity level for performance scores is level 2; senior executives' desensitization intensity level for departments is level 2, and the desensitization intensity level for performance scores is level 1;

[0082] The joint optimization unit obtains the query party's desensitization constraints for each field and prunes the set of joint desensitization strategies, removing those that do not meet the constraints. For example, if a department manager's desensitization constraint is that the performance score's desensitization intensity level is no greater than 2, all joint desensitization strategies that include a department manager's desensitization intensity level of 3 for the performance score will be removed.

[0083] The calculation unit calculates the profit value of each querying party based on the profit calculation algorithm, specifically including:

[0084] Calculate the data availability of each querying party after executing the joint desensitization strategy;

[0085] Calculate the privacy risk value of each querying party after executing the joint desensitization strategy;

[0086] The data availability and privacy risk value of each querying party are weighted and summed to obtain the benefit value for each querying party, where the weight coefficient for data availability is positive and the weight coefficient for privacy risk is negative. The specific weight coefficients are set by those skilled in the art based on actual needs. The benefit value is used to quantify the satisfaction of each querying party with any joint desensitization strategy. Its core goal is to balance the availability of desensitized data with preventing the leakage of sensitive information.

[0087] The revenue calculation algorithm also includes calculating the data availability of any querying party, specifically including:

[0088] The data availability of each field is calculated and summed to obtain the data availability for the querying party. The method for calculating the data availability of any field includes: calculating and normalizing the statistical error after masking the field; the calculation unit is further configured with an error threshold for each field for each querying party; dividing the normalized value of the statistical error for each field by the error threshold for the corresponding field to obtain an error rate; and the data availability of a field is calculated as 1 minus the error rate. The greater the statistical error after masking, the higher the degree of data distortion and the lower the availability. For example, for HR, the error threshold for performance is 0.8; for executives, the error threshold for performance is 0.6. A larger error threshold indicates a greater allowable error introduced by the masking operation, i.e., a lower requirement for data accuracy. If the normalized value of the statistical error after masking performance is 0.2, the calculated data availability for HR is 0.75; for executives, the calculated data availability is 0.67. Therefore, the masked field has a higher availability for HR, indicating that the lower the data accuracy requirement, the higher the data availability. Conversely, a higher statistical error after masking indicates a lower data availability.

[0089] In this embodiment, the method for calculating the statistical error after field desensitization preferably includes:

[0090] For numeric fields, such as performance scores and salaries, the numerical error rate caused by masking is calculated as the statistical error. For example, if the average performance score of the original data is 85 and the expected performance score after masking is 83, the statistical error of the performance score is 85 minus 83, the absolute value of which is then divided by 85.

[0091] For categorical fields, the Jensen-Shannon divergence (JSD) of the category distribution before and after desensitization is calculated as its statistical error. For example, for a department, the original distribution is that the sales department accounts for 0.3 people, the technical department accounts for 0.5 people, and the marketing department accounts for 0.2 people. After desensitizing the departments into regions, the desensitized distribution is that the East China region accounts for 0.55 people and the Western region accounts for 0.45 people. The JSD between the original distribution and the desensitized distribution is calculated as the statistical error of the department after desensitization.

[0092] For text fields, the ratio of the entropy value of the masked field to the entropy value before and after masking is calculated as the statistical error. For example, if the original value of a reward and punishment record is "Outstanding Employee of the Year Award, Quarterly Sales Champion," and the masked value is "Company Award Received," the entropy value before and after masking is calculated using the Shannon entropy formula to determine the statistical error for the field.

[0093] The profit calculation algorithm also includes calculating the privacy risk value of any querying party, specifically including:

[0094] The privacy risk value of each masked field is calculated and summed to obtain the privacy risk value of the querying party. The privacy risk value of any field is the ratio of the querying party's sensitivity to the field to the masking intensity level of the field. The higher the sensitivity, the higher the privacy protection requirements for the relevant field; the higher the masking intensity level, the lower the risk of privacy leakage.

[0095] The joint optimization unit generates the Pareto frontier of the joint desensitization strategy based on the benefit value of each query party, which includes:

[0096] Generate a benefit value sequence for each joint desensitization strategy; the benefit value sequence includes the benefit value of each query party in the joint desensitization strategy;

[0097] For any two payoff value sequences P and Q, if each payoff value in P is not less than the corresponding item in Q, and at least one payoff value in P is greater than the corresponding item in Q, then P is the dominated solution of Q, and P is a non-dominated solution; if there is no dominated solution in P, then P is a non-dominated solution; the joint desensitization strategies corresponding to all non-dominated solutions constitute the Pareto frontier of the joint desensitization strategy.

[0098] The joint optimization unit selects a joint desensitization strategy based on the Pareto front, specifically including:

[0099] Obtain query records of each querying party; based on the query records, count the frequency of each querying party simultaneously querying all fields in the risk field combination, and select the querying party with the highest frequency as the leading querying party of the risk field combination;

[0100] Based on the benefit value sequence of each joint desensitization strategy in the Pareto front, the joint desensitization strategy with the largest benefit value for the dominant query party is located as the joint desensitization strategy selected by the joint optimization unit.

[0101] This embodiment generates a corresponding joint desensitization strategy for each risk field combination through joint optimization of multiple query parties, thereby allocating a desensitization level strength for each field to each query party at one time.

[0102] The existing technology independently calculates the most favorable desensitization strategy for each querying party without considering the needs of other querying parties. Therefore, the same risk field combination needs to generate multiple versions of desensitization strategies for different querying parties, resulting in an exponential increase in storage costs and the need to maintain multiple sets of desensitization processes. Compared with the joint optimization of this embodiment, the system complexity is higher, and the operating efficiency and maintenance costs are high.

[0103] The desensitization execution module is used to execute the joint desensitization strategy and perform joint desensitization of risk field combinations for each querying party.

[0104] The desensitization execution module includes a desensitization algorithm unit and a strategy execution unit;

[0105] The desensitization algorithm unit is configured with a desensitization algorithm corresponding to each desensitization intensity level; the strategy execution unit is used to respond to and execute the joint desensitization strategy, and perform joint desensitization of risk field combinations for each querying party by calling the corresponding desensitization algorithm.

[0106] An example of the partial desensitization algorithm preferred in this embodiment and its joint desensitization of the risk field combination of the querying party is as follows: the querying party is a department manager, who requests to view the distribution of years of service of the top 10% employees in the department, and the risk field combination involved is "performance score + years of service + department"; based on the joint desensitization strategy, the desensitization algorithm is called to jointly desensitize the risk field combination, specifically including: executing a desensitization algorithm with a desensitization intensity level of 3 on the performance score, adding Laplace noise to it to hide the specific value; executing a desensitization algorithm with a desensitization intensity level of 2 on the years of service, generalizing the years of service to "<3 years", "3-5 years", and ">5 years"; executing a desensitization algorithm with a desensitization intensity level of 1 on the department, maintaining its original value, because the querying party's authority allows the department to query.

[0107] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0108] The above describes the embodiments of the present application in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose and scope of protection of this application, all of which are protected by this application.

Claims

1. A dynamic data desensitization system for employee performance appraisal, characterized by: It includes field analysis module, query analysis module, desensitization strategy module, and desensitization execution module; among them: The field analysis module is used to identify risky field combinations between fields in object data; The query analysis module is used to calculate the sensitivity of each querying party to each field in the risk field combination; the query analysis module is also used to obtain the query constraint authority of each querying party and map the query constraint authority to the desensitization constraint condition of the querying party for each field; The desensitization strategy module generates a joint desensitization strategy for each query party based on the sensitivity of each field in the risk field combination and the desensitization constraints of each query party. The desensitization execution module is used to execute the joint desensitization strategy, and perform joint desensitization of risk field combinations for each query party; The desensitization strategy module includes a joint optimization unit and a calculation unit; wherein the calculation unit is configured with a benefit calculation algorithm; the joint optimization unit is used to construct a joint desensitization strategy set; the calculation unit calculates the benefit value of each query party in each joint desensitization strategy based on the benefit calculation algorithm; the joint optimization unit generates a Pareto frontier of the joint desensitization strategy based on the benefit value of each query party, and selects a joint desensitization strategy based on the Pareto frontier; The joint desensitization strategy set includes all joint desensitization strategies; any joint desensitization strategy includes specifying the desensitization risk level of each field in the risk field combination for each querying party; The joint optimization unit obtains the desensitization constraint conditions of each field set by the querying party, and prunes the joint desensitization strategy set to eliminate the joint desensitization strategies that do not meet the desensitization constraint conditions.

2. The dynamic data desensitization system for employee performance appraisal according to claim 1, characterized in that: The field analysis module includes an association combination unit; The association combination unit is configured with an association identification strategy; The association identification strategy is used to identify the association field combination between the fields of the object data; specifically, it includes: Obtain an object dataset; calculate each marginal probability of each field and each joint marginal probability of any m fields based on the object dataset; m is an integer greater than 1; Based on each marginal probability of each field and each joint marginal probability of any m fields, the mutual information value of any m fields is calculated as the correlation value of the corresponding m fields; The association combination unit is further configured with an association threshold; the association identification strategy further includes: if the association value of any m fields is greater than the association threshold, the corresponding m fields are an association field combination.

3. The dynamic data desensitization system for employee performance appraisal according to claim 2, characterized in that: The field analysis module also includes a risk identification unit; The risk identification unit is configured with a risk identification strategy; The risk identification strategy is used to identify risk field combinations in the associated field combinations, specifically including: For any combination of associated fields, count the number of people in all sub-combinations; among all sub-combinations with non-zero number of people, take the number of people in the sub-combination with the least number of people as the anonymity of the corresponding associated field combination; The risk identification unit is further configured with an anonymity threshold; the risk identification strategy further comprises: if the anonymity of any associated field combination is less than the anonymity threshold, marking the corresponding associated field combination as a risky field combination.

4. The dynamic data desensitization system for employee performance appraisal according to claim 1, characterized in that: The query analysis module includes a sensitivity unit; The sensitivity unit is used to calculate the sensitivity of each querying party to each field in the risk field combination; The sensitivity unit calculates the sensitivity of the querying party to each field in the risk field combination in the following manner: obtaining the desensitization operation record of the querying party for each field in each risk field combination, and counting the proportion of the desensitization operation records with a desensitization intensity level of 3 in the desensitization operation records as the sensitivity of the corresponding field; The desensitization intensity levels include level 1, level 2, and level 3. The larger the value of the desensitization intensity level, the higher the desensitization intensity.

5. The dynamic data desensitization system for employee performance appraisal according to claim 4, characterized in that: The query analysis module also includes a desensitization constraint unit; The desensitizing constraint unit is used to obtain the query constraint authority of each query party; the query constraint authority is the authority range of the corresponding query party to view the object data; The desensitization constraint unit is further configured to map the query constraint authority to a desensitization constraint condition for each field of the query party; the desensitization constraint condition includes a constraint condition for the desensitization intensity level of the corresponding field; The desensitizing constraint unit is configured with a mapping rule; based on the mapping rule, the desensitizing constraint unit maps the query constraint authority to the desensitizing constraint condition of the query party for each field.

6. The dynamic data desensitization system for employee performance appraisal according to claim 5, characterized in that: The calculation unit calculates the profit value of each querying party based on the profit calculation algorithm, specifically including: Calculate the data availability of each querying party after executing the joint desensitization strategy; Calculate the privacy risk value of each querying party after executing the joint desensitization strategy; The data availability and privacy risk value of each querying party are weighted and summed to obtain the benefit value of each querying party, wherein the weight coefficient of data availability is positive and the weight coefficient of privacy risk value is negative.

7. The dynamic data desensitization system for employee performance appraisal according to claim 6, characterized in that: The revenue calculation algorithm also includes calculating the data availability of any querying party, specifically including: Calculating the data availability of each field and summing the data to obtain the data availability of the querying party; wherein the method of calculating the data availability of any field includes: calculating the statistical error after desensitization of the field and normalizing it; the calculation unit is further configured with an error threshold for each field of each querying party; dividing the normalized value of the statistical error of each field by the error threshold of the corresponding field to obtain an error rate; the data availability of the field is 1 minus the error rate; The profit calculation algorithm also includes calculating the privacy risk value of any querying party, specifically including: Calculate the privacy risk value of each field after desensitization and sum them up to obtain the privacy risk value of the querying party. The privacy risk value of any field is the ratio of the querying party's sensitivity to the field to the desensitization intensity level of the field.

8. The dynamic data desensitization system for employee performance appraisal according to claim 7, characterized in that: The joint optimization unit generates the Pareto frontier of the joint desensitization strategy based on the benefit value of each query party, which includes: Generate a benefit value sequence for each joint desensitization strategy; the benefit value sequence includes the benefit value of each query party in the joint desensitization strategy; For any two payoff value sequences P and Q, if every payoff value in P is not less than the corresponding item in Q, and at least one payoff value in P is greater than the corresponding item in Q, then P is the dominated solution of Q, and P is a non-dominated solution; if there is no dominated solution in P, then P is a non-dominated solution; the joint desensitization strategies corresponding to all non-dominated solutions constitute the Pareto frontier of the joint desensitization strategy; The joint optimization unit is further configured to obtain query records of each querying party; based on the query records, count the frequency of each querying party simultaneously querying all fields in the risk field combination, and select the querying party with the highest frequency as the leading querying party of the risk field combination; Based on the benefit value sequence of each joint desensitization strategy in the Pareto front, the joint desensitization strategy with the largest benefit value for the dominant query party is located as the joint desensitization strategy selected by the joint optimization unit.

9. The dynamic data desensitization system for employee performance appraisal according to claim 1, characterized in that: The desensitization execution module includes a desensitization algorithm unit and a strategy execution unit; The desensitization algorithm unit is configured with a desensitization algorithm corresponding to each desensitization intensity level; The policy execution unit is used to respond to and execute the joint desensitization policy, and perform joint desensitization of risk field combinations for each querying party by calling a corresponding desensitization algorithm.

Citation Information

Patent Citations

  • Data security system based on big data

    CN116502265A

  • Sensitive data transmission method and device, equipment and storage medium

    CN117834514A

  • Self-adaptive personal information desensitization method and system based on availability evaluation

    CN119622815A