Data medium table-based field-level automatic desensitization method and system
By combining a multi-dimensional sensitivity scoring and grading algorithm and a multi-objective optimization model with an improved NSGA-II algorithm, the problems of low accuracy in sensitive field identification and desensitization strategies in the power sector are solved, achieving efficient and compliant field-level automatic desensitization, which is suitable for power data platform environments.
Patent Information
- Application Number
- CN202510675908.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-10-28
AI Technical Summary
Existing technologies have low accuracy in identifying sensitive fields in power scenarios, and the desensitization strategies are difficult to balance privacy protection and data availability. Furthermore, they lack automation and consistency, and cannot meet the needs of cross-system and cross-platform data platform environments.
By employing a multidimensional sensitivity scoring and grading algorithm and a multi-objective optimization model, and by improving the NSGA-II algorithm to search for the optimal desensitization strategy, and combining uniqueness, information entropy and correlation indicators, the algorithm dynamically adapts to the needs of power business and achieves automatic field-level desensitization.
It improves the comprehensiveness and accuracy of sensitive field identification, generates Pareto optimal solution sets, provides diverse desensitization schemes, is suitable for large-scale power datasets, reduces manual intervention and error risks, and improves data utilization efficiency.
Smart Images

Figure CN120850329A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data security technology, and specifically relates to a field-level automatic de-identification method and system based on a data middle platform. Background Technology
[0002] In the information age, data has become a vital asset for all industries. As data security regulations in various countries and regions become increasingly stringent, strict requirements for the protection of personal information have been imposed. However, sensitive information contained within data, such as personal privacy and trade secrets, also faces the risk of leakage. To ensure data security and protect personal privacy, enterprises need to effectively protect sensitive data.
[0003] Traditional data masking techniques generally suffer from two major pain points: First, static rule bases struggle to adapt to dynamic business scenarios, resulting in coarse-grained masking. This either excessively disrupts data relationships, impacting analytical value, or overlooks the risk of data leakage due to cross-field combinations. Second, existing masking methods lack synergistic optimization of data characteristics and business needs, failing to balance the contradictions between privacy protection strength, data availability, and processing costs. This is particularly problematic in heavily regulated industries like power, where complex scenarios involving intertwined equipment and user behavior data can easily lead to compliance issues or business disruptions with a single masking strategy. Furthermore, cross-system, cross-platform data platform environments require automated and standardized masking processes, while traditional methods relying on manual annotation or fixed rules struggle to meet real-time, consistency, and scalability requirements. Therefore, a field-level automatic masking method based on a data platform is urgently needed to achieve the dual goals of data security and value release through intelligent hierarchical classification and dynamic optimization.
[0004] Chinese patent CN112613069A discloses an automatic data anonymization method based on negative list data resources, including automatic data anonymization and data security protection. The automatic data anonymization includes strategy setting, rule setting, work order management, data scanning, anonymization review, data anonymization, and data verification. The data security protection includes data traceability, tamper management, anomaly query, and statistical query. This invention achieves the filtering of sensitive information by matching and setting de-anonymization strategies and rules to the data applied for in work orders through a series of steps such as scanning, reviewing, anonymizing, and issuing. However, this invention relies on manual review and management, has a low degree of automation, low accuracy in sensitive data identification, and lacks multi-objective optimization capabilities. Summary of the Invention
[0005] The purpose of this invention is to provide a field-level automatic desensitization method and system based on a data platform to solve the problems of low accuracy in identifying sensitive fields in the power sector and difficulty in balancing privacy protection and data availability in desensitization strategies. Through a multi-dimensional sensitivity scoring and grading algorithm and a multi-objective optimization model, it achieves efficient and compliant field-level automatic desensitization that dynamically adapts to the needs of power business.
[0006] The technical solution of the present invention is as follows: On the one hand, this invention provides a field-level automatic de-identification method based on a data middle platform, comprising the following steps: Data on power users and equipment is collected through various power management systems to construct a power dataset, which is then sent to a data platform for centralized and automatic anonymization.
[0007] For all power user and equipment data in the power dataset, a multi-dimensional sensitivity scoring and grading algorithm based on uniqueness, information entropy, correlation, and compliance is used to determine whether the power user and equipment data fields are sensitive fields, and the power user and equipment data corresponding to the sensitive fields are classified according to sensitivity.
[0008] For power user and equipment data and their sensitivity level information that are identified as sensitive fields, a multi-objective optimization model is used to perform multi-objective optimization through an improved NSGA-II algorithm to search for the optimal desensitization strategy. Based on the optimal desensitization strategy, the corresponding data desensitization method is executed for each sensitive field.
[0009] Preferably, the multidimensional sensitivity scoring and grading algorithm calculates the multidimensional sensitivity score based on three-dimensional quantitative indicators: uniqueness, information entropy, and correlation. The uniqueness index is expressed as:
[0010] In the formula, For power user and equipment data Uniqueness indicators for each field; For power user and equipment data The number of unique values in each field is obtained by counting the number of non-repeating values in the statistical data field. For power user and equipment data The total number of records for each field.
[0011] The information entropy index is expressed as:
[0012] In the formula, For power user and equipment data Information entropy index for each field; For power user and equipment data Different values in each field The set of occurrences; value Frequency of occurrence.
[0013] The correlation index is expressed as follows:
[0014] In the formula, For power user and equipment data The correlation index of each field; For power user and equipment data The first field combines power user and equipment data. The number of unique values in each field; For power user and equipment data , The total number of records for each field; To iterate through the electricity dataset except for the first... Other power user and equipment data fields in the first field are the same as the first field. The highest proportion of unique values after combining multiple fields; For power user and equipment data The network weight coefficient of each field measures the association risk of the field in the global data; For power user and equipment data The degree centrality of each field is calculated based on the field association network.
[0015] The multidimensional sensitivity score is expressed as follows:
[0016] In the formula, For power user and equipment data Multidimensional sensitivity scores for each field; , , The scoring weights are as follows: uniqueness index, information entropy index, and relevance index.
[0017] Preferably, the multidimensional sensitivity scoring and grading algorithm determines whether the power user and equipment data fields are sensitive fields, and classifies the power user and equipment data corresponding to the sensitive fields as follows: The compliance index in the sensitivity scoring and grading algorithm uses a binary judgment to determine the data of power users and equipment through an assignment rule. Specifically, for power user and equipment data that are explicitly listed as sensitive by the power industry's explicit regulations or expert experience, the compliance judgment is 1; otherwise, it is 0.
[0018] The multidimensional sensitivity scoring and grading algorithm determines whether a power user and equipment data field is a sensitive field by whether the multidimensional sensitivity score exceeds a preset scoring threshold: if the multidimensional sensitivity score exceeds the preset threshold, it is determined to be a sensitive field; otherwise, it is determined to be a non-sensitive field.
[0019] At the same time, a sensitivity level classification rule is established, which classifies the sensitivity level of power user and equipment data according to the preset multi-dimensional sensitivity score and indicator threshold range. Power user and equipment data with a compliance indicator of 1 and reaching the highest score and indicator threshold range are classified as the highest sensitivity level.
[0020] Preferably, the objective function of the multi-objective optimization model includes a privacy leakage risk function, a data utility loss function, and a processing cost function, wherein the objective function is defined as:
[0021] Where, The objective function is... Minimize the function; Functions that mitigate privacy risks; The data utility loss function; To handle the cost function.
[0022] The privacy breach risk function is defined as follows:
[0023] Where, Total number of fields; For power user and equipment data Sensitivity level of the field; For power user and equipment data The field corresponding to the first The risk leakage coefficient of the desensitization method is assigned a value based on the strength of the desensitization method.
[0024] The data utility loss function is defined as:
[0025] Where, For power user and equipment data Business importance weight of the field; For power user and equipment data The field corresponding to the first Data loss rate caused by anonymization methods.
[0026] The processing cost function is defined as follows:
[0027] Where, For power user and equipment data The field corresponding to the first The computational cost of desensitization methods.
[0028] Preferably, for power user and equipment data and their sensitivity level information identified as sensitive fields, a multi-objective optimization model is used to perform multi-objective optimization through an improved NSGA-II algorithm to search for the optimal desensitization strategy. Specifically: S1: Initialize the population and randomly assign a corresponding desensitization method to each data field under the condition of satisfying the sensitivity level constraint. Associated data fields are assigned the same desensitization method.
[0029] S2: Calculate the objective function value of the multi-objective optimization model for each individual, which is used for subsequent non-dominated sorting.
[0030] S3: Perform non-dominated sorting based on the objective function values of all individuals to delineate the Pareto front. For individuals on the same front, assess individual diversity by calculating crowding.
[0031] S4: Generate offspring populations by performing tournament selection, uniform crossover, and mutation operations on the parent population.
[0032] S5: Merge the parent and offspring populations to form a temporary population. Sort the temporary population using non-dominated sorting and crowding calculation. Use an elite retention strategy to retain the best individuals in the temporary population to form a new generation population.
[0033] S6: Repeat steps S2-S5 until the termination condition of the maximum number of iterations is reached, and output the Pareto solution set of the optimal desensitization strategy.
[0034] Preferably, the congestion level in the improved NSGA-II algorithm is a business-oriented congestion level, expressed as:
[0035] In the formula, For individuals The original level of congestion; For power user and equipment data Business importance weight of the field; For power user and equipment data The field corresponding to the first Data loss rate caused by anonymization methods.
[0036] Preferably, step S4 performs the mutation operation based on a dynamically adjusted mutation probability, wherein the mutation probability is expressed as:
[0037] In the formula, For power user and equipment data The probability of field mutation; Basic mutation probability; For power user and equipment data Sensitivity level of the field; The normalization coefficient is determined by the sensitivity level set.
[0038] On the other hand, the present invention provides a field-level automatic desensitization system based on a data middle platform, including a power data acquisition and centralization module, a data sensitivity assessment and classification module, and a data desensitization strategy formulation and execution module.
[0039] The power data acquisition and centralization module is used to collect power user and equipment data through various power management systems and build power datasets. The power datasets are then sent to the data platform for centralized and automatic de-identification.
[0040] The data sensitivity assessment and classification module is used to determine whether any field in the power user and equipment data is a sensitive field for all power user and equipment data in the power dataset, based on a multi-dimensional sensitivity scoring and classification algorithm that considers uniqueness, information entropy, correlation, and compliance. The module then classifies the power user and equipment data corresponding to the sensitive fields according to their sensitivity.
[0041] The data anonymization strategy formulation and execution module is used to perform multi-objective optimization on power user and equipment data and their sensitivity level information that are identified as sensitive fields, and to search for the optimal anonymization strategy by using a multi-objective optimization model and an improved NSGA-II algorithm. Based on the optimal anonymization strategy, the module then executes the corresponding data anonymization method on each sensitive field.
[0042] In another aspect, the present invention also provides an electronic device, the electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the field-level automatic desensitization method based on a data platform as described in any embodiment of the present invention.
[0043] In another aspect, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the field-level automatic desensitization method based on a data platform as described in any embodiment of the present invention.
[0044] Compared with the prior art, the present invention has the following technical effects: 1. This invention, through the research of field-level anonymization technology using a data platform, enables the sharing and collaboration of sensitive data while protecting privacy, reduces the risk of human intervention and errors, improves data utilization efficiency, promotes business development and innovation, and builds a more open and shared data ecosystem.
[0045] 2. This invention combines the proportion of unique field values, information disorder, and cross-field correlation to effectively distinguish between highly sensitive fields and low-sensitivity fields, avoiding omissions or misjudgments caused by traditional single indicators; and through compliance binary judgment, it forces industry standards (such as GDPR and power data security standards) to be included in the classification logic, ensuring the rationality of sensitivity level classification and significantly improving the comprehensiveness and accuracy of sensitive field identification.
[0046] 3. This invention improves the NSGA-II algorithm, simultaneously optimizing three objectives: privacy leakage risk, data utility loss, and processing cost, to generate a Pareto-optimal solution set. This avoids extreme de-identification strategies (such as over-identification that damages data usability) caused by single-objective optimization, and provides diverse de-identification schemes for decision-making. Furthermore, the improved NSGA-II algorithm, combined with uniform crossover and dynamic mutation operations, can quickly converge to a high-quality solution set in complex solution spaces, making it suitable for large-scale power datasets. Attached Figure Description
[0047] Figure 1 This is an overall flowchart of the field-level automatic desensitization method based on a data middle platform as described in this invention. Detailed Implementation
[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present application and with reference to the accompanying drawings.
[0049] Example 1 This embodiment provides a field-level automatic data masking method based on a data middle platform. (See attached document.) Figure 1 As shown, it includes the following steps: Data on power users and equipment is collected through various power management systems to construct a power dataset, which is then sent to a data platform for centralized and automatic anonymization.
[0050] Specifically, power management systems typically include multiple subsystems, such as electricity consumption monitoring, equipment management, fault detection, and load forecasting. These systems collect data from electricity users and equipment in real time or periodically. This data may include: Electricity user data: including user ID, ID number, address, electricity consumption, load curve, billing information, and electricity consumption time period. Equipment data: including equipment ID, equipment type, operating status, maintenance records, fault history, equipment location, and operation records. Each power management system, based on its subordinate data acquisition modules, periodically collects electricity user and equipment data and then uses a unified communication protocol (such as MQTT, OPC UA, Modbus, etc.) to transmit the data to a data platform for routine data preprocessing (including data cleaning, outlier handling, etc.), forming a unified power dataset.
[0051] For all power user and equipment data in the power dataset, a multi-dimensional sensitivity scoring and grading algorithm based on uniqueness, information entropy, correlation, and compliance is used to determine whether the power user and equipment data fields are sensitive fields, and the power user and equipment data corresponding to the sensitive fields are classified according to sensitivity.
[0052] As a preferred embodiment of this example, the multidimensional sensitivity scoring and grading algorithm calculates the multidimensional sensitivity score based on three-dimensional quantitative indicators: uniqueness, information entropy, and correlation. The uniqueness metric measures a field's ability to directly identify an individual. The higher the uniqueness of a field (such as a user's ID number), the stronger its ability to directly identify an individual, and the higher its sensitivity. This is expressed as:
[0053] Where, For power user and equipment data Uniqueness indicators for each field; For power user and equipment data The number of unique values in each field is obtained by counting the number of non-repeating values in the statistical data field. For power user and equipment data The total number of records for each field.
[0054] Information entropy is used to measure the richness of information in a data field. The higher the information entropy value, the more information the field contains. It is expressed as:
[0055] Where, For power user and equipment data Information entropy index for each field; For power user and equipment data Different values in each field The set of occurrences; value The frequency of occurrence.
[0056] The correlation index is used to assess the sensitivity of data field combinations. If the uniqueness of a certain field is significantly improved when combined with other fields (e.g., equipment location + equipment model), the sensitivity is higher, as expressed as:
[0057] Where, For power user and equipment data The correlation index of each field; For power user and equipment data The first field combines power user and equipment data. The number of unique values in each field; For power user and equipment data , The total number of records in each field; To iterate through the power data except for the first... Other power user and equipment data fields in the first field are the same as the first field. The highest proportion of unique values after combining multiple fields; For power user and equipment data The network weight coefficient of each field measures the association risk of the field in the global data; For power user and equipment data The degree centrality of each field is calculated based on the field association network.
[0058] The multidimensional sensitivity score is expressed as follows:
[0059] Where, For power user and equipment data Multidimensional sensitivity scores for each field; , , The scoring weights for uniqueness, information entropy, and relevance indicators are set based on experience and can be adjusted according to business needs.
[0060] As a preferred embodiment of this example, the multidimensional sensitivity scoring and grading algorithm determines whether the power user and equipment data fields are sensitive fields, and performs sensitivity grading on the power user and equipment data corresponding to the sensitive fields, specifically as follows: The compliance index in the sensitivity scoring and grading algorithm uses a binary judgment to determine the data of power users and equipment through an assignment rule. Specifically, for power user and equipment data that are explicitly listed as sensitive by the power industry's explicit regulations or expert experience, the compliance judgment is 1; otherwise, it is 0.
[0061] The multidimensional sensitivity scoring and grading algorithm determines whether a power user and equipment data field is a sensitive field by whether the multidimensional sensitivity score exceeds a preset scoring threshold: if the multidimensional sensitivity score exceeds the preset threshold, it is determined to be a sensitive field; otherwise, it is determined to be a non-sensitive field. The preset scoring threshold can be adaptively optimized, such as dynamically calculating the scoring threshold based on historical data (determining the top 20% of data fields as sensitive fields).
[0062] At the same time, a sensitivity level classification rule is established, which classifies the sensitivity level of power user and equipment data according to the preset multi-dimensional sensitivity score and indicator threshold range. Power user and equipment data with a compliance indicator of 1 and reaching the highest score and indicator threshold range are classified as the highest sensitivity level.
[0063] Specifically, there is no limit to the number of sensitivity levels; they can be determined based on the actual operational needs of the power system. Below are some simple examples: Sensitivity levels are divided into 5 levels. Data with a score of 0.9 or higher, or a uniqueness index of 1 and a compliance index of 1, is set to the highest sensitivity level, L3. Data with a score less than 0.9 but greater than or equal to 0.7, or an information entropy index greater than or equal to 0.9 or a correlation index greater than or equal to 0.7, is set to the sensitivity level, L2. Data with a score of 0.6 or higher is set to the sensitivity level, L1.
[0064] Furthermore, the sensitivity level can be adjusted through the data association network of power users and equipment: the association network is constructed with data fields as nodes and data association ratio as weights. If the ratio of unique values between data fields is greater than a preset threshold, an edge is formed. The pivotal role of a data field in the association network is measured by calculating the influence of nodes through proximity centrality. If the proximity centrality of a data field is greater than a preset threshold, the sensitivity level is increased to correct the level.
[0065] For power user and equipment data and their sensitivity level information that are identified as sensitive fields, a multi-objective optimization model is used to perform multi-objective optimization through an improved NSGA-II algorithm to search for the optimal desensitization strategy. Based on the optimal desensitization strategy, the corresponding data desensitization method is executed for each sensitive field.
[0066] In a preferred embodiment of this invention, the objective function of the multi-objective optimization model includes a privacy leakage risk function, a data utility loss function, and a processing cost function, wherein the objective function is defined as follows:
[0067] In the formula, The objective function is... Minimize the function; Functions that mitigate privacy risks; The data utility loss function; To handle the cost function.
[0068] The privacy breach risk function is defined as follows:
[0069] In the formula, Total number of fields; For power user and equipment data Sensitivity level of the field; For power user and equipment data The field corresponding to the first The risk leakage coefficient of the desensitization method is assigned a value based on the strength of the desensitization method.
[0070] The data utility loss function is defined as:
[0071] In the formula, For power user and equipment data Business importance weight of the field; For power user and equipment data The field corresponding to the first The data loss rate caused by anonymization methods can be assessed in conjunction with actual business operations.
[0072] The processing cost function is defined as follows:
[0073] In the formula, For power user and equipment data The field corresponding to the first The computational cost of desensitization methods can be obtained through historical data analysis.
[0074] As a preferred embodiment of this example, the power user and equipment data and their sensitivity level information identified as sensitive fields are optimized using a multi-objective optimization model and an improved NSGA-II algorithm to search for the optimal desensitization strategy. Specifically, the optimal desensitization strategy is as follows: S1: Initialize the population and randomly assign an anonymization method to each data field while satisfying the sensitivity level constraints. Associated data fields are assigned the same anonymization method. Specifically, each anonymization strategy is represented as a chromosome, with genes corresponding to the anonymization method selection for each sensitive field. Simultaneously, time-series data fields require a unified anonymization granularity. Device-associated data fields must use the same anonymization method. Different anonymization strengths are matched to different sensitivity levels. For example, the highest anonymization level data is matched with the highest strength anonymization, including but not limited to: mandatory irreversible encryption (AES-256 encryption) or hashing (SHA-3 hashing, HMAC hashing (with salt)), etc., ensuring irreversible anonymization of direct identifiers. Medium-sensitivity level data is matched with medium-strength anonymization, including but not limited to: dynamically masking key parts or partial generalization, allowing limited associations but preventing the restoration of the original value. Low-sensitivity level data is matched with low-strength anonymization, including but not limited to partial masking or lightweight anonymization, meeting basic privacy requirements while maintaining high data utility.
[0075] S2: Calculate the objective function value of the multi-objective optimization model for each individual, which is used for subsequent non-dominated sorting.
[0076] S3: Perform non-dominated ranking based on the objective function values of all individuals to delineate the Pareto front. For individuals on the same front, assess individual diversity by calculating crowding. Specifically, the ranking rule is: if individual A is not inferior to individual B on all objectives and is superior to individual B on at least one objective, then A dominates B.
[0077] S4: Perform tournament selection, uniform crossover, and mutation operations on the parent population to generate the offspring population. Specifically, the size of the tournament selection (the number of individuals compared in each selection operation) is not limited here, and is generally set in the range of 2-5. The mutation operation must meet the sensitivity level constraints. A simple example using the L3 sensitivity level above is that this sensitivity level requires the use of de-identification methods such as encryption or hashing that conform to the sensitivity level.
[0078] S5: Merge the parent and offspring populations to form a temporary population. Sort the temporary population using non-dominated sorting and crowding calculation. Use an elite retention strategy to retain the best individuals in the temporary population to form a new generation population.
[0079] S6: Repeat steps S2-S5 until the termination condition of the maximum number of iterations is reached, and output the Pareto solution set of the optimal desensitization strategy.
[0080] In a preferred embodiment of this invention, the congestion level in the improved NSGA-II algorithm is a business-oriented congestion level, expressed as follows:
[0081] In the formula, For individuals The original level of congestion; For power user and equipment data The business importance weight of a field can be determined by factors such as data usage frequency and correlation analysis requirements. For power user and equipment data The field corresponding to the first Data loss rate caused by anonymization methods.
[0082] For each individual within the frontier, they are sorted in ascending order of their objective function values. The original crowding density is used to measure the distribution density of individuals within the same frontier in the objective space, prioritizing the retention of sparsely distributed individuals, defined as:
[0083] In the formula, For individuals The degree of congestion; The total number of objective functions; For individuals In the objective function The previous individual The objective function value; For individuals In the objective function The previous individual The objective function value; For all individuals in the objective function The maximum value of the objective function; For all individuals in the objective function The objective function is to find the minimum value.
[0084] In a preferred embodiment of this invention, step S4 performs a mutation operation based on a dynamically adjusted mutation probability to ensure the stability of the desensitization method for highly sensitive fields. The mutation probability is expressed as:
[0085] In the formula, For power user and equipment data The probability of field mutation; The basic mutation probability is usually set to 0.1 by default. For power user and equipment data Sensitivity level of the field; The normalization coefficient is determined by the sensitivity level set.
[0086] Furthermore, the data platform can periodically monitor the anonymized data to ensure its security and compliance, including checking data access logs and operation records. It can also periodically audit the anonymization process and generate corresponding audit reports to identify potential data security issues and make timely improvements.
[0087] After automatic data anonymization, the data is integrated into various application systems, such as smart grid monitoring, user management systems, and energy management platforms, through the data interface provided by the data platform. Furthermore, the data platform can use the anonymized data for data mining, trend prediction, and behavioral analysis to generate various power management reports, energy optimization reports, and more.
[0088] Example 2 Accordingly, this embodiment provides a field-level automatic desensitization system based on a data platform. The system is used to implement the field-level automatic desensitization method based on a data platform as described in Embodiment 1 of the present invention, including a power data acquisition and centralization module, a data sensitivity assessment and classification module, and a data desensitization strategy formulation and execution module.
[0089] The power data acquisition and centralization module is used to collect power user and equipment data through various power management systems and build power datasets. The power datasets are then sent to the data platform for centralized and automatic de-identification.
[0090] The data sensitivity assessment and classification module is used to determine whether any field in the power user and equipment data is a sensitive field for all power user and equipment data in the power dataset, based on a multi-dimensional sensitivity scoring and classification algorithm that considers uniqueness, information entropy, correlation, and compliance. The module then classifies the power user and equipment data corresponding to the sensitive fields according to their sensitivity.
[0091] The data anonymization strategy formulation and execution module is used to perform multi-objective optimization on power user and equipment data and their sensitivity level information that are identified as sensitive fields, and to search for the optimal anonymization strategy by using a multi-objective optimization model and an improved NSGA-II algorithm. Based on the optimal anonymization strategy, the module then executes the corresponding data anonymization method on each sensitive field.
[0092] Example 3 This embodiment provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the field-level automatic desensitization method based on a data platform as described in Embodiment 1 of this invention.
[0093] Example 4 This embodiment provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the field-level automatic desensitization method based on a data platform as described in Embodiment 1 of this invention.
[0094] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, A and B simultaneously, or B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.
[0095] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0096] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0097] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0098] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A field-level automatic data masking method based on a data middle platform, characterized in that, Includes the following steps: By collecting power user and equipment data through various power management systems and constructing power datasets, the power datasets are uniformly sent to the data platform for centralized automatic de-identification. For all power user and equipment data in the power dataset, a multi-dimensional sensitivity scoring and grading algorithm based on uniqueness, information entropy, correlation, and compliance is used to determine whether the power user and equipment data fields are sensitive fields, and the power user and equipment data corresponding to the sensitive fields are classified according to sensitivity. For power user and equipment data and their sensitivity level information that are identified as sensitive fields, a multi-objective optimization model is used to perform multi-objective optimization through an improved NSGA-II algorithm to search for the optimal desensitization strategy. Based on the optimal desensitization strategy, the corresponding data desensitization method is executed for each sensitive field.
2. The field-level automatic desensitization method based on a data middle platform according to claim 1, characterized in that, The multidimensional sensitivity scoring and grading algorithm calculates multidimensional sensitivity scores based on three-dimensional quantitative indicators: uniqueness, information entropy, and correlation. The uniqueness index is expressed as: In the formula, For power user and equipment data Uniqueness indicators for each field; For power user and equipment data The number of unique values in each field is obtained by counting the number of non-repeating values in the statistical data field. For power user and equipment data The total number of records for each field; The information entropy index is expressed as: In the formula, For power user and equipment data Information entropy index for each field; For power user and equipment data Different values in each field The set of occurrences; value The frequency of occurrence; The correlation index is expressed as follows: In the formula, For power user and equipment data The correlation index of each field; For power user and equipment data The first field combines power user and equipment data. The number of unique values in each field; For power user and equipment data , The total number of records for each field; To iterate through the electricity dataset except for the first... Other power user and equipment data fields in the first field are the same as the first field. The highest proportion of unique values after combining multiple fields; For power user and equipment data The network weight coefficient of each field measures the association risk of the field in the global data; For power user and equipment data The degree centrality of each field is calculated based on the field association network. The multidimensional sensitivity score is expressed as follows: In the formula, For power user and equipment data Multidimensional sensitivity scores for each field; , , The scoring weights are as follows: uniqueness index, information entropy index, and relevance index.
3. The field-level automatic desensitization method based on a data middle platform according to claim 1, characterized in that, The multidimensional sensitivity scoring and grading algorithm determines whether a power user and equipment data field is a sensitive field, and then categorizes the power user and equipment data corresponding to the sensitive fields into sensitivity levels as follows: The compliance index in the sensitivity scoring and grading algorithm uses a binary judgment to determine the data of power users and equipment through an assignment rule. Specifically, for power user and equipment data that are explicitly listed as sensitive by the power industry's explicit regulations or expert experience, the compliance judgment is 1; otherwise, it is 0. The multidimensional sensitivity scoring and grading algorithm determines whether a power user and equipment data field is a sensitive field by whether the multidimensional sensitivity score exceeds a preset scoring threshold: if the multidimensional sensitivity score exceeds the preset threshold, it is determined to be a sensitive field; Otherwise, it is considered a non-sensitive field; At the same time, a sensitivity level classification rule is established, which classifies the sensitivity level of power user and equipment data according to the preset multi-dimensional sensitivity score and indicator threshold range. Power user and equipment data with a compliance indicator of 1 and reaching the highest score and indicator threshold range are classified as the highest sensitivity level.
4. The field-level automatic desensitization method based on a data middle platform according to claim 1, characterized in that, The objective function of the multi-objective optimization model includes a privacy leakage risk function, a data utility loss function, and a processing cost function. The objective function is defined as follows: In the formula, The objective function is... Minimize the function; Functions that mitigate privacy risks; The data utility loss function; To handle the cost function; The privacy breach risk function is defined as follows: In the formula, Total number of fields; For power user and equipment data Sensitivity level of the field; For power user and equipment data The field corresponding to the first The risk leakage coefficient of the desensitization method is assigned a value based on the strength of the desensitization method; The data utility loss function is defined as: In the formula, For power user and equipment data Business importance weight of the field; For power user and equipment data The field corresponding to the first Data loss rate caused by anonymization methods; The processing cost function is defined as follows: In the formula, For power user and equipment data The field corresponding to the first The computational cost of desensitization methods.
5. The field-level automatic desensitization method based on a data middle platform according to claim 4, characterized in that, For power user and equipment data and their sensitivity level information identified as sensitive fields, a multi-objective optimization model is used to perform multi-objective optimization through an improved NSGA-II algorithm to search for the optimal desensitization strategy. Specifically: S1: Initialize the population, and randomly assign the corresponding desensitization method to each data field under the condition of satisfying the sensitivity level constraint. Assign the same desensitization method to related data fields. S2: Calculate the objective function value of the multi-objective optimization model for each individual, which will be used for subsequent non-dominated ranking; S3: Perform non-dominated ranking based on the objective function values of all individuals to delineate the Pareto front. For individuals on the same front, assess individual diversity by calculating crowding. S4: Generate offspring populations by performing tournament selection, uniform crossover, and mutation operations on the parent population; S5: Merge the parent and offspring populations to form a temporary population. Sort the temporary population using non-dominated sorting and crowding calculation. Use an elite retention strategy to retain the best individuals in the temporary population to form a new generation population. S6: Repeat steps S2-S5 until the termination condition of the maximum number of iterations is reached, and output the Pareto solution set of the optimal desensitization strategy.
6. The field-level automatic desensitization method based on a data middle platform according to claim 5, characterized in that, The congestion level in the improved NSGA-II algorithm is a business-oriented congestion level, expressed as: In the formula, For individuals The original level of congestion; For power user and equipment data Business importance weight of the field; For power user and equipment data The field corresponding to the first Data loss rate caused by anonymization methods.
7. The field-level automatic desensitization method based on a data middle platform according to claim 5, characterized in that, Step S4 performs a mutation operation based on a dynamically adjusted mutation probability, whereby the mutation probability is expressed as: In the formula, For power user and equipment data The probability of field mutation; Basic mutation probability; For power user and equipment data Sensitivity level of the field; The normalization coefficient is determined by the sensitivity level set.
8. A field-level automatic data masking system based on a data middle platform, characterized in that, The system is used to implement the field-level automatic desensitization method based on a data platform as described in any one of claims 1-7, including a power data acquisition and centralization module, a data sensitivity assessment and classification module, and a data desensitization strategy formulation and execution module; The power data acquisition and centralization module is used to collect power user and equipment data through various power management systems and build power datasets, and then send the power datasets to the data platform for centralized automatic de-identification. The data sensitivity assessment and classification module is used to determine whether any field in the power user and equipment data is a sensitive field for all power user and equipment data in the power dataset, based on a multi-dimensional sensitivity scoring and classification algorithm that considers uniqueness, information entropy, correlation, and compliance. The module then classifies the power user and equipment data corresponding to the sensitive fields based on sensitivity. The data anonymization strategy formulation and execution module is used to perform multi-objective optimization on power user and equipment data and their sensitivity level information that are identified as sensitive fields, and to search for the optimal anonymization strategy by using a multi-objective optimization model and an improved NSGA-II algorithm. Based on the optimal anonymization strategy, the module then executes the corresponding data anonymization method on each sensitive field.
9. An electronic device, the electronic device comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the computer program, it implements the field-level automatic desensitization method based on a data platform as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the field-level automatic desensitization method based on the data platform as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Automatic desensitization method based on negative manifest data resources
CN112613069A
Cited By
Vehicle visual sentry optimization method and device, storage medium and program product
CN121354064A