Data desensitization method, electronic equipment and medium

By performing hierarchical privacy processing and risk feature extraction on the original data, and combining the parameters of the management platform for desensitization and inverse desensitization, the contradiction between data privacy protection and risk identification in the existing technology is solved, and the reversible restoration of data and the balance between risk identification capabilities is achieved.

CN120449194AActive Publication Date: 2025-08-08HANGZHOU PINGPONG INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510954398.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-08-08
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

While protecting user data, existing data privacy protection technologies have low risk identification capabilities and irreversible desensitization operations, which cannot meet the compliance needs of restoring original data in authorized scenarios.

Method used

Provide a data desensitization method, which can privacy the original data through multiple privacy processing parameters set, extract risk characteristics and desensitization, and use the parameters of the management platform to perform inverse desensitization operations, combining hierarchical desensitization and dynamic risk assessment to ensure the balance between data privacy protection and risk identification.

Benefits of technology

It realizes the precise identification of potential risk data while protecting data privacy, supports reversible data restoration, and meets compliance needs in authorized scenarios. It is suitable for scenarios that require both data security and business analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449194A_ABST
    Figure CN120449194A_ABST
Patent Text Reader

Abstract

The invention discloses a data desensitization method, electronic equipment and a computer readable medium, and relates to the technical field of data protection. Performing privacy processing on the original data by using a plurality of set privacy processing parameters to obtain privacy processing data, and sending the privacy processing parameters to a management platform, the management platform having a management authority; extracting risk features in the privacy processing data to obtain risk feature representation data of the privacy processing data, and sending related parameters of the risk feature representation data to a management platform; and performing desensitization processing on the original data according to the risk feature representation data, outputting desensitized data, and performing inverse desensitization operation on the desensitized data through privacy processing parameters stored in the management platform and related parameters of the risk feature representation data to obtain the original data. The method has the advantages and characteristics that data privacy protection and risk identification are balanced, and the desensitized data are controllably and reversibly desensitized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data protection technology, and in particular to a data desensitization method, electronic equipment, and medium. Background Art

[0002] While financial institutions conduct in-depth analysis of user data to identify potential risks, they also need to protect the privacy of that data. Privacy protection requires limiting the processing and analysis of user data, while risk identification requires comprehensive and in-depth analysis of user data.

[0003] Although current data privacy protection technologies (such as static data masking technology, K-anonymity technology, traditional differential privacy technology, and federated learning technology) can protect user data to varying degrees, the data practicality and availability are poor, which leads to low risk identification capabilities. Summary of the Invention

[0004] The present invention aims to solve one of the technical problems in the related art to a certain extent. To this end, the present invention provides a data desensitization method, an electronic device for performing the data desensitization method, and a computer-readable medium, which have the advantages of balancing data privacy protection and risk identification.

[0005] In order to achieve the above object, as a first aspect of the present invention, a data desensitization method is provided, wherein the data desensitization method comprises: Get the original data; Performing privacy processing on the original data using a plurality of set privacy processing parameters to obtain privacy-processed data, and sending the privacy processing parameters to a management platform, which has management authority; Extracting risk features from the privacy-processing data, obtaining risk feature representation data of the privacy-processing data, and sending relevant parameters of the risk feature representation data to a management platform; The original data is desensitized according to the risk characteristic representation data, and the desensitized data is output, wherein the desensitized data can be reverse-desensitized by using the privacy processing parameters and relevant parameters of the risk characteristic representation data stored in the management platform to obtain the original data.

[0006] Optionally, performing desensitization processing on the original data according to the risk feature representation data and outputting the desensitized data includes: Determining a risk level and a desensitization rule based on the data complexity of the risk feature representation data; wherein the desensitization rule corresponds to the risk level in a one-to-one manner; Using the risk level as a guide condition, desensitizing the original data according to a desensitization rule corresponding to the risk level to obtain desensitized data; Determine the desensitization quality score based on the information entropy of the desensitized data and the information entropy of the original data; When the desensitization quality score is greater than or equal to a preset desensitization quality threshold, the desensitized data is output.

[0007] Optionally, the privacy processing parameters include an initial privacy budget parameter, an adjusted privacy budget parameter, and a noise parameter, and the initial privacy budget parameter corresponds one-to-one to the sensitivity level of the original data; The step of performing privacy processing on the original data using the set multiple privacy processing parameters to obtain privacy processed data includes: determining an adjusted privacy budget parameter based on the sensitivity level and the initial privacy budget parameter corresponding thereto; determining a noise parameter based on the sensitivity level; Noise processing is performed on the original data using the noise parameter, the adjusted privacy budget parameter, and the sensitivity level to obtain privacy-processed data.

[0008] Optionally, the noise parameter includes a noise type, and determining the noise parameter according to the sensitivity level includes: In a case where the sensitivity level is high sensitivity, the noise type includes Laplace noise; When the sensitivity level is medium sensitivity, the noise type includes Laplace noise or Gaussian noise; In a case where the sensitivity level is low sensitivity, the noise type includes Gaussian noise.

[0009] Optionally, performing noise processing on the original data using the noise parameter, the adjusted privacy budget parameter, and the sensitivity level to obtain privacy-processed data includes: In a case where the noise parameter is Laplace noise, inputting the adjusted privacy budget parameter and the sensitivity level into a Laplace noise function, adding the Laplace noise function to the original data, and obtaining privacy-processed data after Laplace noise addition; When the noise parameter is Gaussian noise, the adjusted privacy budget parameter, the sensitivity level, and the calibration constant are input into a Gaussian noise function, and the Gaussian noise function is added to the original data to obtain privacy-processed data after Gaussian noise addition.

[0010] Optionally, the noise parameter further includes an amplitude of the noise type and a standard deviation of the noise type, and the data desensitization method further includes: determining a noise range ratio according to a standard deviation of the noise type and a preset data range; Determining a noise precision ratio according to a standard deviation of the noise type and a preset data precision; The magnitude of the noise type is adjusted according to the noise range ratio and the noise precision ratio so that the noise range ratio is within a preset range factor and the noise precision ratio is within a preset precision factor.

[0011] Optionally, the risk feature representation data includes a fused risk feature vector and a risk prediction value, and extracting the risk features from the privacy-processed data to obtain the risk feature representation data of the privacy-processed data includes: Preprocessing the privacy-processing data according to preset standardized processing rules to obtain standard privacy-processing data; Inputting the standard privacy-processed data into feature engineering to obtain multi-dimensional enhanced features; Inputting the multi-dimensional enhanced features into the traditional risk feature model and the deep learning risk feature model respectively to obtain traditional risk features and potential risk features; The traditional risk features and the potential risk features are fused to generate a fused risk feature vector and a risk prediction value.

[0012] Optionally, the fusing of the traditional risk features and the potential risk features to generate a fused risk feature vector and a risk prediction value further includes: Determine the feature quality score of the fusion risk feature according to the preset feature evaluation rules; When the feature quality score is less than or equal to the preset feature quality threshold, the feature quality score is input into the feature engineering to guide the feature engineering to re-extract the feature quality score. Multi-dimensional enhanced features corresponding to quasi-privacy processed data; When the feature quality score is greater than a preset feature quality threshold, a fused risk feature vector and a risk prediction value are output.

[0013] As a second aspect of the present invention, there is provided an electronic device, comprising: one or more processors; A memory having one or more computer programs stored thereon, wherein when the one or more computer programs are executed by the one or more processors, the one or more processors implement the data desensitization method provided according to the first aspect of the present invention.

[0014] In addition, as a third aspect of the present invention, a computer-readable medium is provided, on which a computer program is stored, wherein when the computer program is executed by a processor, the data desensitization method provided by the first aspect of the present invention is implemented.

[0015] The data desensitization method provided by the present invention performs different degrees of privacy processing on the original data according to different privacy processing parameters, and then extracts the risk characteristics of the privacy-processed data after privacy processing, and uses the risk characteristics as conditions to accurately desensitize the original data. During the entire desensitization process, on the one hand, the privacy of the original data is protected, and on the other hand, it is further guaranteed that the potential risk data is identified in a targeted manner without excessive use of data. In addition, the entire data desensitization method will synchronously send the operating parameters and characteristic data of each processing link to a management platform with management authority functions, and the desensitized data can be reverse-demensitized according to the privacy processing parameters and risk characteristic representation parameters stored in the management platform to restore the original data. The data desensitization method provided by the present invention combines hierarchical desensitization processing with dynamic risk assessment, which not only meets the requirements of data privacy protection, but also retains the analytical value of the data to the greatest extent; at the same time, the perfect operation record mechanism provides a reliable guarantee for data recovery and auditing, and is particularly suitable for scenarios that need to take into account both data security and business analysis.

[0016] These features and advantages of the present invention will be further disclosed in the following detailed description and accompanying drawings. The preferred embodiments and means of the present invention will be fully illustrated in conjunction with the accompanying drawings, but are not intended to limit the technical solutions of the present invention. Furthermore, although multiple features, elements, and components may be present in each of the following text and accompanying drawings, they may be labeled with different symbols or numbers for convenience, but all represent components with the same or similar structure or function. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention will be further described below in conjunction with the accompanying drawings: Figure 1 A flowchart of a data desensitization method provided by the present invention; Figure 2 A flowchart of an implementation of step S140 of the data desensitization method provided by the present invention; Figure 3 This is a flow chart of another implementation of step S120 of the data desensitization method provided by the present invention; Figure 4 This is a flow chart of another implementation of step S122 of the data desensitization method provided by the present invention; Figure 5 This is a flow chart of another implementation of step S123 of the data desensitization method provided by the present invention; Figure 6 This is a flow chart of another implementation of step S120 of the data desensitization method provided by the present invention; Figure 7 This is a flow chart of another implementation of step S130 of the data desensitization method provided by the present invention; Figure 8This is a flow chart of another implementation of step S134 of the data desensitization method provided by the present invention; Figure 9 This is a system framework diagram corresponding to the data desensitization method provided by the present invention; Figure 10 Flowchart of the hierarchical differential privacy module of the system framework diagram provided by the present invention; Figure 11 Flowchart of the risk feature extraction and representation learning module of the system framework diagram provided by the present invention; Figure 12 This is a flowchart of the risk analysis and privacy protection balance module of the system framework diagram provided by the present invention; Figure 13 A module diagram of an electronic device provided by the present invention; Figure 14 A schematic diagram of a computer-readable medium provided by the present invention.

[0018] Description of Reference Numerals Among them, 101 is a processor; 102 is a memory; 103 is an I / O interface; and 104 is a bus. DETAILED DESCRIPTION

[0019] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described in the embodiments are intended to explain the present invention and are not to be construed as limiting the present invention.

[0020] References in this specification to "one embodiment," "an example," or "an example" mean that a particular feature, structure, or characteristic described in connection with the embodiment itself can be included in at least one embodiment disclosed herein. The appearances of the phrase "in one embodiment" in various places in the specification are not necessarily all referring to the same embodiment.

[0021] At present, existing data privacy protection technologies usually use unified privacy parameters to mechanically protect data, resulting in some data with practical value and weak correlation with user privacy being over-protected, further leading to a decline in risk identification capabilities, or reducing the intensity of privacy protection in order to improve risk identification accuracy; in addition, the desensitization operations performed on data by existing data privacy protection technologies are mostly irreversible operations, which cannot meet the compliance requirements of restoring original data in authorized scenarios.

[0022] In view of this, in order to solve the technical contradiction between data privacy protection and risk identification, and the problem that desensitized data cannot be restored to original data, as a first aspect of the present invention, a data desensitization method is provided, such as Figure 1 As shown, the data desensitization method includes: In step S110, original data is obtained; In step S120, the original data is privacy-processed using the set multiple privacy processing parameters to obtain privacy-processed data, and the privacy processing parameters are sent to a management platform, which has management authority; In step S130, risk features are extracted from the privacy-processing data to obtain risk feature representation data of the privacy-processing data, and relevant parameters of the risk feature representation data are sent to the management platform; In step S140, the original data is desensitized according to the risk characteristic representation data, and the desensitized data is output, wherein the desensitized data can be reverse-desensitized by using the privacy processing parameters and relevant parameters of the risk characteristic representation data stored in the management platform to obtain the original data.

[0023] The data desensitization method provided by the present invention performs different degrees of privacy processing on the original data according to different privacy processing parameters, and then extracts the risk characteristics of the privacy-processed data after privacy processing, and uses the risk characteristics as conditions to accurately desensitize the original data. During the entire desensitization process, on the one hand, the privacy of the original data is protected, and on the other hand, it is further guaranteed that the potential risk data is identified in a targeted manner without excessive use of data. In addition, the entire data desensitization method will synchronously send the operating parameters and characteristic data of each processing link to a management platform with management authority functions, and the desensitized data can be reverse-demensitized according to the privacy processing parameters and risk characteristic representation parameters stored in the management platform to restore the original data. The data desensitization method provided by the present invention combines hierarchical desensitization processing with dynamic risk assessment, which not only meets the requirements of data privacy protection, but also retains the analytical value of the data to the greatest extent; at the same time, the perfect operation record mechanism provides a reliable guarantee for data recovery and auditing, and is particularly suitable for scenarios that need to take into account both data security and business analysis.

[0024] The above desensitization process is described in detail as an optional implementation of step S140. Figure 2 As shown, the desensitizing process is performed on the original data according to the risk feature representation data, and the desensitized data is output, including: In step S141, the risk level and desensitization rules are determined according to the data complexity of the risk feature representation data; wherein the desensitization rules correspond to the risk levels one by one; In step S142, the risk level is used as a guide condition, and the original data is desensitized according to the desensitization rule corresponding to the risk level to obtain desensitized data; In step S143, a desensitization quality score is determined based on the information entropy of the desensitized data and the information entropy of the original data; In step S144, when the desensitization quality score is greater than or equal to the preset desensitization quality threshold, the desensitized data is output.

[0025] Steps S141-S144 above illustrate the process for desensitizing raw data. Prior to desensitization, the risk level and desensitization rules can be determined based on the data complexity of the risk signature data. Data complexity can be calculated based on factors such as the number of features in the risk signature data (more features, higher data complexity; fewer features, lower data complexity) and the data distribution of the risk signature data (data distribution is determined by calculating entropy; higher entropy indicates a more complex data distribution; lower entropy indicates a simpler data distribution). In practical applications, calculation methods include, but are not limited to, the complexity of the number of features and the complexity of the data distribution, and require flexible adjustment based on computing resources, computational time complexity, and accuracy of the data complexity. Desensitization rules correspond to risk levels. For high-risk data, key-based symmetric encryption, format-preserving encryption, reversible tokenization, and partial-preserving desensitization with differential privacy enhancement can be used; for medium-risk data, partial-preserving desensitization with differential privacy enhancement can be used; and for low-risk data, general encryption algorithms can be used. It should be noted that the desensitization rules for different risk levels described above can be flexibly adjusted or other desensitization rules can be adopted in practical applications. However, after desensitizing the data using the desensitization rules, it is necessary to evaluate the desensitization effect (such as calculating the change in information entropy of the data before and after desensitization in step S143) and output desensitized data that meets the desensitization quality requirements.

[0026] De-massaging operations related to the above-mentioned de-massaging operations must be performed with explicit authorization. Authorization can be granted by a management platform with administrative privileges, ensuring the high privacy of de-massaging operations, thereby further enhancing the security of data de-massaging and de-massaging.

[0027] Similarly, in order to protect the original data in a targeted manner and reduce the situation where data is over-protected or under-protected in traditional data protection methods, the present invention provides a layered differential privacy protection mechanism to determine privacy processing parameters, wherein the privacy processing parameters include an initial privacy budget parameter, an adjusted privacy budget parameter, and a noise parameter. The initial privacy budget parameter corresponds to the sensitivity level of the original data in a one-to-one manner; as an implementation method of step S120, Figure 3 As shown, the privacy processing of the original data using the set multiple privacy processing parameters to obtain privacy-processed data includes: In step S121, an adjusted privacy budget parameter is determined according to the sensitivity level and the initial privacy budget parameter corresponding thereto; In step S122, a noise parameter is determined according to the sensitivity level; In step S123 , the original data is subjected to noise processing using the noise parameter, the adjusted privacy budget parameter, and the sensitivity level to obtain privacy-processed data.

[0028] First, the original data is classified and then the sensitivity level is determined based on the classification results. Different privacy processing parameters are used for data of different sensitivity levels. Finally, the original data is processed in a targeted manner based on the privacy processing parameters. It is necessary to specifically explain the classification. As a classification method, the original data can be divided into personal identity information (such as name, ID number, contact information), financial transaction information (such as account balance, transaction record, credit limit), behavioral characteristic information (such as consumption habits, location information, device information), and derivative risk information (such as credit score, risk level, default probability, etc.). Similarly, a method for determining a sensitivity level can be determined according to the following formula (1): (1) In the above formula (1), It is a data identification indicator used to measure whether the data can uniquely identify an individual, with a value range of [0,1]; It is a data sensitivity index used to measure the potential harm caused by data leakage, with a value range of [0,1]; It is a data timeliness indicator used to measure the effective time length of data, with a value range of [0,1]; It is a data correlation index, which is used to measure the degree of correlation between data and other data, and its value range is [0,1]. It is a regulatory requirement indicator used to measure the data protection requirements in regulations, with a value range of [0,1]; to is the corresponding weight and satisfies S is the sensitivity score, which is used to determine the sensitivity level based on the sensitivity score. Data with S ≥ 0.7 is considered highly sensitive and requires the strictest privacy protection; data with 0.3 ≤ S < 0.7 is considered moderately sensitive and requires moderate privacy protection; and data with low sensitivity S < 0.3 is considered lowly sensitive and requires basic privacy protection. Regarding the above-mentioned determination of sensitivity levels based on sensitivity scores, it is important to note that weights can be automatically adjusted based on external changes, sensitivity levels can be dynamically adjusted based on data usage frequency and scenarios, and data sensitivity needs to be regularly reassessed to adapt to data changes. Furthermore, machine learning algorithms can be used to analyze historical data usage patterns to predict trends in sensitivity changes.

[0029] After the original data is divided into three levels of high sensitivity fields, medium sensitivity fields, and low sensitivity fields through the above classification operation, different initial privacy budget parameters are assigned to different sensitivity levels. The corresponding relationship between the initial privacy budget parameters and the sensitivity levels is formula (2): (2) In formula (2), Corresponding to high-sensitivity data, medium-sensitivity data and low-sensitivity data respectively; is the initial privacy budget parameter; for highly sensitive data, a smaller initial privacy budget parameter is set. ; For medium-sensitive data, set a medium initial privacy budget parameter, ; For low-sensitivity data, set a larger initial privacy budget parameter, In addition, the initial privacy budget parameters can also be set based on the sensitivity level, data type, and usage scenario.

[0030] After setting the initial privacy budget parameters, they need to be dynamically adjusted according to the sensitivity level to obtain the adjusted privacy budget parameters. For highly sensitive fields, conservative adjustments are made according to formula (3): (3) in, is the initial privacy budget parameter corresponding to the highly sensitive field, is the data distribution characteristic factor, is the inter-field correlation factor, is the adjusted privacy budget parameter corresponding to the highly sensitive field; conservative adjustment according to formula (3) can not only fine-tune according to data characteristics but also ensure the strictness of privacy protection. For medium-sensitive fields, balanced adjustment is performed according to formula (4): (4) in, is the initial privacy budget parameter corresponding to the sensitive field, is the business importance factor, is the data distribution characteristic factor, is the adjusted privacy budget parameter corresponding to the medium-sensitive field; a balanced adjustment is performed according to formula (4) to achieve a balance between privacy protection and business needs and data availability. For low-sensitivity fields, a loose adjustment is performed according to formula (5): (5) in, is the initial privacy budget parameter corresponding to the low-sensitivity field, is the risk tolerance factor, is the adjusted privacy budget parameter corresponding to the low-sensitivity field; according to formula (5), a loose adjustment can be made to prioritize the availability of data based on the acceptable risk level.

[0031] In order to accurately protect the privacy of the original data, it is necessary to first determine the noise type and then inject the noise into the original data. As an implementation method of step S122, Figure 4 As shown, the noise parameter includes a noise type, and determining the noise parameter according to the sensitivity level includes: In step S122a, when the sensitivity level is high sensitivity, the noise type includes Laplace noise; In step S122b, when the sensitivity level is medium sensitivity, the noise type includes Laplace noise or Gaussian noise; In step S122c, when the sensitivity level is low sensitivity, the noise type includes Gaussian noise.

[0032] After determining the type of noise to be injected according to the sensitivity level, the original data is subjected to noise processing using the noise parameter, the adjusted privacy budget parameter, and the sensitivity level to obtain privacy-processed data. As an optional implementation of step S123, Figure 5 Shown, including: In step S123a, if the noise parameter is Laplace noise, the adjusted privacy budget parameter and the sensitivity level are input into a Laplace noise function, and the Laplace noise function is added to the original data to obtain privacy-processed data after Laplace noise addition; In step S123b, when the noise parameter is Gaussian noise, the adjusted privacy budget parameter, the sensitivity level, and the calibration constant are input into a Gaussian noise function, and the Gaussian noise function is added to the original data to obtain privacy-processed data after Gaussian noise addition.

[0033] It is necessary to explain the noise addition process in detail: Laplace noise injection and Gaussian noise injection are as follows: (6) (7) In formulas (6) and (7), S is the sensitivity, ε is the adjusted privacy budget parameter, c is the calibration constant, X is the original data, Process the data for privacy after adding noise.

[0034] In order to further protect the original data, the noise parameters also include the amplitude of the noise type and the standard deviation of the noise type, such as Figure 6 As shown, the data desensitization method also includes: In step S122a, a noise range ratio is determined according to the standard deviation of the noise type and a preset data range; In step S122b, a noise precision ratio is determined according to the standard deviation of the noise type and a preset data precision; In step S122c, the amplitude of the noise type is adjusted according to the noise range ratio and the noise precision ratio, so that the noise range ratio is within a preset range factor and the noise precision ratio is within a preset precision factor.

[0035] Special attention should be paid to the range factor and the difference of precision. The range factor ensures that the noise amplitude is appropriate for the data range. Its calculation methods include, but are not limited to, the relative noise ratio method: noise ratio = noise standard deviation / data range. When the noise ratio is < 0.01, the noise may be too low, and privacy protection is insufficient. When the noise ratio is 0.01 ≤ noise ratio ≤ 0.1, it is generally considered appropriate. When the noise ratio is > 0.1, the noise may be too high, and the data's usefulness is impaired. It is generally recommended to control the noise ratio between 0.01 and 0.1, depending on privacy requirements. The difference of precision ensures that noise does not excessively affect data accuracy. Its calculation methods include, but are not limited to, the precision threshold comparison method: precision impact ratio = noise standard deviation / data precision requirement. When the precision impact ratio is < 0.5, the noise does not significantly affect accuracy. When the precision impact ratio is 0.5 ≤ precision impact ratio < 1, the noise is close to affecting accuracy. When the precision impact ratio is ≥ 1, the noise has excessively affected accuracy.

[0036] The completion of noise injection marks the completion of the layered differential privacy protection mechanism, and the privacy-processed data is output for subsequent operations. The privacy-processed data has the same data format as the original data. During the layered differential implementation process, information such as privacy parameters and noise type is recorded and sent to the management platform.

[0037] Extracting risk feature information of privacy-processing data requires the use of a parallel traditional risk feature model and a deep learning risk feature model, as an optional implementation of step S130, such as Figure 7 As shown, the risk feature representation data includes a fusion risk feature vector and a risk prediction value, and the risk feature representation data of the privacy-processed data is obtained by extracting the risk feature from the privacy-processed data, including: In step S131, the privacy-processing data is pre-processed according to a preset standardization processing rule to obtain standard privacy-processing data; In step S132, the standard privacy-processed data is input into feature engineering to obtain multi-dimensional enhanced features; In step S133, the multi-dimensional enhanced features are input into the traditional risk feature model and the deep learning risk feature model respectively to obtain traditional risk features and potential risk features; In step S134, the traditional risk features and the potential risk features are fused to generate a fused risk feature vector and a risk prediction value.

[0038] In order to better extract risk feature information, the standardized processing rules of step S131 include but are not limited to operations such as missing value filling, outlier processing, and data format unification to perform data processing for subsequent features; the preprocessed data enters step S132 for multi-dimensional feature transformation and enhancement, and the feature engineering process includes but is not limited to one or more combined operations such as time series feature extraction, relationship feature construction, interaction feature generation, category feature encoding, and feature binning. The output results of the feature engineering are simultaneously fed into the traditional risk feature model and the deep learning risk feature model. In step S133, the traditional risk feature model and the deep learning risk feature model are used as parallel models to simultaneously obtain the output results of the feature engineering. The traditional risk feature model extracts highly explanatory features, including but not limited to principal component analysis, independent component analysis, linear discriminant analysis, etc.; the deep learning risk feature model extracts highly abstract and compact potential feature representations to capture nonlinear patterns and high-order relationships in the data. The deep learning risk feature model includes but is not limited to: variational autoencoder construction, latent space regularization, adversarial training enhancement, attention mechanism integration, representation compression and noise reduction operations. Step S134 fuses traditional features and deep learning features to generate a comprehensive risk feature vector. Fusion methods include but are not limited to: weighted fusion, stacked fusion, attention fusion, multi-view consistency learning, heterogeneous graph fusion, etc.

[0039] In order to improve the accuracy of risk prediction, it is necessary to perform a quality assessment on the risk feature vector before using the fused risk feature vector for risk identification and prediction. As an implementation of step S134, Figure 8 As shown, the fusion of the traditional risk features and the potential risk features to generate a fused risk feature vector and a risk prediction value also includes: In step S134a, the feature quality score of the fused risk feature is determined according to a preset feature evaluation rule; In step S134b, if the feature quality score is less than or equal to the preset feature quality threshold, the feature quality score is input into feature engineering to guide feature engineering to re-extract multi-dimensional enhanced features corresponding to the standard privacy-processed data; In step S134c, when the feature quality score is greater than a preset feature quality threshold, the fused risk feature vector and the risk prediction value are output.

[0040] It is necessary to explain the quality assessment rules in detail, which include but are not limited to information gain assessment, stability test, multicollinearity detection, robustness detection, and interpretability assessment.

[0041] In addition, the multiple processing links of the data desensitization method provided by the present invention can be integrated into multiple modules respectively, and the association relationship between the multiple modules also constitutes a system corresponding to the data desensitization method provided by the present invention. Figure 9A system diagram of the data desensitization method provided by the present invention is given. In this system, the original customer data is first sent to the data classification and sensitivity assessment module, which is classified and identified according to the preset sensitivity level determination method, and the sensitivity classification result (customer data with sensitivity identification) is output; the customer data with sensitivity identification is input into the hierarchical differential privacy framework module, which applies the differential privacy protection algorithm of corresponding strength according to different sensitivity levels and outputs differential privacy processed data; the data after differential privacy processing does not carry customer privacy information and can be accessed and used normally; the data after differential privacy processing is input into the risk feature extraction and representation learning module, which extracts the differential privacy information from the differential privacy protection algorithm. Key risk features are extracted from the privacy-processed data and feature representations are generated. The risk feature representation results are output. The risk feature representation results also do not contain customer privacy information and can be accessed and used normally. The risk feature representation results and the original data are input into the controllable and reversible desensitization module. The original data can be desensitized according to the sensitivity level. The risk feature representation results can also be used as a guiding condition to guide the desensitization of the original data, and the desensitized data is output. It should be noted that all operations involving the desensitization of customer sensitive information and the reversal of desensitization operations can only be performed under the authorization of the security authorization and control module. Every operation during the execution process is recorded to ensure the compliance, traceability and data security of the desensitization operation. The desensitized data and its risk feature prediction accuracy results are input into the risk analysis and privacy protection balance module to evaluate the risk identification ability and privacy protection level under the current parameter settings. The optimized feedback parameters are output to the hierarchical differential privacy framework module and the risk feature extraction and representation learning module to adjust the differential privacy parameters and feature representation parameters respectively to achieve a dynamic balance between risk identification ability and privacy protection level. The security audit module monitors the entire processing flow, records the operation logs of each module, ensures that all data processing activities in the system are traceable and auditable, and generates compliance audit reports. This module adopts a distributed log collection architecture to capture the operational events of each functional module of the system in real time, including but not limited to key activities such as user login / logout behavior, data access requests, desensitizing operation execution, permission changes, and system configuration modifications. The collected audit data is tamper-proofed and stored in a dedicated security audit database, and a data integrity verification mechanism is applied to ensure the authenticity and integrity of the log data. This module can automatically generate multi-dimensional audit reports, including data processing compliance reports, privacy protection assessment reports, security incident analysis reports, etc., and supports visual display and export functions. All audit records are set with a retention period in accordance with relevant regulations, and strict access control is implemented to ensure the security and confidentiality of the audit data itself. With the system Figure 9In one embodiment, financial institutions can achieve a balance between customer data privacy protection and risk identification capabilities through the following steps: The system receives raw customer data containing personal identity information, transaction records, and behavioral characteristics. Using a data classification and sensitivity assessment module, the data is classified into three sensitivity levels: high, medium, and low. For high-sensitivity data (such as ID numbers and bank account numbers), the system applies strict differential privacy parameters; for medium-sensitivity data (such as age ranges and occupational categories), medium-strength privacy protection parameters are used; and for low-sensitivity data (such as consumer preferences and transaction frequency), relaxed privacy protection parameters are used. The processed data is then processed using a deep autoencoder network to extract risk features, retaining over 95% of risk prediction capability while also achieving data dimensionality reduction. The system then processes the data using a controllable and reversible desensitization mechanism based on homomorphic encryption, ensuring that the original information can be restored through a distributed key management system in authorized scenarios (such as regulatory inspections and judicial investigations). The desensitized data and feature prediction accuracy information are fed into a risk analysis and privacy protection balancing module, which dynamically adjusts differential privacy parameters using a multi-objective optimization algorithm. Ultimately, a balance is achieved between customer privacy protection and risk identification capabilities. The detailed flow charts of the hierarchical differential privacy module, risk feature extraction and representation learning module, and risk analysis and privacy protection balance module in the system diagram are as follows: Figure 10 、 11 、12.

[0042] As a second aspect of the present invention, there is provided an electronic device, such as Figure 13 Shown, including: One or more processors 101; The memory 102 stores one or more computer programs. When the one or more computer programs are executed by the one or more processors 101, the one or more processors 101 implement the data desensitization method provided according to the first aspect of the present invention.

[0043] The tool may further include one or more I / O interfaces 103 connected between the processor 101 and the memory 102 and configured to implement information exchange between the processor 101 and the memory 102 .

[0044] Among them, the processor 101 is a device with data processing capabilities, including but not limited to the central processing unit 101 (CPU); the first memory 102 is a device with data storage capabilities, including but not limited to random access memory 102 (RAM, more specifically such as SDRAM, DDR, etc.), read-only memory 102 (ROM), electrically erasable programmable read-only memory 102 (EEPROM), flash memory (FLASH); the I / O interface 103 (read-write interface) is connected between the processor 101 and the memory 102, and can realize information exchange between the processor 101 and the memory 102, including but not limited to the data bus 104 (Bus), etc.

[0045] In some embodiments, the processor 101 , the memory 102 , and the I / O interface 103 are connected to each other via a bus 104 , and further connected to other components of the computing device.

[0046] In addition, as a third aspect of the present invention, a computer readable medium is provided, on which a computer program is stored. Figure 14 As shown, when the computer program is executed by a processor, the data desensitization method provided by the first aspect of the present invention is implemented.

[0047] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. Accordingly, the computer program can be stored in a non-volatile computer-readable storage medium, and when the computer program is executed, it can implement the method of any of the above-mentioned embodiments. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0048] The above are only specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Those skilled in the art should understand that the present invention includes but is not limited to the contents described in the drawings and the above specific embodiments. Any modifications that do not deviate from the functional and structural principles of the present invention are intended to be included within the scope of the claims.

Claims

1. A data desensitization method, characterized in that: The data desensitization method includes: Get the original data; Performing privacy processing on the original data using a plurality of set privacy processing parameters to obtain privacy-processed data, and sending the privacy processing parameters to a management platform, which has management authority; Extracting risk features from the privacy-processing data, obtaining risk feature representation data of the privacy-processing data, and sending relevant parameters of the risk feature representation data to a management platform; The original data is desensitized according to the risk characteristic representation data, and the desensitized data is output, wherein the desensitized data can be reverse-desensitized by using the privacy processing parameters and relevant parameters of the risk characteristic representation data stored in the management platform to obtain the original data.

2. The data desensitization method according to claim 1, characterized in that: The desensitizing the original data according to the risk feature representation data and outputting the desensitized data includes: Determining a risk level and a desensitization rule based on the data complexity of the risk feature representation data; wherein the desensitization rule corresponds to the risk level in a one-to-one manner; Using the risk level as a guide condition, desensitizing the original data according to a desensitization rule corresponding to the risk level to obtain desensitized data; Determine the desensitization quality score based on the information entropy of the desensitized data and the information entropy of the original data; When the desensitization quality score is greater than or equal to a preset desensitization quality threshold, the desensitized data is output.

3. The data desensitization method according to claim 1, characterized in that: The privacy processing parameters include an initial privacy budget parameter, an adjusted privacy budget parameter, and a noise parameter, wherein the initial privacy budget parameter corresponds to the sensitivity level of the original data. The step of performing privacy processing on the original data using the set multiple privacy processing parameters to obtain privacy processed data includes: determining an adjusted privacy budget parameter based on the sensitivity level and the initial privacy budget parameter corresponding thereto; determining a noise parameter based on the sensitivity level; Noise processing is performed on the original data using the noise parameter, the adjusted privacy budget parameter, and the sensitivity level to obtain privacy-processed data.

4. The data desensitization method according to claim 3, characterized in that: The noise parameter includes a noise type, and determining the noise parameter according to the sensitivity level includes: In a case where the sensitivity level is high sensitivity, the noise type includes Laplace noise; When the sensitivity level is medium sensitivity, the noise type includes Laplace noise or Gaussian noise; In a case where the sensitivity level is low sensitivity, the noise type includes Gaussian noise.

5. The data desensitization method according to claim 3, characterized in that: The performing noise processing on the original data using the noise parameter, the adjusted privacy budget parameter, and the sensitivity level to obtain privacy-processed data includes: In a case where the noise parameter is Laplace noise, inputting the adjusted privacy budget parameter and the sensitivity level into a Laplace noise function, adding the Laplace noise function to the original data, and obtaining privacy-processed data after Laplace noise addition; When the noise parameter is Gaussian noise, the adjusted privacy budget parameter, the sensitivity level, and the calibration constant are input into a Gaussian noise function, and the Gaussian noise function is added to the original data to obtain privacy-processed data after Gaussian noise addition.

6. The data desensitization method according to any one of claims 3 to 5, characterized in that: The noise parameters also include the amplitude of the noise type and the standard deviation of the noise type. The data desensitization method also includes: determining a noise range ratio according to a standard deviation of the noise type and a preset data range; Determining a noise precision ratio according to a standard deviation of the noise type and a preset data precision; The magnitude of the noise type is adjusted according to the noise range ratio and the noise precision ratio so that the noise range ratio is within a preset range factor and the noise precision ratio is within a preset precision factor.

7. The data desensitization method according to claim 1, characterized in that: The risk feature representation data includes a fused risk feature vector and a risk prediction value, and extracting the risk features from the privacy-processed data to obtain the risk feature representation data of the privacy-processed data includes: Preprocessing the privacy-processing data according to preset standardized processing rules to obtain standard privacy-processing data; Inputting the standard privacy-processed data into feature engineering to obtain multi-dimensional enhanced features; Inputting the multi-dimensional enhanced features into the traditional risk feature model and the deep learning risk feature model respectively to obtain traditional risk features and potential risk features; The traditional risk features and the potential risk features are fused to generate a fused risk feature vector and a risk prediction value.

8. The data desensitization method according to claim 7, characterized in that: The fusing of the traditional risk features and the potential risk features to generate a fused risk feature vector and a risk prediction value further includes: Determine the feature quality score of the fusion risk feature according to the preset feature evaluation rules; If the feature quality score is less than or equal to a preset feature quality threshold, input the feature quality score into feature engineering to guide feature engineering to re-extract multi-dimensional enhanced features corresponding to the standard privacy-processed data; When the feature quality score is greater than a preset feature quality threshold, a fused risk feature vector and a risk prediction value are output.

9. An electronic device, characterized in that: include: one or more processors; A memory having one or more computer programs stored thereon, wherein when the one or more computer programs are executed by the one or more processors, the one or more processors implement the data desensitization method according to any one of claims 1 to 8.

10. A computer-readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the data desensitization method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Data processing method and device, electronic equipment and storage medium

    CN111737750A

  • Data blood relationship tracking method, system and device for privacy security protection

    CN118536164A

  • Privacy protection data exchange method and system under zero-trust network architecture

    CN119316239A

  • User data intelligent protection method and system based on differential privacy

    CN119720263A

  • Dynamic desensitization processing method and processing device for biological medicine classification variable data

    CN120257342A

Cited By

  • Log desensitization method and device, computer equipment and readable storage medium

    CN122153969A