Data processing method and device based on differential privacy protection
By performing structured processing on the original data and generating mirrored false data, and adding noise in differential privacy protection technology to form a privacy-preserving dataset, the balance problem between privacy protection and data availability in existing technologies is solved, and a balance between data security and accuracy is achieved.
Patent Information
- Application Number
- CN202510665375.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-09-19
AI Technical Summary
While existing differential privacy technologies protect data privacy, they also lead to a decline in data quality and availability, making it difficult to find a balance between privacy protection and data availability.
The original data is pre-processed to generate structured data, and differential noise is generated for each data unit based on differential privacy protection technology. At the same time, multiple mirrored false data are generated based on the data structure, and differential noise is added to the original data and the mirrored false data to form a privacy-preserving dataset.
It solves the problem of single piece of information distortion in the original data after differential noise processing, while ensuring data security, maintaining data accuracy and improving the efficiency and accuracy of data use.
Smart Images

Figure CN120671174A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer information processing, and more specifically, to a data processing method and device based on differential privacy protection. Background Art
[0002] In today's information age, data security and authenticity are increasingly important issues, and data privacy protection is a key topic. Differential privacy, a theoretical model that quantifies and constrains the risk of privacy leakage, has been widely used in various data publishing and analysis scenarios. It protects data privacy to a certain extent by injecting a certain amount of noise into the original data to mask personal information.
[0003] However, differential privacy technology has an inherent trade-off: greater amounts of injected noise enhance privacy protection, but this also leads to decreased data quality and availability. Finding a balance between privacy protection and data availability has always been a key issue in this field.
[0004] On the one hand, the core of differential privacy technology lies in the proper selection of noise parameters, such as the privacy budget ε. Choosing this parameter requires a deep understanding of the balance between privacy risks and data utilization value. Improper parameter settings can lead to overprotection (reduced data value) or privacy leakage (increased privacy risks). This places high technical demands on practitioners using differential privacy technology and requires a deep understanding of privacy protection theory.
[0005] On the other hand, while differential privacy improves the privacy security of individuals, statistical analysis of overall data still maintains the accuracy of the original data. This raises a new issue: the overall statistical data processed using differential privacy still faces the risk of being exploited by attackers. Attackers may exploit background knowledge and other information to infer individual privacy. This requires further exploration of combining other privacy protection methods to better prevent targeted privacy breaches.
[0006] The above information disclosed in this Background section is only for enhancement of understanding of the background of the application and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0007] In view of this, the present application provides a data processing method and device based on differential privacy protection, which can solve the problem of distortion of single information of original data after differential noise processing, ensure data security while ensuring data accuracy, and improve the efficiency and accuracy of data use.
[0008] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.
[0009] According to one aspect of the present application, a data processing method based on differential privacy protection is proposed, which includes: performing structured preprocessing on original data to generate structured data; generating a set of differential noise for each data unit in the structured data based on differential privacy protection technology; generating multiple mirror false data according to the data structure of the original data; adding differential noise to the original data and the multiple mirror false data to generate original noise data and multiple mirror noise data; combining the original noise data and the multiple mirror noise data to form a data group, and marking the data group to form a privacy-protected data set.
[0010] In an exemplary embodiment of the present application, it also includes: extracting the original noise data and the corresponding multiple mirror noise data contained in the data group according to the label of each data group in the privacy protection data set; calculating the mean value of the original noise data and the corresponding data units in the multiple mirror noise data based on the numerical values to eliminate differential noise; and reconstructing the original data according to the calculated mean value.
[0011] In an exemplary embodiment of the present application, the original data is subjected to structured preprocessing to generate structured data, including: performing word segmentation processing on the original data to identify semantic units in the data; performing part-of-speech tagging and named entity recognition on the semantic units to extract structured fields; and generating structured data based on the structured fields.
[0012] In an exemplary embodiment of the present application, a set of differential noise is generated for each data unit in the structured data based on differential privacy protection technology, including: generating a set of differential noise for each data unit in the structured data based on Laplace distribution or Gaussian distribution.
[0013] In an exemplary embodiment of the present application, a set of differential noise is generated for each data unit in the structured data based on the Laplace distribution or the Gaussian distribution, including: determining differential privacy parameters and calculation sensitivity; generating a set of differential noise for each data unit in the structured data based on the Laplace distribution or the Gaussian distribution and the differential privacy parameters; or generating a set of differential noise for each data unit in the structured data based on the Laplace distribution or the Gaussian distribution, the differential privacy parameters and the calculation sensitivity.
[0014] In an exemplary embodiment of the present application, multiple mirrored false data are generated according to the data structure of the original data, including: parsing the structural features of the original data to identify key data units therein; generating multiple mirrored data for each key data unit; and combining the generated multiple mirrored data according to the original data structure to form multiple mirrored false data.
[0015] In an exemplary embodiment of the present application, multiple false data are generated for each key data unit, including: generating multiple false data for each key data unit based on a large model according to the data structure and calculation sensitivity of each key data unit.
[0016] In an exemplary embodiment of the present application, differential noise is added to the original data and the multiple mirrored false data respectively to generate original noise data and multiple mirrored noise data, including: adding the differential noise value to the data unit corresponding to the original data to form the original noise data; adding the differential noise value to the corresponding data unit in the multiple mirrored false data to form multiple mirrored noise data; wherein the average of the multiple differential noise values in each data unit is 0.
[0017] In an exemplary embodiment of the present application, reconstructing the original data based on the calculated mean includes: identifying the same data group consisting of the original data and its corresponding multiple mirror noise data based on the tag information in the data group; performing statistical mean calculation on the numerical values of the corresponding fields in the data group to obtain estimated values of each field; and reversely mapping the estimated values to the original structured fields according to a preset mapping relationship, thereby reconstructing the original data.
[0018] According to one aspect of the present application, a data processing device based on differential privacy protection is proposed, which includes: a structuring module for performing structured preprocessing on original data to generate structured data; a differential module for generating a set of differential noise for each data unit in the structured data based on differential privacy protection technology; a mirror module for generating multiple mirror false data according to the data structure of the original data; an adding module for adding differential noise to the original data and the multiple mirror false data respectively to generate original noise data and multiple mirror noise data; a data module for combining the original noise data and the multiple mirror noise data to form a data group, and marking the data group to form a privacy-protected data set.
[0019] According to one aspect of the present application, an electronic device is proposed, which includes: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described above.
[0020] According to one aspect of the present application, a computer-readable medium is provided, on which a computer program is stored. When the program is executed by a processor, the method described above is implemented.
[0021] According to the data processing method and device based on differential privacy protection of the present application, structured data is generated by performing structured preprocessing on the original data; a set of differential noise is generated for each data unit in the structured data based on differential privacy protection technology; multiple mirror false data are generated according to the data structure of the original data; differential noise is added to the original data and the multiple mirror false data respectively to generate original noise data and multiple mirror noise data; the original noise data and the multiple mirror noise data are combined to form a data group, and the data group is marked to form a privacy-protected data set. This method can solve the problem of distortion of a single piece of information in the original data after differential noise processing, ensure data accuracy while ensuring data security, and improve efficiency and accuracy in the data use process.
[0022] It should be understood that the foregoing general description and the following detailed description are merely illustrative and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The above and other objects, features, and advantages of the present application will become more apparent by describing in detail exemplary embodiments thereof with reference to the accompanying drawings. The drawings described below are merely some embodiments of the present application, and it is apparent to those skilled in the art that other drawings can be derived from these drawings without inventive effort.
[0024] Figure 1 The figure is a flowchart of a data processing method based on differential privacy protection according to an exemplary embodiment.
[0025] Figure 2 The figure is a flowchart of a data processing method based on differential privacy protection according to an exemplary embodiment.
[0026] Figure 3 is a flowchart of a data processing method based on differential privacy protection according to another exemplary embodiment.
[0027] Figure 4 is a schematic diagram of a data processing method based on differential privacy protection according to another exemplary embodiment.
[0028] Figure 5 The figure is a block diagram of a data processing device based on differential privacy protection according to an exemplary embodiment.
[0029] Figure 6It is a block diagram of an electronic device according to an exemplary embodiment.
[0030] Figure 7 It is a block diagram of a computer-readable medium according to an exemplary embodiment. DETAILED DESCRIPTION
[0031] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. Like reference numerals in the drawings represent like or similar parts, and thus repetitive description thereof will be omitted.
[0032] In addition, described feature, structure or characteristic can be combined in one or more embodiments in any suitable manner.In the following description, many specific details are provided so as to provide a full understanding of the embodiments of the present application. However, it will be appreciated by those skilled in the art that the technical scheme of the present application can be put into practice without one or more of the specific details, or other methods, components, devices, steps etc. can be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the application.
[0033] The block diagrams shown in the accompanying drawings are merely functional entities and do not necessarily correspond to physically separate entities. That is, these functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0034] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, while others may be combined or partially combined. Therefore, the actual execution order may vary depending on the actual situation.
[0035] It should be understood that although the terms first, second, third, etc. may be used herein to describe various components, these components should not be limited by these terms. These terms are used to distinguish one component from another. Thus, the first component discussed below could be referred to as the second component without departing from the teachings of the present invention. As used herein, the term "and / or" includes any one and all combinations of one or more of the associated listed items.
[0036] Those skilled in the art will understand that the drawings are merely schematic diagrams of example embodiments, and the modules or processes in the drawings are not necessarily necessary for implementing the present application, and therefore cannot be used to limit the scope of protection of the present application.
[0037] In order to solve the problems existing in the prior art, this application optimizes the shortcomings of differential privacy enhancement technology in data privacy protection and proposes a method for generating mirrored false data information to balance the problem of reduced accuracy caused by adding differential noise to the original data. In ordinary differential privacy enhancement technology, there are two major problems. One is that the algorithm designer needs to have a deep understanding of the meaning of the original data set and determine different data sensitivities for different data types, so as to formulate reasonable noise parameters to prevent over-protection and break the balance between data accuracy and privacy protection. The other is that this method undermines the credibility of a single piece of data and can only analyze statistical results.
[0038] To address these issues, the technical solution of this application designs a mirrored fake dataset corresponding to the original dataset for noise removal. This ensures that the original real data can be restored after internal noise removal, while also preventing attackers from stealing privacy. First, because the information group generated from a single piece of data can ultimately be restored to the original data through calculation, in-depth analysis of the original data's sensitivity is not required when adding noise, reducing the developer's understanding of data from different fields.
[0039] Furthermore, since the mirrored dataset can be used for final noise removal, the choice of noise model during differential noise insertion is more flexible. For example, there's no difference between noise based on a Gaussian function and noise based on a Laplacian function in the final restored dataset. However, the noise function will affect the final statistical analysis results (obtained by the attacker), thereby better protecting privacy. In fact, noise interference with the original data is only the initial stage of data security protection. After preprocessing, enterprise-specific keys are generated, and multi-layer encryption and distributed storage are implemented, along with a comprehensive security audit mechanism, providing more comprehensive and robust security protection for private data.
[0040] The technical content of this application is described in detail below with the help of specific embodiments.
[0041] Figure 1 FIG1 is a flowchart of a data processing method based on differential privacy protection according to an exemplary embodiment. The data processing method 10 based on differential privacy protection includes at least steps S102 to S110.
[0042] like Figure 1As shown, in S102, the original data is pre-processed to generate structured data. The original data can be segmented to identify semantic units in the data; part-of-speech tagging and named entity recognition are performed on the semantic units to extract structured fields; and structured data is generated based on the structured fields.
[0043] Structured datasets can make data models clearer and facilitate storage, management, and retrieval. Compared to unstructured data, structured data is easier to classify, filter, and analyze. Furthermore, in terms of data analysis and modeling, structured data can better support various data analysis and machine learning modeling tasks, such as statistical analysis and predictive modeling. These tasks require data to be well organized and interpretable.
[0044] There are many specific structuring methods, such as tokenization, part-of-speech tagging, named entity recognition (NER), text normalization, sentiment analysis, topic modeling, text summarization, and text embedding. Choose the appropriate data structuring method based on the actual data type.
[0045] In S104, a set of differential noises is generated for each data unit in the structured data based on differential privacy protection technology. A set of differential noises can be generated for each data unit in the structured data based on Laplace distribution or Gaussian distribution.
[0046] In practice, word segmentation can be performed on raw data. For example, natural language processing (NLP) techniques are used to identify word boundaries and extract semantic units. These semantic units are then subjected to part-of-speech tagging and named entity recognition (NER) to identify structured entity information such as names of people, places, organizations, and time. A data structure in the form of key-value pairs is then constructed based on the extracted structured fields (such as name, address, and amount) to generate structured data. For example, for a user's financial transaction record, fields such as "transaction time," "transaction amount," and "transaction type" can be extracted after processing to form standardized structured data.
[0047] More specifically, differential privacy parameters and computational sensitivity can be determined; based on the Laplace distribution or Gaussian distribution, a set of differential noise is generated for each data unit in the structured data using the differential privacy parameters; or based on the Laplace distribution or Gaussian distribution, a set of differential noise is generated for each data unit in the structured data using the differential privacy parameters and the computational sensitivity.
[0048] The commonly used mechanisms are the Laplace mechanism or the Gaussian mechanism. The former is suitable for situations where ε-differential privacy is satisfied, while the latter is suitable for situations where (ε, δ)-differential privacy requirements are satisfied.
[0049] The following operations may be further included: first, determine the differential privacy parameter ε (and δ, if a Gaussian mechanism is used) and the sensitivity Δf of each data field (i.e., the maximum change in the statistical results of the field when a single record changes); then, use the formula Lap(Δf / ε) or Gauss(Δf / ε,δ) to generate multiple groups of noise samples for each data unit, so that the noise distribution meets the desired privacy constraints and maintains a numerical mean of 0 to support subsequent reverse denoising.
[0050] In S106, multiple pieces of mirrored false data are generated based on the data structure of the original data. For example, the original data may be analyzed for structural features to identify key data units; multiple pieces of mirrored data are generated for each key data unit; and the generated multiple pieces of mirrored data are combined according to the original data structure to form multiple pieces of mirrored false data.
[0051] It can parse the original structured data and identify key data units with high sensitivity or high query frequency, such as ID number, account balance, etc.; then, based on a large language model or rule template, it generates multiple fake data items with similar statistical characteristics for these key data units to ensure that they are consistent with the real data in syntax, structure and distribution; finally, the mirror data units are recombined according to the field order of the original data to construct multiple complete mirror fake data records.
[0052] In S108, differential noise is added to the original data and the plurality of mirrored false data to generate original noise data and a plurality of mirrored noise data. For example, the differential noise value may be added to a data unit corresponding to the original data to generate original noise data; and the differential noise value may be added to a data unit corresponding to the plurality of mirrored false data to generate a plurality of mirrored noise data; wherein the average of the plurality of differential noise values in each data unit is 0.
[0053] A noise value sampled from the corresponding distribution is added to each data unit in the original data to obtain the original noise data. The corresponding noise sample is added to the data unit at the same position in the mirrored false data to obtain multiple mirrored noise data. The overall mean of the added noise is 0, so that it can effectively offset the noise in the subsequent statistical mean processing.
[0054] In S110, the original noise data and the multiple mirror noise data are combined to form a data group, and the data group is labeled to form a privacy-preserving dataset. The original noise data and the mirror noise data are combined to form a data group, and labeling information is added to distinguish between the original and mirror data, for example, by setting the label field "source" to "original" or "mirror." These data groups constitute a privacy-preserving dataset and can be used for model training or data analysis, protecting individual privacy while improving the statistical efficiency of the overall data.
[0055] According to the data processing method based on differential privacy protection of the present application, structured data is generated by performing structured preprocessing on the original data; a set of differential noise is generated for each data unit in the structured data based on differential privacy protection technology; multiple mirror false data are generated according to the data structure of the original data; differential noise is added to the original data and the multiple mirror false data respectively to generate original noise data and multiple mirror noise data; the original noise data and the multiple mirror noise data are combined to form a data group, and the data group is marked to form a privacy-protected data set. This method can solve the problem of distortion of a single piece of information in the original data after differential noise processing, ensure data accuracy while ensuring data security, and improve efficiency and accuracy in the data use process.
[0056] Figure 2 The figure is a flowchart of a data processing method based on differential privacy protection according to an exemplary embodiment. Figure 2 The process 20 shown is Figure 1 Supplementary description of the process shown.
[0057] like Figure 2 As shown, in S202, according to the label of each data group in the privacy protection data set, the original noise data and the corresponding multiple mirror noise data contained in the data group are extracted.
[0058] Specifically, each record in the data set contains source tag information (such as "original" or "mirror"). Based on this tag, the system can identify a set of records belonging to the same data group, that is, a set consisting of one original noise data and its corresponding multiple mirror noise data.
[0059] In S204 , based on the values of corresponding data units in the original noise data and the plurality of mirror noise data, their averages are calculated respectively to eliminate differential noise.
[0060] More specifically, for all values recorded in the same field position in the data group, the data unit value at that position is extracted and an averaging operation is performed. The characteristic of "noise mean is 0" is used to cancel out the differential noise, thereby restoring an estimated value that is closer to the real data.
[0061] In S206, the original data is reconstructed based on the calculated mean. Based on the tag information in the data group, the same data group consisting of the original data and its corresponding multiple mirror noise data can be identified; statistical mean calculation is performed on the values of the corresponding fields in the data group to obtain estimated values for each field; and the estimated values are reverse-mapped to the original structured fields according to a preset mapping relationship to reconstruct the original data.
[0062] First, based on the data set's label information and the corresponding relationships between structured fields, the values of the corresponding fields in the original noisy data and its mirrored data are located and aggregated. A statistical mean operation is then performed on each field to obtain an estimated value for each field. Finally, based on a preset field mapping dictionary or decoding table, the estimated value is restored to a valid representation of the original structured field (such as a classification label, enumeration value, or actual value), thus completing the reconstruction of the original data. In practical applications, this mean reconstruction method can effectively improve data availability while maintaining differential privacy.
[0063] Continuing with the previous example, during data restoration, a dataset generated from each piece of original information can be restored based on the dataset number of each piece of information stored in the final database. The original information can be restored by averaging the different data structures. Still using the data in 3.2.3 as an example, the name sequence number is averaged to 6, and the original name obtained from the name database is "Wei Moumou"; the quantity is averaged to 5; and the noun sequence number is averaged to 10, and the original noun obtained from the noun database is "Cefoperazone Sodium." Therefore, the original information can be restored to "Wei Moumou bought 5 units of Cefoperazone Sodium."
[0064] It should be clearly understood that this application describes how to form and use specific examples, but the principles of this application are not limited to any details of these examples. On the contrary, based on the teaching of the content disclosed in this application, these principles can be applied to many other embodiments.
[0065] Figure 3 is a flowchart of a data processing method based on differential privacy protection according to another exemplary embodiment. Figure 3 The process 30 shown is Figure 2A detailed description of “generating multiple mirrored false data according to the data structure of the original data” in the process shown.
[0066] like Figure 3 As shown, in S302, the structured features of the raw data are parsed to identify key data units. Specifically, the raw structured data can be parsed at the field level to analyze the role and sensitivity of each field in the data semantics. For example, fields containing user identity, behavior trajectory, and transaction information can be identified as key data units. This step can be combined with natural language processing technology, sensitivity calculation methods, or a predefined list of key fields for auxiliary identification, ensuring that subsequent processing focuses on the perturbation and falsification of sensitive information.
[0067] In S304, multiple mirror data are generated for each key data unit. Specifically, rule generation, random sampling, or a forgery method based on a large model (such as a language model or a generative model) can be used to generate a number of mirror data that are consistent with the original data structure but have fictitious content according to the type of the key data unit (such as numerical type, categorical type, text type, etc.). For example, for the "place of residence" field, multiple city names that exist but are not the user's actual address can be generated; for the "amount of consumption" field, multiple approximate values can be generated by adding small perturbations based on the original value. Each generated mirror data should maintain consistency with the original data in semantics and format.
[0068] In one embodiment, multiple false data can be generated for each key data unit based on the data structure and calculated sensitivity of each key data unit based on the large model. The generation of corresponding false data can be selected based on each piece of processed original information, its different data structure, and the level of sensitivity. For example, the following example:
[0069] Original information: Wei (6) bought 5 units of cefoperazone sodium (10)
[0070] "Wei Moumou" is the data with sequence number 6 in the name structure, "5" is a quantity and can be directly added with noise, and "Cefoperazone Sodium" is the data with sequence number 10 in the noun structure. Assuming that four false information needs to be generated, then after adding the modified original data, the result is as follows:
[0071] Wei (6) bought 7 doses of amoxicillin (3) [revised original information]
[0072] Mr. Jing (1) bought 4 units of cefoperazone sodium (10) [generated false information 1]
[0073] Tao (3) bought 9 aspirin (15) [generated false information 2]
[0074] Meng Moumou (8) bought 0 acyclovir capsules (8) [generated false information 3]
[0075] Lou Moumou (12) bought 5 units of aminomethylbenzoic acid (14) [generated false information 4]
[0076] Each name in the name structure set has a corresponding serial number, which contains both true and false information. Similarly, the noun structure set also contains true and false information, each with its own number. Considering the modified original information and the four false information as a data set (generated from a single message), the mean value for each structure is the value of the original structure. For example, in the name structure, the mean of the serial numbers of the five names in the example is 6, which is the serial number of "Wei Moumou", corresponding to the original information. The same is true for the noun structure. In the quantity structure, the mean value can be directly calculated.
[0077] The examples shown here are for ease of understanding. In practice, simply storing names in a name database can also lead to personal information leakage (for example, an attacker in a medical database could determine that a person actually appears in the database). Privacy can be further protected by storing each character of a name separately or encrypted. After labeling the dataset in which each piece of information resides, all generated fake information is combined to form a mirrored fake dataset. This is then mixed with the original dataset and stored.
[0078] In S306, the generated multiple mirrored data are combined according to the original data structure to form multiple mirrored false data records. Specifically, the mirrored values of each key data unit and non-key fields are combined in the same field order and data type as the original data, leaving them unchanged or generating appropriate default values, to construct a complete and consistent mirrored false data record. This ensures that the mirrored data is structurally indistinguishable, improving the anonymity and attack resistance of the differentially private dataset.
[0079] Figure 4 is a schematic diagram of a data processing method based on differential privacy protection according to another exemplary embodiment. Figure 4 This paper describes the data encryption preprocessing stage in data privacy protection, using structured data sets, differential noise addition, mirrored false data set construction and other technologies to optimize the data differential privacy protection process, thereby adding another layer of protection for subsequent data encryption, and finally obtaining the original real information through the noise elimination algorithm.
[0080] First, the raw data is preprocessed, and information is extracted from the unstructured dataset. Through techniques such as tokenization, part-of-speech tagging, and named entity recognition (NER), a structured standard dataset is obtained to facilitate subsequent data processing. Differential noise is then added to the structured dataset to obtain a dataset enhanced with differential privacy protection in the general sense. While adding noise to each piece of data, corresponding mirrored false data is generated. A denoising algorithm is then used to verify that the data set generated from each piece of raw data can be restored to the true data.
[0081] To add differential noise, we first need to determine the differential privacy parameters ε and δ. The former controls the strength of privacy protection; a smaller value indicates stronger privacy protection; the latter represents the upper bound on the probability of privacy protection failure.
[0082] Calculating sensitivity defines the maximum difference between any two adjacent records in the database and is primarily used to determine the amount of noise to add. Because image denoising technology maintains consistency between data quality and privacy, this step can be selectively implemented based on specific hardware requirements.
[0083] Generates noise that follows a Laplace or Gaussian distribution based on ε, δ, and sensitivity. The variance of Laplace noise can be 2*sensitivity / ε, and the variance of Gaussian noise can be 2*sensitivity^2*ln(1.25 / δ) / ε^2.
[0084] However, unlike conventional differential algorithms, this application requires generating multiple pieces of false information for each piece of processed information. Therefore, rather than generating just one piece of noise with a certain structure for each piece of processed information, a set of noises with the mean value of the original data can be generated. Specifically, by modifying the expected value of the probability distribution to the initial data, a set of noises that meets the requirements can be generated.
[0085] Finally, it is necessary to add differential noise to the false information generated in the next step to complete the entire noise addition process.
[0086] This application addresses the irreconcilable contradiction between privacy protection and data quality in differential privacy technology, solves the problem of distortion of single information in original data after differential noise processing, exchanges space for protection benefits, ensures data accuracy while ensuring data security, and makes data managers more efficient in the data processing process.
[0087] Those skilled in the art will appreciate that all or part of the steps implementing the above embodiments can be implemented as a computer program executed by a CPU. When executed by the CPU, the computer program performs the functions defined in the above method provided herein. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disk.
[0088] Furthermore, it should be noted that the aforementioned figures are merely illustrative of the processes included in the methods according to exemplary embodiments of the present application and are not intended to be limiting. It is readily understood that the processes illustrated in the aforementioned figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.
[0089] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.
[0090] Figure 5 FIG is a block diagram of a data processing device based on differential privacy protection according to an exemplary embodiment. Figure 5 As shown, the data processing device 50 based on differential privacy protection includes: a structuring module 502 , a difference module 504 , a mirroring module 506 , an adding module 508 , and a data module 510 . The data processing device 50 based on differential privacy protection may further include: a reconstruction module 512 .
[0091] The structuring module 502 is used to perform structural preprocessing on the original data to generate structured data; the structuring module 502 is also used to perform word segmentation processing on the original data to identify semantic units in the data; perform part-of-speech tagging and named entity recognition on the semantic units to extract structured fields; and generate structured data based on the structured fields.
[0092] The differential module 504 is used to generate a set of differential noise for each data unit in the structured data based on differential privacy protection technology; the differential module 504 is also used to generate a set of differential noise for each data unit in the structured data based on Laplace distribution or Gaussian distribution.
[0093] The mirror module 506 is used to generate multiple mirror false data according to the data structure of the original data; the mirror module 506 is also used to parse the structural characteristics of the original data and identify the key data units therein; generate multiple mirror data for each key data unit; and combine the generated multiple mirror data according to the original data structure to form multiple mirror false data.
[0094] The adding module 508 is used to add differential noise to the original data and the multiple mirrored false data respectively to generate original noise data and multiple mirrored noise data; the adding module 508 is also used to add the differential noise value to the data unit corresponding to the original data to form original noise data; and add the differential noise value to the corresponding data unit in the multiple mirrored false data to form multiple mirrored noise data; wherein the average of the multiple differential noise values is 0 in each data unit.
[0095] The data module 510 is configured to combine the original noise data and the plurality of mirror noise data to form a data group, and mark the data group to form a privacy-preserving data set.
[0096] The reconstruction module 512 is used to extract the original noise data and the corresponding multiple mirror noise data contained in the data group based on the label of each data group in the privacy-preserving dataset; calculate the mean value of the corresponding data units in the original noise data and the multiple mirror noise data respectively to eliminate differential noise; and reconstruct the original data based on the calculated mean value.
[0097] According to the data processing device based on differential privacy protection of the present application, structured data is generated by performing structured preprocessing on the original data; a set of differential noise is generated for each data unit in the structured data based on differential privacy protection technology; multiple mirror false data are generated according to the data structure of the original data; differential noise is added to the original data and the multiple mirror false data respectively to generate original noise data and multiple mirror noise data; the original noise data and the multiple mirror noise data are combined to form a data group, and the data group is marked to form a privacy-protected data set. This method can solve the problem of distortion of a single piece of information in the original data after differential noise processing, ensure data accuracy while ensuring data security, and improve efficiency and accuracy in the process of data use.
[0098] Figure 6 It is a block diagram of an electronic device according to an exemplary embodiment.
[0099] Refer to the following Figure 6 hereinafter, an electronic device 600 according to this embodiment of the present application is described. Figure 6 The electronic device 600 shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0100] like Figure 6As shown, electronic device 600 is implemented as a general-purpose computing device. Components of electronic device 600 may include, but are not limited to, at least one processing unit 610, at least one storage unit 620, a bus 630 connecting various system components (including storage unit 620 and processing unit 610), a display unit 640, and the like.
[0101] The storage unit stores program codes, which can be executed by the processing unit 610, so that the processing unit 610 performs the steps described in this specification according to various exemplary embodiments of the present application. For example, the processing unit 610 can perform the following steps: Figure 1 , Figure 2 , Figure 3 Follow the steps shown in .
[0102] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access memory unit (RAM) 6201 and / or a cache memory unit 6202 , and may further include a read-only memory unit (ROM) 6203 .
[0103] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205, such program modules 6205 including but not limited to: an operating system, one or more application programs, other program modules and program data, each of which or some combination may include an implementation of a network environment.
[0104] Bus 630 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0105] The electronic device 600 can also communicate with one or more external devices 600' (e.g., a keyboard, a pointing device, a Bluetooth device, etc.), devices that allow a user to interact with the electronic device 600, and / or any device that allows the electronic device 600 to communicate with one or more other computing devices (e.g., a router, a modem, etc.). This communication can occur via an input / output (I / O) interface 650. Furthermore, the electronic device 600 can communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 via the bus 630. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 600, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0106] Through the above description of the embodiments, it is easy for those skilled in the art to understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Figure 7 As shown, the technical solution according to the embodiment of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes a number of instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiment of the present application.
[0107] The software product can be any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, a system, device or component of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0108] The computer-readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein the readable program code is carried. The data signal propagated may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable storage medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the readable storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.
[0109] The program code for performing the operations of the present application can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0110] The above-mentioned computer-readable medium carries one or more programs. When the above-mentioned one or more programs are executed by a device, the computer-readable medium implements the following functions: performing structured preprocessing on the original data to generate structured data; generating a set of differential noise for each data unit in the structured data based on differential privacy protection technology; generating multiple mirror false data according to the data structure of the original data; adding differential noise to the original data and the multiple mirror false data respectively to generate original noise data and multiple mirror noise data; combining the original noise data and the multiple mirror noise data to form a data group, and marking the data group to form a privacy-protected data set.
[0111] Those skilled in the art will appreciate that the modules described above can be distributed in the device according to the description of the embodiment, or can be modified accordingly to be used in one or more devices that are different from the embodiment. The modules of the above embodiment can be combined into one module or further divided into multiple submodules.
[0112] Through the description of the above embodiments, it is easy for those skilled in the art to understand that the example embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solution according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.
[0113] While the exemplary embodiments of the present application have been specifically illustrated and described above, it should be understood that the present application is not limited to the detailed structures, configurations, or implementations described herein; rather, the present application is intended to encompass various modifications and equivalent configurations within the spirit and scope of the appended claims.
Claims
1. A data processing method based on differential privacy protection, characterized in that: include: Perform structured preprocessing on the original data to generate structured data; generating a set of differential noise for each data unit in the structured data based on differential privacy protection technology; Generate multiple mirrored false data according to the data structure of the original data; Adding differential noise to the original data and the plurality of mirrored false data respectively to generate original noise data and a plurality of mirrored noise data; The original noise data and the plurality of mirror noise data are combined to form a data group, and the data group is marked to form a privacy-preserving data set.
2. The method according to claim 1, wherein Also includes: Extracting original noise data and corresponding multiple mirror noise data contained in each data group according to the label of the data group in the privacy-preserving dataset; Calculating the mean values of the original noise data and the corresponding data units in the plurality of mirror noise data to eliminate differential noise; Reconstruct the original data based on the calculated mean.
3. The method according to claim 1, wherein Perform structured preprocessing on raw data to generate structured data, including: Perform word segmentation on the original data to identify the semantic units in the data; Performing part-of-speech tagging and named entity recognition on the semantic units to extract structured fields; Structured data is generated according to the structured fields.
4. The method according to claim 1, wherein A set of differential noise is generated for each data unit in the structured data based on differential privacy protection technology, including: A set of differential noises is generated for each data unit in the structured data based on a Laplace distribution or a Gaussian distribution.
5. The method according to claim 4, wherein Generating a set of differential noises for each data unit in the structured data based on a Laplace distribution or a Gaussian distribution, including: Determine differential privacy parameters and computational sensitivity; generating a set of differential noise for each data unit in the structured data based on Laplace distribution or Gaussian distribution and the differential privacy parameter; or A set of differential noises is generated for each data unit in the structured data based on a Laplace distribution or a Gaussian distribution, the differential privacy parameter, and the computational sensitivity.
6. The method according to claim 1, wherein Generate multiple mirrored false data according to the data structure of the original data, including: Analyze the structural features of raw data and identify key data units; Generate multiple mirror data for each key data unit; The generated multiple mirror data are combined according to the original data structure to form multiple mirror false data.
7. The method according to claim 6, wherein Generate multiple false data for each key data unit, including: Based on the large model, multiple false data are generated for each key data unit according to the data structure and calculation sensitivity of each key data unit.
8. The method according to claim 1, wherein Adding differential noise to the original data and the plurality of mirrored false data to generate original noise data and a plurality of mirrored noise data, respectively, comprises: Adding the differential noise value to the data unit corresponding to the original data to form original noise data; adding the differential noise value to corresponding data units in the plurality of mirrored false data to form a plurality of mirrored noise data; The average value of the multiple differential noise values is 0 in each data unit.
9. The method according to claim 2, wherein Reconstruct the original data based on the calculated mean, including: Based on the marking information in the data group, identifying the same data group consisting of the original data and the corresponding multiple mirror noise data; Performing statistical mean calculation on the values of the corresponding fields in the data group to obtain an estimated value of each field; According to a preset mapping relationship, the estimated value is reversely mapped into the original structured field, thereby reconstructing the original data.
10. A data processing device based on differential privacy protection, characterized in that: include: Structuring module, used to perform structural preprocessing on raw data and generate structured data; A differential module, configured to generate a set of differential noise for each data unit in the structured data based on differential privacy protection technology; A mirror module, configured to generate a plurality of mirror false data according to the data structure of the original data; An adding module, configured to add differential noise to the original data and the plurality of mirrored false data to generate original noise data and a plurality of mirrored noise data; A data module is configured to combine the original noise data and the plurality of mirror noise data to form a data group, and mark the data group to form a privacy-preserving data set.