Methods, systems, and computer readable media for applying pairwise differential privacy to variables in dataset
By specifying random instance seed values and adaptive sensitivity parameters, additive noise is generated, which solves the problem that differential privacy technology in the prior art is difficult to deal with highly correlated data, and achieves efficient data privacy protection and data utility balance.
Patent Information
- Application Number
- CN202380073814.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-19
- Filing Date
- 2023-10-25
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to effectively apply differential privacy technologies without compromising data utility, especially when processing highly correlated data in the field of life sciences.
By specifying random instance seed values to variables in the dataset and generating additive noise based on the identified high correlation and adaptive sensitivity parameters, the Laplace mechanism is applied to generate pseudonymous datasets.
It realizes that the level of data privacy protection is improved without damaging the data utility, especially when processing highly correlated biochemical data, effectively retaining the statistical properties of the data.
Smart Images

Figure CN120092239A_ABST
Abstract
Description
[0001] Priority Claim
[0002] This application claims priority to U.S. Patent Application Serial No. 17 / 969,507, filed October 19, 2022, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0003] The subject matter described herein relates to data privacy and related noise generation techniques. More particularly, the subject matter described herein relates to methods, systems, and computer-readable media for applying pairwise differential privacy to variables in a dataset. Background Art
[0004] Differential privacy is a general term for mathematical methods that implement the functionality of the ε-differential privacy definition. It defines a quantitative measure of privacy that exists in a relational database. A method that provides ε = 1 privacy translates to a privacy guarantee that each entry in the dataset has approximately the same level of privacy as if its corresponding data were completely removed. One mathematical method for implementing ε-differential privacy is the Laplace mechanism that generates additive noise that is pseudo-randomly applied to successive values in the dataset.
[0005] Utilizing ε-differential privacy in real-world situations has been quickly determined to be infeasible in most cases because ε = 1 privacy distorts data excessively in many cases, reducing the utility of the modified data beyond an acceptable and / or useful state. Efforts to relax the expectations set by this definition have resulted in "Epsilon Delta" differential privacy (i.e., (ε,Δ)-differential privacy), where an additional parameter Δ is added to estimate the maximum probability that a privacy leak occurs. If the probability is characterized as "low" for a particular event, then the privacy requirements are more relaxed. Thus, the (ε,Δ)-differential privacy Laplace mechanism will produce less additive noise compared to its more stringent ε = 1 counterpart.
[0006] Notably, other proposed methods involve cases that operate on one or more values of the same variable. In the field of life sciences, which often produces biochemical and physical measurements, unique requirements arise. Specifically, there is a need for mutual correlation or relationship between two or more variables. Measurements obtained from the same sample (e.g., a blood sample analyzed by a mass spectrometer) can have relationships that are produced by the complex processes of the human body. For example, these processes can interact or interfere during the measurement process. To produce pseudonymization that is complex enough for life science use, the mutual correlation of the original data needs to be addressed.
[0007] Accordingly, there is a need for improved methods and systems for applying pairwise differential privacy to variables in a dataset. SUMMARY OF THE INVENTION
[0008] A method for applying pairwise differential privacy to variables in a dataset includes assigning a random instance seed value to a first dataset variable in an original dataset. The method further includes assigning the random instance seed value to the at least one additional dataset variable if a high correlation is identified between the first dataset variable and at least one additional dataset variable in the original dataset. The method further includes determining an adaptive sensitivity parameter corresponding to the first dataset variable. The method further includes generating additive noise by a noise generation manager using two or more of the first dataset variable, the random instance seed value, and / or the adaptive sensitivity parameter and applying the additive noise to the first dataset variable to produce a pseudonymized variable for inclusion in a pseudonymized dataset associated with the original dataset.
[0009] In another aspect of the subject matter described herein, the method of applying pairwise differential privacy to variables in a dataset is repeated for each remaining dataset value included in the original dataset.
[0010] In another aspect of the subject matter described herein, the high correlation is identified by an operator.
[0011] In another aspect of the subject matter described herein, the high correlation includes either a high positive correlation or a high negative correlation.
[0012] In another aspect of the subject matter described herein, the first dataset variable and the at least one additional dataset variable are biochemical data variables associated with a common subject sample.
[0013] In another aspect of the subject matter described herein, the adaptive sensitivity parameter scales with a numerical measurement value associated with the first dataset variable.
[0014] In another aspect of the subject matter described herein, the adaptive sensitivity parameter indicates the distribution range of the additive noise applied to the first dataset variable.
[0015] In another aspect of the subject matter described herein, the adaptive sensitivity parameter is used to establish the magnitude of the additive noise applied to the first dataset variable.
[0016] In another aspect of the subject matter described herein, the noise generation manager is a Laplace transform mechanism.
[0017] In another aspect of the subject matter described herein, the original data set includes a relational database. In another aspect of the subject matter described herein, a system for applying pairwise differential privacy to variables in a data set is provided. The system includes a computing platform that includes at least one processor and a memory. The system also includes a pairwise differential privacy (PDP) engine that includes a correlation manager and a noise generation manager (NGM) and is stored in the memory, and when executed by the at least one processor, the engine is configured to: assign a random instance seed value to a first data set variable in the original data set using the correlation manager; if a high correlation is identified between the first data set variable and at least one additional data set variable in the original data set, then assign the random instance seed value to the at least one additional data set variable using the correlation manager; determine an adaptive sensitivity parameter corresponding to the first data set variable using the correlation manager; and generate additive noise using two or more of the first data set variable, the random instance seed value, and / or the adaptive sensitivity parameter via the noise generation manager and apply the additive noise to the first data set variable to produce a pseudonymized variable to be included in a pseudonymized data set associated with the original data set.
[0018] In another aspect of the subject matter described herein, the correlation manager and the noise generation manager are configured to repeat each action for each remaining data set value included in the original data set.
[0019] In another aspect of the subject matter described herein, the high correlation is identified by an operator.
[0020] In another aspect of the subject matter described herein, the high correlation includes a high positive correlation or a high negative correlation.
[0021] In another aspect of the subject matter described herein, the first data set variable and the at least one additional data set variable are biochemical data variables associated with a common subject sample.
[0022] In another aspect of the subject matter described herein, the adaptive sensitivity parameter scales with a numerical measurement value associated with the first data set variable.
[0023] In another aspect of the subject matter described herein, the adaptive sensitivity parameter indicates the distribution range of the additive noise applied to the first data set variable.
[0024] In another aspect of the subject matter described herein, the adaptive sensitivity parameter is used to establish the magnitude of the additive noise applied to the first data set variable.
[0025] In another aspect of the subject matter described herein, the noise generation manager uses a Laplace transform mechanism.
[0026] According to another aspect of the subject matter described herein, a non-transitory computer-readable medium having executable instructions stored thereon that, when executed by a processor of a computer, control the computer to perform a method that includes assigning a random instance seed value to a first dataset variable in an original dataset. The method further includes assigning the random instance seed value to the at least one additional dataset variable if a high correlation is identified between the first dataset variable and at least one additional dataset variable in the original dataset. The method further includes determining an adaptive sensitivity parameter corresponding to the first dataset variable. The method further includes generating additive noise by a noise generation manager using two or more of the first dataset variable, the random instance seed value, and / or the adaptive sensitivity parameter and applying the additive noise to the first dataset variable to produce a pseudonymized variable to include in a pseudonymized dataset associated with the original dataset.
[0027] Exemplary computer-readable media suitable for implementing the subject matter described herein include non-transitory devices such as disk storage devices, chip memory devices, programmable logic devices, and application specific integrated circuits. Additionally, the computer-readable media implementing the subject matter described herein can be located on a single device or computing platform or can be distributed across multiple devices or computing platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The subject matter described herein will now be explained with reference to the drawings, in which:
[0029] Figure 1 illustrates an exemplary system configured to apply pairwise differential privacy to variables in a dataset;
[0030] Figure 2 illustrates exemplary pseudocode for applying pairwise differential privacy to variables in a dataset;
[0031] Figure 3 illustrates an exemplary database table generated and utilized by a pairwise differential privacy engine or algorithm;
[0032] Figure 4 illustrates a graph depicting sensitivity value testing with experimental data;
[0033] Figure 5 illustrates a graph depicting adaptive 3-level sensitivity versus fixed sensitivity;
[0034] Figure 6 illustrates a graph depicting 3-level sensitivity versus percentage sensitivity;
[0035] Figure 7 illustrates a chart depicting example nuchal translucency (NT) multiple of the median (MOM) summaries for various data sources;
[0036] Figure 8 A graph depicting the NT MoM summary of the raw data and the ε = 1 data without gestational age (GA) pseudonymization is shown;
[0037] Figure 9 A graph depicting the association between the raw pregnancy-associated plasma protein A (PAPPA) and PAPPA MoM is shown; and
[0038] Figure 10 is a flowchart illustrating an exemplary process for applying pairwise differential privacy to variables in a dataset. DETAILED DESCRIPTION
[0039] With the emergence of new data security / confidentiality regulations and standards affecting software solutions provided to enterprise customers, methods for enabling software service providers to ensure compliance have become increasingly important. This is especially true in cases where employees of software service providers are provided access to private data belonging to customers. Generally, when a customer shares their data (e.g., for customer support, research and development activities, etc.), added value is generated, but this data sharing process often poses a potential threat of leakage of private customer data to software service providers.
[0040] For example, accidental data leakage may occur when email files are shared or when an employee's unlocked laptop is stolen. In these cases, conventional data security methods cannot protect the confidential data of customer users (e.g., patients) when access to private data is obtained in an improper manner. There is a need to address reconstruction, database linking, and re-identification attacks on confidential patient data in software products.
[0041] The present subject matter discloses a pairwise differential privacy method that allows for the retention of data utility without compromising the underlying source data (e.g., patient data). In particular, the disclosed subject matter relates to pairwise (ε, Δ) differential privacy techniques that include a method for creating a pseudonymized dataset from a highly correlated raw dataset. Currently, there is no existing method for retaining patient data privacy without affecting the utility of the data. In some embodiments, the pseudorandom decisions involved in the method are instantiated for each individual observation. If an observation contains variables known to be correlated (e.g., based on or identified by domain knowledge), then the randomness of the noise applied is fixed to a common constant value for each of these correlated variables.
[0042] A second aspect of the disclosed subject matter is an extension of a sensitivity parameter, which is typically derived from domain knowledge. In contrast, the disclosed system utilizes an adaptive sensitivity parameter that scales with the value being operated on. In some embodiments, this functionality can be achieved using a normalized percentage.
[0043] Reference will now be made in detail to various embodiments of the subject matter described herein, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numerals will be used throughout the drawings to refer to the same or like parts.
[0044] Figure 1 An exemplary system configured to apply pairwise differential privacy to variables in a dataset is illustrated. That is, Figure 1 System 100 is illustrated, which includes a client-side host 102 and a data management host 112 (e.g., a cloud-based software-as-a-service provider) communicatively coupled via a network 130 (e.g., the Internet). In some embodiments, host 102 can represent any one or more suitable entities (e.g., one or more nodes, devices, or computing platforms) for performing various aspects associated with applying pairwise differential privacy to variables in a dataset. In some embodiments, host 102 can represent or include a host server, a personal computer, a laptop computer, a tablet device, or any other suitable computing device accessible to a user (e.g., a patient or clinician user). As Figure 1 shown, host 102 includes one or more processors 104 and a memory 106. In some embodiments, the (one or more) processors 104 include a microprocessor, such as a central processing unit (CPU), or any other hardware-based processor unit. The processor 104 is configured to execute and / or utilize software to execute a plurality of software components stored in the memory 106. Similarly, the memory 106 can include non-transitory computer-readable media, such as disk storage devices, chip memory devices, programmable logic devices, and application specific integrated circuits. Similarly, the data management host 112 (similar to host 102) can also include a host server, a personal computer, a laptop computer, a tablet device, or any other suitable computing device including one or more processors 114 (not unlike processor 104). Likewise, the data management host 112 includes a memory 116 comparable to the memory 106 in host 102.
[0045] In Figure 1In [the figure], the memory 106 includes a paired differential privacy (PDP) engine 108 and one or more application programming interface (API) elements 110. In some embodiments, the PDP engine 108 includes a software program or algorithm executed by the processor 104 and stored in the memory 106. The PDP engine 108 may include hardware, software, and / or firmware components for implementing the management and execution of applying paired differential privacy to variables in a dataset as described herein. In an exemplary implementation, the PDP engine 108 includes functionality for adding domain-related noise to an unchanged and / or original matrix-like dataset (e.g., dataset R) in a pseudo-random manner. Notably, a resulting derived dataset (e.g., dataset D) is generated from the original dataset that has undergone paired differential privacy processing and / or additive noise by the PDP engine 108. In some embodiments, the PDP engine 108 includes a correlation manager (CM) 107 and a noise generation manager (NGM) 109. Notably, each of the correlation manager 107 and the NGM 109 may include a software program portion or component (e.g., an algorithm, pseudocode, etc.) implemented as the entire software program of the PDP engine 108. Additionally, each of the correlation manager 107 and the NGM 109 may also be stored in the memory and executed by the processor.
[0046] Notably, while the customer owns the original, unchanged data R, an employee of the software service company receives the D dataset from the customer to perform any data-related tasks. The noise is carefully customized for the specific task, so the pseudonymization via differential privacy does not affect the conclusions drawn from the task. As used herein, pseudonymization may refer to a de-identification procedure by which the personally identifiable information fields in a data record are replaced with one or more artificial identifiers or pseudonyms. Thus, even if a third party gains access to the dataset D, this data cannot be automatically linked to the confidential true dataset D owned by the customer, thereby increasing data privacy in a passive manner. In the dataset D, instead of changing each value, only some of the observations contain the changed values, making the DP method performed by the PDP engine 108 (and / or the correlation manager 107) pseudo-random.
[0047] In some embodiments, the host 102 may process the customer's original, unchanged data R to derive the D dataset, which may be securely transmitted and / or transferred to a cloud-based service provider host 112 suitable for performing any and / or specific data-related tasks.
[0048] In some embodiments, the noise is carefully customized to be task - specific by the Noise Generation Manager (NGM) 109 of the PDP engine 108. In particular, the PDP engine 108 and / or the NGM 109 are configured to generate noise in a manner that does not affect the conclusions drawn from the task for pseudonymization via differential privacy. Even if a third party gains access to the dataset D generated by the PDP engine 108, this data cannot be automatically linked to the confidential real dataset R that the customer still owns, thereby increasing data privacy in a passive manner.
[0049] Once the dataset D is transferred from the host 102 to the host 112 (e.g., via one or more APIs 110), the host 112 is configured to use the dataset D as an input to the Pseudo - Random Data (PD) manager 118, which is configured to generate statistical reports. Notably, the statistical reports generated using the pseudonymized data of the dataset D will yield the same conclusions as those generated by the PD manager 118 using the original dataset R.
[0050] Figure 2 An exemplary algorithm or software component that can be executed by the PDP engine 108 (and / or the Correlation Manager 107 and NGM 109) is illustrated as example pseudocode 200. It can be understood that the pseudocode 200 can be implemented using any computer code or language (e.g., the C# programming language) without departing from the scope of the disclosed subject matter. In some embodiments, the pseudocode 200 can be used (e.g., by the PDP engine 108 and / or its Correlation Manager 107 depicted in Figure 1 ) to process the raw customer data represented in the data table 301(t 1 ) and the correlation - preserving data represented in the data table 302(t 2 ). More specifically, the raw data table 301 includes multiple data entries (i.e., rows) of customer data values associated with multiple variables (e.g., x 1 –x 4 columns). Notably, the data table 301 contains the raw data (e.g., private customer data or R dataset) to be pseudorandomized by the PDP engine (e.g., to produce the dataset D). In some embodiments, the correlation - preserving data table 302 is created by a separate analysis and / or supplied by the domain knowledge of an expert or system administrator. Notably, the correlation - preserving data table 302 can include a table accessible by the system administrator that specifies which variables included in the data table 301 are relevant. For example, the first row in the data table 302 indicates that the variable x 1 is relevant to x 2 . The second row in the data table 302 indicates that the variable x 1 is relevant to x 4is relevant. The third row in the data table 302 indicates the variable x 2 is relevant to x 4 is relevant. Thus, relevance is retained where two (or more) variables exist in a row of the data table 302 indicating that the variables have a high negative or positive correlation generated by domain knowledge or an expert, a system administrator, or other sources of knowledge about the relationships between data variables in a particular domain.
[0051] Figure 3 also includes a random seed value table 303 (t 3 ), which includes a plurality of random seed values (for the variables in table 301) generated by the PDP engine (and / or its relevance manager). In some embodiments, the PDP engine and / or the relevance manager may be configured to use the pseudocode 200 to generate the loop seed value data in the data table 303. For example, entries for the random seed value data table 303 are created in lines 7 - 8 of the pseudocode 200. Moreover, Figure 3 the data table 304 in is the differential result table, which contains differential private values (i.e., pseudorandomized data values) corresponding to the original data values included in the original data table 301. In some embodiments, the PDP engine and / or the relevance manager may provide the pseudorandomized data (at least a portion thereof) in table 304 to a data management host (e.g., a software - as - a - service provider) to generate a statistical report that produces the same conclusions as those produced when using the underlying original data (in table 301).
[0052] Returning to Figure 2 , in line 0 of the pseudocode 200, the PDP engine and / or the relevance manager may be configured to (e.g., by a system administrator) set or define each of the epsilon value, the delta value, and the parameter sensitivity percentage value to be used in the execution of the pseudocode 200. The data table 303 is also reset and / or cleared in line 0. In line 1 of the pseudocode 200, the PDP engine (and / or the relevance manager) is configured to execute a loop function that iterates over each row of the data table 301. In line 2, the PDP engine includes a pseudorandom generator component that is configured to generate a random instance seed value (i.e., "instance_seed") initially stored in a local buffer.
[0053] In lines 3-4 of the pseudocode 200, the PDP engine and / or the correlation manager perform iterative data processing using the raw data (e.g., private patient data) in the data table 301 and the correlation retention data (e.g., variable correlations that can be predefined by the system administrator) in the data table 302. Notably, the PDP engine (and / or the correlation manager) executes lines 3-4 together with line 5 of the pseudocode 200 to determine whether the raw data variables included in the data table 301 are similarly included or contained in the correlation retention data table 302. If the PDP engine (and / or the correlation manager) finds a matching entry in the data table 302, then the PDP engine (and / or the correlation manager) designates the matching variable as a relevant variable (i.e., line 5), thereby determining to retain a certain level of correlation. For example, after processing line 5 of the pseudocode 200, the PDP engine (and / or the correlation manager) can determine that x is found in the first two rows of table 302. 1 parameter. Then, the PDP engine (and / or the correlation manager) calculates and / or correlates variable pairs containing the x 1 variable, and in this example, the variable pair includes relevant data = [x 2 , x 4 .
[0054] In line 6 of the pseudocode 200, the PDP engine (and / or the correlation manager) tests whether the table 303 contains a seed related to the parameter x 1 . During the first iteration of the pseudocode 200, before executing lines 6-8, the data table 303 is initially empty (i.e., there are no data entries in the random seed value data table). In lines 7-8 of the pseudocode 200, the PDP engine (and / or the correlation manager) is configured to append new vectors (or value entries) to the data table 303. Once lines 6-8 of the pseudocode 200 are processed, the PDP engine (and / or the correlation manager) adds one or more rows to the data table 303. For example, the PDP engine adds [x 1 , instance_seed], [x 2 , instance_seed], and [x 4 , instance_seed] as row entries to table 303. Notably, a random instance seed value was previously generated by the pseudorandom generator (e.g., see line 2). In this way, when iterating over each of the variables x 2 and x 4 , their associated seed values are respectively selected from the table 303 (e.g., selected by the PDP engine and / or the correlation manager), such that x 2 and x 4 use the same random instance seed value as x 1 (i.e., because x1 previously determined to be related to x 2 and x 4 (related). Notably, in pseudocode lines 7 - 8, variable and seed value pairs can be appended to the temporary table 303 so that these values can be checked and selected by the IF ELSE construct of the pseudocode 200.
[0055] In lines 9 - 10 of the pseudocode 200, the PDP engine (and / or the relevance manager) determines that the data table 303 contains the variable currently being processed and sets the instance seed value to the instance seed previously determined from the table 303 (if there already exists an instance seed associated with the variable). Notably, this ELSE construct in lines 9 - 10 is included because if the previous IF construct tests as "FALSE" (i.e., a record of the iterated variable or any of its related variables already exists in the table 303), then the associated random seed is selected from the table 303.
[0056] In line 11 of the pseudocode 200, the PDP engine calculates the noisy result (using the previously defined (e.g., see line 0) epsilon value, delta value, and percentage value parameter values as inputs to its noise generation manager). In some embodiments, the noise generation manager can include a Laplace mechanism, such as the "EpsilonDeltaLaplaceNoise" function, which is configured to apply deterministic noise determined by a random seed value in addition to inputs including the value x, epsilon value, delta value, sensitivity value (i.e., produce a noisy output value). Notably, the PDP engine also includes a "RelativeSensitivity" function that receives the "x value" as an input and returns a "y percentage" value representing the sensitivity level as an output. This calculated sensitivity value is further used by the PDP engine as an input for determining the noisy output value result as mentioned above. In some embodiments, the sensitivity level can be defined by using an adaptive sensitivity parameter that scales with the value being operated on. In some embodiments, a normalized percentage can be used by the PDP engine to achieve this functionality.
[0057] In some embodiments, the PDP engine adds the noisy x 1 value to the "row_result vector" (e.g., see line 12 of the pseudocode 200), which is subsequently appended as a vector entry to the first row of the differential result data table 304 (e.g., see line 13) (t 4)。At this stage (e.g., line 14), the PDP engine resets the random seed value data table 303 so that the data table 303 can be repopulated for the second row of the original data table 301. In line 15 of the pseudocode 200, a statistical report representing the selected data in table 304 is generated.
[0058] Referring to the differential result data table 304, it should be noted that since each of the variables x 1 , x 2 and x 4 uses the same random seed as the noise generation manager of the PDP engine (e.g., the EpsilonDeltaLaplaceNoise function), the noise added to each of these variables is the same relative to the values being differentiated. The "upward arrow" shown in the data table 304 indicates that the same amount of noise has been applied to each of the variables x 1 , x 2 and x 4 .
[0059] Returning to Figure 1 , the PDP engine 108 can select at least a portion of the resulting pseudorandomized data generated in the data table 304 and send it to the data management host 112 using the API 110. Similarly, the host 112 includes a pseudorandomized data (PD) manager 118 configured to receive the pseudorandomized data generated by the PDP engine 108. In particular, the PD manager 118 is configured to generate a statistical report using the pseudorandomized data received from the host 102 (e.g., via the API 120). It is worth noting that the statistical report can include the same information and / or conclusions as the information and / or conclusions generated when the PD manager 118 uses the original data processed by the PDP engine 108.
[0060] The disclosed subject matter also pertains to adding differential privacy functionality to the customer service process of software products, such as laboratory data management and statistical services. For example, the disclosed subject matter can be optimally utilized under the following conditions:
[0061] · First trimester screening data is sent from a customer laboratory (e.g., a client-side host computer) to a statistical team (e.g., a software service provider host computer) for evaluation. This sending of screening data can be triggered in response to a customer's inquiry about the quality of the (private) biomarker data they possess.
[0062] · The customer is primarily concerned with trends, such as changes or level differences over time, and less concerned with individual observations. One task of the statistician is to investigate these population-level summary statistics and provide a statistical report to the customer.
[0063] · Although the "statistical export" version of the received data does not include easily identifiable information (e.g., customer name and social security number), the data often contains numerical information that may be identifiable. This is the information that differential privacy is designed to anonymize in a sufficient manner.
[0064] · The cloud capacity of the laboratory data management software enables the operator to implement a differential privacy solution for the client applications of existing customers (e.g., the PDP engine at the client-side host) for local pseudonymization and push the anonymized data to the cloud storage device associated with the software service provider host.
[0065] Methods and Materials
[0066] In some embodiments, the "statistical export" file designed for private raw data collection may include 10,000 observation batches containing multiple variables, some of which are used in early pregnancy risk prediction, while others constitute additional information. The statistical export file also contains different data types, which is important in terms of differential privacy because different mechanisms can be used for different data types. Moreover, changing some variables has a greater impact on the conclusions of a statistical survey than others. For example, changing the categorical coding of race (e.g., "1" for Caucasian can be changed to "2" for East Asian) for adding pseudonymization has a significant impact on how risk modeling is performed, thus significantly changing the overall risk score. This is not the case for numerical variables such as the median multiple (MoM) of a biomarker, where a deviation of appropriate size does not significantly change the risk scoring result. In some cases, MoM represents the biomarker concentration / median of the patient population or a more complex formula considering gestational age (GA). It should be noted that a specific biomarker has a specific MoM formula.
[0067] Therefore, only the following numerical variables are considered for differential privacy pseudonymization in this example:
[0068] · Pregnancy-related measurements
[0069] ○ GA (Gestational Age)
[0070] ○ BPD (Biparietal Diameter)
[0071] ○ BPD2 (Twins)
[0072] ○ CRL (Crown-Rump Length)
[0073] ○ CRL2 (Twins)
[0074] ○ HC (Head Circumference)
[0075] ○ HC2 (Twins)
[0076] ○LMP (Last Menstrual Period)
[0077] ○Maternal age
[0078] ○Maternal weight
[0079] · Biomarker and biophysical measurements
[0080] ○AFP (Alpha-Fetoprotein)
[0081] ○hCG (Human Chorionic Gonadotropin)
[0082] ○hCGb (Human Chorionic Gonadotropin, beta subunit)
[0083] ○uE3 (Unconjugated Estriol)
[0084] ○PAPPA (Pregnancy-Associated Plasma Protein A)
[0085] ○Inhibin-A
[0086] ○NT (Nuchal Translucency)
[0087] ○NT2 (Twins)
[0088] ○uE3Upd (Unconjugated Estriol)
[0089] ○PlGF (Placental Growth Factor)
[0090] ○MAP (Mean Arterial Pressure)
[0091] ○UTPi (Uterine Artery Pulsatility Index)
[0092] ○sFlt1 (Soluble fms-like tyrosine kinase 1)
[0093] ○AFP MoM
[0094] ○HCG MoM
[0095] ○HCGb MoM
[0096] ○uE3 MoM
[0097] ○PAPPA MoM
[0098] ○Inhibin-A MoM
[0099] ○NT MoM
[0100] ○uE3Upd MoM
[0101] ○PlGF MoM
[0102] ○MAP MoM
[0103] ○UTPiMoM
[0104] ○sFlt1 MoM
[0105] A viable method for implementing differential privacy for continuous variables is to use the Laplace mechanism, where noise generated from the Laplace distribution can be added to the values. It follows the definition of the differential privacy mechanism. Differential privacy in its current form “Epsilon-delta differential privacy” utilizes three parameters: the Δ value parameter, the ε value parameter, and iii) the adaptive sensitivity value parameter (or percentage). These parameters should be selected based on their applicability to tasks related to pseudonymization, which in this scenario is the laboratory data statistics service. Delta or Δ can represent the (estimated) probability of data leakage in the system and can be specified in some embodiments as:
[0106] Δ = 1 / data observations,
[0107] It implements a more practical (ε,Δ) differential privacy rather than the more restrictive (ε,0) differential privacy that has limited applications in the real world. Thereafter, for all experiments, Δ was fixed at 0.0001. Epsilon or ε directly affects the amount of anonymity preserved, as ε can represent the available privacy budget (and / or privacy budget ceiling). In many embodiments, an ε equal to 1 can be used. Thus, this parameter can be fixed to be equal to 1. The sensitivity, the parameter that determines the relative amount of noise added, can also be determined iteratively during the study. In some embodiments, the statistical analysis software R and RStudio can be used to generate various outputs.
[0108] Experimental Results
[0109] Given a simple toy data of a continuous variable, the first iteration of the code was written. An additive Laplace noise function was written in R that supports the (ε,Δ) differential privacy parameters ε, Δ, and sensitivity. In some embodiments, ε and Δ are fixed, so the initial tests involve determining the appropriate value of the sensitivity parameter. Figure 4 Different magnitudes of sensitivity values were shown, where a sensitivity of 10 was considered too extreme to be used for any type of biomarker data as it produces too much variation between the true and derived values. In particular, Figure 4Graphs 401 - 404 are illustrated, each graph depicting the use of different sensitivity value tests. For example, graph 401 illustrates the sensitivity value test where the sensitivity value is set to 10. Similarly, graph 402 illustrates the use of a sensitivity value equal to 1, graph 403 illustrates the use of a sensitivity value equal to 0.1, and graph 404 illustrates the use of a sensitivity value equal to 0.01. It is worth noting that the lower the sensitivity value used, the smaller the deviation exhibited by the resulting data.
[0110] At this point, the design limitations of different variables have been fully realized because, compared to biomarker concentration and demographic information, MoM values approximately in the range (0, 20] have greater limitations on noise. For example, a MoM value near 1 is considered normal compared to the patient median, values greater than 1 are considered elevated. Also, compared to the patient median, values less than 1 and greater than zero are considered decreased. Thus, on the "positive" side (e.g., greater than 1), MoM behavior is linear in a way that deviates from the patient median, while on the "negative" side of 0 < x < 1, due to the division in the MoM formula, it exhibits non - linearity. This means that the additive noise mechanism needs to address this issue and does not use a fixed sensitivity parameter. Additionally, due to the added noise, a transition from the "positive" side to the "negative" side is not allowed. It is worth noting that MoM values in (0,.., 1) cannot be changed to > 1, and (1,.., +∞] cannot be changed to < 1, but values in [0.95,.., 1) and (1,.., 1.05] can be rounded to 1. This information indicates that the amount of positive or negative noise should be related to the value being operated on.
[0111] The second iteration can include a conditional structure for sensitivity, where:
[0112] 1. If the value is less than 10, then sensitivity = 0.05. Otherwise:
[0113] 2. If the value >= 10 and < 1000, then the sensitivity is 0.5. Otherwise:
[0114] 3. If the value >= 1000, then the sensitivity is 5.
[0115] Compared to the fixed noise as shown in Figure 5 This results in a more adaptive noise addition. In particular, graphs 501 - 504 illustrate an example of adaptive 3 - level sensitivity compared to a fixed sensitivity of 0.5. As the sensitivity of 0.5 starts to be too large for values near 1, the deviation from the actual value increases to an unacceptable proportion, while the 3 - level uses a sensitivity of 0.05 for values near 1. For example, Figure 5 shows an improvement in the reduction of deviation near 1, but it introduces a rough step in terms of deviation, which is inFigure 5 In the “Sensitivity = Adaptive, level 3, logarithmic scale, up to 50” sub - graph (i.e., graph 502), it is obvious within the range after an input of 10. To increase adaptability and universality (for example, an IF ELSE structure may not be applicable to other products), the sensitivity will be determined as the relative percentage of the value to be anonymized.
[0116] Figure 6 It shows that at the first two step - lengths of the level 3 mechanism, the sensitivity with respect to the value behaves similarly, but after a value of 10 or higher, the amount of added noise increases. Notably, adding 5% noise is more intuitive compared to any IF ELSE structure, and it can be shown that it does not violate the rules related to the MoM value. Refer to Figure 6 , the level 3 sensitivity is represented by row 601, while the 5% percentage sensitivity is represented by row 602. In Figure 6 , the y - axis is represented on a logarithmic scale because it demonstrates the more dynamic behavior of the percentage mechanism.
[0117] At this point during the experiment, the sensitivity mechanism is feasibly implemented on the false / experimental data set, so the first - round experiment on the real data set has been completed. In some embodiments, subjective evaluation by domain experts can be used to evaluate the differential privacy method. For example, statisticians can generate reports with real and false data and investigate whether the same conclusions can be drawn using these two data sets. The ε parameter is of concern at this stage, so the statisticians generated three (3) reports with different ε values:
[0118] · Statistical report with original data
[0119] · Statistical report with derived data, using ε = 2
[0120] · Statistical report with derived data, using ε = 4
[0121] Notably, gestational age (GA) can be used to group biomarker results, but when differentiating it, this creates groups that did not originally exist in the data set. Figure 7 This is demonstrated using NT MoM, and the risk recalculation is highly affected by this deviation from the original data. Figure 7 Graphs 701 - 703 are illustrated, which provide a summary of NT MoM for the original data (e.g., graph 701) and the differentiated data with ε = 2 (e.g., graph 702) and ε = 3 (e.g., graph 703). Notably, pseudonymization creates new GA - week groups that did not exist in the real data.
[0122] In particular, the following variables are set to be non-anonymous: "BPD", "BPD2", "CRL", "CRL2", "HC", "HC2", "gestational age", and "LMP". The statistical report is recalculated, now with ε = 1, since ε = 2 is still viable, and thus the limits of this parameter are also studied. Figure 8 The figure illustrates the differences between the new iteration and the report completed with the original data. The summaries by GA week are almost identical, and it is confirmed that the same results are achieved for all biomarkers using the original and anonymized datasets. Also, ε = 1 is considered suitable for this experiment. In Figure 8 Figure 801 illustrates the NT MoM summary of the ε = 1 data without GA pseudonymization (i.e., more stringent than 2 or 4), while Figure 802 depicts the NT MoM summary of the original data. Notably, the summaries are almost identical.
[0123] In some embodiments, the sensitivity mechanism can be reworked to generate an additional noise of +-3%. This is mainly due to the requirements set by changing the MoM values, and thus the transition of MoM = 1 from the positive side to the negative side (and vice versa) cannot be calculated. For example, MoM = 1.05 and 0.95 (small positive and negative effects) cannot be rounded to 1 after a 3% differentiation. The report is recalculated and it is checked that the conclusions have not changed.
[0124] At this stage of the experiment, it is noted that biomarkers can be represented as multiple variables, concentrations, and MoM results (and derivatives of MoM, such as Log MoM). The noise generation manager (e.g., the Laplace mechanism) is not aware of this relationship and theoretically can generate a situation where 3% positive noise is added to the patient's concentration result and 3% negative noise is added to the MoM result. This breaks the association that the concentration and MoM values can have. The correction for this is to use paired differentiation, where the same random seed is used for two values of the same patient, thus generating a deviation in the phasor magnitude and direction.
[0125] A new derived dataset is created using the paired mechanism, and the statistical report is recreated. The overall conclusions have not changed, which is expected since the concentration is not checked in the reporting protocol. After this check, the correlation is used to see if the paired mechanism preserves the relationship between biomarker concentration and MoM. In Figure 9 Figure 803, the method that does not consider the association reduces the correlation between pregnancy-associated plasma protein A (PAPPA) and PAPPA MoM by 7.4%, while the disclosed paired mechanism (e.g., the PDP engine) reduces this correlation by a negligible 0.2%. The results demonstrate the feasibility of the disclosed subject matter. In particular, Figure 9Illustrated is a comparison of the original PAPPA concentration versus the PAPPA MoM correlation (Chart 901) with naive differentiation (Chart 902) and paired differentiation (Chart 903). Notably, paired differentiation does not significantly change the correlation (0.2%), while the naive method changes the correlation by 7.4%.
[0126] In some embodiments, the disclosed subject matter (e.g., the PDP engine 108 in Figure 1 can be configured to identify and support both positive and negative correlations. Notably, the type of correlation employed or identified depends on the analysis required by the use case or the problem the operator is solving (e.g., whether it is desirable and / or preferred to retain the negative or positive correlation of two or more variables). For example, one application may require retaining a highly positive correlation between BMI and a certain biomarker, while another application may want to retain a highly negative correlation between the age of a subject / patient and a designated biomarker).
[0127] Exemplary Embodiments
[0128] In some embodiments, program parameters (e.g., parameters for the pseudocode 200 and / or the PDP engine) can be stored in a JSON file. As indicated above, the program parameters can include Δ, ε, and sensitivity (e.g., either including three values and two thresholds or including a single percentage value). The parameters can also include a list of differential privacy column names and a list of differential privacy groups, such that each group contains the names of those differential privacy columns to which the same noise percentage is to be applied. A differential privacy column can belong to at most one differential privacy group, but it does not need to belong to any group. A separate non-differential privacy column can be designated as the primary key column and all its values must be unique. For example, its column name does not appear in the list of differential privacy column names).
[0129] In some embodiments, a program (e.g., the pseudocode 200 and / or the PDP engine) can read a CSV file line by line. The first line in the file contains the column names for mapping the configuration data to the column indices. For each line, the value of the primary key column is retrieved and its SHA256 hash value is calculated (e.g., by the PDP engine). The first few bytes of the hash are converted by the PDP engine into an integer, which can be used as a seed for a random number generator for this line. Notably, this primary key column can later be removed or its value can be replaced with its hash value. Thus, the customer site can repeat the anonymization and obtain the same result, but cannot retrieve the original value from the result).
[0130] In some embodiments, all differentially private columns (sum values) are processed by the PDP engine in the order in which the columns appear in the JSON file. If a column value is null or cannot be converted to a double-precision value, then the value remains unchanged as it may contain, for example, the string "N / A" to indicate a missing value. If the conversion is successful, then the original string representation of the column value is examined to determine whether i) the value is an integer value without a decimal point, ii) a real value expressed in exponential notation, or iii) a real value with a decimal point and a fractional part but without an exponent. For a real number without an exponent, the PDP engine determines the number of digits in its fraction. Similarly, for a real number with an exponent, the precision is determined. After the noise is applied to the value by the PDP engine (and / or its noise generation manager), the pseudorandom value is converted to a string so that it has the same format as the associated original data value.
[0131] If a differentially private column does not belong to any group, then the random number generator instance of the row being processed is used to generate a noise percentage using the configured Δ, ε, and sensitivity. In some embodiments, the sensitivity can be calculated by the PDP engine as a fixed percentage of the input value. However, if a differentially private column belongs to a group, then the column is first examined to determine whether a noise percentage has already been calculated for this group. In some embodiments, each row in the CSV file can have a dictionary of group names and their corresponding noise percentages. If a noise percentage has not been calculated for this group and row, then a new noise percentage is calculated by the PDP engine (and / or the noise generation manager) in the same manner as for columns that do not belong to any group. Then, the PDP engine (and / or the noise generation manager) can add this noise value to the dictionary so that the values of other columns in this group can be found when processing this row.
[0132] It is noted that after the noise is applied, the differentially private column value cannot become zero or negative. If this occurs, then a new random noise percentage is calculated by the PDP engine (and / or the noise generation manager) until the result is positive.
[0133] Figure 10 is a flowchart illustrating an example process for applying pairwise differential privacy to variables in a dataset in accordance with embodiments of the subject matter described herein. In some embodiments, Figure 10 The method 1000 depicted in Figure 2 can be an algorithm, program, pseudocode (e.g., the pseudocode 200 in ), or script stored in a memory that, when executed by a processor, performs the steps set forth in blocks 1002 - 1008. In some embodiments, the method 1000 represents a list of steps implemented via software code programming and / or the logic of the PDP engine.
[0134] In block 1002, method 1000 includes assigning a random instance seed value to a first dataset variable in an original dataset. In some embodiments, the PDP engine (and / or its correlation manager) is configured to calculate the random instance seed value. For example, the PDP engine can generate a random instance seed value for each row of an original data table (e.g., the table 301 in Figure 3 ).
[0135] In block 1004, method 1000 includes assigning the random instance seed value to the at least one additional dataset variable if a high correlation is identified between the first dataset variable and at least one additional dataset variable in the original dataset. In some embodiments, the PDP engine (and / or correlation manager) accesses a correlation retention table (e.g., the table 302 in Figure 3 ), which includes data entries indicating correlations between different dataset variables (e.g., vector parameters). For example, the correlation retention table 302 can be configured by a system administrator who can access domain knowledge allowing an administrator to define vector parameter correlations. In some embodiments, a high correlation can include a highly positive or negative correlation identified by an operator. When the PDP engine or correlation manager finds two or more variables in the same row of the correlation retention table 302, the PDP engine or correlation manager can identify that there is a high correlation between the variables. Thus, determining whether there is a high correlation between variables can include the PDP engine or correlation manager performing a lookup in the correlation retention table 302 to determine whether the first dataset variable identified in step 1002 exists in any row of the correlation retention table 302 and is thus highly correlated with any other variable in the same row.
[0136] In block 1006, method 1000 includes determining an adaptive sensitivity parameter corresponding to the first dataset variable. In some embodiments, the PDP engine (and / or correlation manager) is configured to determine the adaptive sensitivity parameter by scaling a sensitivity value relative to the magnitude of the value being manipulated. In some embodiments, the PDP engine (and / or correlation manager) can be configured to determine the adaptive sensitivity parameter as a normalized percentage value and ε and Δ values for each vector variable being processed. In some embodiments, the adaptive sensitivity parameter is quantified as a percentage.
[0137] In block 1006, method 1000 includes generating additive noise by a noise generation manager (and / or the PDP engine) using two or more of the first dataset variable, the random instance seed value, and / or the adaptive sensitivity parameter and applying it to the first dataset variable to produce a pseudonymized variable to include in a pseudonymized dataset associated with the original dataset.
[0138] In some embodiments, blocks 1002 - 1008 are repeatedly executed by the PDP engine to generate all of the pseudorandomized data supplied in a differential data table (e.g., data table 304 in Figure 3 ). Once the pseudorandomized data is generated, at least a portion of the data is sent by the PDP engine to the data management server. It is noted that the data management server can subsequently use the received pseudorandomized data to generate a statistical report.
[0139] As described above, the disclosed subject matter enables the PDP engine to produce pseudorandomized data that can be used by a data management entity to securely generate a statistical report. It is noted that when compared to the original data, using the pseudorandomized data in this manner can generate the same conclusions. However, with differential privacy, the amount of privacy risk within the data set is significantly reduced. It is noted that the disclosed subject matter provides a sensitivity mechanism applicable to other products and situations, thereby creating an appropriate set of default parameters for any differential privacy implementation.
[0140] It will be understood that various details of the presently disclosed subject matter may be changed without departing from the scope of the presently disclosed subject matter. Additionally, the foregoing description is for illustrative purposes only and not for purposes of limitation.
Claims
1. A method for applying pairwise differential privacy to variables in a dataset, the method comprises: assigning a random instance seed value to a first dataset variable in an original dataset; if a high correlation is identified between the first dataset variable and at least one additional dataset variable in the original dataset, then assigning the random instance seed value to the at least one additional dataset variable; determining an adaptive sensitivity parameter corresponding to the first dataset variable; and generating additive noise by a noise generation manager using two or more of the first dataset variable, the random instance seed value, and / or the adaptive sensitivity parameter and applying the additive noise to the first dataset variable to produce a pseudonymized variable for inclusion in a pseudonymized dataset associated with the original dataset.
2. The method according to claim 1, wherein the method is repeated for each remaining dataset value included in the original dataset.
3. The method according to claim 1 or 2, wherein the high correlation is identified by an operator.
4. The method according to claim 3, wherein the high correlation includes either a high positive correlation or a high negative correlation.
5. The method according to any one of the preceding claims, wherein the first dataset variable and the at least one additional dataset variable are biochemical data variables associated with a common subject sample.
6. The method according to any one of the preceding claims, wherein the adaptive sensitivity parameter scales with a numerical measurement value associated with the first dataset variable.
7. The method according to any one of the preceding claims, wherein the adaptive sensitivity parameter indicates a distribution range of the additive noise applied to the first dataset variable.
8. The method according to claim 7, wherein the adaptive sensitivity parameter is used to establish a magnitude of the additive noise applied to the first dataset variable.
9. The method according to any one of the preceding claims, wherein the noise generation manager is a Laplace transform mechanism.
10. The method according to any one of the preceding claims, wherein the original dataset includes a relational database.
11. A system for applying pairwise differential privacy to variables in a dataset, the system comprises: a computing platform including at least one processor and a memory; and a pairwise differential privacy (PDP) engine including a correlation manager and a noise generation manager (NGM) and stored in the memory, and when executed by the at least one processor, the pairwise differential privacy (PDP) engine is configured to: assign a random instance seed value to a first dataset variable in an original dataset using the correlation manager; if a high correlation is identified between the first dataset variable and at least one additional dataset variable in the original dataset, then assign the random instance seed value to the at least one additional dataset variable using the correlation manager; determine an adaptive sensitivity parameter corresponding to the first dataset variable using the correlation manager; and The noise generation manager generates additive noise using two or more of the first data set variable, the random instance seed value, and / or the adaptive sensitivity parameter and applies the additive noise to the first data set variable to produce a pseudonymized variable for inclusion in a pseudonymized data set associated with the original data set.
12. The system of claim 11, wherein the correlation manager and the noise generation manager are configured to repeat each action for each remaining data set value included in the original data set.
13. The system of claim 11 or 12, wherein a high correlation is identified by an operator.
14. The system of claim 13, wherein a high correlation includes a high positive correlation or a high negative correlation.
15. The system of any one of claims 11 to 14, wherein the first data set variable and the at least one additional data set variable are biochemical data variables associated with a common subject sample.
16. The system of any one of claims 11 to 15, wherein the adaptive sensitivity parameter scales with a numerical measurement value associated with the first data set variable.
17. The system of any one of claims 11 to 16, wherein the adaptive sensitivity parameter indicates a distribution range of the additive noise applied to the first data set variable.
18. The system of claim 17, wherein the adaptive sensitivity parameter is used to establish an amount value of the additive noise applied to the first data set variable.
19. The system of any one of claims 11 to 18, wherein the noise generation manager uses a Laplace transform mechanism.
20. A non-transitory computer-readable medium having executable instructions stored thereon, the instructions when executed by a processor of a computer control the computer to execute a method, the method comprising: specifying a random instance seed value to a first data set variable in an original data set; if a high correlation is identified between the first data set variable and at least one additional data set variable in the original data set, then specifying the random instance seed value to the at least one additional data set variable; determining an adaptive sensitivity parameter corresponding to the first data set variable; and generating additive noise using two or more of the first data set variable, the random instance seed value, and / or the adaptive sensitivity parameter by a noise generation manager and applying the additive noise to the first data set variable to produce a pseudonymized variable for inclusion in a pseudonymized data set associated with the original data set.