System and process for anonymizing data based on the risks of each piece of data
The data anonymization system addresses the inadequacies of existing methods by assessing and transforming high-risk data to meet strict anonymization criteria, ensuring compliance and data quality through a risk-based approach.
Patent Information
- Application Number
- FR2023013604
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2043-12-05
AI Technical Summary
Existing data anonymization methods, such as pseudonymization and synthetic data, are inadequate for ensuring compliance with strict anonymization criteria while preserving data quality and usability, as they either allow for easy re-identification or require time-consuming model creation, and current risk assessment tools lack guidance for implementing these criteria effectively.
A data anonymization system that assesses individualization, correlation, and inference risks using a characterization means, exposure levels, severity scales, and exploitability scores, allowing targeted transformation of high-risk data to minimize degradation and maintain data quality.
The system provides a precise risk assessment and minimizes data degradation by transforming only high-risk data, ensuring compliance with anonymization criteria and maintaining data usability and quality, with a digital tool for analyzing large volumes of data and generating reports for compliance.
Smart Images

Figure 00000020_0000 
Figure 00000021_0000 
Figure 00000022_0000
Abstract
Description
Title of the invention: System and method for anonymizing data based on the risks of each piece of data
[0001] The invention relates to the field of systems and methods for processing and anonymizing personal data and means of re-identifying persons.
[0002] Personal data have become a valuable element in science and technology. They are particularly useful for conducting clinical studies, testing and validating computer applications, and are crucial in the field of machine learning in artificial intelligence.
[0003] Unfortunately, the use of personal data is currently hampered by the implementation of legislation relating to the protection of personal data (for example: GDPR, CCPA, HIPAA, etc.). This legislation specifies, in particular, that in order to use personal data for the purposes described above, it must be anonymized, that is to say, transformed into data that is no longer personal.
[0004] Thus, several measures have been proposed to anonymize data, but none are entirely satisfactory. A common strategy is pseudonymization (of which data masking is a variant), which consists of replacing one attribute with another, thereby aiming to limit the risk of identifying an individual. In other words, directly or indirectly identifying data (such as name, surname, email address, postal address) is replaced by pseudonyms (such as an alias, a number, etc.).
[0005] Another strategy is the use of so-called synthetic data instead of real data in order to protect people. Synthetic data is data generated based on the original data, using artificial intelligence models; it has the particularity of retaining the properties of real data while not containing real information.
[0006] Unfortunately, pseudonymization is not entirely satisfactory because it is considered reversible. If a pseudonymized dataset is removed, the author may be able to reconstruct the personal data relatively easily. Thus, pseudonymized data is still considered personal data under regulations. Furthermore, synthetic data is difficult to use because it is generally difficult to produce datasets that truly reflect the original data, resulting in a lack of reliability in synthetic data. In addition, the generation of Synthetic data is a time-consuming operation that requires creating a new model for each dataset.
[0007] Anonymization in the strict sense is defined at the European level by the Article 29 Working Party on Data Protection (WP29). According to this group, an anonymization solution must be developed on a case-by-case basis and adapted to the intended uses. To help evaluate a good anonymization solution, the WP29 proposes three criteria: - Individualization: Is it still possible to isolate an individual after anonymization? - Correlation: Is it possible to link together distinct sets of data concerning the same individual? - Inference: can we deduce information about an individual?
[0008] Thus, according to the Article 29 Working Party (WP29), a dataset for which it is not possible to individualize, correlate, or infer is a priori anonymous. Furthermore, a dataset for which at least one of the three criteria is not met can only be considered anonymous following a detailed risk analysis of re-identification. The tools for implementing the WP29 recommendations are not provided in the prior art.
[0009] Unfortunately, the criteria presented in the G29 opinion are too strict and unusable as they stand, because they systematically produce excessively high risks, leading to the application of overly rigorous anonymization, which tends to considerably destroy the data by rendering it unusable in an anonymized form.
[0010] In light of this situation, the CNIL proposes two ways to assess the risks: either strictly adhere to these criteria, or conduct a re-identification risk analysis. However, the CNIL provides very little guidance on how to carry out this re-identification risk analysis.
[0011] The main difficulties are 1) being able to calculate re-identification scores according to the 3 criteria of the G29; 2) being able to explain the scores calculated according to re-identification risks; 3) integrating them into a risk analysis approach following the EBIOS model, recommended by the CNIL.
[0012] Thus, one objective of the present invention is to remedy the shortcomings of the prior art, and in particular to offer a data anonymization solution that significantly limits the risk of re-identification while allowing the use of the data in an anonymized form. The invention is based on a risk assessment tool for individualization, correlation, and inference, which, depending on the risk level, allows only the data involving the greatest risk to individuals to be transformed. Therefore, the tool allows for minimal degradation of the dataset in order to preserve its quality in an anonymized form.
[0013] To achieve these objectives, the invention proposes a data anonymization system taking as input at least one dataset from an organization, the anonymization system comprising: - a means of characterizing data on the basis of a “data schema”, allowing to define for each data, a name, a level of exposure (allowing to evaluate the possibility of accessing the data according to whether it can be accessed from outside the organization or from inside), and preferably a type of data sensitivity; - a means of identifying feared events allowing for the definition, for each data point, of a severity scale preferably broken down into a type of feared event and / or an impact scale of this event, evaluated on the basis of at least one given criterion; - a means of evaluating the legal criteria for anonymization, allowing us to define for each piece of data, an individualization score, a correlation score, and an inference score; - a means of evaluation of a level of exploitability determined on the basis of a combination of exposure levels and individualization, correlation and inference scores; - a means of assessing the overall risk of the dataset based on risk assumptions constructed using data with a significant level of usability and a significant severity scale; and - preferably a means of reducing risks including proposals for countermeasures or transformations to limit the level of exploitability and / or the severity scale of the most at-risk data.
[0014] Advantageously, the invention makes it possible to assess the risk of the different data in a dataset and to transform part of the data; those involving a significant risk, so that the dataset is anonymized without significantly degrading its quality.
[0015] The invention is also a digital tool for analyzing data and for proving that the risks inherent in the data used have been validly assessed. The invention allows for simplified proof that the risk assessment has been carried out and is applicable to large volumes of data (hundreds or even thousands of variables).
[0016] According to one variant, the characterization means makes it possible to assign each data point an exposure level based on variables, the number of which can vary between 2 and 10, preferably four variables, more preferably from among: - a “restricted internal” level if the data is accessible to a limited number of people internal to the user's organization; - an “extended internal” level if the data is accessible to any people internal to the user's organization; - a “restricted external” level if the data is accessible to a limited number of people external to the user's organization; - an “extended external” level if the data is accessible to any people outside the user's organization.
[0017] This allows for a different risk assessment depending on third-party access to the data, resulting in a more precise risk assessment and improved quality of anonymized data. The level of exposure enables a more accurate calculation of the correlation criterion and helps identify the data at highest risk.
[0018] According to one variant, the characterization means makes it possible to assign to each data a type of sensitivity on the basis of three variables, preferably among: - a sensitive data type if the data can have personal impacts on the person concerned; - a type of data perceived as sensitive if the data is perceived as sensitive; - a common data type if the data is routine and not sensitive.
[0019] This makes it possible to discriminate the sensitivities of the different data in order to better assess the risk and improve the quality of the anonymized data.
[0020] According to one variant, the means of identifying feared events makes it possible to assign at least one severity scale to four variables, preferably: minor, significant, serious, and critical.
[0021] This makes it easy to assess the severity of the disclosure of each data point and of the dataset. It also makes it possible to record these events for each data point and to forecast them in a report.
[0022] According to one variant, the means of evaluating the legal criteria for anonymization allows for the assignment of scores whose number of levels can vary between 3 and 5, preferably: low, moderate, high, very high.
[0023] This simplifies the assessment of the risk of re-identification.
[0024] According to one variant, the exploitability level assessment means makes it possible to assign exploitability levels whose number of values can vary between 3 and 5, preferably very difficult, difficult, easy, very easy.
[0025] This simplifies the assessment of the risk of data exploitability by a third party.
[0026] According to one variant, the system further includes a means for generating a colour code with an importance scale for at least one evaluated level.
[0027] This allows for a quick visualization of the danger of an assessed risk.
[0028] According to one variant, the system further includes a means of editing a report including, among other things, the identified risks and / or countermeasures.
[0029] This makes it possible to prove to a third party or institution that the risks have been assessed and measures have been taken to minimize them.
[0030] Another object of the invention relates to a method for anonymizing personal data in at least one user dataset, the anonymization method comprising: - a data characterization step based on a data schema in which each piece of data is assigned a name, an exposure level allowing assessment of the possibility of accessing the data depending on whether it can be accessed from outside the organization or from inside, and preferably a type of data sensitivity; - a stage of identifying feared events in which each piece of data is assigned a severity scale preferably broken down into a type of feared event and / or an impact scale of that event, assessed on at least one of the material, physical, moral aspects; - an evaluation stage of the legal criteria for anonymization in which each piece of data is assigned an individualization score, a correlation score, and an inference score; - an assessment step of a level of exploitability based on a combination of exposure levels and individualization, correlation, and inference scores; - a step of assessing the overall risk of the dataset based on risk assumptions constructed using data with a significant level of usability and a significant severity scale; and - a risk reduction step including proposals for countermeasures or transformations to limit the level of exploitability and / or the severity scale of the most at-risk data.
[0031] The invention further relates to a computer program comprising program code instructions for executing the steps of a data anonymization process according to the invention, when said program is running on a computer.
[0032] The invention will be further detailed by describing non-limiting embodiments, and based on the accompanying figures illustrating preferred embodiments of the invention, in which: - [Fig.l] schematically illustrates an anonymization system according to a preferred embodiment of the invention, and an edition of a report which can be physical or electronic; - [Fig.2] schematically illustrates an anonymization process according to one implementation method; - [Fig. 3] schematically illustrates a first example of data risk analysis of a dataset using the invention, before and after application of a maximum countermeasures according to the invention against an internal or external hacker; and - [Fig.4] schematically illustrates a second example of data risk analysis similar to that of [Fig.4] with minimal countermeasures according to the invention.
[0033] The invention relates to a system and a method for anonymizing personal data in at least one user dataset.
[0034] The invention is implemented by computer means, for example via a computer, a server, a tablet, a smartphone or other or a combination of at least two of these elements.
[0035] The anonymization system includes several hardware and software means for loading and evaluating the data, preferably transforming it or proposing countermeasures to improve the security of the dataset.
[0036] The anonymization system includes a means for loading at least one dataset to be analyzed. In particular, it includes a module for reading a file containing said dataset. The dataset may be in any file format, for example in a spreadsheet format (known as ".csv").
[0037] The anonymization system further includes a means for identifying a data schema. The data schema is illustrated in Table 1. In this data schema, each piece of data is assigned a name, preferably a label, a level of exposure to persons internal or external to the user's organization, and preferably a data sensitivity type.
[0038] Table 1: Example of a data schema # Name Label Exposure Sensitivity 1 Antibiotic Susceptibility Test Molecule X 1-Restricted Internal Current 2 Bacterial Species Species Y 1-Restricted Internal Current 3 Sample Type 8 modalities 1-Restricted Internal Current 4 Sample Date Shifted by a random number 2-Extended Internal Current 5 MALDLT Spectrum oF Single Vector 1000 values 1-Restricted Internal Current 6 Date of Birth Rounded by 5 years 4-Extended External Current 7 Sex Two categories 4-Extended External Current
[0039] In the field of health or clinical studies, the name of the data is, for example, antibiogram; bacterial species; type of sample; date of sample collection; such and such a test name, for example MALDI-ToF spectrum; date of birth, sex. The “Name” may refer to an identification number, a driver's license number, a points system in the field of road safety, or the name of any test in another technical field.
[0040] In addition to the “Name”, it is possible to define the “Label” to enter a summary description of the data.
[0041] The “Exposure” level of an attribute assesses how easy it is for a hacker to obtain the information in question from another dataset. For example, it is probably more difficult to obtain a person’s “Type of Sample” from another dataset than to find their “Date of Birth” or “Sex”.
[0042] The level of “Exposure” is differentiated according to internal third parties within the organization using the dataset, and those internal or external to that organization who do not have the right to access the data. Furthermore, the level of exposure is preferably identified by fewer than ten discrete variables, preferably four variables.
[0043] The level of exposure can be classified into four main categories in ascending order of level: - internal-restricted; - internal-enlarged; - external-restricted; and - extreme-widened.
[0044] Regarding the internally restricted level: attributes with an internally restricted exposure level are accessible only to a limited number of authorized individuals. These attributes may include medical information, sensitive financial data, or sensitive personal information.
[0045] In particular, we will speak of a "restricted internal" level if the data is accessible to a limited number of people internal to the user's organization, for example, people belonging to a particular department.
[0046] A traffic light-type color code can be generated according to the exposure level. The restricted internal level is, for example, green (V) because data in this category are less likely to be found in other datasets.
[0047] For example, bacterial species, analysis results, banking operations may have a restricted internal level, and a green color code V.
[0048] Regarding the internal-extended level: attributes with an internal-extended exposure level are accessible within the organization, by employees belonging to several departments or by all employees. These attributes may include employee identification data, internal activity reports, but also data of relevant individuals (e.g., patients, clients, etc.) passing from one department to another (e.g., client / patient IDs, heights, weights, blood pressures, etc.).
[0049] In particular, data has this level of exposure if it is accessible to any persons within the organization or to several different departments. The color code is, for example, yellow J.
[0050] Regarding the external-restricted level: attributes with an external-restricted exposure level are accessible to third parties, but require specific searches or specific data sources to access them. They may include information shared with business partners, survey data, or industry-specific information.
[0051] In particular, data is considered to have this level if it is accessible to a limited number of people outside the user's organization due to the complexity required to collect it. This includes people who may already know the information in question or find it by conducting extensive research. For example, an admission date or an exit date may have a restricted external access level. The color code is, for example, orange O (or dark orange).
[0052] Regarding the extremely broad level: attributes exhibiting an extremely broad level of exposure are easily accessible and can be obtained from external sources without much difficulty. These attributes may include publicly available information, such as a name, surname, age, postal address, email address, or general professional data.
[0053] In particular, data is at this level if it is accessible to any persons external to the user's organization, for example from social networks and search engines. The color code is, for example, red R.
[0054] Data breach depends, among other things, on the ability to cross-reference different datasets and therefore on the ability to find in another dataset data present in the anonymized dataset.
[0055] In the context of the invention, it is desirable to be able to assign a score between 1 and 4 which quantify the degree of exposure of an attribute in the dataset. He quite naturally came to propose the following scores: extreme-wide with a score equal to 4 (critical), external-restricted with a score equal to 3 (high), internal-wide with a score equal to 2 (medium) then internal-restricted with a score equal to 1 (low).
[0056] A high exposure score may indicate that randomization should be applied to this attribute. This could make it more difficult to find suitable values from external sources.
[0057] In addition to the exposure level, the identification method allows each piece of data to be assigned a sensitivity type based on three variables. These can be referred to as CNIL data types. We can distinguish: - a sensitive data type if the data relates to racial or ethnic origin, political opinions, religious or philosophical beliefs or trade union membership, as well as the processing of genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person's sex life or sexual orientation; - a type of data perceived as sensitive if the data is perceived as sensitive, for example banking data, biometric data, a social security number; - a common data type if the data is neither sensitive nor perceived as sensitive.
[0058] The risk of re-identification depends on this sensitivity.
[0059] The anonymization system further includes a means of identifying feared events. This information demonstrates that this aspect was assessed before the data was used. To this end, each piece of data is assigned a severity scale to determine whether a minor, significant, serious, or critical event would occur if the data were lost. Preferably, the scale contains fewer than 10 levels, more preferably fewer than 6, and even more preferably, 4 variables are used for the severity scale. Furthermore, the severity preferably includes a type of feared event, for example, the disclosure of a patient's illness, and / or an impact scale for this event, assessed on at least one of three criteria: material, physical, or psychological.
[0060] Table 2: Example of feared events Sensitive Attribute Feared Event Material Physical Moral Severity Bacterial Species Disease Disclosure Minor Serious Critical Critical
[0061] For better visibility and quick understanding, the feared events are assigned to a traffic light type colour code, with a green colour V for events of no severity, a yellow colour J and / or orange O for events of intermediate severity, and a red colour R for a significant severity.
[0062] The anonymization system further includes a means for evaluating the criteria for re-identifying individuals, in which each piece of data is assigned an individualization score, a correlation score, and an inference score. In each of these cases, the score is preferably evaluated on a scale of fewer than 10 values, more preferably between 3 and 5 values, and more preferably according to the values: Low, medium, high, and critical. These scores can be assessed collegially, conventionally, or statistically for each data point.
[0063] Table 3: Risk assessment of re-identification # Name Individualization Correlation Inference 1 Antibiotic susceptibility testing 1-Low 1-Low 2-Moderate 2 Bacterial species 3-High 1-Low 4-Very high 3 Sample type 1-Low 1-Low 1-Low 4 Sample date 4-Very high 3-High 4-Very high 5 MALDLT spectrum of 3-High 1-Low 3-High 6 Date of birth 2-Moderate 3-High 2-Moderate 7 Sex 1-Low 3-High 1-Low
[0064] Similarly, a traffic light-type colour code can be used. Critical risk is red R; high risk, orange O; medium risk, yellow J; and low risk, green V.
[0065] The individualization score can be assessed or calculated. Individualization assesses the possibility of isolating an individual within the dataset. It refers to the level of detail or specificity of the information contained in each attribute. Some attributes can be very granular, providing precise information such as the full date of birth or the complete address. Other attributes can be more general, with a lower individualization score, such as the year of birth or the city of residence. The individualization score of the attributes can influence the potential for disclosure and the associated level of risk.
[0066] A high individualization score can be remedied by applying generalization. However, it remains important to consider the loss of data utility that may accompany generalization.
[0067] The risk of inference can be assessed or calculated. Inference can be likened to the statistical concept of discrimination rate. This is a measure that assesses the ability of an attribute to distinguish or discriminate between an individual and others. It is often used to assess the risk of re-identifying individuals from anonymized data. An attribute with a high discrimination rate can provide information enabling
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074] to identify or reveal specific personal characteristics, which increases the risk of a breach of confidentiality. As with the rx ratio obtained for the granularity measure, the discrimination rate is a value in the interval [0, 1]. Denoting this metric as DR, we can define the associated SDR score in a manner analogous to the granularity score: ^DR — |"4 • DR] with ^4 the upper integer part, also called the “ceiling”. The correlation criterion is assessed based on the level of exposure and the individualization criterion. The anonymization system also includes a means of assessing exploitability levels, comprising at least a combination of exposure and risk levels related to individualization, correlation, and inference. Essentially, when the scores are high, then exploitability is easy, and a hacker can easily access personal data; and vice versa. Table 4: Risk assessment of exploitability # Individualization Correlation Inference Operability 1 1-Low 1-Low 2-Moderate 1-Very difficult 2 3-High 1-Low 4-Very high 2-Difficult 3 1-Low 1-Low 1-Low 1-Very difficult 4 4-Very high 3-High 4-Very high 4-Very easy 5 3-High 1-Low 3-High 2-Difficult 6 2-Moderate 3-High 2-Moderate 3-High 7 1-Low 3-High 1-Low 2-Difficult A possible exploitability assessment scale is shown in Table 5. Table 5: Method for assessing exploitability risks Category Individualization Correlation Inference Exploitability Expanded External (Sex) 1-Low 2-Moderate 1-Low 1-Very difficult Expanded Internal (Date of Sample) 4-Very high 3-High 4-Very high 4-Very easy Restricted Internal (Antibiography) 4-Very high 1-Low 4-Very high 2-Difficult A color code can be assigned to the exploitability. The Very easy level is red R, difficult is yellow J, and very difficult is green V.
[0075] The anonymization system further includes a means of overall risk assessment of the dataset in which data with a significant level of exploitability and / or a significant level of severity are identified.
[0076] The exploitability level and the severity level are preferably represented in a two-dimensional graph (one dimension for each level) to allow for better visualization of the risks and easy comparison of anonymized games. The graph preferably includes a traffic light-type color code. This type of graph is illustrated in [Fig. 3] and [Fig. 4].
[0077] The anonymization system further includes a means of correcting the dataset comprising proposals, preferably automatic, of countermeasures or transformations of the data having a significant level of exploitability and / or a significant level of severity so as to reduce said level.
[0078] The system can alert on the risks of identification at high exploitability levels, identify vulnerabilities via initial vulnerability modules noted D1, D3, etc., and propose corresponding countermeasures: - a Dl module requires generalizing as much as possible the variables involving a high level of exploitability (easy) for example extended internal variables such as sampling dates. - Module D3 identifies the risks of re-identification related to restricted internal variables.
[0079] The system can further analyze compliance with the GDPR by means of at least one of the following specific vulnerability modules: - a Cl module requiring compliance with the principle of data minimization (because only the data strictly necessary for the study should be used according to the regulations, it will then be necessary to define a list of data and to justify the use which is made; - a C2 module requesting the deletion of data for special persons; - a C3 module requiring the definition of a reasonable data retention period according to the purpose of the data processing; - a C4 module requiring a data purging mechanism to be provided at the end of data processing; - a C5 module prohibiting the export of data outside the system; - a C6 module requiring that the persons concerned be informed that the data will be anonymized; - a C7 module alerting on the location of the data, because the data must be located in the EU zone, otherwise standard or binding contractual clauses with the cloud network provider must be signed; - a C8 module requiring the stakeholder to commit to not trying to re-identify people, it can be coupled with module D3; - a C9 module requiring a high-security password, for example according to the recommendations of ANSSI.
[0080] The anonymization system further includes a means for generating an R report summarizing the identified risks and countermeasures. The R report automatically includes the risk analysis elements and serves as evidence to demonstrate to the authorities that the risks have been assessed and minimized. The R report can be in physical or electronic format.
[0081] Example RI:
[0082] By hypothesis, a hacker exfiltrates insufficiently anonymized data after exploiting a weakness in authentication, and re-identifies the persons concerned on the basis of the expanded internal variables.
[0083] The system of the invention analyzes the data and identifies the vulnerabilities of the modules D1, D3, C2 and C8 with a high risk of re-identification (for example a level 4- very easy (red).
[0084] A first set of countermeasures is proposed, for D1, a CM1 countermeasure: generalize the expanded internal variables; for C8, a CM3 countermeasure: define a password policy using ANSSI standards; for C2, a CM5 countermeasure: delete the data of special persons; for D3, CM6 countermeasures: generalize the restricted internal variables.
[0085] The countermeasures limit the risk, which goes from a level 4-Critical to a level 2-Medium.
[0086] On an initial exploitability scale of 3 / 4, and a severity scale of 4 / 4 directly visible in the system (top of [Fig.3]), the countermeasures lower exploitability to 1 / 4, and severity to 3 / 4 as illustrated by the bottom of [Fig.3].
[0087] The system can propose minimum countermeasures to be taken into account, here CM1 and CM3. The countermeasures maintain the risk at level 2-Medium. In this case, the severity remains at 4 / 4, but the exploitability decreases to 1 / 4 as illustrated in [Fig. 4].
[0088] Example R2:
[0089] By hypothesis, a researcher re-identifies the persons concerned on the basis of the restricted internal variables.
[0090] The system of the invention analyzes the data and identifies the vulnerabilities of the modules D1, D3, C2 and C7 with a high risk of re-identification (for example a level 4- very easy (red).
[0091] A first set of countermeasures is proposed: for D1, a CM1 countermeasure: generalize the expanded internal variables; for D3 and C7, CM6 countermeasures: generalize the restricted internal variables and CM4 countermeasures: plan measures contractual for the operator who undertakes not to attempt to re-identify the persons; for C2, a countermeasure CM5: delete the data of special persons.
[0092] The countermeasures limit the risk which reaches a level 2-Medium.
[0093] On an initial exploitability scale of 3 / 4, and a severity scale of 4 / 4 directly visible in the system, the countermeasures lower exploitability to 1 / 4, and severity to 3 / 4 as illustrated in [Fig.3].
[0094] The system preferably proposes minimal countermeasures to be taken into account, here CM1 and CM4. The countermeasures maintain the risk at level 2-Medium. In this case, the severity remains at 4 / 4, but the exploitability drops to 1 / 4 as illustrated in [Fig. 4].
[0095] Preferably, the level of correlation is determined according to the level of exposure and the level of individualization based on the following table:
[0096] Table 6: Evaluation of the level of correlation Exposure Individualization Internal restricted Internal broadened External restricted External broadened Very high 1-Low 3-High 4-Very high 4-Very high High 1-Low 2-Moderate 4-Very high 4-Very high Moderate 1-Low 2-Moderate 3-High 3-High Low 1-Low 1-Low 2-Moderate 2-Moderate
[0097] Preferably, the level of exploitability is determined according to the level of correlation and the level of inference based on the following table:
[0098] Table 7: Assessment of the level of exploitability Exploitability Inference 1-Low 2-Moderate 3-High 4-Very High Very High 2-Difficult 3-Easy 4-Very Easy 4-Very Easy High 2-Difficult 2-Difficult 3-Easy 4-Very Easy Moderate 1-Very difficult 2-Difficult 3-Easy 3-Easy Low 1-Very difficult 1-Very difficult 2-Difficult 3-Easy
[0099] Preferably, a contextual exploitability level is determined based on the exploitability level and the remarks and countermeasures above, on the basis of the following table:
[0100] Table 8: Assessment of the level of exploitability Contextual exploitability Exploitability (Contextual) 1-Low 2-Moderate 3-High 4-Very High Very Easy 2-Difficult 3-Easy 4-Very Easy 4-Very Easy Easy 2-Difficult 2-Difficult 4-Very Easy 4-Very Easy Difficult 1-Very Difficult 2-Difficult 3-Easy 3-Easy Very Difficult 1-Very Difficult 1-Very Difficult 2-Difficult 3-Easy
[0101] Preferably, the risk of re-identification (legal criteria for anonymization) is determined according to the level of severity and the level of exploitability (preferably contextual exploitability) based on the following table:
[0102] Table 9: Assessment of the level of exploitability Severity Exploitability Negligible Limited Significant Maximum Very easy 2-Medium 3-High 4-Critical 4-Critical Easy 2-Medium 2-Medium 3-High 4-Critical Difficult 1-Low 2-Medium 3-High 3-High Very difficult 1-Low 1-Low 2-Medium 2-Medium
[0103] The invention further relates to a method of anonymizing personal data in at least one user dataset, based on a system as described above.
[0104] The process includes steps for implementing the different modules.
[0105] The process is implemented by computer.
[0106] The invention also relates to a computer program for implementing the invention via computer means.
Claims
Demands
1. A data anonymization system implemented by computer means, taking as input at least one dataset in an organization, the anonymization system comprising: - a means of characterizing the data (M1) on the basis of a "data schema" allowing to define for each data, a name, a level of exposure on the basis of a number of variables, allowing to assess the possibility of accessing the data according to whether it can be accessed from outside the organization or from inside, and preferably a type of sensitivity of the data on the basis of three variables; - a means of identifying feared events (M2) allowing to define for each data, a severity scale of fewer than 10 variables preferably broken down into a type of feared event and / or an impact scale of this event, assessed on the basis of at least one given criterion;- a means of evaluating the legal criteria for anonymization (M3) allowing the definition, for each data point, of an individualization score on a scale of fewer than 10 values, a correlation score on a scale of fewer than 10 values, and an inference score on a scale of fewer than 10 values; - a means of evaluating a level of usability (M4) based on a number of values, determined on the basis of a combination of exposure levels and individualization, correlation, and inference scores; - a means of evaluating the overall risk (M5) of the dataset based on risk assumptions constructed using data with a significant level of usability and a significant severity scale; and - a means of risk reduction (M6) including proposals for countermeasures or transformations allowing the limitation of the level of usability and / or the severity scale of the most at-risk data.
2. An anonymization system according to any one of the preceding claims, characterized in that the characterization means (M1) makes it possible to assign to each data point a level of exposure on the variable base, the number of which can vary between 2 and 10, preferably four variables, more preferably from: - a "restricted internal" level if the data is accessible to a restricted number of people internal to the user's organization; - an "extended internal" level if the data is accessible to any people internal to the user's organization; - a "restricted external" level if the data is accessible to a restricted number of people external to the user's organization; - an "extended external" level if the data is accessible to any people external to the user's organization.
3. An anonymization system according to any one of the preceding claims, characterized in that the characterization means (M1) makes it possible to assign to each data a type of sensitivity on the basis of three variables, preferably among: - a sensitive data type if the data may have personal impacts on the person concerned; - a data type perceived as sensitive if the data is perceived as sensitive; - a common data type if the data is usual and not sensitive.
4. An anonymization system according to any one of the preceding claims, characterized in that the means for identifying feared events (M2) makes it possible to assign at least one severity scale to four variables, preferably minor, significant, serious, and critical.
5. An anonymization system according to any one of the preceding claims, characterized in that the means for evaluating the legal criteria for anonymization (M3) allows for the assignment of scores whose number of levels can vary between 3 and 5, preferably low, moderate, high, very high.
6. An anonymization system according to any one of the preceding claims, characterized in that the exploitability level assessment means (M4) allows exploitability levels to be assigned, the number of values of which can vary between 3 and 5, preferably very difficult, difficult, easy, very easy.
7. An anonymization system according to any one of the preceding claims, characterized in that it further comprises a means generating a colour code (V, Y, O, R) with an importance scale for at least one assessed level.
8. An anonymization system according to any one of the preceding claims, characterized in that it further comprises a means for editing a report (R) including among other things the identified risks and / or countermeasures (CM1, CM3, CM4, CM5, CM6).
9. A method for anonymizing personal data in at least one user dataset, the anonymization method being implemented by computer means, and comprising: - a data characterization step based on a data schema in which each piece of data is assigned a name, an exposure level based on a number of variables, allowing the possibility of accessing the data to be assessed according to whether it can be accessed from outside the organization or from inside, and preferably a type of data sensitivity based on three variables; - a feared event identification step in which each piece of data is assigned a severity scale of fewer than 10 variables preferably broken down into a type of feared event and / or an impact scale of this event, assessed on at least one of the material, physical, or moral aspects;- a step of evaluating the legal criteria for anonymization in which each data point is assigned an individualization score on a scale of fewer than 10 values, a correlation score on a scale of fewer than 10 values, and an inference score on a scale of fewer than 10 values; - a step of assessing the level of usability based on a number of values, determined on the basis of a combination of exposure levels and individualization, correlation, and inference scores; - a step of assessing the overall risk of the dataset based on risk assumptions constructed using data with a significant level of usability and a significant severity scale; and - a risk reduction step including proposals for countermeasures or transformations to limit the; 19 level of exploitability and / or the severity scale of the data most at risk.
10. A computer program comprising program code instructions for performing the steps of a data anonymization process according to claim 9, when said program is running on a computer.