System and process for anonymizing data based on the risks of each data item

The system addresses the limitations of existing data anonymization methods by assessing and mitigating re-identification risks through a targeted approach, ensuring effective anonymization with minimal data degradation.

FR3156220A1Active Publication Date: 2025-06-06COACHMESEC CONSULTING
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
FR2023013604
Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-06
Estimated Expiration
2043-12-05

AI Technical Summary

Technical Problem

Existing data anonymization methods, such as pseudonymization and synthetic data generation, are not fully satisfactory as they are either reversible or difficult to produce reliably, leading to overly rigorous anonymization that degrades data quality.

Method used

A system and method for anonymizing data by assessing the risks of individualization, correlation, and inference using a tool that evaluates data items based on exposure levels, sensitivity, and severity, allowing for targeted transformations to minimize re-identification risks while preserving data quality.

Benefits of technology

The solution effectively reduces the risk of re-identification while minimizing data degradation, allowing for the use of anonymized data in a reliable and efficient manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000020_0000
    Figure 00000020_0000
  • Figure 00000021_0000
    Figure 00000021_0000
  • Figure 00000022_0000
    Figure 00000022_0000
Patent Text Reader

Abstract

The invention relates to a data anonymization system comprising:- a means for identifying the data (M1) and assigning a level of exposure to persons inside or outside the user's organization;- a means for identifying feared events (M2) and their severity;- a means for assessing the risks of re-identification of persons (M3);- a means for assessing the level of exploitability (M4);- a means for assessing the overall risk (M5) of the data set in which the data having a significant level of exploitability and / or a significant level of severity are identified; and- a means for correcting (M6) the data and for countermeasures (CM1, CM3, CM4, CM5, CM6) or for transforming the data having a significant level of exploitability and / or severity so as to limit said level. The invention also relates to a program based on such a method. Abstract figure: Fig. 2
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: System and method for anonymizing data based on the risks of each data item

[0001] The invention relates to the field of systems and methods for processing and anonymizing personal data and means for re-identifying people.

[0002] Personal data has become a valuable asset in science and technology. It is particularly useful for conducting clinical studies, testing and validating computer applications, and is crucial in the field of machine learning in artificial intelligence.

[0003] Unfortunately, the use of personal data is, nowadays, hampered by the implementation of legislation relating to the protection of personal data (for example: GDPR, CCPA, HIPAA, ...). This legislation specifies in particular that to use personal data for the purposes described above, it must be anonymized, that is to say, transformed into data that is no longer personal.

[0004] Thus, several measures have been proposed to anonymize data, but are not fully satisfactory. A common strategy is pseudonymization (of which data masking is a variant) which consists of replacing one attribute with another, thus aiming to limit the risk of identifying an individual. In other words, directly or indirectly identifying data (such as name, first name, email address, postal address) are replaced by pseudonyms (such as an alias, a number, etc.).

[0005] Another strategy is the use of so-called synthetic data instead of real data in order to protect people. Synthetic data is data generated on the basis of the original data, relying on artificial intelligence models; it has the particularity of retaining the properties of real data while not containing real information.

[0006] Unfortunately, pseudonymization is not fully satisfactory because it is considered reversible. In the event of subtraction of a pseudonymized data set, it is possible that the author can reconstruct the personal data relatively easily. Thus, pseudonymized data are still considered personal data by the regulations. Furthermore, synthetic data are difficult to use because it is generally difficult to produce data sets that truly reflect the original data, which leads to a lack of reliability of the synthetic data. In addition, the generation of Synthetic data is a time-consuming operation that requires creating a new model for each dataset.

[0007] Anonymization in the strict sense is defined at the European level by the Article 29 Working Party on Data Protection (abbreviated as WP29). According to this group, an anonymization solution must be constructed on a case-by-case basis and adapted to the intended uses. To help evaluate a good anonymization solution, WP29 proposes three criteria: - individualization: is it still possible to isolate an individual after anonymization? - correlation: is it possible to link together distinct sets of data concerning the same individual? - inference: can we deduce information about an individual?

[0008] Thus, for G29, a data set for which it is not possible to individualize, correlate, or infer is a priori anonymous. Furthermore, a data set for which at least one of the three criteria is not respected can only be considered anonymous following a detailed analysis of the risks of re-identification. The tools for implementing the recommendations of G29 are not proposed in the prior art.

[0009] Unfortunately, the criteria presented in the G29 opinion are too strict and unusable as they stand, because they systematically produce risks that are too high, leading to the application of anonymization that is too rigorous, which tends to considerably destroy the data by making them unusable in an anonymized form.

[0010] Faced with this observation, the CNIL proposes two ways of assessing the risks: either scrupulously respect these criteria, or carry out a re-identification risk analysis. However, the CNIL gives very little guidance on how to carry out this re-identification risk analysis.

[0011] The main difficulties are 1) being able to calculate the re-identification scores according to the 3 criteria of G29; 2) being able to explain the scores calculated according to re-identification risks; 3) integrating them into a risk analysis approach following the EBIOS model, recommended by the CNIL.

[0012] Thus, an objective of the present invention is to remedy the defects of the prior art, and in particular to propose a solution for anonymizing data significantly limiting the risk of re-identification while allowing the use of the data in an anonymized form. The invention is based on a tool for assessing the risks of individualization, correlation and inference, making it possible, depending on the value of the risk, to transform only the data involving the greatest risks for people. Thus, the tool allows for little degradation of the data set in order to preserve its qualities in an anonymized form.

[0013] To achieve these objectives, the invention proposes a data anonymization system taking as input at least one data set in an organization, the anonymization system comprising: - a means of characterizing data based on a “data schema”, making it possible to define for each data item, a name, a level of exposure (making it possible to assess the possibility of access to the data depending on whether it can be accessed from outside the organization or from inside), and preferably a type of sensitivity of the data; - a means of identifying feared events making it possible to define for each data item, a severity scale preferably broken down into a type of feared event and / or a scale of impact of this event, evaluated on the basis of at least one given criterion; - a means of evaluating the legal anonymization criteria, making it possible to define for each data item, an individualization score, a correlation score, and an inference score; - a means of evaluation of an exploitability level determined on the basis of a combination of exposure levels and individualization, correlation and inference scores; - a means of assessing the overall risk of the dataset based on risk hypotheses constructed using data having a significant level of exploitability and a significant severity scale; and - preferably a means of risk reduction including proposals for countermeasures or transformations to limit the level of exploitability and / or the severity scale of the most at-risk data.

[0014] Advantageously, the invention makes it possible to evaluate the risk of the different data in a data set and to transform a part of the data; those involving a significant risk, so that the data set is anonymized without significantly degrading its quality.

[0015] The invention is furthermore a digital tool for analyzing data, and for proving that the risks inherent in the data used have been validly assessed. The invention makes it possible to prove in a simplified manner that the risk assessment has been carried out, and is applicable to large volumes of data (hundreds, even thousands of variables)

[0016] According to a variant, the characterization means makes it possible to assign to each data item an exposure level on the basis of variables the number of which can vary between 2 and 10, preferably four variables, more preferably among: - a “restricted internal” level if the data is accessible to a limited number of people within the user’s organization; - an “extended internal” level if the data is accessible to any persons within the user’s organization; - a “restricted external” level if the data is accessible to a limited number of people outside the user’s organization; - an “extended external” level if the data is accessible to any persons outside the user’s organization.

[0017] This makes it possible to assess the risk differently depending on the accessibility of the data by a third party in order to be more precise in the risk assessment, and to improve the quality of the anonymized data. The level of exposure makes it possible to calculate the correlation criterion more precisely and to identify the data most at risk.

[0018] According to a variant, the characterization means makes it possible to attribute to each data item a type of sensitivity on the basis of three variables, preferably among: - a sensitive data type if the data item can have personal impacts on the person concerned; - a data type perceived as sensitive if the data is perceived as sensitive; - a common data type if the data is usual and not sensitive.

[0019] This makes it possible to discriminate the sensitivities of different data to better assess the risk and improve the quality of anonymized data.

[0020] According to a variant, the means for identifying feared events makes it possible to assign at least one severity scale to four variables, preferably: minor, significant, serious, and critical.

[0021] This makes it easy to assess the severity of disclosure of each data and dataset. It also makes it possible to note these events for each data and predict them in a report.

[0022] According to a variant, the means of evaluating the legal anonymization criteria makes it possible to assign scores whose number of levels can vary between 3 and 5, preferably: low, moderate, high, very high.

[0023] This makes it possible to simplify the assessment of the risk of re-identification.

[0024] According to a variant, the means for evaluating the level of exploitability makes it possible to assign levels of exploitability, the number of values ​​of which can vary between 3 and 5, preferably very difficult, difficult, easy, very easy.

[0025] This makes it possible to simplify the assessment of the risk of data exploitability by a third party.

[0026] According to a variant, the system further comprises means for generating a color code with a scale of importance for at least one evaluated level.

[0027] This makes it possible to quickly visualize the danger of an assessed risk.

[0028] According to a variant, the system further comprises a means for editing a report including, among other things, the identified risks and / or the countermeasures.

[0029] This makes it possible to prove to a third party or an institution that the risks have been assessed and measures have been taken to minimize them.

[0030] Another object of the invention relates to a method for anonymizing personal data in at least one set of data of a user, the anonymization method comprising: - a data characterization step based on a data schema in which each piece of data is assigned a name, a level of exposure allowing the possibility of accessing the data to be assessed depending on whether it can be accessed from outside the organization or from inside, and preferably a type of data sensitivity; - a step of identifying feared events in which each piece of data is assigned a severity scale preferably broken down into a type of feared event and / or a scale of impact of this event, evaluated on at least one of the material, bodily and moral levels; - a step of evaluating the legal anonymization criteria in which each data item is assigned an individualization score, a correlation score, and an inference score; - a step of evaluating a level of exploitability based on a combination of exposure levels and individualization, correlation, and inference scores; - a step of overall risk assessment of the data set based on risk hypotheses constructed using data with a significant level of exploitability and a significant severity scale; and - a risk reduction step including proposals for countermeasures or transformations to limit the level of exploitability and / or the severity scale of the most at-risk data.

[0031] The invention further relates to a computer program comprising program code instructions for executing the steps of a data anonymization method according to the invention, when said program operates on a computer.

[0032] The invention will be further detailed by the description of non-limiting embodiments, and on the basis of the appended figures illustrating preferred embodiments of the invention, in which: - [Fig.l] schematically illustrates an anonymization system according to a preferred embodiment of the invention, and an edition of a report which can be physical or electronic; - [Fig.2] schematically illustrates an anonymization method according to one embodiment; - [Fig.3] schematically illustrates a first example of data risk analyses of a data set using the invention, before and after application of a maximum countermeasures according to the invention for an internal hacker or an external hacker; and - [Fig.4] schematically illustrates a second example of data risk analyses similar to that of [Fig.4] with minimal countermeasures according to the invention.

[0033] The invention relates to a system and a method for anonymizing personal data in at least one data set of a user.

[0034] The invention is implemented by computer means, for example via a computer, a server, a tablet, a smartphone or other or a combination of at least two of these elements.

[0035] The anonymization system comprises several hardware and software means for loading and evaluating the data, preferably transforming it or proposing countermeasures to improve the security of the data set.

[0036] The anonymization system comprises a means for loading at least one data set to be analyzed. This is in particular a module for reading a file comprising said data set. The data set may be in any file format, for example in computer file format (called “.csv”).

[0037] The anonymization system further comprises a means for identifying a data schema. The data schema is illustrated in Table 1. In this data schema, each data item is assigned a name, preferably a label, a level of exposure to people inside or outside the user's organization, and preferably a type of sensitivity of the data.

[0038] Table 1: Example of data schema # Name Label Exposure Sensitivity 1 Antibiogram Molecule X 1-Internal restricted common 2 Bacterial species Species Y 1-Internal restricted common 3 Type of sampling 8 modalities 1-Internal restricted common 4 Date of sampling Shifted by a random number 2-Internal expanded common 5 MALDLT oF spectrum Single vector 1000 values ​​1-Internal restricted common 6 Date of birth Rounded by 5 years 4- External expanded common 7 Sex Two categories 4- External expanded common

[0039] In the field of health or clinical studies, the name of the data is for example antibiogram; bacterial species; type of sample; date of sample; such and such name of test, for example MALDI-ToF spectrum; date of birth, sex. The “Name” can concern an identification number, a license number, a number of points in the field of road safety, or the name of any test in another technical field.

[0040] In addition to the “Name”, you can define the “Label” to enter a summary description of the data.

[0041] The “Exposure” level of an attribute evaluates how easy it is for a hacker to obtain the information in question from another dataset. For example, it is likely more difficult to obtain a person’s “Sample Type” from another dataset than to find their “Date of Birth” or “Sex”.

[0042] The level of “Exposure” is discriminated according to the third parties internal to the organization using the data set, and those internal or external to this organization not having the right to access the data. Furthermore, the level of exposure is preferably identified by less than ten discrete variables, preferably four variables.

[0043] The level of exposure can be classified into four main categories in ascending order of level: - internal-restricted; - internal-enlarged; - external-restricted; and - externally expanded.

[0044] Regarding the internal-restricted level: Attributes with an internal-restricted exposure level are accessible only to a limited number of people authorized to view them. These attributes may include medical information, sensitive financial data, or sensitive personal information.

[0045] In particular, we will speak of a “restricted internal” level if the data is accessible to a restricted number of people internal to the user’s organization, for example people belonging to a particular department.

[0046] A traffic light-type color code can be generated depending on the exposure level. The restricted internal level is, for example, green V because data in this category is less likely to be found in other data sets.

[0047] For example, the bacterial species, analysis results, banking operations may have a restricted internal level, and a green V color code.

[0048] Regarding the internal-extended level: attributes having an internal-extended exposure level are accessible within the organization, by employees belonging to several departments or by all employees. These attributes may include employee identification data, internal activity reports, but also data of data subjects (e.g., patients, customers, etc.) passing from one department to another (e.g., customer / patient identifiers, heights, weights, blood pressures, etc.)

[0049] In particular, data has this level of exposure if it is accessible to any persons within the organization or several different departments. The color code is for example yellow J.

[0050] Regarding the external-restricted level: Attributes with an external-restricted exposure level are accessible to third parties, but require specific research or specific data sources to access them. They may include information shared with business partners, survey data, or industry-specific information.

[0051] In particular, data has this level if it is accessible to a limited number of people external to the user's organization due to the complexity required to collect it. These include people who may know the information in question or find it by conducting in-depth research. For example, an admission date, a discharge date may have a restricted external level. The color code is, for example, orange O (or dark orange).

[0052] Regarding the external-extended level: Attributes with an external-extended exposure level are easily accessible and can be obtained from external sources without much difficulty. These attributes may include publicly available information, such as a name, first name, age, postal address, email address, or general professional data.

[0053] In particular, data is at this level if it is accessible to any persons external to the user's organization, for example from social networks and search engines. The color code is for example red R.

[0054] Data breach depends, among other things, on the ability to cross-reference different data sets and therefore the ability to find in another data set, data present in the anonymized data set.

[0055] In the context of the invention, it is desired to be able to assign a score between 1 and 4 which quantifies the degree of exposure of an attribute in the dataset. It comes quite naturally to propose the following scores: external-extended with a score equal to 4 (critical), external-restricted with a score equal to 3 (high), internal-extended with a score equal to 2 (medium) then internal-restricted with a score equal to 1 (low).

[0056] A high exposure score may be a sign that randomization should be applied to this attribute. This could make it more difficult to find adequate values ​​in external sources.

[0057] In addition to the exposure level, the identification means makes it possible to assign to each piece of data a type of sensitivity with three variables. We can speak of CNIL data type. We can distinguish: - a sensitive data type if the data concerns racial or ethnic origin, political opinions, religious or philosophical beliefs or trade union membership, as well as the processing of genetic data, biometric data for the purpose of uniquely identifying a natural person, data concerning health or data concerning a natural person's sex life or sexual orientation; - a type of data perceived as sensitive if the data is perceived as sensitive, for example banking data, biometric data, a social security number; - a common data type if the data is neither sensitive nor perceived as sensitive.

[0058] The risk of re-identification depends on this sensitivity.

[0059] The anonymization system further comprises a means for identifying feared events. This information makes it possible to prove that this aspect has been evaluated before the use of the data. For this purpose, each data item is assigned a severity scale to determine a minor, significant, serious or critical event to be expected in the event of loss of the data. Preferably the scale contains less than 10 levels, more preferably less than 6, more preferably 4 variables are used for the severity scale. Furthermore, the severity preferably comprises a type of feared event, for example the disclosure of a patient's illness, and / or a scale of impact of this event, evaluated on at least one of the three criteria: material, physical, moral.

[0060] Table 2: example of feared events Sensitive Attribute Event re doubted Bodily Material Moral Severity Bacterial Species Disease Disclosure Minor Severe Critical Critical

[0061] For better visibility and rapid understanding, the feared events are assigned a traffic light type color code, with a green color V for events of no seriousness, a yellow color J and / or orange color O for events of intermediate seriousness, and a red color R for significant seriousness.

[0062] The anonymization system further comprises a means for evaluating the criteria for re-identifying people in which each data item is assigned an individualization score, a correlation score, an inference score. In each of these cases, the score is preferably evaluated on a scale of less than 10 values, more preferably between 3 and 5 values, more preferably according to the values: low, medium, high, and critical. These scores can be evaluated collegially, conventionally, or statistically for each data item.

[0063] Table 3: Re-identification risk assessment # Name Individualization Correlation Inference 1 Antibiogram 1-Low 1-Low 2-Moderate 2 Bacterial species 3-High 1-Low 4-Very high 3 Sample type 1-Low 1-Low 1-Low 4 Sample date 4-Very high 3-High 4-Very high 5 MALDLT oF spectrum 3-High 1-Low 3-High 6 Date of birth 2-Moderate 3-High 2-Moderate 7 Sex 1-Low 3-High 1-Low

[0064] Similarly, the traffic light type color code can be used. Critical risk is red R; high risk, orange O; medium risk, yellow J; and low risk, green V.

[0065] The individualization score can be assessed or calculated. Individualization assesses the ability to isolate an individual in the dataset. It refers to the level of detail or specificity of the information contained in each attribute. Some attributes may be very granular, providing precise information such as the full date of birth or full address. Other attributes may be more general, with a lower individualization score, such as the year of birth or city of residence. The individualization score of the attributes may influence the potential for disclosure and the associated level of risk.

[0066] A high individualization score can be addressed by applying generalization. However, it remains important to consider the loss of data utility that may accompany generalization.

[0067] The risk of inference can be assessed or calculated. Inference can be likened to the statistical concept of discrimination rate. This is a measure that assesses the ability of an attribute to distinguish or discriminate an individual from others. It is often used to assess the risk of re-identification of individuals from anonymized data. An attribute with a high discrimination rate can provide information that allows

[0068]

[0069]

[0070]

[0071]

[0072]

[0073]

[0074] to identify or reveal specific personal characteristics, which increases the risk of breach of confidentiality. As for the rx ratio obtained for the granularity measurement, the discrimination rate is a value in the interval [0, 1]. By denoting DR this metric, we can define the SDR score associated in a similar way to the granularity score: ^DR — |"4 • DR]with ^4 the upper integer part, also called “ceiling”. The correlation criterion is assessed on the basis of the exposure level and the individualization criterion. The anonymization system further includes a means of assessing the level of exploitability comprising at least a combination of the levels of exposure and risks of individualization, correlation, and inference. In essence, when the scores are high, then exploitability is easy, and a hacker can easily gain access to personal data; and vice versa. Table 4: Exploitability risk assessment # Individualization Correlation Inference Operability 1 1-Low 1-Low 2-Moderate 1-Very difficult 2 3-High 1-Low 4-Very high 2-Difficult 3 1-Low 1-Low 1-Low 1-Very difficult 4 4-Very high 3-High 4-Very high 4-Very easy 5 3-High 1-Low 3-High 2-Difficult 6 2-Moderate 3-High 2-Moderate 3-Strong 7 1-Low 3-High 1-Low 2-Difficult An exploitability assessment scale can be as follows in Table 5. Table 5: Means of assessing exploitability risks Category Individualization Correlation Inference Exploitability Expanded external (sex) 1-Low 2-Moderate 1-Low 1-Very difficult Expanded internal (sampling date) 4-Very high 3-High 4-Very high 4-Very easy Restricted internal (Antibiogr.) 4-Very high 1-Low 4-Very high 2-Difficult Exploitability can be color-coded. Very easy is in red R, difficult in yellow J, and very difficult in green V.

[0075] The anonymization system further comprises a means for evaluating the overall risk of the data set in which the data having a significant level of exploitability and / or a significant level of severity are identified.

[0076] The exploitability level and the severity level are preferably represented in a two-dimensional graph (one for each level), to allow better visualization of the risks and easy comparison of anonymized games. The graph preferably includes a traffic light type color code. This type of graph is illustrated in [Fig.3] and 4.

[0077] The anonymization system further comprises a means for correcting the data set comprising proposals, preferably automatic, for countermeasures or transformations of the data having a significant level of exploitability and / or a significant level of severity so as to reduce said level.

[0078] The system can alert on identification risks on high exploitability levels, identify vulnerabilities via first vulnerability modules noted D1, D3, etc., and propose corresponding countermeasures: - a Dl module requires generalizing as much as possible the variables involving a high level of exploitability (easy) for example the expanded internal variables such as the sampling dates. - module D3 identifies the risks of re-identification linked to restricted internal variables.

[0079] The system may further analyze GDPR compliances using at least one of the following specific vulnerability modules: - a Cl module requiring compliance with the principle of data minimization (because only data strictly necessary for the study must be used according to the regulations, it will then be necessary to define a list of data and justify the use made; - a C2 module requesting to delete data of special persons; - a C3 module requiring the definition of a reasonable data retention period depending on the purpose of the data processing; - a C4 module requiring the provision of a data purging mechanism at the end of data processing; - a C5 module prohibiting the export of data outside the system; - a C6 module requiring information to the persons concerned that the data will be subject to anonymization; - a C7 module alerting on the location of the data, because the data must be located in the EU zone, failing which standard or binding contractual clauses with the cloud network provider must be signed; - a C8 module requiring the stakeholder to undertake not to try to re-identify people, it can be coupled with module D3; - a C9 module requiring a high-security password, for example according to ANSSI recommendations.

[0080] The anonymization system further comprises a means for editing an R report containing the identified risks and countermeasures. The R report automatically contains the risk analysis elements and serves as proof to show the authorities that the risks have been assessed and minimized. The R report can be physical or electronic.

[0081] Example RI:

[0082] Hypothetically, a hacker exfiltrates insufficiently anonymized data after exploiting a weakness in authentication, and reidentifies the persons concerned on the basis of the expanded internal variables.

[0083] The system of the invention analyzes the data and identifies the vulnerabilities of the modules D1, D3, C2 and C8 with a high risk of re-identification (for example a level 4 - very easy (red).

[0084] A first set of countermeasures is proposed, for D1, a countermeasure CM1: generalize the expanded internal variables; for C8, a countermeasure CM3: define a password policy using ANSSI standards; for C2, a countermeasure CM5: delete the data of special people; for D3, countermeasures CM6: generalize the restricted internal variables.

[0085] Countermeasures limit the risk from level 4-Critical to level 2-Medium.

[0086] On an initial exploitability scale of 3 / 4, and a severity scale of 4 / 4 directly visible in the system (top of [Fig.3]), the countermeasures lower the exploitability to 1 / 4, and the severity to 3 / 4 as illustrated at the bottom of [Fig.3].

[0087] The system can propose minimal countermeasures to be taken into account, here CM1 and CM3. The countermeasures maintain the risk at level 2-Medium. In this case, the severity remains at 4 / 4, but the exploitability drops to 1 / 4 as illustrated in [Fig.4].

[0088] Example R2:

[0089] Hypothetically, a researcher re-identifies the individuals concerned on the basis of the restricted internal variables.

[0090] The system of the invention analyzes the data and identifies the vulnerabilities of the modules D1, D3, C2 and C7 with a high risk of re-identification (for example a level 4 - very easy (red).

[0091] A first set of countermeasures is proposed, for Dl, a countermeasure CM1: generalize the expanded internal variables; for D3 and C7, countermeasures CM6: generalize the restricted internal variables and CM4: provide measures contractual for the operator who undertakes not to attempt to re-identify individuals; for C2, a CM5 countermeasure: delete the data of special individuals.

[0092] Countermeasures limit the risk to a level 2-Medium.

[0093] On an initial exploitability scale of 3 / 4, and a severity scale of 4 / 4 directly visible in the system, the countermeasures lower the exploitability to 1 / 4, and the severity to 3 / 4 as illustrated in [Fig.3].

[0094] The system preferably proposes minimal countermeasures to be taken into account, here CM1 and CM4. The countermeasures maintain the risk at level 2-Medium. In this case, the severity remains at 4 / 4, but the exploitability drops to 1 / 4 as illustrated in [Fig.4],

[0095] Preferably, the level of correlation is determined as a function of the level of exposure and the level of individualization on the basis of the following table:

[0096] Table 6: Evaluation of the level of correlation Exposure Individualization Internal restricted Internal expanded External restricted External expanded Very high 1-Low 3-High 4-Very high 4-Very high High 1-Low 2-Moderate 4-Very high 4-Very high Moderate 1-Low 2-Moderate 3-High 3-High Low 1-Low 1-Low 2-Moderate 2-Moderate

[0097] Preferably, the level of exploitability is determined based on the level of correlation and the level of inference based on the following table:

[0098] Table 7: Evaluation of the level of exploitability Exploitability Inference 1-Low 2-Moderate 3-High 4-Very High Very High 2-Difficult 3-Easy 4-Very Easy 4-Very Easy High 2-Difficult 2-Difficult 3-Easy 4-Very Easy Moderate 1-Very Difficult 2-Difficult 3-Easy 3-Easy Low 1-Very Difficult 1-Very Difficult 2-Difficult 3-Easy

[0099] Preferably, a contextual exploitability level is determined based on the exploitability level and the above remarks and countermeasures, based on the following table:

[0100] Table 8: Evaluation of the level of exploitability Contextual exploitability Contextual exploitability 1-Low 2-Moderate 3-High 4-Very high Very easy 2-Difficult 3-Easy 4-Very easy 4-Very easy Easy 2-Difficult 2-Difficult 4-Very easy 4-Very easy Difficult 1-Very difficult 2-Difficult 3-Easy 3-Easy Very difficult 1-Very difficult 1-Very difficult 2-Difficult 3-Easy

[0101] Preferably, the risk of re-identification (legal anonymization criteria) is determined based on the severity level and the exploitability level (preferably contextual exploitability) based on the following table:

[0102] Table 9: Evaluation of the level of exploitability Severity Exploitability Negligible Limited Significant Maximum Very easy 2-Medium 3-High 4-Critical 4-Critical Easy 2-Medium 2-Medium 3-High 4-Critical Difficult 1-Low 2-Medium 3-High 3-High Very difficult 1-Low 1-Low 2-Medium 2-Medium

[0103] The invention further relates to a method for anonymizing personal data in at least one data set of a user, based on a system as described above.

[0104] The method comprises steps of implementing the different modules.

[0105] The method is implemented by computer.

[0106] The invention also relates to a computer program for implementing the invention via computer means.

Claims

Claims

1. Data anonymization system implemented by computer means, taking as input at least one data set in an organization, the anonymization system comprising: - a means of characterizing the data (Ml) on the basis of a "data schema" making it possible to define for each data item, a name, a level of exposure on the basis of a number of variables, making it possible to evaluate the possibility of access to the data depending on whether it can be accessed from outside the organization or from inside, and preferably a type of sensitivity of the data on the basis of three variables; - a means of identifying feared events (M2) making it possible to define for each data item, a severity scale with less than 10 variables preferably broken down into a type of feared event and / or a scale of impact of this event, evaluated on the basis of at least one given criterion;- a means of evaluating the legal anonymization criteria (M3) making it possible to define for each data item, an individualization score on a scale of less than 10 values, a correlation score on a scale of less than 10 values, and an inference score on a scale of less than 10 values; - a means of evaluating a level of exploitability (M4) on the basis of a number of values, determined on the basis of a combination of the exposure levels and the individualization, correlation, and inference scores; - a means of evaluating the overall risk (M5) of the data set on the basis of risk hypotheses constructed based on data having a significant level of exploitability and a significant severity scale; and - a means of reducing risks (M6) comprising proposals for countermeasures or transformations making it possible to limit the level of exploitability and / or the severity scale of the data most at risk.

2. Anonymization system according to any one of the preceding claims, characterized in that the characterization means (Ml) makes it possible to assign to each data item a level of exposure on the base of variables, the number of which may vary between 2 and 10, preferably four variables, more preferably among: - a “restricted internal” level if the data is accessible to a restricted number of people internal to the user’s organization; - an “extended internal” level if the data is accessible to any people internal to the user’s organization; - a “restricted external” level if the data is accessible to a restricted number of people external to the user’s organization; - an “extended external” level if the data is accessible to any people external to the user’s organization.

3. Anonymization system according to any one of the preceding claims, characterized in that the characterization means (Ml) makes it possible to attribute to each data a type of sensitivity on the basis of three variables, preferably from among: - a sensitive data type if the data can have personal impacts on the person concerned; - a data type perceived as sensitive if the data is perceived as sensitive; - a current data type if the data is usual and not sensitive.

4. Anonymization system according to any one of the preceding claims, characterized in that the means for identifying feared events (M2) makes it possible to assign at least one severity scale to four variables, preferably minor, significant, serious, and critical.

5. Anonymization system according to any one of the preceding claims, characterized in that the means for evaluating the legal anonymization criteria (M3) makes it possible to assign scores whose number of levels can vary between 3 and 5, preferably low, moderate, high, very high.

6. Anonymization system according to any one of the preceding claims, characterized in that the exploitability level evaluation means (M4) makes it possible to assign exploitability levels, the number of values of which can vary between 3 and 5, preferably very difficult, difficult, easy, very easy.

7. Anonymization system according to any one of the preceding claims, characterized in that it further comprises means generation of a color code (V, J, 0, R) with an importance scale for at least one level evaluated.

8. Anonymization system according to any one of the preceding claims, characterized in that it further comprises means for editing a report (R) including, among other things, the identified risks and / or the countermeasures (CM1, CM3, CM4, CM5, CM6).

9. Method for anonymizing personal data in at least one set of data of a user, the anonymization method being implemented by computer means, and comprising: - a step of characterizing the data on the basis of a data schema in which each data item is assigned a name, a level of exposure on the basis of a number of variables, making it possible to evaluate the possibility of access to the data depending on whether it can be accessed from outside the organization or from inside, and preferably a type of sensitivity of the data on the basis of three variables; - a step of identifying feared events in which each data item is assigned a severity scale with less than 10 variables preferably broken down into a type of feared event and / or a scale of impact of this event, evaluated on at least one of the material, bodily, moral plane;- a step of evaluating the legal anonymization criteria in which each data item is assigned an individualization score on a scale of less than 10 values, a correlation score on a scale of less than 10 values, and an inference score on a scale of less than 10 values; - a step of evaluating the level of exploitability of evaluating a level of exploitability on the basis of a number of values, determined on the basis of a combination of the exposure levels and the individualization, correlation, and inference scores; - a step of evaluating the overall risk of the data set on the basis of risk hypotheses constructed based on data having a significant level of exploitability and a significant severity scale; and - a step of reducing risks including proposals for countermeasures or transformations to limit the; 19 exploitability level and / or severity scale of the most at-risk data.

10. A computer program comprising program code instructions for performing the steps of a data anonymization method according to claim 9, when said program is running on a computer.

Citation Information

Patent Citations

  • Method and device for anonymizing data stored in a database

    US11068619B2

  • Simulated risk contribution

    US20210049282A1

  • Computer-implemented privacy engineering system and method

    US20230359770A1