Privacy index system based on privacy attribute inline relation
By constructing a multi-dimensional privacy assessment indicator module, the problem of lack of relevance and dynamic response in the existing privacy assessment system is solved, realizing a full-process, multi-dimensional privacy protection assessment, improving the timeliness and accuracy of the assessment, and providing precise privacy decision support.
Patent Information
- Application Number
- CN202511217484.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-11-07
AI Technical Summary
The existing privacy assessment system lacks the correlation of multi-dimensional indicators, which leads to the disconnect between assessment results and actual risks, making it difficult to achieve dynamic response and global optimal solution, and failing to balance privacy protection and data availability.
A multi-dimensional privacy assessment indicator module is constructed, including indicator system decomposition and algorithm adaptation, indicator association mapping, and dynamic update module. The relationship between indicators is established through the analytic hierarchy process and association rule mining to achieve real-time dynamic assessment and optimization.
It achieves a comprehensive, multi-dimensional privacy protection assessment covering the entire process, improving the timeliness and accuracy of the assessment, providing precise privacy decision support, and reducing the risk of privacy leaks.
Smart Images

Figure CN120910394A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of data privacy protection, and particularly relates to a privacy index system based on internal relations of privacy attributes through information security evaluation, dynamic system modeling and multi-dimensional index linkage analysis. BACKGROUND
[0002] With the rapid development of data economy, privacy protection has become a core demand in the fields of artificial intelligence, big data analysis, cloud computing, and Internet of Things. Under the wave of digital transformation, enterprises and institutions collect, process, and analyze user data more frequently, and the application scenarios cover intelligent recommendation, precision marketing, medical health data sharing, and other fields. Privacy leakage not only brings economic losses and information misuse risks to users, but also may make enterprises face regulatory penalties and reputation crises. Under this background, traditional privacy evaluation methods (such as single index independent evaluation, static rule verification, etc.) have significant limitations such as the isolation of single index, the lack of dynamic response capability, and the lack of systematic evaluation.
[0003] The existing privacy evaluation system usually uses independent indexes for single-point evaluation, without considering the internal correlation between indexes. This isolated evaluation mode breaks the integrity of the privacy protection system, leading to a serious disconnection between the evaluation results and the actual risks. In intelligent recommendation systems, excessive optimization of privacy protection strength may cause the recommendation algorithm to fail due to data distortion; in medical data analysis scenarios, noise interference may even lead to an increase in misdiagnosis rates. On the other hand, if data usability is excessively pursued, the anonymization processing or access permission restrictions on data are relaxed, which may accidentally increase the risk of user identification, allowing attackers to restore sensitive user information through data correlation analysis. This "trade-off" optimization method is prone to cause "index optimization imbalance" problems, making enterprises trapped in a dilemma between privacy protection and business development.
[0004] In technology development, it is common to adjust technical parameters such as differential privacy parameters and anonymization levels, but traditional evaluation methods lack tools to quantify the impact of real-time parameter modification on multi-dimensional indexes, relying on manual experience and trial-and-error, which is inefficient and prone to systemic risks. For example, when adjusting the differential privacy strength, the privacy strength and data accuracy changes need to be manually calculated, and multiple experiments are required to verify, which may ignore related indexes such as system performance and compliance, leading to privacy and data usability conflicts (such as excessive response time affecting experience, forcing to relax privacy standards).
[0005] Privacy protection is essentially a multi-objective optimization problem, which needs to balance privacy security, data utility, compliance, user experience and other multiple objectives. However, the existing evaluation scheme cannot support dynamic simulation of "pulling a hair and moving the whole body", and it is difficult to achieve the global optimal solution in the process of technical iteration. There is a game relationship between the objectives, such as increasing the encryption strength to enhance the security but reducing the processing efficiency, strictly minimizing the data to meet the compliance but limiting the data value. Traditional methods lack systematic modeling and cannot depict the interaction between objectives, resulting in a lack of scientific basis for enterprise strategy formulation and easily falling into the dilemma of over-conservatism (sacrificing innovation) or ignoring risks (facing penalties). Especially in complex scenarios such as cross-border data and emerging technologies, traditional schemes cannot predict the comprehensive influence of multiple factors, making it difficult to develop sustainable privacy strategies. SUMMARY
[0006] The purpose of the present application is to provide a privacy evaluation system based on index correlation dynamic analysis, which can comprehensively consider the dimensions of privacy evaluation indicators and perform linkage analysis and dynamic updating when the indicators change, solving the problems of single-dimensional evaluation, lack of overall analysis and dynamic response capability of existing evaluation systems.
[0007] The technical scheme of the present application covers data lifecycle indicators at each link by constructing a multi-dimensional privacy evaluation indicator module, and establishes an index correlation model to analyze the internal logical relationship. When a certain indicator changes, the system triggers a dynamic updating mechanism according to the pre-set correlation rules and algorithms, automatically recalculates and adjusts the affected indicator values and evaluation results, and realizes real-time dynamic and comprehensive evaluation of privacy status.1. A privacy indicator system based on attribute internal relationship, characterized by comprising a multi-dimensional indicator calculation module, an index correlation mapping module, a dynamic updating module, a rule verification and optimization module;
[0008] The multi-dimensional indicator calculation module includes an indicator system disassembly and algorithm adaptation unit and a multi-dimensional evaluation value aggregation calculation unit;
[0009] The indicator system disassembly and algorithm adaptation unit performs hierarchical analysis on the privacy protection effect evaluation index system, and adapts the calculation algorithm for different types of indicators;
[0010] The multi-dimensional evaluation value aggregation calculation unit uses the analytic hierarchy process to construct the index hierarchy, first compares the indicators two by two, and constructs the judgment matrix; then uses the eigenvalue method to solve the matrix and obtains the weight of each indicator;
[0011] The index correlation mapping module includes an association relationship mining and modeling unit and a rule processing and matrix construction unit;
[0012] The association relationship mining and modeling unit constructs an initial association rule library as static rules through the known logical relationship between indexes, and then analyzes historical evaluation data and simulation test data through an association rule mining algorithm to mine potential associations as dynamic rules, thereby supplementing and perfecting the association rule library.
[0013] The rule processing and matrix construction unit generates a comprehensive rule set by comprehensively combining static rules and dynamic rules, and constructs an index association matrix according to the obtained dynamic rules.
[0014] The dynamic updating module includes a fluctuation benchmark determination unit, a user demand conversion unit and an index updating unit.
[0015] The fluctuation benchmark determination unit calculates the fluctuation benchmark of an index in different scenarios according to the business characteristics and historical data in the different scenarios.
[0016] The user demand conversion unit converts the user's description of the enhancement of an index into a change amplitude in units of the fluctuation benchmark when receiving the user's demand to enhance a certain privacy index.
[0017] The index updating unit extracts the influence coefficient of a specified index on other indexes from the index association matrix, calculates the linkage change value of other indexes caused by the update of the specified index, and updates the benchmark values of other indexes based on the linkage change value.
[0018] The rule verification and optimization module performs rule validity verification on the updated index system and updates and iterates the index association matrix.
[0019] Further, the index evaluation system covers four first-level indexes: data characteristics, compliance, availability and security; and 22 second-level indexes.
[0020] The calculation algorithm is adapted to different types of indexes as follows: the basic statistical index traverses the data set to calculate the characteristic value; the privacy model index substitutes the data into the model to calculate the privacy protection parameter under the model constraint; and the risk assessment index calculates the privacy leakage risk probability value by constructing an attack model.
[0021] Further, the multi-dimensional evaluation value aggregation calculation unit is as follows:
[0022] First, determine the hierarchical structure and input parameters: clearly define the decision goal as the weight allocation of the second-level indexes under each first-level dimension, the criterion layer as the second-level indexes to be calculated, and construct an index system hierarchical structure diagram to clearly define the hierarchical relationship; input the judgment matrix T=[a ij ]∈R n×n , where a ij represents the relative weight of index m i to index m jThe relative importance of the matrix; matrix dimension n and the random consistency index RI;
[0023] Calculate the normalized value of each element in a column of the judgment matrix, which is the ratio of that element to the sum of all elements in its column, thus standardizing the column dimensions of the matrix. For the column-normalized matrix, calculate the sum of the elements in each row to obtain the row sum N. i ; Combine each row with N i After normalization, the weight values of each indicator are obtained, forming a weight vector V = [v1, v2, ..., v n ] T , where v i ,i=1,2,…,n are the weights corresponding to the secondary indicators;
[0024] To obtain the weight vector V = [v1, v2, ..., v] for each secondary indicator... n ] T Then, it needs to be combined with the original indicator data to calculate the comprehensive aggregate value for each evaluation object. The specific calculation formula is as follows:
[0025] A 聚合 =∑A i *v i
[0026] Where A 聚合 A represents the aggregated index value. i This is a secondary indicator;
[0027] Finally, a consistency check is performed to complete the construction of the judgment matrix.
[0028] Furthermore, the relationship mining and modeling unit is specifically as follows:
[0029] The association relationship includes static rules and dynamic rules. First, static rules are constructed based on publicly available academic literature and industry standards to quantify and score the logical relationship between each indicator. Then, the association rule is represented by a four-tuple structure, namely (indicator A, association type, indicator B, weight). The association type includes positive correlation, negative correlation, and non-linear correlation. The weight is calculated by the analytic hierarchy process.
[0030] Then the dynamic rules are constructed; first, the historical evaluation data are acquired and the simulation test data are set, and the data are standardized; the numerical indicators and discrete indicators are defined as different items respectively, the Apriori algorithm is used for frequent item set mining, the minimum support threshold, the minimum confidence threshold and the maximum item set length are set, wherein the support represents the probability of the occurrence of a single item, and the confidence represents the probability of the occurrence of B at the same time under the condition of the occurrence of A; the rules of "if the value of indicator A belongs to a certain category, then the value of indicator B is likely to belong to a certain category" are extracted from the frequent item set, and the confidence is quantified; the lift of the extracted relationship is calculated to verify the non-accidentalness of the specific rules, the correlation of the specific rules is verified and extracted by using the chi-square test+Cramer's V algorithm; the lift takes indicator B as an example, and the calculation formula is as follows:
[0031] Lift(A→B)=Confidence(A→B) / Support(B)
[0032] Wherein, Confidence(A→B) represents the confidence of indicator B relative to indicator A, and Support(B) represents the support of indicator B;
[0033] The dynamic strength is calculated by comprehensively considering the support, the confidence and the lift.
[0034] Dynamic strength=support×confidence×lift / (1+support×confidence×lift)
[0035] The dynamic rules are generated in the form of four-tuple (indicator A, association type, indicator B, dynamic strength).
[0036] Further, the rule processing and matrix construction unit is as follows:
[0037] If there is only static rule or only dynamic rule between two attributes, they are directly merged; if there is both static rule and dynamic rule between two attributes, the total strength is calculated according to the following formula:
[0038] Comprehensive strength=0.6×dynamic strength+0.4×static strength
[0039] Wherein, the dynamic strength is the weight in the dynamic rule set, and the static strength is the weight in the static rule;
[0040] According to the obtained dynamic rules, the "privacy indicator" is taken as an identifier, and the indicator association matrix M=R n×n is constructed, wherein n represents the number of privacy indicators, M i,j represents the influence coefficient of indicator i on indicator j, and the value range is [-1, 1], the absolute value represents the influence strength, and the sign represents the influence direction.
[0041] Further, the fluctuation benchmark determination unit is specifically as follows: for an index with historical data, the mean and standard deviation are calculated by a statistical method, and the standard deviation is taken as the fluctuation benchmark; for a new index or a scene lacking historical data, an initial baseline is set through expert research combined with industry best practices and business objectives, and the initial baseline is dynamically calibrated according to actual data feedback in system operation;
[0042] In the user demand conversion unit, the demand is divided into qualitative demand and quantitative demand; for the qualitative demand, conversion is performed according to a preset industry conventional adjustment amplitude conversion rule, and for the quantitative demand, a percentage is converted into a multiple of a fluctuation benchmark according to a historical data distribution feature of the index, and finally a change value in units of the fluctuation benchmark is output.
[0043] Further, the index updating unit specifically updates the index through the following steps:
[0044] An influence coefficient M of the index i on the index j is extracted from the correlation matrix M i,j A linkage change value Δ of the index j due to the index i updating j The calculation formula is: Δ j = Δ i × M i,j ;
[0045] The current benchmark value of the index j is k j , and the updated benchmark value k j ' is: k j '= k j + Δ j .
[0046] Further, the rule verification and optimization unit is specifically as follows:
[0047] Rule validity verification: historical business scene data is selected, Δ i converted by the user demand and the correlation matrix M are substituted into the system, the index linkage updating process is simulated, the deviation of the updating result from the historical actual business performance is compared, and the deviation rate is calculated:
[0048] Deviation rate = (|actual change value-simulated change value| / simulated change value) x 100%
[0049] If the deviation rate exceeds a preset threshold, it is marked as not passing the verification, and the influence coefficient of the correlation matrix M or the conversion logic of Δ i needs to be checked back;
[0050] Correlation matrix iteration: based on the deviation data in the verification stage, the influence coefficient M i,j of the correlation matrix M is updated regularly: if the actual influence of the index i on j is greater than the simulation value, M i,j is adjusted upward by the deviation proportion; if the actual influence is less than the simulation value, M is adjusted downwardi,j ; for the emerging business scenario, supplement the index correlation in the scene, mark the initial M i,j And through the subsequent data verification iteration.
[0051] The beneficial effects of the present application are as follows:
[0052] 1. Breakthrough of the integrity of the evaluation system: the existing privacy evaluation focuses on a single dimension, and the present application relies on a system covering four dimensions of data characteristics, compliance, usability and security and subdivided indicators, realizing full-process and multi-dimensional coverage of privacy protection effect. From data collection (such as quasi-identifier dimension, sensitivity dimension) to use and destruction (such as sensitive attribute re-identification risk), the privacy protection state is quantified comprehensively, solving the one-sidedness problem of traditional evaluation and providing a more complete privacy protection effect portrait for users or institutions.
[0053] 2. Technical advantages of dynamic correlation analysis: by establishing the correlation mapping and dynamic updating rules between indicators, the traditional evaluation "static fragmentation" limitation is broken. When any indicator (such as data distinguishability, entropy leakage) changes, the associated indicators can be automatically triggered for linkage analysis, reflecting the chain effect of indicator changes on the system in real time, allowing the evaluation to be upgraded from "post-static statistics" to "in-process dynamic perception", accurately capturing the privacy risk transmission path and improving the timeliness and accuracy of the evaluation.
[0054] 3. Precise support for privacy decision and optimization: providing clear guidance for privacy protection strategy optimization: on the one hand, the dynamically updated evaluation results can quantitatively present the effect of strategy adjustment (such as how the "entropy leakage" and "privacy gain" indicators change after optimizing the encryption algorithm); on the other hand, based on the correlation rules, the key indicators affecting the privacy protection effect can be quickly located (such as finding the "sensitive attribute re-identification risk" anomaly, which can be traced back to the associated indicators such as data record anonymity rate and K-Anonymity), helping users or institutions to optimize privacy protection measures, reduce privacy leakage risk, and improve the scientificity and efficiency of privacy governance. BRIEF DESCRIPTION OF DRAWINGS
[0055] Figure 1 System design of privacy evaluation system.
[0056] Figure 2 Design of privacy evaluation index system.
[0057] Figure 3 Visualization of privacy evaluation index system.
[0058] Figure 4 Visualization of privacy index system.
[0059] Figure 5 Comparison chart of predicted values of privacy index system. DETAILED DESCRIPTION
[0060] The present application will be further described below in conjunction with the accompanying drawings and specific embodiments, the schematic embodiments of the present application and the description are used to explain the present application, but not as a limitation of the present application.
[0061] A privacy attribute interrelation-based privacy index evaluation method, examples of use in the communication industry include Figure 1 As shown, comprising: a multi-dimensional index calculation module, quantifying the privacy protection effect through a multi-dimensional index calculation model; an index association mapping module, according to the index association mapping relationship, to break through the linkage analysis link; a dynamic updating module, designing a dynamic updating rule set, realizing system adaptive adjustment; a rule verification and optimization module, through simulation test, historical data backtracking verification, testing rule effectiveness, and optimizing rule parameters. Next, each module is described.
[0062] The multi-dimensional index calculation module includes an index system disassembly and algorithm adaptation unit and a multi-dimensional evaluation value aggregation calculation unit, and the specific index evaluation system designed by the present application covers four first-level indexes (data characteristics, compliance, availability, and security) and 22 second-level indexes.
[0063] The index system disassembly and algorithm adaptation unit performs hierarchical analysis on the privacy protection effect evaluation index system, as shown in Figure 2 The four first-level dimensions of data characteristics, compliance, availability, and security are clear, the definitions and calculation boundaries of the subordinate 22 second-level indexes such as "quasi-identifier dimension", "sensitive attribute dimension", and "K-Anonymity" are combed, and the calculation algorithm is adapted for different types of indexes; Specifically:
[0064] The basic statistical index (such as "minimum equivalent class size" and "average generalization degree") uses statistical counting, mean calculation, etc. to traverse the data set to calculate the characteristic value;
[0065] The privacy model index (such as "inherent privacy" and "average generalization degree") calls the privacy algorithm library and substitutes the privacy protection parameters under the data calculation model constraint;
[0066] The risk assessment index (such as "sensitive attribute re-identification risk" and "overall re-identification risk") calculates the privacy leakage risk probability value by constructing an attack model (such as simulating a re-identification algorithm).
[0067] As a further example, the specific indexes are as follows:
[0068] Quasi-identifier dimension (Dimen QI ): Quasi-identifier dimension refers to the number of quasi-identifier attributes in the data set. The quasi-identifier of individual u is represented as QID uQuasi-identifier dimensions: Quasi-identifier dimensions are those attributes that may not uniquely identify an individual when viewed in isolation, but may indirectly reveal identity when combined with multiple attributes. The size of the quasi-identifier dimension affects the complexity and effectiveness of privacy protection methods. If the quasi-identifier dimension is low, the data is easier to anonymize, but may provide less information. If the quasi-identifier dimension is high: the data is more likely to be re-identified, requiring more powerful anonymization techniques.
[0069] Dimen QI =|QID u |
[0070] Sensitive attributes: Sensitive attributes refer to information that needs to be protected and is not intended to be directly associated with an individual's identity. These attributes are usually the core content of subsequent research and public data sets, so they are generally not anonymized. For example: diseases, medical history, treatment information in medical data sets; income, credit score, bank account balance in financial data sets; religious beliefs, political leanings in population survey data, etc.
[0071] Sensitive attribute dimension (Dimen SA ): The sensitive attribute dimension refers to the number of sensitive attributes (SA) for individual u. The number of sensitive attributes for individual u is SA u .
[0072] Dimen SA =|SA u |
[0073] Minimum equivalence class size (priv minE ): The size of the equivalence class indicates how many records are grouped together, while the minimum equivalence class size indicates the smallest number of records in all equivalence classes in the data set. The size of the minimum equivalence class reflects the maximum K-anonymity that can be achieved in the data set, representing the lower limit of data privacy protection. Although in many cases the data set has been effectively de-identified, most equivalence classes contain a large number of records, but there are a small number of equivalence classes with few records. These small equivalence classes may become potential attack vulnerabilities, with user data in them facing a higher risk of exposure. The size of the equivalence class is represented as |E|.
[0074] priv minE =min(|E|)
[0075] Inherent privacy (priv IP ): Inherent privacy refers to the length of the uniform distribution interval corresponding to the uncertainty of the probability distribution of the data set. Inherent privacy is measured by entropy, which measures the level of privacy protection of a random variable (here referring to equivalence classes) in front of an attacker. It represents the theoretical number of binary questions an attacker needs to answer to re-identify the equivalence class. The larger the value, the greater the diversity of equivalence classes in the data set, and the lower the security. The entropy of the data set is represented as H(X).
[0076] priv IP = 2 H(X)
[0077] K-Anonymity: K-Anonymity means dividing a dataset into multiple equivalence classes, each containing at least K records. Specifically, if a dataset satisfies K-Anonymity, then for any given record, there are at least K-1 other records that share the same quasi-identifier values.
[0078] priv K ≡ k, where
[0079] L-Diversity: L-Diversity is often used to enhance the effectiveness of K-Anonymity methods, ensuring that each equivalence class in a dataset not only satisfies K-Anonymity in quasi-identifier values but also satisfies a certain diversity in sensitive value distribution. Specifically, if a dataset satisfies L-Diversity, then for any equivalence class, it contains at least L different sensitive values.
[0080] priv L ≡ l, where
[0081] T-Closeness: To prevent adversaries who know the global distribution of sensitive attributes from inferring information, T-Closeness requires limiting the distribution of sensitive values within equivalence classes. Specifically, for an equivalence class and a sensitive attribute, if the distribution of the sensitive attribute within the equivalence class does not differ from its distribution in the entire dataset by more than a threshold t, then the equivalence class satisfies T-Closeness.
[0082] priv T ≡ t, where
[0083] (α, k)-Anonymity: (α, k)-Anonymity introduces additional constraints on equivalence classes in a dataset based on the concept of K-Anonymity. Specifically, it requires that the sensitive attribute values in an equivalence class cannot exceed a certain proportion, represented by a threshold α.
[0084] priv AK ≡ (α, k), where
[0085] Data distinguishability: Data distinguishability refers to the degree to which data records remain distinguishable after privacy protection techniques are applied. A higher distinguishability value indicates that data records remain relatively distinguishable, while a lower value indicates that the differences between records are reduced. For a given record t, if there are s other records in the published dataset that have the same non-sensitive attributes as t, then the information loss for t is recorded as s. For a record l that has been hidden in the dataset, its information loss is equal to the total number of records in the original dataset T. The formula is as follows:
[0086]
[0087] where CDM represents the total amount of information loss, E represents the equivalence class, |E| represents the number of tuples in the equivalence class, n s represents the number of hidden tuples, and |T| represents the total number of tuples in the dataset.
[0088] Data record anonymity rate: The data record anonymity rate is defined as the proportion of records that are hidden or deleted during the privacy protection process. During the implementation of privacy protection, some tuples that do not meet privacy conditions are usually deleted, resulting in information loss. These deleted records are usually referred to as hidden records. The data record anonymity rate quantifies the degree of data removal by calculating the ratio of the number of hidden records to the total number of records in the original dataset, expressed as:
[0089]
[0090] where n s represents the number of hidden records, and |T| represents the total number of tuples in the dataset.
[0091] Average generalization degree: The average generalization degree is the total number of tuples in the dataset divided by the number of equivalence classes, indicating the average number of records in each equivalence class. By calculating the average generalization degree, we can determine how many records with the same quasi-identifier each record is hidden in, thus reflecting the overall anonymization level of the dataset. The total number of records in the dataset is represented as |T|, and the number of equivalence classes is represented as |num E |.
[0092]
[0093] Data loss degree: Data loss degree quantifies the uncertainty introduced in the privacy protection process by calculating the attribute penalty value of each tuple. It measures the degree of data modification required to achieve privacy protection. The calculation involves determining the normalized certainty penalty of each attribute in the tuple and combining the attribute weights to calculate the weighted certainty penalty. Finally, the total weighted certainty penalty serves as an indicator to evaluate the overall uncertainty in the dataset.
[0094]
[0095] where d is the number of attributes in QID (i.e., dimension). T represents equivalence class. is a numerical or string type attribute, with weight w i where ∑w i = 1.
[0096] When A i is a numerical attribute, the corresponding formula calculates the ratio of the range of values in the current equivalence class and the range of values in the entire dataset. For example, if the range of values of the age attribute in the entire dataset is 0 to 100, and in a given equivalence class, all users are summarized to be 30-50 years old, then the NCP(T) of this equivalence class on the age attribute is
[0097] Similarly, when A i is a string type attribute, the formula calculates the ratio of the length of the part hidden due to summarization and the total length of the attribute values. For example, if only the middle four digits of the phone number are hidden, then the NCP(T) of this equivalence class on the phone number attribute is 4 / total length of the phone number.
[0098] Entropy-based average data loss degree: The entropy-based average data loss degree measures the difference between the initial privacy level and the updated privacy level. Specifically, the data loss can be calculated as (initial entropy - current entropy) / initial entropy. The closer this value is to 1, the greater the data loss; the closer it is to 0, the smaller the data loss. The initial entropy represents the entropy of the original data, reflecting the unpredictability of the original data, assuming that each record in the original data is unique. The current entropy, on the other hand, represents the entropy of the processed data, reflecting the distribution of equivalence classes after generalization. The formula is as follows:
[0099]
[0100] where |E i | represents the size of the equivalence class, and |T| represents the size of the dataset. p(E i ) is the proportion of the equivalence class in the dataset, where and
[0101] Unique record fraction: The unique record fraction refers to the proportion of unique records in a dataset. The simplest definition is that if there is only one record in an equivalence class, then that record is considered unique. A unique record means that there is no other record in the dataset with the same quasi-identifier value. When an attacker identifies this equivalence class, they can be 100% certain of the sensitive attribute value of the individual under attack. Therefore, when the unique record fraction is not 0, the dataset has a very high risk of re-identification. A low unique record fraction means that the dataset has been effectively anonymized to some extent, with each record hidden among other similar records, which helps to maintain the privacy protection of the data. However, too low a unique record fraction can also lead to information loss, affecting the usability of the data. Because privacy protection techniques may over-generalize or delete data, reducing the accuracy and analysis value of the data.
[0102]
[0103] where |E| represents the size of the equivalence class, and |T| represents the size of the dataset.
[0104] Distribution leakage: Distribution leakage can be defined as determining the probability distribution of a certain sensitive attribute equivalence class before and after publication, and calculating the Euclidean distance between the two. Distribution leakage can be seen as a measure of the overall divergence of the attribute value distribution from one state to another. For each given equivalence class, measure the leakage of the sensitive attribute distribution in the original dataset and the published dataset, and take the maximum value as the distribution leakage of the data. The formula is as follows:
[0105]
[0106] where A represents the selected sensitive attribute, E represents the selected equivalence class, and DLeakage(E,S) represents the distribution leakage of equivalence class E for sensitive attribute S. E(S i ) represents the distribution of sensitive value S i on the equivalence class, and T(S i ) represents the distribution of sensitive value S i on the entire dataset. The above is only the distribution leakage of a certain equivalence class, and when we calculate the distribution leakage of all equivalence classes, we take the maximum value as the distribution leakage of the dataset.
[0107] Entropy leakage: The main idea of entropy leakage is: for a certain sensitive attribute, determine an equivalence class and the probability distribution of the original data before and after publication, and measure the privacy leakage of individuals in the equivalence class by the difference between the initial entropy of the original distribution and the equivalence class entropy. The formula is as follows:
[0108]
[0109] where S denotes the selected sensitive attribute, E denotes the selected equivalence class, ELeakage(E, S) denotes the entropy leakage of equivalence class E for sensitive attribute S. E(S i ) denotes the distribution of sensitive value S i over the equivalence class. T(S i ) denotes the distribution of sensitive value S i over the entire dataset. The above is the entropy leakage of a certain equivalence class, when we calculate the entropy leakage of all equivalence classes, we take the maximum value as the distribution leakage of the dataset.
[0110] Privacy Gain: Privacy gain is usually used to measure the effectiveness of data anonymization or perturbation techniques in protecting privacy. It quantifies the change in the inferability of sensitive information before and after data publication. Intuitively, privacy gain represents the benefit or improvement in the protection of sensitive attributes after applying privacy protection methods. In this paper, the privacy gain of a specific sensitive attribute value is defined as the sum of the sizes of the equivalence classes containing that sensitive attribute value divided by the number of records containing that sensitive attribute value.
[0111]
[0112] where S i denotes the sensitive attribute value, |S i | denotes the number of records containing the sensitive attribute, |E| is the size of the equivalence class containing the sensitive attribute, PG(S i ) is the privacy gain of sensitive attribute value S i , and the privacy gain of the sensitive attribute is the average of all PG(S i ).
[0113] KL-Divergence: KL-Divergence is a measure used to measure the distance between two frequency distributions. In this paper, it is used to measure the distance between the distribution of sensitive attributes in equivalence classes and the distribution of that sensitive attribute in the entire dataset. The KL-Divergence formula for sensitive attribute value S i is as follows:
[0114]
[0115] where S i is the sensitive attribute value, P(S i ) is the probability distribution of the sensitive attribute value in the equivalence class, and Q(S i ) is the probability distribution of the sensitive attribute value in the dataset. KL-Divergence measures the logarithmic difference between the discrete probability distributions of P and Q. After calculating the KL-Divergence of all equivalence classes, we take the maximum value as the KL-Divergence of the dataset.
[0116] Sensitive Attribute Re-identification Risk: The risk of sensitive attribute re-identification is quantified by calculating the maximum entropy and relative entropy after the application of privacy protection technologies. The maximum entropy value represents the uncertainty of information, and the relative entropy value represents the impact of privacy protection technologies on data security. By assessing the risk of sensitive attribute re-identification, we can better understand the privacy risks after data publication and take appropriate measures to protect them.
[0117] Maximum entropy is considered the maximum amount of information required to determine whether sensitive data is related to identifier data. It represents the maximum uncertainty in the information corresponding to an event that occurs with a certain probability. Maximum entropy H max and equivalence class H E The entropy value is calculated as follows:
[0118]
[0119] Where S is the sensitive attribute, s i For sensitive attribute values, n is s i The quantity, P(s) i ) represents the sensitive attribute value s i The proportion of the class in equivalence class E.
[0120] Through maximum entropy H max and relative entropy D KL Calculate S using (E||max) i SAR re-identification risk i .
[0121]
[0122] After calculating the re-identification risk for each equivalence class, we take the maximum value as the re-identification risk for the entire dataset, i.e.
[0123] Overall re-identification risk: The SAR above only considers the re-identification risk of a single sensitive attribute, while the overall re-identification risk considers all sensitive attributes as a whole and calculates their re-identification risk accordingly. When the dataset contains multiple sensitive attributes, the formulas for TR and SAR remain unchanged, but H... max and H E The definition of H will be different, in which case the calculation of H... max and H E All sensitive attributes must be considered. For example, in billing data, the sensitive attributes are (extra charges, monthly package fee, account balance). Therefore, when calculating TR, (22.45, 29, 124.5) and (17.2, 29, 244) represent different sensitive attribute values. It's worth noting that because all sensitive attributes are considered simultaneously, H... max and HE There can be an increase on both TRs. Thus, there is no clear size relationship to prove which one is larger.
[0124] Quasi-identifier re-identification risk: The difference between quasi-identifier re-identification risk and the above two indicators is that it does not consider the re-identification risk from the perspective of equivalence groups, but considers the re-identification risk of data from the perspective of specific quasi-identifier attributes within the equivalence group. That is, it focuses on a single quasi-identifier attribute and assesses the re-identification risk that an attacker can cause after the data is published. Like TR, QIR considers the risk of all sensitive attributes, so the H max The same. However, QIR no longer calculates the probability of re-identification through equivalence values, but re-identifies private information based on a quasi-identifier attribute. QIR needs to calculate H Q The formula is as follows:
[0125]
[0126] Where Q is a quasi-identifier attribute, P(Q, s i ) is the proportion of sensitive attribute values s i in equivalence class Q divided by the proportion of quasi-identifier attribute values q i .
[0127] The formula for QIR is as follows:
[0128]
[0129] The multi-dimensional evaluation value aggregation calculation unit constructs an index hierarchy using the analytic hierarchy process, determines the relative importance of each secondary index through expert scoring, and then calculates the subjective weight to weight and integrate secondary indexes in the same field;
[0130] First, compare the indexes pairwise to construct a judgment matrix; then solve the matrix using the eigenvalue method to obtain the weight of each index, providing a quantitative basis for evaluation. To ensure the rationality of the weight, the judgment matrix needs to be verified through consistency test to avoid the influence of subjective judgment bias on the accuracy of the result. Specifically as follows:
[0131] 1) Determine the hierarchy structure and input parameters: Clearly define the decision goal as the weight distribution of the secondary indexes under each primary dimension, and the criterion layer as the secondary indexes whose weights are to be calculated. Construct an index system hierarchy diagram to clarify the hierarchical relationship; input the judgment matrix T = [a ij ] ∈ R n×n , where a ij represents the relative importance of index m i relative to index m j ; the matrix dimension n and the random consistency index RI.
[0132] 2) Column normalization judgment matrix: in the judgment matrix T = [a ij ] ∈R n×n , "column" has a clear reference base meaning. Among them, the column index j (j ∈ [1, n]) corresponds to the jth secondary index, and all elements a 1j , a 2j , a 3j ,..., a nj of this column are based on m j to describe the relative importance of other indicators to m j . For example, a ij represents the importance of indicator m i relative to indicator m j , and a jj is always 1, because the importance ratio of the same indicator itself is 100%. For each column j (j ∈ [1, n]) of the judgment matrix, the normalized value of each element in the column is calculated, that is, the ratio of the element to the sum of all elements in the column, to achieve the standardization processing of the column dimension.
[0133] 3) Calculate row normalization sum and weight vector: for the column normalized matrix, calculate the sum of elements of each row i (i ∈ [1, n]), to get the row sum N i ; normalize each row sum N i (i.e. each N i is divided by the sum of all row sums) to get the weight value of each indicator, which constitutes the weight vector V = [v1, v2,..., v n ] T .
[0134] 4) Index aggregation value calculation: after obtaining the weight vector V = [v1, v2,..., v n ] T of each secondary index, it needs to be combined with the original index data to calculate the comprehensive aggregation value of each evaluation object, providing a quantitative basis for the final decision. The specific calculation formula is: A 聚合 = ∑A i *v i , where A 聚合 is the aggregated index value, A i is the secondary index, and v i , i = 1, 2,..., n is the weight corresponding to the secondary index. The visualization of the established privacy index system is shown in Figure 4 .
[0135] 5) Consistency check: calculate the maximum eigenvalue λ max of the judgment matrix, calculate the consistency index CI by the formula CI = (λ max -n) / (n-1), and then obtain the consistency ratio CR according to CR = RI / CI.
[0136] 6) Result output: if CR≤0.1, it is judged that the consistency of the matrix is passed, and the normalized weight vector V and the consistency ratio CR are returned; if CR>0.1, the consistency is not passed, and an adjustment prompt is returned: please re-construct the judgment matrix to improve the consistency.
[0137] The index association mapping module comprises an association relationship mining and modeling unit and a rule processing and matrix construction unit.
[0138] The association relationship mining and modeling unit is based on the index slam-dunk rule library, introduces an association rule mining algorithm, analyzes historical evaluation data and simulation test data, mines potential associations, and supplements and perfects the association relationship network, as follows:
[0139] Firstly, static association analysis is performed, based on expert knowledge and field research, the known logical relationships between indexes are sorted out, and an initial association rule library is constructed.
[0140] 1) Based on the published academic literature and industry standards, the logical relationships between indexes are quantitatively scored (the score range is 1-5 points, and 5 points represent strong association).
[0141] 2) A four-tuple structure is used to represent the association rule, i.e. (index A, association type, index B, weight), wherein the association type includes positive correlation, negative correlation, nonlinear correlation, etc., and the weight is calculated by the analytic hierarchy process (AHP). An example rule expression is: (data record anonymity rate, negative correlation, sensitive attribute re-identification risk, 0.85), wherein the weight 0.85 represents the importance degree of the association relationship in the overall evaluation system.
[0142] Then, dynamic association mining is performed: an association rule mining algorithm is introduced, historical evaluation data and simulation test data are analyzed, potential associations are mined, and the association relationship network is supplemented and perfected.
[0143] 1) Data preprocessing. Collect the original records and derived risk quantification data (sample size not less than 1000) related to identity, consumption and behavior of users in the telecommunication / telecom business in the past three years, and set simulation test data (sample size not less than 500 groups) based on these data, and perform standardization processing on the data, using the Z-score method to calculate: Wherein, μ is the mean, and σ is the standard deviation.
[0144] 2) Definition of numerical item set. The discretized index category is regarded as "item", and the numerical index and discrete index are defined as different items respectively. For example: the discretized item of index A data distinguishability: A1 (≤ 50%), A2 (51%-65%), A3 (66%-80%), A4 (81%-90%), A5 (> 90%). The discretized item of index B (sensitive attribute re-identification risk): B1 (low), B2 (lower), B3 (medium), B4 (higher), B5 (high). Then "item set" refers to the combination of these categories (such as {A5, B1} represents "anonymous rate > 90% and low risk").
[0145] 3) Frequent item set mining using Apriori algorithm, setting the minimum support threshold not less than 0.3, the minimum confidence threshold not less than 0.7, and the maximum item set length not more than 3.
[0146] 4) Extracting the rule "if the value of index A belongs to a certain category, then the value of index B is likely to belong to a certain category" from the frequent item set, and quantifying it with confidence: the confidence of rule A5→B1 = (the number of A5 and B1) / (the total number of A5) = 280 / 300 ≈ 0.93 ≥ 0.7, which means "when the anonymous rate is > 90%, the probability of low risk is 93%"
[0147] 5) Verify the effectiveness of the extracted relationship. Calculate the lift of the extracted relationship to verify the non-accidental nature of the specific rule, and use the chi-square test + Cramer's V algorithm to test and extract the correlation of the specific rule.
[0148] a. Calculate the support. Support (Support) represents the probability of the occurrence of a certain item alone, that is
[0149] Support(A) = the number of A appearing / the total number of experiments
[0150] b. Calculate the confidence. Confidence (Confidence) represents the probability of the occurrence of B at the same time under the condition of the occurrence of A, that is
[0151] Confidence(A→B) = Support(A∩B) / Support(A)
[0152] c. Calculate the lift. The calculation formula of lift is as follows:
[0153] Lift(A→B) = Confidence(A→B) / Support(B)
[0154] Keep the association rules with lift greater than 1.2 to ensure that the rules have practical predictive value.
[0155] d. Verify significance by chi-square test. Set the contingency table of two categorical variables as N = (n ij ), the matrix dimension is r x c (r is the number of categories of row variable, c is the number of categories of column variable). The formula of chi-square test is χ 2 =∑(O-E) 2 / E. Where O represents the actual observed frequency, i.e. the actual data in each cell of the contingency table. E represents the expected frequency, i.e. the "theoretically deserved frequency" of each cell assuming the two variables are independent. ∑ represents the sum of the calculation results of all cells in the contingency table. The degree of freedom df = (r-1) x (c-1) is calculated. In the chi-square distribution critical value table with degree of freedom df, the probability corresponding to the calculated value of chi-square statistic is found to obtain the p value, i.e. p = P(χ 2 ≥ calculated value), which represents the probability of observing the current or more extreme chi-square statistic under the condition that the null hypothesis is true. Only the items with p value less than 0.05 are retained, which are recognized as statistically significant.
[0156] e. Calculate Cramer's V coefficient. Total sample size: Cramer's V coefficient is a standardization of chi-square statistic, which eliminates the influence of sample size N and category number, and the formula is:
[0157]
[0158] where χ 2 is the chi-square statistic calculated above; N is the total sample size (the sum of all elements in the matrix); min(r-1, c-1) is the minimum of (r-1) and (c-1) (the standardization term of matrix dimension)
[0159] 6) Calculate dynamic strength by combining support, confidence and lift:
[0160] Dynamic strength = support x confidence x lift / (1 + support x confidence x lift)
[0161] 7) Generate dynamic rules. Generate dynamic rules in the form of four-tuple (indicator A, association type, indicator B, dynamic strength).
[0162] Rule processing and matrix construction unit integrates static rules and dynamic rules to generate a comprehensive rule set and construct an indicator association matrix;
[0163] If there is only static rule or only dynamic rule between two attributes, they are directly merged. If there are both static rule and dynamic rule between two attributes, the total strength is calculated according to the following formula:
[0164] Comprehensive strength = 0.6 x dynamic strength + 0.4 x static strength
[0165] wherein the dynamic strength is the weight in the dynamic rule set, and the static strength is the weight in the static rule.
[0166] According to the obtained dynamic rule, taking the "privacy index" as the identifier, an index correlation matrix M=R is constructed n×n , n represents the number of privacy indexes, which is 22 in the present application. M i,j represents the influence coefficient of index i on index j, which takes the value range [-1, 1], and the absolute value represents the influence strength, and the sign represents the influence direction. The attribute correlation represented by the attribute correlation graph is shown in Figure 3 .
[0167] The dynamic updating module includes a fluctuation benchmark determination unit, a user demand conversion unit, and an index updating unit.
[0168] The fluctuation benchmark determination unit calculates the fluctuation benchmark of the index in each scene according to the business characteristics and historical data in different scenes such as data collection, storage, and use, for the 22 privacy indexes. Specifically, for the indexes with historical data, the mean and standard deviation are calculated by statistical methods, and the standard deviation is taken as the fluctuation benchmark; for new indexes or scenes lacking historical data, the initial baseline is set through expert discussion combined with industry best practices and business objectives, and is dynamically calibrated according to actual data feedback in system operation;
[0169] Taking the "entropy leakage" index as an example, in the data transmission scene, the index data of past multiple data transmission tasks in this scene is collected, the mean and standard deviation σ are calculated by statistical methods, and the fluctuation benchmark is set as σ current = σ.
[0170] When the user receives the demand for enhancing a certain privacy index, the user demand conversion unit converts the user's description of index enhancement into a change amplitude in the unit of fluctuation benchmark; wherein for qualitative demand, the conversion is performed according to the preset industry conventional adjustment amplitude conversion rule, and for quantitative demand, the percentage is converted into a multiple of the fluctuation benchmark according to the historical data distribution characteristics of the index, and finally the change value in the unit of fluctuation benchmark is output;
[0171] For example, when the user proposes the demand for enhancing the data record anonymity rate, the system converts the user's description of index enhancement (such as "significantly improve the data record anonymity rate" "increase the data record anonymity rate by 10%") into a change amplitude in the unit of σ current :
[0172] 1) If the user description is a qualitative demand (such as "significantly improve" "significantly enhance"), the conversion rule is defined in combination with the conventional adjustment amplitude of the industry for the index (for example: slight change ±10% σcurrent Moderate variation is ±20% σ current The variation is significant, ranging from ±30% σ. current (The specific threshold is preset by experts based on business characteristics);
[0173] 2) If the user's description is a quantitative requirement (e.g., "improve compliance by 10%)", convert the percentage to σ by utilizing the distribution characteristics of historical data for the indicator. current The multiple (e.g., a 10% improvement in compliance corresponds to an actual numerical change of Δx, through...) The corresponding σ is calculated. current Multiples, such as Δx = 3%σ current This translates to "increase σ by 3%" current ”).
[0174] 3) The final output is in σ current The change in units is Δ = k × σ current (k is the coefficient after conversion, such as "improving compliance by 30% σ") current That is, k = +0.3).
[0175] The indicator update unit completes the change value Δ on the user-specified privacy indicator (denoted as indicator i). i =k×σ current After transformation, based on the generated indicator correlation matrix M, dynamic linkage updates of other indicators (denoted as indicator j, j≠i) are achieved. Specifically, the influence coefficient of the specified indicator on other indicators is extracted from the correlation matrix, the linkage change value of other indicators caused by the update of the specified indicator is calculated, and the baseline value of other indicators is updated based on the linkage change value. The specific process is as follows:
[0176] 1) Calculation of correlation influence transmission: Extract the influence coefficient M of indicator i on indicator j from the correlation matrix M. i,j (Value range [-1, 1], positive numbers indicate positive impact, negative numbers indicate negative impact, and absolute value indicates the intensity of impact). The linkage change Δ caused by the update of index j to index i. j The calculation formula is: Δ j =Δ i ×M i,j (For example: if indicator i is "data record anonymity rate", its change value Δ) i =0.3×σ current_i And M in the matrix i,j =0.5 (j is the "average generalization"), then Δ j =0.3×σ current_i ×0.5=0.15×σ current_j This means that the "average generalization" is positively boosted by 0.15 × σ. current_j .
[0177] 2) Privacy Index Update: Let the current baseline value of index j be k j , the updated baseline value k j ' be: k j ' = k j + Δ j . For example, if the current baseline value of "average generalization degree" is k j = 80%, Δ j = 0.15σ current_j , and σ current_j = 0.05 = 5%, then Δ j = 0.15 x 5% = 0.75%, and the updated baseline value k j ' = 80% + 0.75% = 80.75% = 0.8075. An example of privacy index update is shown in Figure 5 .
[0178] Rule Verification and Optimization Unit: After completing the dynamic linkage update of the privacy index, it is necessary to verify and continuously optimize the rules to ensure that the updated index system and associated rules meet the actual business needs, industry compliance standards, and system operation stability. The specific process is as follows:
[0179] 1) Rule Validity Verification: Select historical business scenario data (such as privacy index records in the past 3 months), and input the user demand converted Δ i and the associated matrix M into the system to simulate the index linkage update process. Compare the deviation between the update result and the actual business performance (such as whether the actual change of "average generalization degree" is consistent with the simulated Δ j after the "data record anonymization rate" is enhanced). Calculate the deviation rate:
[0180] Deviation rate = (|actual change value - simulated change value| / simulated change value) x 100%
[0181] If the deviation rate exceeds the preset threshold (such as 15%), it is marked as not passing the verification, and the influence coefficient of the associated matrix M or the conversion logic of Δ i needs to be checked back.
[0182] 2) Associated Matrix Iteration: Based on the deviation data in the verification stage, regularly update (such as every quarter) the influence coefficient M i,j of the associated matrix M: if the actual influence of index i on j is greater than the simulated value, increase M i,j by the deviation rate (such as a deviation rate of +20%, then M i,j is updated to the original coefficient x 1.2); if the actual influence is less than the simulated value, decrease M i,j (by a deviation rate of -15%, then M i,jUpdate to original coefficient x 0.85). For new business scenarios (such as the addition of a "cross-border data transmission" scenario), supplement the index association relationship under this scenario, and mark the initial M i,j and verified iteratively through subsequent data.
[0183] It is to be understood that the application is described by way of example only, and that modifications or alterations can be made to the features and embodiments described without departing from the spirit or scope of the application as set out in the claims. In addition, modifications can be made to the features and embodiments described to adapt them to specific circumstances and materials without departing from the spirit and scope of the application. The application is therefore not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims are intended to be within the scope of the application.
Claims
1. A privacy indicator system based on attribute interlinkages, characterized in that, The application comprises a multi-dimensional index calculation module, an index correlation mapping module, a dynamic updating module, and a rule verification and optimization module. The multi-dimensional index calculation module comprises an index system disassembly and algorithm adaptation unit and a multi-dimensional evaluation value aggregation calculation unit. The index system disassembly and algorithm adaptation unit performs hierarchical analysis on the privacy protection effect evaluation index system, and adapts the calculation algorithm according to different types of indexes. The multi-dimensional evaluation value aggregation calculation unit uses the analytic hierarchy process to construct an index hierarchy structure, first compares indexes two by two to construct a judgment matrix, and then solves the matrix using the eigenvalue method to obtain the weight of each index. The index correlation mapping module comprises an association relationship mining and modeling unit and a rule processing and matrix construction unit. The association relationship mining and modeling unit constructs an initial association rule base as a static rule based on the known logical relationship between indexes, and then analyzes historical evaluation data and simulation test data to mine potential associations as dynamic rules to supplement and improve the association rule base. The rule processing and matrix construction unit generates a comprehensive rule set by comprehensively processing static rules and dynamic rules, and constructs an index correlation matrix according to the obtained dynamic rules. The dynamic updating module comprises a fluctuation benchmark determination unit, a user demand conversion unit, and an index updating unit. The fluctuation benchmark determination unit calculates the fluctuation benchmark of each index in different scenarios according to the business characteristics and historical data in different scenarios. The user demand conversion unit converts the user's description of the enhanced index into a change amplitude in units of fluctuation benchmarks when receiving the user's demand to enhance a certain privacy index. The index updating unit extracts the influence coefficient of the specified index on other indexes from the index correlation matrix, calculates the linkage change value of other indexes caused by the update of the specified index, and updates the benchmark values of other indexes based on the linkage change value. The rule verification and optimization module verifies the effectiveness of the updated index system and iteratively updates the index correlation matrix.
2. The privacy indicator system based on attribute interlinkages according to claim 1, wherein, The index evaluation system covers four first-level indexes: data characteristics, compliance, availability, and security, and 22 second-level indexes. The calculation algorithm for different types of indexes is as follows: the basic statistical index traverses the data set to calculate the characteristic value; the privacy model index substitutes the data into the privacy protection parameter under the model constraint; and the risk assessment index calculates the privacy leakage risk probability value by constructing an attack model.
3. The privacy indicator system based on attribute interlinkages according to claim 2, wherein, The multi-dimensional evaluation value aggregation calculation unit is as follows: First, determine the hierarchy and input parameters: clear decision goal is the weight allocation of the second-level indicators under each first-level dimension, the criterion layer is the second-level indicators whose weights are to be calculated, and the index system hierarchy diagram is constructed to clarify the hierarchical relationship; input the judgment matrix T=[a ij ]∈R n×n , where a ij represents the relative importance of indicator m i relative to indicator m j ; the matrix dimension n and the random consistency index RI; The normalized value of each element in the judgment matrix array is calculated, that is, the ratio of the element to the sum of all elements in the column, to realize the standardization processing of the matrix array dimension; the sum of the elements of each row is calculated after the column normalization, to obtain the row sum N i ; the row sum N i is normalized to obtain the weight value of each index, to form a weight vector V=[v1,v2,...,v n ] T , wherein v i ,i=1,2,...,n is the weight corresponding to the secondary index. The weight vector V = [v1, v2,..., vn] of each secondary index is obtained, and the weight vector V = [v1, v2,..., vn] of each secondary index is obtained. n ] T Then, the comprehensive aggregation value of each evaluation object is calculated by combining the original index data, and the specific calculation formula is as follows: A 聚合 =∑A i *v i wherein A 聚合 is the index value after polymerization, A i is a secondary index; Finally, consistency test is performed to complete the construction of the judgment matrix.
4. The privacy indicator system based on attribute interlinkages according to claim 3, wherein, The association relationship mining and modeling unit is as follows: The association relationship includes static rules and dynamic rules. First, static rules are constructed, the logical relationship between indexes is quantitatively scored based on published academic literature and industry standards, and then a four-tuple structure is used to represent the association rule, i.e., (index A, association type, index B, weight), wherein the association type includes positive correlation, negative correlation, and nonlinear correlation, and the weight is calculated by the analytic hierarchy process. Then the dynamic rules are constructed; first, the historical evaluation data are acquired and the simulation test data are set, and the data are standardized; the numerical indicators and discrete indicators are defined as different items respectively, the Apriori algorithm is used for frequent item set mining, the minimum support threshold, the minimum confidence threshold and the maximum item set length are set, wherein the support represents the probability of the occurrence of a single item, and the confidence represents the probability of the occurrence of B at the same time under the condition of the occurrence of A; the rules of "if the value of indicator A belongs to a certain category, then the value of indicator B is likely to belong to a certain category" are extracted from the frequent item set, and the confidence is quantified; the lift of the extracted relationship is calculated to verify the non-accidentalness of the specific rules, the chi-square test+Cramer's V algorithm is used to test and extract the correlation of the specific rules; the lift takes indicator B as an example, and the calculation formula is as follows: Lift(A→B)=Confidence(A→B) / Support(B) Wherein, Confidence(A→B) represents the confidence of indicator B relative to indicator A, and Support(B) represents the support of indicator B; The dynamic strength is calculated by comprehensively considering the support, the confidence and the lift: Dynamic strength=support×confidence×lift / (1+support×confidence×lift) The dynamic rules are generated in the form of four-tuple (indicator A, association type, indicator B, dynamic strength).
5. The privacy indicator system based on attribute interlinkages according to claim 4, wherein, The rule processing and matrix construction unit is specifically as follows: If there is only a static rule or only a dynamic rule between two attributes, they are directly combined; if there are both static rules and dynamic rules between two attributes, the total strength is calculated according to the following formula: Comprehensive strength=0.6×dynamic strength+0.4×static strength Wherein, the dynamic strength is the weight in the dynamic rule set, and the static strength is the weight in the static rule; According to the obtained dynamic rules, taking the "privacy index" as the identifier, an index correlation matrix M=R is constructed n×n , n represents the number of privacy indexes, M i,j represents the influence coefficient of index i on index j, and its value range is [-1, 1], the absolute value represents the influence strength, and the sign represents the influence direction.
6. The privacy indicator system based on attribute interlinkages according to claim 5, wherein, The fluctuation benchmark determination unit is specifically as follows: for the indicators with historical data, the mean and the standard deviation are calculated by statistical methods, and the standard deviation is taken as the fluctuation benchmark; for new indicators or scenes lacking historical data, the initial baseline is set through expert discussion combined with industry best practices and business objectives, and the baseline is dynamically calibrated according to the actual data feedback in the system operation; In the user demand conversion unit, the demand is divided into qualitative demand and quantitative demand; for the qualitative demand, the conversion rule is converted according to the preset industry conventional adjustment range, and for the quantitative demand, the percentage is converted into the multiple of the fluctuation benchmark according to the distribution characteristics of the historical data of the indicators, and finally the change value in the unit of the fluctuation benchmark is output.
7. The privacy indicator system based on attribute interlinkages according to claim 6, wherein, The index updating unit specifically updates the indicators through the following steps: An influence coefficient M of index i on index j is extracted from the correlation matrix M i,j , a linkage change value Δ of index j due to index i update j The calculation formula is: Δ j = Δ i × M i,j ; Let the current reference value of index j be k j , the updated reference value k j ′ is: k j ′ = k j + Δ j .
8. The attribute interlinking based privacy indicator system of claim 7, wherein, The rule verification and optimization unit is specifically as follows: Rule validity verification: Select historical business scenario data, convert user demand Δ i and correlation matrix M into the system, simulate the index linkage update process, compare the deviation of the update result and the actual business performance, and calculate the deviation rate: Deviation rate=(|actual change value-simulation change value| / simulation change value)×100% If the deviation rate exceeds the preset threshold, it is marked as verification failure, and the influence coefficient or Δ of the correlation matrix M needs to be checked back i conversion logic; Correlation matrix iteration: based on the deviation data in the verification stage, the influence coefficient M of the correlation matrix M is updated regularly i,j : if the actual influence of index i on j is greater than the simulation value, M is adjusted upwards by the deviation ratio i,j ; if the actual influence is less than the simulation value, M is adjusted downwards i,j ; for new business scenarios, supplement the index correlation relationship under the scene, mark the initial M i,j and iterate through subsequent data verification.
9. The attribute interlinking based privacy indicator system of claim 8, wherein, The secondary indicators include: Quasi-identifier dimension: the number of quasi-identifier attributes in the data set; Sensitive attribute: information that needs to be protected; Sensitive attribute dimension: the number of sensitive attributes; Minimum equivalence class size: the smallest number of records in all equivalence classes in the data set; Inherent privacy: the length of the uniform distribution interval corresponding to the uncertainty of the probability distribution of the data set; K-Anonymity: Each equivalence class contains at least K records; L-Diversity: If a dataset satisfies L-Diversity, then for any equivalence class, it contains at least L different sensitive values; T-Closeness: Limit the distribution of sensitive values within the equivalence class; (α, k)-Anonymity: Requires that the proportion of sensitive attribute values in an equivalence class does not exceed a certain proportion; Data distinguishability: The degree to which data records can still be distinguished after applying privacy protection techniques; Data record anonymity rate: The proportion of records that are hidden or deleted during the privacy protection process; Average generalization degree: The total number of tuples in the dataset divided by the number of equivalence classes, representing the average number of records in each equivalence class; Data loss degree: Calculate the attribute penalty value of each tuple to quantify the uncertainty introduced in the privacy protection process Average data loss degree based on entropy: Calculate the difference between the initial privacy level and the updated privacy level to measure the proportion of unique records in the dataset; Distribution leakage: Determine the equivalence class of a certain sensitive attribute and the probability distribution of the original data before and after publication, and calculate the Euclidean distance between them; Entropy leakage: For a certain sensitive attribute, determine an equivalence class and the probability distribution of the original data before and after publication, and measure the privacy leakage of individuals in the equivalence class by the difference between the initial entropy of the original distribution and the entropy of the equivalence class Privacy gain: Measures the effectiveness of data anonymization or perturbation techniques in protecting privacy KL-Divergence: Measures the distance between the distribution of a sensitive attribute in an equivalence class and the distribution of that sensitive attribute in the entire dataset; Sensitive attribute re-identification risk: Calculate the maximum entropy and relative entropy after applying privacy protection techniques to quantify the re-identification risk of sensitive attributes; Overall re-identification risk: Consider all sensitive attributes as a whole and calculate their re-identification risk; Quasi-identifier re-identification risk: Consider the re-identification risk of data from the quasi-identifier attributes within the equivalence group.
Citation Information
Cited By
Data processing method and system for anonymous person attribute abundance optimization
CN121211510A
A data processing method and system for anonymous person attribute abundance optimization
CN121211510B
Method, device and system for automatically identifying sensitive data to carry out de-identification processing
CN121980617A