Information leakage risk assessment method and system in multi-trust-domain access data aggregation based on granular association
Through granularization and risk assessment of data in multiple trust domains based on particle association, the risk of over-authorization access to sensitive information during data aggregation is solved, and the risk of over-authorization of information leakage is effectively reduced and the formulation of data access strategies is achieved.
Patent Information
- Application Number
- CN202510053439.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-05-06
AI Technical Summary
During the data aggregation process of multi-trust domains, there is a risk of overreach of sensitive information, resulting in an increase in the possibility of information leakage.
The whole-domain data is granulated by a granular method based on particle association, the association rules between data attributes are extracted, and the sensitivity level of the target data object is predicted through the correlation attribute sensitivity level fuzzy set probability measurement, and whether there is a risk of override access by user cross-domain access.
It effectively reduces the risk of information leakage during data aggregation, helps formulate data access policies, controls the analysis of related data, prevents information leakage, and ensures the security of multiple information systems.
Smart Images

Figure CN119939662A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data security access, and in particular to a method and system for assessing information leakage risk in multi-trust domain access data aggregation based on granular association. Background Art
[0002] With the continuous emergence of new technologies such as big data, cloud computing, and blockchain, the trend of digital economy has been set off. The degree of intelligence and digitization in various industries has been increasing, which has accelerated the demand for data collection, aggregation, and sharing on the Internet. Independent trust domains have been established between enterprises. However, good data interaction and sharing can break the information barriers between trust domains and tap the deeper value of data. However, there may be problems such as unequal user permissions and implicit associations between data between multiple information systems, which may cause information leakage, and the confidentiality of sensitive data has been greatly challenged. Specifically, different organizations have differences in security requirements and data management levels, resulting in an imbalance in data collection and processing permissions. Some departments have extensive data collection permissions and high data processing permissions, while other departments may only need limited data access permissions, thus forming a difference between high-security domains and low-security domains. This difference may cause the high-security domain to open data that it does not have access to to the low-security domain, that is, sensitive data exposure. It is of great research value to analyze global data, identify implicit associations, infer the possibility of information leakage caused by high-sensitivity information from the aggregation of associated data, and evaluate the access strategy of users across information systems. These operations are essential to maintaining the normal operation of multiple information systems. .
[0003] The derivation of highly sensitive data information caused by data aggregation belongs to the logical reasoning problem of data. This is a complex weak formalization problem. Logical reasoning occurs when a user leaks some protected data information without direct access, where direct access means that the data allows a certain user to access or query. The aggregated information obtained by the user enables the information system to infer specific privacy and confidential information without special protection measures. There are two protection strategies to protect the security of data information when users access the information system, namely responsive protection and preventive protection. The responsive protection method mainly blocks the access connection or sends an alarm to the administrator when the information system detects abnormal data access patterns or some sensitive data in the database is illegally obtained. The preventive protection method focuses on taking measures to reduce risks before the information leakage incident occurs. From the perspective of computing consumption, the responsive protection method is more expensive and has higher requirements on the system's storage and computing capabilities, while the preventive protection method is more convenient for users to use. Through the rules and protection mechanisms set in advance, complex real-time analysis and processing can be performed before the security incident occurs, and there will be no problem of excessive resource consumption.
[0004] Due to the explosive growth of data volume, we have entered the era of big data. Data from different sources may contain information related to the same subject, which means that there is more data to be mined and utilized, which further increases the risk of data leakage. In addition, based on the similarity of data, classification prediction models and cluster analysis techniques are used to discover the potential characteristics and patterns in the data through the statistical characteristics of a large amount of data, and then infer protected information. Data sharing and third-party cooperation between enterprises or organizations have led to a lack of strict privacy protection measures in the data sharing process, so that directly accessible data is used to infer protected information beyond the scope of sharing, thereby exposing users' personal privacy information.
[0005] The emergence of artificial intelligence has enabled association mining to reach a higher level of data processing, while also posing a greater threat to data security. Some users can use the neural network architecture of deep learning models to extract data features and aggregate them with related data to infer protected information; they can also integrate multi-source data to build knowledge graphs and perform reasoning, comprehensively analyze medical, social, consumer and other data, and mine deeper protected information. Using artificial intelligence technology to mine associated data to discover the problem of deriving highly sensitive information when accessing data aggregation has played an important role in mining different types of data features and processing multi-dimensional data features, but there are still many defects. In practical applications, it is difficult for users to determine whether the features mined by the model are reasonable, and it is difficult to effectively evaluate and review the decision-making basis of the model. In addition, the training and updating of artificial intelligence models requires a lot of computing resources and time. With the dynamic update of data, the model needs to be retrained, resulting in high computing costs and limiting the model's ability to adapt to dynamic changes in data. The integration of granular computing into the problem of sensitive information leakage caused by logical reasoning of data information can achieve a higher level of dynamic data analysis and reasoning, and can deeply analyze the inherent characteristics and associations of data from a microscopic granularity level. For example, in existing particle analysis, data particle sets are formed based on data attributes, similar particle clouds are established using clustering, and the possibility of information leakage is deduced based on the possibility measure of attribute fuzzy sets and the contribution of particles to sensitive particle clouds. The risk of information leakage is reduced by analyzing similar data, but the related data is not analyzed, so there is still a risk of unauthorized access when data is aggregated. Summary of the invention
[0006] To this end, the present invention provides a method and system for assessing the risk of information leakage in multi-trust domain access data aggregation based on granular association, which solves the problem of the risk of unauthorized access to sensitive information when aggregating big data.
[0007] According to the design scheme provided by the present invention, on the one hand, a method for assessing information leakage risk in data aggregation in multiple trust domains accessed based on granular association is provided, comprising:
[0008] Granulate the global data according to the logical deduction relationship between the global data in the target domain, extract the association rules between the granulated data attributes and obtain the target data objects that meet the association condition in the global data, wherein the global data in the target domain includes several trust domains, and each trust domain stores data of a sensitivity level corresponding to the user access rights;
[0009] The sensitivity level of the target data object is predicted by the possibility measure of the fuzzy set of associated attribute sensitivity levels, and whether there is a risk of unauthorized access when the user accesses the target object across domains is determined based on the predicted sensitivity level. The fuzzy set of associated attribute sensitivity levels is used to describe the data association particles corresponding to the specified sensitivity level in the global data attribute space, and the possibility measure of the fuzzy set of associated attribute sensitivity levels is calculated using the fuzzy set membership function of the fuzzy set of associated attribute sensitivity levels and the attribute-related possibility distribution function.
[0010] As the information leakage risk assessment method in the multi-trust domain access data aggregation based on granular association of the present invention, further, the global data is granulated according to the logical deduction relationship between the global data of the target domain, including:
[0011] Preprocessing the data in the global data, and initially granulating the global data according to the conditional attributes and decision attributes of the data to obtain an attribute granule set and a decision granule set, wherein the attribute granule set includes a plurality of attribute granules, and the decision granule set includes a plurality of decision granules;
[0012] According to the attribute particle set and the decision particle set, the upper and lower approximate sets are obtained, and each particle is divided by using the upper and lower approximate sets to obtain the positive domain and boundary domain of the particle, so as to construct the particle feature matrix and eigenvalue matrix by using the positive domain and boundary domain of the particle;
[0013] Based on the particle feature matrix and the eigenvalue matrix and according to the minimum identification attribute set, the attribute importance of the attribute in the distribution simplification is obtained, and the simplification of the conditional attribute set is calculated according to the attribute importance. Based on the simplification, the association relationship between attribute particles and between attribute particles and decision particles is obtained, and the association rules with no less than the minimum support and minimum confidence are recorded into the association rule set according to the association relationship.
[0014] As the information leakage risk assessment method in the multi-trust domain access data aggregation based on granular association of the present invention, further, obtaining the target data object that meets the association condition in the global domain data includes:
[0015] The weight of each association rule is calculated according to the information content of the attribute particle and the quantitative value of the highest sensitivity level of the attribute particle clock attribute, and the association degree between the data objects in the global data is obtained by using the association rule weight and the confidence between the attribute particle association rules;
[0016] The association degree between data objects is set to be greater than or equal to a minimum association degree threshold value to satisfy the association degree condition, the target data object in the global data is obtained according to the association degree condition, and the target data object is recorded in the target data object set.
[0017] As the information leakage risk assessment method in the multi-trust domain access data aggregation based on granular association of the present invention, further, obtaining the target data object that meets the association condition in the global domain data also includes:
[0018] According to the attribute set contained in the conditional attribute set simplification, the newly added data objects in the global space are screened, and the new data objects with the same attribute values as those in the conditional attribute set simplification are taken as existing attribute granules. The new data objects with different attribute values from those in the conditional attribute set simplification are granulated and the granule feature matrix and eigenvalue matrix are updated, so as to use the updated granule feature matrix and eigenvalue matrix to update the association rules of the target data objects.
[0019] As the information leakage risk assessment method in the multi-trust domain access data aggregation based on granular association of the present invention, further, the sensitivity level in the target data object is predicted by the fuzzy set possibility measurement of the associated attribute sensitivity level, including:
[0020] Determine the maximum sensitivity level of the target data object;
[0021] The mutual information between target data objects is obtained according to the information entropy of the target data objects, and the weight of the associated attribute set corresponding to the target data object is calculated using the mutual information, so as to obtain the possibility of the target data object aggregation to derive a sensitivity level higher than the maximum sensitivity level by using the possibility measure of the associated attribute set on the fuzzy set higher than the maximum sensitivity level and the weight of the associated attribute set;
[0022] If the possibility is greater than the information possibility threshold, it is determined that there is information to derive data of a higher sensitivity level when the target data object is aggregated, so as to determine that there is a risk of unauthorized access when the user accesses the target object across domains.
[0023] On the other hand, the present invention also provides a system for assessing information leakage risk in data aggregation based on granular association and multiple trust domains, comprising: a data granulation module and a risk assessment module, wherein:
[0024] A data granulation module is used to granulate the global data according to the logical deduction relationship between the global data in the target domain, extract the association rules between the granulated data attributes and obtain the target data objects that meet the association conditions in the global data, wherein the global data in the target domain includes several trust domains, and each trust domain stores data of a sensitivity level corresponding to the user's access rights;
[0025] The risk assessment module is used to predict the sensitivity level in the target data object through the fuzzy set possibility measure of the associated attribute sensitivity level, and judge whether there is a risk of unauthorized access when the user accesses the target object across domains based on the predicted sensitivity level. The fuzzy set of the associated attribute sensitivity level is used to describe the data association particles corresponding to the specified sensitivity level in the global data attribute space, and the fuzzy set possibility measure of the associated attribute sensitivity level is calculated using the fuzzy set membership function of the associated attribute sensitivity level and the attribute related possibility distribution function.
[0026] Beneficial effects of the present invention:
[0027] The present invention analyzes the correlation between the global data in the target field, mines out highly correlated data objects based on the dependency relationship of data attributes, and then deduces the possibility of deriving highly sensitive information by data aggregation when users access multiple information systems based on the sensitivity level fuzzy set possibility measure of the associated attributes of the data objects. This helps to formulate data access strategies for users, control the analysis of associated data, and reduce the risk of information leakage. It is of great significance for controlling the restricted access of users of multiple information systems to associated data, formulating access strategies for cross-domain users, and preventing information leakage. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic diagram of the information leakage risk assessment process in the multi-trust domain access data aggregation based on granular association in the embodiment;
[0029] Figure 2 This is an illustration of the data aggregation problem in the embodiment;
[0030] Figure 3 It is a schematic diagram of the structure of the particle containing the relationship attribute in the embodiment;
[0031] Figure 4 It is a schematic diagram of the structure of the derived relationship attribute particles in the embodiment;
[0032] Figure 5 The following is a schematic diagram of the algorithm execution efficiency when the data scale changes in the embodiment;
[0033] Figure 6 It is a schematic diagram of the algorithm execution efficiency when the threshold value changes in the embodiment;
[0034] Figure 7 This is a schematic diagram of the algorithm accuracy in the embodiment;
[0035] Figure 8 Schematic diagram of the algorithm error rate in the embodiment. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solutions and advantages of the present invention clearer and more understandable, the present invention is further described in detail below in conjunction with the accompanying drawings and technical solutions.
[0037] In order to solve the problem of sensitive information leakage caused by big data aggregation, the embodiment of the present invention, see Figure 1 As shown, a method for assessing information leakage risk in multi-trust domain access data aggregation based on granular association is provided, comprising:
[0038] S101. Granulate the global data according to the logical deduction relationship between the global data of the target domain, extract the association rules between the granular data attributes and obtain the target data objects that meet the association conditions in the global data, wherein the global data of the target domain includes several trust domains, and each trust domain stores data of a sensitivity level corresponding to the user's access rights.
[0039] like Figure 2 As shown in the figure, each trust domain stores data with different sensitivity levels: P represents public data, S represents secret data, and C represents confidential data. Their sensitivity levels gradually increase. The permissions of users in the domain are different. Assume that users between trust domain A and trust domain B can access information in each other's domain. Data with sensitivity levels of P and S in domain A After merging together, data with a sensitivity level of C can be derived Data with sensitivity level P in trust domain B When aggregated, data with information sensitivity level S can be derived This means that these data are strongly correlated. If the data are aggregated together, it is possible to deduce data with a sensitivity level that exceeds the user's access rights. They should not be allowed to be accessed together. When a user accesses data in a domain space, it is possible to deduce data that is higher than the user's rights, which can cause information leakage and security issues. Therefore, the possibility of aggregating global data to derive information with a higher sensitivity level should be deduced to adjust the user's access to data within the authority. Prevent users from making unauthorized access when accessing data and reduce the risk of sensitive information leakage.
[0040] A large amount of data in various forms is stored in the multi-information system. The logical deduction relationship between the data is mainly manifested as follows:
[0041] i.data i →data j ;
[0042] ii.
[0043] iii.
[0044] iv. wait.
[0045] In view of the logical deduction relationship between the above data, in the embodiment of this case, the data is granulated, and then the implicit association relationship between the attributes of the granulated data is analyzed.
[0046] Among them, the global data is granulated according to the logical deduction relationship between the global data in the target field, which can be designed to include:
[0047] Preprocessing the data in the global data, and initially granulating the global data according to the conditional attributes and decision attributes of the data to obtain an attribute granule set and a decision granule set, wherein the attribute granule set includes a plurality of attribute granules, and the decision granule set includes a plurality of decision granules;
[0048] According to the attribute particle set and the decision particle set, the upper and lower approximate sets are obtained, and each particle is divided by using the upper and lower approximate sets to obtain the positive domain and boundary domain of the particle, so as to construct the particle feature matrix and eigenvalue matrix by using the positive domain and boundary domain of the particle;
[0049] Based on the particle feature matrix and the eigenvalue matrix and according to the minimum identification attribute set, the attribute importance of the attribute in the distribution simplification is obtained, and the simplification of the conditional attribute set is calculated according to the attribute importance. Based on the simplification, the association relationship between attribute particles and between attribute particles and decision particles is obtained, and the association rules with no less than the minimum support and minimum confidence are recorded into the association rule set according to the association relationship.
[0050] Let U be the global space of objects in a certain domain, U = {x1,x2,...,x |U|}, AT is the attribute space, AT={a1,a2,...,a |AT|}, the attribute set AT consists of two parts: the condition attribute set C and the decision attribute set D, and C∪D=AT,
[0051] Let a∈C, vc be the value of attribute a on an object in C, m() be the meaning set function on C, (a, vc) be the atomic formula on C, denoted by a vc , a vc The combination of is a well-formulated formula on C, denoted as but It is called a condition granule cg (condition granules) on U. That is, all the satisfying The decision attribute D divides the domain U into U / IND(D)={D1,D2,...,D n}, D j (1≤j≤n) is a decision granule dg (decision granules) of decision attribute D on U. Decision granule is a special attribute granule, that is, By attribute particle cg i (cg i ∈U / C1) and attribute particle cg j (cg j ∈U / C2) is the association relationship formed by cg i →cg j, the association relationship formed by the attribute particle cg (cg∈U / C) and the decision particle dg (dg∈U / D) is cg→dg.
[0052] make cg i =(α i ,m(α i )), then cgs={cg i ,L,∠} is the attribute granular structure. Among them, L is the level of granular structure, and ∠ is the partial order relationship between granules in granular layers.
[0053] The attribute granule structure is a hierarchical structure. The partial order relations between granule layers include two types: one is the attribute inclusion relation, and the other is the attribute derivation relation. There are also two types of relations between attribute granules in the same layer: one is the similarity relation, and the other is the mutual independence relation. There is no similarity relation between mutually independent attribute granules.
[0054] like Figure 3 As shown in the figure, it is a granular structure of inclusion relationship, which is usually a hierarchical structure composed of single-attribute granules and multi-attribute granules. If the granules in the same layer are similar, that is, cg α ,cg β Similar, the upper layer attribute particles cg α||β is the conjunction of the lower attribute particles, i.e. cg α||β =cg α ∧cg β =(α∧β,m(α)∧m(β)), indicating the common property characteristics and significance of similar particles. If the particles in the same layer are independent particles, that is, The upper attribute particle is the disjunction of the lower data particle, that is, cg α||β =cg α ∨cg β =(α∨β,m(α)∨m(β)), which represents the fusion characteristics and significance of independent attribute particles.
[0055] like Figure 4 As shown, it is the granular structure of the derivation relationship, which is a hierarchical structure built on the basis of the granular derivation relationship. Figure 4 In (a), if That is cg α Derived from the dependence λ1, cg γ ,cg β Derived from the dependence λ2 γ ,So It represents the characteristics and significance of deriving a single attribute particle by aggregating multiple attribute particles. Figure 4 In (b), there is That is cg α Derived cg with dependence λ1 and λ2 respectively γ ,cg ρ, cg β Derived cg with dependency λ3 and λ4 respectively γ ,cg ρ ,So Right now same, Right now Then it must exist, Figure 4 (b) in the figure shows the characteristics and significance of dispersed attribute particles derived by aggregation of multiple attribute particles.
[0056] Let the item of the attribute granule rule set be <cg i →cg j ,sup,conf>, where sup and conf are attribute granule association rules cg i →cg j The support and confidence of can be expressed as follows:
[0057]
[0058] Here, |·| represents the basis of the set.
[0059] In order to improve the efficiency of discovering the association relationship between attribute granules, the embodiment of this case proposes a data association relationship mining algorithm based on granule association, and uses the attribute simplification method to mine the association relationship between attribute granules and between attribute granules and decision granules.
[0060] Let bD be the data set under the same trust domain, Among them, bd i ={x1,x2,...,x n}, S={s1,s2...,s k} is a sensitive set, set bd i is a data list, each list stores a set of data information and contains multiple data objects. Let x i The attribute set is ω ic = <a i1 ,a i2 ,...,a ik ,d i >,d i For x i The decision attribute of ω, the attribute value corresponding to the attribute in ω is v i =<v i1 ,v i2 ,...,v ik ,v id >, data concept object xg i Expressed as
[0061] U / C={E1,E2,...,El},U / D={D1,D2,...,D k}. E i The eigenvector of For E i The eigenvalue vector of is the feature index, obj i is a collection of data objects xg, obj i ={xg k |x k ∈E i}.like That is, the positive domain of the equivalence relation partition, then reg i =P. If That is, the boundary domain divided by the equivalence relation, then reg i =B. The particle feature matrix and eigenvalue matrix can be expressed as follows:
[0062]
[0063] Setting Att min is the minimum identifiable attribute set, Att min ={Att0,Att1,...,Att r}, where Meet Att i ∈D(E i ,E j ), and for Meet Att i ∈D(E i ,E j ). The elements in the minimum discriminative attribute set are different from each other, and there is no inclusion relationship between them. Let Important i Represents attribute a i The attribute importance in the distribution reduction, the attribute importance matrix can be expressed as:
[0064]
[0065] Among them, AT represents the attribute vector of attribute importance, and IM represents the importance vector.
[0066] Data attribute association relationship mining based on granular association is shown in Algorithm 1:
[0067]
[0068]
[0069]
[0070] The algorithm is mainly divided into three parts: the first part (lines 1-21) pre-processes the data objects in the global space, performs the initial granular division according to the conditional attributes and decision attributes of the data, and obtains the granular set It includes multiple attribute particles. The second part (lines 22-35) further divides each particle according to the upper and lower approximate sets to obtain the positive domain POS and boundary domain BND of the particle, and then constructs the particle feature matrix D E and the eigenvalue matrix M EC The third part (lines 36-64) obtains the difference matrix M based on the distribution reduction idea of the minimum identification attribute set D And the reduction Red of the conditional attribute set, get the association relationship of the attribute particles, and record the association rules that are not less than the minimum support min_sup and the minimum confidence min_conf into the association rule set R.
[0071] In order to solve the problem of highly sensitive information leakage caused by the aggregation of associated data, in the embodiment of this case, the association between data objects is mined for the attribute particles of the precondition of the association relationship between the attribute particles mined by Algorithm 1, so as to obtain highly associated data objects through the association deduction relationship.
[0072] Specifically, the target data object that meets the relevance condition in the global data is obtained, which can be designed to include:
[0073] The weight of each association rule is calculated according to the information content of the attribute particle and the quantitative value of the highest sensitivity level of the attribute particle clock attribute, and the association degree between the data objects in the global data is obtained by using the association rule weight and the confidence between the attribute particle association rules;
[0074] The association degree between data objects is set to be greater than or equal to a minimum association degree threshold value to satisfy the association degree condition, the target data object in the global data is obtained according to the association degree condition, and the target data object is recorded in the target data object set.
[0075] Let k be the number of rules in R, R i is the i-th rule, data object xg i and xg j The correlation degree is assoc(xg i ,xg j ), which can be expressed as:
[0076]
[0077] Among them, w i For R i The weight is expressed as:
[0078]
[0079] Among them, Info(cg i ) is the information content of the attribute particle, S max is the quantitative value of the highest sensitivity level of the attribute in the attribute particle. Let minassoc be the minimum association threshold of the data object. If assoc(xg i ,xg j )≥minassoc, then the data object xg i and xg j It is called a strongly associated data object, or associated object for short.
[0080] In the embodiment of this case, a highly associated data object discovery algorithm based on associated inference relationship is also designed to identify highly associated data objects, as shown in the specific pseudo code Algorithm 2.
[0081]
[0082] The algorithm traverses all data objects, calculates the association between data objects according to the association relationship obtained by Algorithm 1, and records data objects with an association higher than min_assoc into the high-association data object set Q.
[0083] As data objects grow dynamically, more data attributes are added, and attribute granules change accordingly, the associations between data objects will also change with the changes in attribute granules, including the degree of association between data object attributes and the possible emergence of new attribute associations.
[0084] To this end, in the embodiment of this case, the newly added data objects in the global space are screened according to the attribute set included in the conditional attribute set simplification, and the new data objects with the same attribute values as those in the conditional attribute set simplification are used as existing attribute granules, and the new data objects with different attribute values from those in the conditional attribute set simplification are granulated and the granule feature matrix and eigenvalue matrix are updated, so as to use the updated granule feature matrix and eigenvalue matrix to update the association rules of the target data objects.
[0085] Dynamic updating of associated data can better maintain the association of global data. This embodiment of the case proposes a dynamic updating algorithm for associated attributes based on mining the association relationship of data attributes. The specific pseudo code is shown in Algorithm 3.
[0086]
[0087]
[0088] The algorithm first screens the newly added data according to the attribute set contained in Red, directly converts the xg with the same attribute value as Red into the existing attribute granules, and re-granulates the unsatisfactory ones to update The final updated association rule set R t+1 .
[0089] S102. Predict the sensitivity level of the target data object through the fuzzy set possibility measure of the associated attribute sensitivity level, and judge whether there is a risk of unauthorized access when the user accesses the target object across domains based on the predicted sensitivity level. The fuzzy set of the associated attribute sensitivity level is used to describe the data association particles corresponding to the specified sensitivity level in the global data attribute space, and the fuzzy set possibility measure of the associated attribute sensitivity level is calculated using the fuzzy set membership function of the associated attribute sensitivity level and the attribute related possibility distribution function.
[0090] Specifically, the sensitivity level of the target data object is predicted by the fuzzy set possibility measure of the associated attribute sensitivity level, which can be designed to include:
[0091] Determine the maximum sensitivity level of the target data object;
[0092] The mutual information between target data objects is obtained according to the information entropy of the target data objects, and the weight of the associated attribute set corresponding to the target data object is calculated using the mutual information, so as to obtain the possibility of the target data object aggregation to derive a sensitivity level higher than the maximum sensitivity level by using the possibility measure of the associated attribute set on the fuzzy set higher than the maximum sensitivity level and the weight of the associated attribute set;
[0093] If the possibility is greater than the information possibility threshold, it is determined that there is information to derive data of a higher sensitivity level when the target data object is aggregated, so as to determine that there is a risk of unauthorized access when the user accesses the target object across domains.
[0094] Let U be the global space, AT be the attribute space on U, V be the value range of AT, and let is the associated attribute sensitive fuzzy set on AT, The sensitivity is s i Let X be a variable on AT, the probability distribution related to X is Πx, and the probability distribution function of Πx is π x , and numerically defined as F s The membership degree of
[0095] Let F be the sensitivity Fuzzy sets on AT, Π X is the probability distribution associated with variable X, and X takes values in AT, then F The likelihood measure is defined as:
[0096]
[0097] Among them, F( c ) is the membership function of F, πX ( c ) is the probability distribution function related to X.
[0098] The possibility of inferring higher-sensitivity information by aggregating related objects P , depends on the sum of the product of the possibility measure of the associated attribute set on the fuzzy set higher than the highest sensitivity level and its weight. The specific calculation process can be expressed as:
[0099]
[0100] Among them, k is the set of strongly associated aggregate objects Q i Poss() is the fuzzy set possibility measure with the highest sensitivity among all associated attributes; w i is the weight of the associated particle, and the calculation formula is as follows:
[0101]
[0102] Among them, I(Q i 1 ;Q i 2 ) is Q i 1 ,Q i 2 The mutual information is defined as:
[0103] I(Q i 1 ;Q i 2 )=H(Q i 1 )-H(Q i 1 |Q i 2 ),
[0104] Among them, H(P) is the information entropy that defines P, P is an equivalent partition of the entire domain U, and the information entropy is defined as:
[0105]
[0106] If P ≥ τ, it is considered that the aggregation of associated objects may lead to the derivation of information with a higher sensitivity level. Here, τ refers to the sensitivity level S derived from the associated objects. max The information possibility threshold (the highest sensitivity level of the associated data object) means that when data objects that meet this condition are accessed simultaneously, there may be security issues such as unauthorized access and information leakage. It is necessary to restrict the joint access to this data and improve the user's access policy.
[0107] Based on the above content, an algorithm for deducing the sensitivity level of associated data aggregation information is designed in the embodiment of this case. The pseudo code of the algorithm is shown in Algorithm 4 as follows.
[0108]
[0109]
[0110] The algorithm is based on the Q obtained in Algorithm 2. s , Q s As a subset of Q, we first determine the maximum sensitivity level of the data objects in the set, and then use the possibility measure of the sensitive fuzzy set of associated attributes in the data objects with high association to evaluate Q. s The possibility of deriving a higher sensitivity level can be calculated and obtained as <Q s ,S i ,P>.
[0111] Further, based on the above method, an embodiment of the present invention also provides a system for assessing information leakage risk in data aggregation based on granular association in multiple trust domains, comprising: a data granulation module and a risk assessment module, wherein:
[0112] A data granulation module is used to granulate the global data according to the logical deduction relationship between the global data in the target domain, extract the association rules between the granulated data attributes and obtain the target data objects that meet the association conditions in the global data, wherein the global data in the target domain includes several trust domains, and each trust domain stores data of a sensitivity level corresponding to the user's access rights;
[0113] The risk assessment module is used to predict the sensitivity level in the target data object through the fuzzy set possibility measure of the associated attribute sensitivity level, and judge whether there is a risk of unauthorized access when the user accesses the target object across domains based on the predicted sensitivity level. The fuzzy set of the associated attribute sensitivity level is used to describe the data association particles corresponding to the specified sensitivity level in the global data attribute space, and the fuzzy set possibility measure of the associated attribute sensitivity level is calculated using the fuzzy set membership function of the associated attribute sensitivity level and the attribute related possibility distribution function.
[0114] In order to verify the effectiveness of this solution, the following is a further explanation based on experimental data:
[0115] The simulation experiment environment is: python 3.7.16, Intel(R) Core(TM) i7-10875H CPU@2.30GHz, 32.0GB memory, and the system environment is Windows 10. The main purpose of this experiment is to explore the problem that low-sensitivity related data may cause users to deduce high-sensitivity data and cause information leakage. In order to meet the experimental requirements for multiple attribute relationships between data, scientific research data with easy-to-divid attribute characteristics in laboratories and the Internet and artificially synthesized data are used. According to the characteristics of the experimental data, data attributes are extracted, and data attributes and data attribute values are quantified to construct a priori knowledge base of data association relationships, the possible distribution of attributes and attribute sets, and fuzzy sets at various sensitivity levels. The experimental evaluation indicators include the execution efficiency of the algorithm and the accuracy of the algorithm deduction.
[0116] On the basis of obtaining the association relationship of the attribute sets of all data objects, the association relationship map between the attribute sets and the association relationship map between the data objects with high association are constructed. The minimum support minsup=0.04, the confidence minconf=0.5, and the association minassoc=0.12 are set to simulate the data with a data scale of 1000 and generate the association relationship map. In the association relationship map between attribute sets, a node is an attribute or an attribute set, and each node can represent a particle, that is, a data object or a data object set. The size of the node is used to reflect its importance, that is, the support. The directed edges between nodes represent the association relationship between nodes, and the depth of the color of the directed edges reflects the connection between nodes, that is, the confidence. In the association relationship map between data objects with high association, each node represents a data object, the connection between nodes represents the association relationship between nodes, and the depth of color can be used to reflect the association between data objects.
[0117] For the above selected data information, the performance of the data association mining algorithm and the high-sensitivity level information possibility deduction algorithm is evaluated from the following two aspects.
[0118] (1) Algorithm execution efficiency when data scale changes
[0119] According to the change of the number of selected data, the execution efficiency of the algorithm is evaluated when minsup=0.04, minconf=0.5, minassoc=0.12. The comparison results are as follows Figure 5As shown in the figure, under the same data scale, the efficiency of this scheme is higher than that of the existing object aggregation information level deduction method based on attribute association. The reference [5] refers to the existing object aggregation information level deduction method based on attribute association, which guides the implementation of network boundary access control policy by studying the association between objects. As the data scale increases, the association complexity between data objects increases, which increases the execution time of the algorithm. The execution efficiency of the algorithm increases, which is related to the time complexity of the algorithm. When the number of data for experimental test is 600, the size of its attribute space is 3000, and the execution time of the algorithm is 110.95s. At this time, the execution speed of the algorithm is relatively fast, the number of rules mined is 227, and the scale of associated objects with high association is 15133. When the number of data is 1500, the size of the attribute space is 7500, the execution time of the algorithm is 919.09s, the number of rules mined is 159, and the scale of associated objects is 35316. At this time, the execution speed of the algorithm is significantly increased compared with the time when the number of data is 600, which shows that with the increase of data scale, the number of associated objects obtained is more and the complexity is greater. Due to the increase of data in this experiment, the data attributes become more complex, resulting in changes in the support and confidence of the association rules. Due to the unified setting of the threshold, the increase of data scale reduces the number of association rules mined.
[0120] (2) Algorithm execution efficiency under threshold changes
[0121] According to the changes in parameters such as minimum support minsup and confidence minconf, the number of experimental data n is set to 1400, and the association degree minassoc = 0.12. The algorithm is evaluated under these conditions. The execution efficiency of the algorithm is as follows Figure 6 As shown in the figure. From the experimental results, it can be seen that with the decrease of minsup and minconf, the execution time of the algorithm increases. When minsup=0.12 and minconf=0.9 are set for experimental testing, the number of association rules mined is 16, and the scale of data object relationships with high association is 74311. When minsup=0.04 and minconf=0.5 are set for experimental testing, the number of association rules mined is 163, and the scale of data object relationships with high association is 28944. The reduction of the threshold setting increases the number of association rules, that is, the increase of associations between attributes, which will lead to an increase in associations between data objects, but the support and confidence values of the mined association rules are low, resulting in a decrease in the association between data objects, that is, the scale of data objects with high association becomes smaller, but the algorithm's running complexity increases, and its execution time also increases accordingly.
[0122] In order to illustrate the execution effect of the algorithm and verify its accuracy, the measurement indicators used are the accuracy rate (CR), error rate (WR) and error rate (ER) of the deduction algorithm. Let T be the experimental data set with safety deduction problems derived from the algorithm of this case, and N is the data set with safety deduction problems under standard conditions. The indicator calculation formula is as follows:
[0123] (1) Accuracy rate (CR): CR = |T∩N| / |N|, that is, the proportion of data objects contained in T in N to the total number of experimental data;
[0124] (2) Error rate (WR): WR = |T-(T∩N)| / |N|, i.e., the proportion of data with inference errors: the proportion of data objects in T that are not included in N to the total number of experimental data;
[0125] (3) Error rate (ER): ER = |N-(T∩N)| / |N|, which is the proportion of experimental data that are not derived and have safety deduction problems.
[0126] When the threshold value τ of sensitive level data deduction is set to 0.8, 0.7, 0.6, and 0.5, experimental simulations are performed on experimental data of scales 1000, 1200, and 1400 to aggregate and deduce the possibility of higher-level sensitive information. The accuracy and error rate of the algorithm are statistically analyzed. Minsup=0.04, minconf=0.5, and minassoc=0.12 are set. The experimental results are as follows Figure 7 and 8 shown.
[0127] From the simulation results, we can see that the higher the possibility threshold τ, the more data of related objects are ignored during the deduction process, so the accuracy of the algorithm is lower and more sensitive data is omitted; the lower τ means that more related objects that may cause security problems can be deduced. At this time, the accuracy of the algorithm is higher, but it is easy to misjudge normal data, resulting in an increase in the error rate. From the above analysis, it can be seen that a reasonable setting of the possibility threshold τ is the key to ensuring the accuracy of the deduction. The data obtained from the experimental simulation show that when τ = 0.6, the accuracy of the algorithm deduction is roughly maintained at more than 90%, and the error rate of the algorithm does not exceed 2%. From the actual effect, it can better detect the existence of security problems.
[0128] The above experimental data show that the solution in this case can calculate the possibility of deriving higher-sensitivity data due to data aggregation, effectively avoid information leakage and security issues, and then adjust user access to authorized data, prevent unauthorized access caused by cross-domain access by users, and reduce the risk of sensitive information leakage.
[0129] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0130] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0131] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.
[0132] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits, and accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of combination of hardware and software.
[0133] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A method for assessing information leakage risk in data aggregation in multiple trust domains based on granular association, which is used to prevent users from unauthorized access to sensitive data during cross-domain access, characterized in that: Include: Granulate the global data according to the logical deduction relationship between the global data in the target domain, extract the association rules between the granulated data attributes and obtain the target data objects that meet the association condition in the global data, wherein the global data in the target domain includes several trust domains, and each trust domain stores data of a sensitivity level corresponding to the user access rights; The sensitivity level of the target data object is predicted by the possibility measure of the fuzzy set of associated attribute sensitivity levels, and whether there is a risk of unauthorized access when the user accesses the target object across domains is determined based on the predicted sensitivity level. The fuzzy set of associated attribute sensitivity levels is used to describe the data association particles corresponding to the specified sensitivity level in the global data attribute space, and the possibility measure of the fuzzy set of associated attribute sensitivity levels is calculated using the fuzzy set membership function of the fuzzy set of associated attribute sensitivity levels and the attribute-related possibility distribution function.
2. The method for assessing information leakage risk in data aggregation based on granular association in multiple trust domains according to claim 1 is characterized in that: Granulate the global data according to the logical deduction relationship between the global data in the target domain, including: Preprocessing the data in the global data, and initially granulating the global data according to the conditional attributes and decision attributes of the data to obtain an attribute granule set and a decision granule set, wherein the attribute granule set includes a plurality of attribute granules, and the decision granule set includes a plurality of decision granules; According to the attribute particle set and the decision particle set, the upper and lower approximate sets are obtained, and each particle is divided by using the upper and lower approximate sets to obtain the positive domain and boundary domain of the particle, so as to construct the particle feature matrix and eigenvalue matrix by using the positive domain and boundary domain of the particle; Based on the particle feature matrix and the eigenvalue matrix and according to the minimum identification attribute set, the attribute importance of the attribute in the distribution simplification is obtained, and the simplification of the conditional attribute set is calculated according to the attribute importance. Based on the simplification, the association relationship between attribute particles and between attribute particles and decision particles is obtained, and the association rules with no less than the minimum support and minimum confidence are recorded into the association rule set according to the association relationship.
3. The method for assessing information leakage risk in data aggregation based on granular association in multiple trust domains according to claim 2 is characterized in that: The constructed particle feature matrix is expressed as: The eigenvalue matrix is expressed as: in, For data object E i The feature vector of i is the feature index, obj i is a data object collection, reg i is the positive domain of the particle, δ i is the boundary region of the particle, For E i The eigenvalue vector, e im For data object x i The property a m The characteristic value of .
4. The method for assessing information leakage risk in data aggregation based on granular association in multiple trust domains is characterized in that: Get the target data object that meets the relevance condition in the global data, including: The weight of each association rule is calculated according to the information content of the attribute particle and the quantitative value of the highest sensitivity level of the attribute particle clock attribute, and the association degree between the data objects in the global data is obtained by using the association rule weight and the confidence between the attribute particle association rules; The association degree between data objects is set to be greater than or equal to a minimum association degree threshold value to satisfy the association degree condition, the target data object in the global data is obtained according to the association degree condition, and the target data object is recorded in the target data object set.
5. The method for assessing information leakage risk in data aggregation based on granular association in multiple trust domains according to claim 4 is characterized in that: Data object xg i ,xg j The calculation formula of the correlation degree between them is expressed as: Among them, w i is the i-th association rule R i The weight of k is the number of association rules used to obtain the target data object, Info(cg i ) is the attribute particle cg i The amount of information, S max is the quantitative value of the highest sensitivity level of the attribute in the attribute particle, conf(R i ) is the confidence of the association rule.
6. The method for assessing information leakage risk in data aggregation based on granular association in multiple trust domains according to claim 4 is characterized in that: Obtaining the target data object that meets the relevance condition in the global data also includes: According to the attribute set contained in the conditional attribute set simplification, the newly added data objects in the global space are screened, and the new data objects with the same attribute values as those in the conditional attribute set simplification are taken as existing attribute granules. The new data objects with different attribute values from those in the conditional attribute set simplification are granulated and the granule feature matrix and eigenvalue matrix are updated, so as to use the updated granule feature matrix and eigenvalue matrix to update the association rules of the target data objects.
7. The method for assessing information leakage risk in data aggregation based on granular association in multiple trust domains according to claim 1 is characterized in that: Predicting the sensitivity level in the target data object by the possibility measure of the fuzzy set of associated attribute sensitivity level, including: determining the maximum sensitivity level of the target data object; The mutual information between target data objects is obtained according to the information entropy of the target data objects, and the weight of the associated attribute set corresponding to the target data object is calculated using the mutual information, so as to obtain the possibility of the target data object aggregation to derive a sensitivity level higher than the maximum sensitivity level by using the possibility measure of the associated attribute set on the fuzzy set higher than the maximum sensitivity level and the weight of the associated attribute set; If the possibility is greater than the information possibility threshold, it is determined that there is information to derive data of a higher sensitivity level when the target data object is aggregated, so as to determine that there is a risk of unauthorized access when the user accesses the target object across domains.
8. The method for assessing information leakage risk in data aggregation based on granular association in multiple trust domains according to claim 7 is characterized in that: The calculation formula for the probability of the target data object aggregation deriving a sensitivity level higher than the maximum sensitivity level is expressed as: Where k is the number of target data objects, P oss() is the fuzzy set possibility measure with the highest sensitivity among all associated attributes, w i is the weight of the associated attribute set, F s is the associated attribute sensitive fuzzy set on the attribute space, X is the value variable on the attribute space, F is the fuzzy set on the attribute space corresponding to the sensitivity level, S() is the sensitivity level, and a1 and a2 are the target data objects.
9. A system for assessing information leakage risk in data aggregation based on granular association in multiple trust domains, used to prevent users from unauthorized access to sensitive data during cross-domain access, characterized in that: Contains: data granulation module and risk assessment module, among which, A data granulation module is used to granulate the global data according to the logical deduction relationship between the global data in the target domain, extract the association rules between the granulated data attributes and obtain the target data objects that meet the association conditions in the global data, wherein the global data in the target domain includes several trust domains, and each trust domain stores data of a sensitivity level corresponding to the user's access rights; The risk assessment module is used to predict the sensitivity level in the target data object through the fuzzy set possibility measure of the associated attribute sensitivity level, and judge whether there is a risk of unauthorized access when the user accesses the target object across domains based on the predicted sensitivity level. The fuzzy set of the associated attribute sensitivity level is used to describe the data association particles corresponding to the specified sensitivity level in the global data attribute space, and the fuzzy set possibility measure of the associated attribute sensitivity level is calculated using the fuzzy set membership function of the associated attribute sensitivity level and the attribute related possibility distribution function.
10. An electronic device, characterized in that: include: at least one processor, and a memory coupled to the at least one processor; The memory stores a computer program, and the computer program can be executed by the at least one processor to implement the method according to any one of claims 1 to 8.