Context-Aware Method, System and Data Access Control Method Based on Weighted GraphSAGE

By adopting the weighted GraphSAGE model in the big data system, automatically perceive and model context information, the problem of dynamic access control in the big data system is solved, and the active control and efficient access management of user access requests is realized.

CN115658979BActive Publication Date: 2025-06-10Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211213490.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-06-10
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

It is difficult for the prior art to realize dynamic access control in big data systems, especially when user permissions need to be changed accordingly according to dynamically changing access scenarios.

Method used

The context-aware method based on weighted GraphSAGE is adopted to transform the access control context-aware problem into a graph neural network node learning problem, and the context information is automatically perceived through the graph neural network and access control rules are automatically generated.

Benefits of technology

It realizes dynamic control of user access requests, improves the flexibility and efficiency of big data access control, can automatically model and infer contextual relationships, and reduces manual intervention and workload.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115658979B_ABST
    Figure CN115658979B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of data access control, and particularly relates to a context-aware method, system and data access control method based on weighted GraphSAGE. The method includes obtaining original context information describing access control entity data from a user access record dataset; constructing a weighted graph of the context information, and modeling the access control context-aware problem as an inductive learning clustering problem of graph node representations; using a GraphSAGE graph neural network model to learn the graph node representations and obtain node embedding information, where the node embedding information at least includes node features, node relationship features and node relationship strengths; and using a clustering algorithm to perform inductive clustering on the node embedding information and generate scenario information composed of a set of context information, and automatically generating access control rules for dynamically controlling access requests from the scenario information. The present invention can realize dynamic control of user access requests based on context information, which is convenient for application in actual big data access control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data access control, and particularly relates to a context awareness method, system and data access control method based on weighted GraphSAGE. Background Art

[0002] With the rapid development of big data, the demand for its resource management and control has become increasingly prominent. Access control technology is still the key technology to protect the secure sharing of big data resources and prevent unauthorized access. However, compared with the data resources in traditional information systems, big data resources have new characteristics such as large data volume, dynamic generation, multi-source heterogeneity, and multi-element correlation. Users usually need to request access to resources (information, services, etc.) anytime and anywhere, but the access scenarios are dynamically changing. In such a dynamic environment, the user's permissions are not only related to static elements such as their identity, role, and inherent attributes, but are also affected by dynamic elements such as time, location, entity relationship, access intention or purpose. Big data systems urgently need to be able to perform authorization and access control according to relevant access scenarios, and be able to change the user's permissions accordingly with the dynamic changes of the access scenarios, and dynamically control the user's access. For example, users are not allowed to access system data resources during non-working hours, and when the accessing user and the data resource are not within the same location range, their continued access to the resource content is refused, etc. However, traditional classic access control models based on static policies, such as DAC, MAC, RBAC, etc., are difficult to directly meet the requirements of dynamic access control of big data resources.

[0003] To meet the requirement of dynamic access control according to changes in access scenarios, the Context aware Access Control (CAAC) model flexibly and dynamically controls the behavior of users accessing resources based on different "context" information scenarios. Context Information is used to describe the status of access control entities such as users, resources, or the environment, as well as the relationships between different entities. For example, information such as user type, resource category, time, location, relationships between entities, access intent, and access behavior history can all be regarded as context information available for access control. In the CAAC model, Context Aware is the prerequisite and foundation for dynamic access control using context information, mainly including steps such as context acquisition, modeling, and reasoning. However, in existing literature, context-aware methods for dynamic access control have the following problems: 1. They do not support automatic perception of context types and value ranges. When formulating context-based access control rules, it is necessary to pre-perceive the categories and values of the context in advance to determine the change range of the corresponding context information in the rules, so as to achieve dynamic control of the subject's access. However, currently, the CAAC model uses an artificial method by administrators to specify the types of context information in the rules and their corresponding change ranges. This artificial method has a large workload, low efficiency, and it is difficult to pre-define complete context information. 2. There is a lack of automatic modeling and reasoning of context relationships. There are many types of context involved in big data systems, and each type may include multiple constituent elements. There are various associations between the elements according to the access scenario. These associations are very important for context awareness and the formulation of context-based access control policies. Therefore, when formulating context-based access control rules, it is necessary to infer composite and implicit context information and the relationships between contexts from basic context information, such as inferring that the subject and the object are in the same place through their context information. However, currently, most CAAC models use an artificial method to define ontologies to model contexts and administrators formulate relevant reasoning rules. They cannot automatically establish the connections between context information, and the established context relationships are fixed and cannot be changed with the change of the context status. Therefore, there is an urgent need for an automatic and efficient context-aware method to obtain the context information of big data system access in a non-artificial way and be able to model and reason about the association relationships between these context information, providing necessary support for intelligent decision-making in big data access control. Summary of the Invention

[0004] To this end, the present invention provides a context-aware method, system and data access control method based on weighted GraphSAGE, which transforms the access control context awareness problem into a graph neural network node learning problem, and uses the graph neural network model to automatically sense context information to automatically generate context-aware access control rules, so as to realize the dynamic control of user access requests based on context information, which is convenient for application in actual big data access control.

[0005] According to the design scheme provided by the present invention, a context-aware method based on weighted GraphSAGE is provided for big data access control, including the following content:

[0006] Obtain the original context information describing the access control entity data from the user access record dataset;

[0007] Construct a weighted graph of context information, and model the access control context awareness problem as an inductive learning clustering problem of graph node representation. Among them, in the weighted graph, the context information in the user access record is used as the graph node, and the relationship between the context information is used as the edge. Each node is composed of multiple elements, and each element is the same as the corresponding position element of the access record, and the edge weight is set according to the node elements;

[0008] Use the GraphSAGE graph neural network model to learn the graph node representation and obtain the node embedding information, where the node embedding information at least includes node features, node relationship features and node relationship strength; and use the clustering algorithm to perform inductive clustering on the node embedding information and generate scenario information composed of a set of context information, and automatically generate access control rules for dynamically controlling access requests from the scenario information.

[0009] As the context-aware method based on weighted GraphSAGE in the present invention, further, the context information is represented as a multi-tuple composed of access control entity data elements, where the access control entity data elements at least include subject status information, object status information, environmental status information, operation request information, relationship status information and access purpose status information.

[0010] As the context-aware method based on weighted GraphSAGE in the present invention, further, when modeling the access control context awareness problem, a multi-graph G=(V, E) is used to model the context information features and relationships, where each user access record X i is used as a node V of the graph i , and each graph node V i is composed of m elements. An element edge is established between two nodes whose element distances of the corresponding elements in the node are less than the preset element distance threshold. Each element is associated with an element weight, and the edge weight is the sum of the corresponding element weights.

[0011] As the context-aware method based on weighted GraphSAGE in the present invention, further, when using the GraphSAGE graph neural network model to obtain node embedding information, first, the neighbor sampling function is used to sample the nodes in the graph from the inside out, and then the sampled nodes in the graph are aggregated one by one from the outer layer to the inner layer, and the embedding information of the nodes is generated. Among them, the probability of each neighbor node being selected in the sampling function is set as the proportion of the edge weight of the corresponding node in the sum of the edge weights of all neighbor nodes.

[0012] As the context-aware method based on weighted GraphSAGE in the present invention, further, when using the sampling function to perform neighbor sampling on the nodes in the graph, first, an initial node set is set, and all first-order neighbor nodes of the initial nodes are obtained; then, the first-order neighbor nodes are sorted according to the node edge weights between the nodes and the first-order neighbor nodes, and weighted sampling is performed from the neighbor nodes; then, the sampled nodes are weighted to sample their respective neighbor nodes, and weighted sampling is performed layer by layer in turn according to the preset sampling layers to obtain a weighted neighbor node set.

[0013] As the context-aware method based on weighted GraphSAGE in the present invention, further, in the weighted aggregation operation, first, the representation information of the outermost neighbor nodes of each node is multiplied by the node weight, and the products are aggregated; then, the aggregated value of the neighbor nodes is concatenated with the current representation information of the second outermost layer nodes, and the representation information of the second outermost layer nodes is obtained through the transformation of the non-linear activation function; then, the representation information of the second outermost layer neighbor nodes is used to continue weighted aggregation until the aggregation operation reaches the innermost layer nodes.

[0014] As the context-aware method based on weighted GraphSAGE in the present invention, further, the K-means clustering algorithm is used to perform clustering operations on the feature vectors of the node embedding information, and the node clusters are used to cluster the access records with the same or similar context information, so that each node cluster after clustering corresponds to a context scenario, and the feature information of the node cluster represents the feature information of the corresponding context scenario.

[0015] As the context-aware method based on weighted GraphSAGE of the present invention, further, the access control rules automatically generated from the scenario information are represented as <RuleID, UserClass, Contextscenario, ClusterCenter, ControlResult>, where RuleID represents the rule number, and the rule number is the same as the context scenario serial number, UserClass represents the user classification identifier represented by the user role and characteristic attributes, Contextscenario represents the description of the clustering result of the corresponding context information, ClusterCenter represents the clustering center feature vector, and ControlResult represents the access control result.

[0016] Further, the present invention also provides a context-aware system based on weighted GraphSAGE for big data access control, including: a data acquisition module, a problem modeling module, and an analysis and clustering module, where

[0017] The data acquisition module is used to acquire the original context information describing the access control entity data from the user access record dataset;

[0018] The problem modeling module is used to construct a weighted graph of the context information and model the access control context-aware problem as an inductive learning clustering problem represented by graph nodes. In the weighted graph, the context information in the user access record is used as graph nodes, the relationship between the context information is used as edges, each node consists of multiple elements, and each element is the same as the corresponding position element in the access record, and the edge weights are set according to the node elements;

[0019] The analysis and clustering module is used to use the GraphSAGE graph neural network model to learn the graph node representation and obtain the node embedding information, where the node embedding information at least includes node features, node relationship features, and node relationship strengths; and use the clustering algorithm to perform inductive clustering on the node embedding information and generate scenario information composed of a set of context information, and automatically generate access control rules for dynamically controlling access requests from the scenario information.

[0020] Further, the present invention also provides a data access control method, which uses the above context-aware method to generate access control rules, and makes an access control permission decision by matching the user access request with the generated access control rules. The matching process includes: first, converting the user access request information into a context vector and obtaining the context information; then, calculating the distances between the context information and each clustering center; then, sorting the multiple distance values, selecting the clustering center with the closest distance as the matching item, and using the access control rule corresponding to the matching item to respond to the user access request.

[0021] Advantages of the present invention:

[0022] The present invention applies a graph neural network to solve the problem of context information perception in access control. It performs graph modeling on the access record data set, transforms the access control context perception problem into a graph neural network node learning problem, uses the context information in the user access record as graph nodes, and the relationships between the context information as edges to construct a weighted graph of the context information. Through the graph neural network, it realizes graph node representation learning and obtains node embedding information including node features, node relationship features, and node relationship strengths, etc. Correspondingly, context information including context features, context relationships, and their strengths is obtained. The graph neural network model WGraphSAGE can use the multi-dimensional vectors of access records as input to obtain the original context information, perform graph modeling on the context information through multi-graph construction and weighted graph conversion, learn the node embedding of the context information through weighted neighbor sampling and weighted aggregation functions, infer the context scenario information through a clustering algorithm, and automatically generate access control rules from the context scenario, thereby realizing the dynamic control of access requests. The original GraphSAGE model adopted equivalent random sampling and aggregation calculation methods, ignoring the different impacts of neighbor nodes with different weights on node representation. In the solution of this case, weighted neighbor sampling and weighted aggregation algorithms are used to sample and aggregate calculate neighbors according to different weight information. The obtained graph node embedding information not only contains the feature information of the node itself and neighbor nodes, but also integrates the respective different weight information of relevant neighbor nodes, thus realizing the automatic modeling and reasoning of context relationships, being more flexible and facilitating applications in dynamic access control of big data resources. Description of the Drawings

[0023] Figure 1 Schematic diagram of the context awareness process based on weighted GraphSAGE in the embodiment;

[0024] Figure 2 Schematic diagram of the principle framework of the WGraphSAGE model in the embodiment;

[0025] Figure 3 For the data set D and sample X in the embodiment i Schematic diagram of the data structure;

[0026] Figure 4 Schematic diagram of a simple graph and a multi-graph in the embodiment;

[0027] Figure 5 Schematic diagram of converting a multi-graph to a weighted graph in the embodiment;

[0028] Figure 6 Example of normalizing node edge weights of a weighted graph in the embodiment;

[0029] Figure 7 Schematic illustration of clustering of graph node embeddings in the embodiment;

[0030] Figure 8 Example of weighted neighbor sampling process in the embodiment;

[0031] Figure 9 Example of weighted aggregation in the embodiment;

[0032] Figure 10 Graph of node quantity - time relationship in the embodiment;

[0033] Figure 11 Graph of iteration times - loss function value relationship in the embodiment;

[0034] Figure 12 Graph of clustering quantity - metric relationship in the embodiment;

[0035] Figure 13 Schematic illustration of clustering results in the embodiment;

[0036] Figure 14 Schematic illustration of access control decision results in the embodiment;

[0037] Figure 15 Schematic illustration of ablation experiment analysis in the embodiment. Detailed implementation manners

[0038] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and technical solutions.

[0039] Context information is a key element for implementing dynamic access control of big data. In traditional access control models, such as DAC, MAC, RBAC, etc., the permissions assigned to users are pre-granted and static. However, the access process in big data systems involves different types of and dynamically changing context information. Therefore, these models cannot dynamically change permissions based on the context information at the time of access, nor can they adapt to the continuously changing context scenarios. To support dynamic access control decisions, many literatures have extended the traditional access control models to add support for context. For example, elements such as time information, space information, and the combination of time and location information have been added to the RBAC model. According to the time, space, and other information at the time of access, the role and its corresponding permissions are dynamically activated. However, these improved models can only attach specific types of context information to the original model, supporting fewer types of context information, and do not consider other types of environmental factors or context information such as entity relationships and access intentions. The ABAC access control model performs dynamic control through global "environmental conditions". Conceptually, the environmental conditions of ABAC also belong to the category of context information. However, the environmental conditions of ABAC only include some common context conditions, such as time, location, etc., and cannot well support the dynamic changes in access permissions caused by the context conditions of the subject and object themselves, such as the subject state, the relationship between the subject and the object, etc. Context aware Access Control can implement dynamic control of access behavior according to different scenarios. In the existing methods, it cannot support the automatic perception of context types and value ranges, and lacks automatic modeling and reasoning of context relationships, which will thus affect the learning results and execution efficiency of subsequent context-aware access control rules. Therefore, the embodiments of the present invention provide a context-aware method based on weighted GraphSAGE for big data access control, including:

[0040] S101. Obtain the original context information describing the access control entity data from the user access record dataset;

[0041] S102. Construct a weighted graph of context information, and model the access control context awareness problem as an inductive learning clustering problem represented by graph nodes. Among them, in the weighted graph, the context information in the user access record is used as graph nodes, the relationship between context information is used as edges, each node consists of multiple elements, each element is the same as the corresponding position element in the access record, and the edge weights are set according to the node elements;

[0042] S103. Use the GraphSAGE graph neural network model to learn the graph node representation and obtain node embedding information, where the node embedding information at least includes node features, node relationship features, and node relationship strength; and use a clustering algorithm to inductively cluster the node embedding information and generate scenario information composed of a set of context information, and automatically generate access control rules for dynamically controlling access requests from the scenario information.

[0043] Convert the user access record set into a graph structure, design a weighted neighbor sampling and weighted aggregation algorithm, learn the weighted representation method of graph nodes, and then cluster the learned node embedding information to obtain context scenario information, and the context scenario can automatically generate access control rules, so as to realize the dynamic control of user access according to context information. Among them, context information: information used to describe the state of access control-related entities such as users, resources, or environments, the relationships between different entities, access purposes or intentions, etc.

[0044] The GraphSAGE model using an inductive framework transforms "learning the representation of each node" into "learning the representation method of nodes". The information of a node is obtained by randomly sampling its neighbor nodes, and then using the selected aggregation function to obtain the feature representation of the node, transforming transductive learning into inductive learning, avoiding the situation where the features of nodes need to be retrained each time, and supporting incremental features; moreover, by introducing neighbor sampling, the node representation corresponds to the inductive representation of nodes with multiple local structures, which can effectively prevent overfitting in training and enhance the generalization ability. The disadvantages of GraphSAGE are: (1) It cannot handle weighted graphs and can only perform equal-weight aggregation on neighbor nodes; (2) Random equal-length sampling will cause the loss of important local information of some nodes, and the features of the same node embedding in the inference stage are unstable.

[0045] See Figure 2 As shown, in the embodiment of this case, the weighted graph neural network model WGraphSAGE is used to transform the context awareness problem into a graph representation learning problem, introduce "edge weights" in the graph neural network, and accurately characterize the context information features and relationships, overcoming the problems of the original GraphSAGE model's "equal-weight" calculation that cannot distinguish the importance of neighbor nodes and does not support multiple types of node features. Compared with the existing CAAC model, WGraphSAGE can realize the automatic perception of context types and value ranges, and realize the automatic modeling and reasoning of context relationships.

[0046] As a preferred embodiment, further, the context information is represented as a multi-tuple composed of access control entity data elements, where the access control entity data elements at least include subject status information, object status information, environmental status information, operation request information, relationship status information, and access purpose status information.

[0047] In the embodiments of this case, the context information can be represented as <u, o, e, op, r, p>, where u represents subject status information, o represents object status information, e represents environmental status information, op represents operation request information, r represents relationship status information, and p represents access purpose status information. Each element can further include multiple sub-elements. For example, the environmental status information e can further include time t, location l, a combination of time and location <t, l>, etc. To facilitate learning of the node representation, the values of each element or sub-element in the context information are set to discrete values or can be null values. Context scenario: A set of similar context information. The context information within the same context scenario is closely related, while the context information between different context scenarios has no association or is distantly related. Context aware: The process of obtaining, modeling, and reasoning about context scenarios is called context aware.

[0048] Through the weighted graph neural network model WGraphSAGE, context scenarios are perceived from the user access record dataset. The problem definition can be expressed as follows:

[0049] WG: D → CS (1)

[0050] Where D is the access record set, CS is the context scenario set, and WG is the WGraphSAGE model. The goal is to learn WGraphSAGE to obtain high-quality context scenarios, that is, the context scenarios retain the characteristics of the context information in the access record dataset and the relationships between the context information.

[0051] A multi-graph is used to model the context information features and relationships, and the multi-graph is defined as G = (V, E). In order to more accurately model the access record set through the multi-graph. Among them, the graph node (Graph Node) takes each record X i as a node V of the graph i , that is, V i = X i (0 <= i <= n). In the element of the graph node (Element of Graph Node), each graph node V i is composed of m elements, and each element x of the graph node ij is the same as the corresponding element in the access record X i .

[0052] The Minkowski distance between the eigenvectors of the corresponding elements of two nodes in the graph is represented by the Distance of Elements (DoE). distence j is the Minkowski distance function of the eigenvector of element j, and node V p 、V q The j-th element distance is expressed as:

[0053]

[0054] The Threshold for DoE (TfD) is the critical value for determining whether there is an edge based on the corresponding element distance. TfD j represents the distance threshold of the j-th element. In the Edge for Elements (EfE), an element can be regarded as a child node of a node. If the element distance between the corresponding positions of two nodes does not exceed the threshold TfD, there is an element edge between the corresponding elements of the two nodes, indicating that there is a certain relationship between them.

[0055] In the Sign for EfE (s), s pqj represents the flag bit indicating whether there is an element edge at the j-th element position between two nodes V p 、V q If the element distance does not exceed the threshold, the corresponding elements of the two nodes in the graph are similar or close, and there is an edge between them. The flag bit s pqj = 1; if the element distance exceeds the threshold, the corresponding elements of the two nodes in the graph are not related, and there is no edge between them. The flag bit s pqj = 0. The calculation formula for the element edge flag bit is as follows:

[0056]

[0057] The Sign Vector for EfE (SV): The element edge flag bits of nodes V p 、V q form an m-dimensional vector SV pq = [s pq1 , s pq2 , …, s pqj , …, s pqm . For example, in Figure 5 , the element edge vector SV between V p 、V q is: [0, 1,...1,..., 1]. The number of element edges between two nodes is

[0058] The edge between graph nodes is called the node edge EfN (Edge for Nodes). In the element weight WoE (Weight of Elements), each element of the graph node is associated with a weight value, and this weight value is called the element weight. The element weights at the same position of any node are the same, and the column vector of the element weight is WoE = [w 1 , w 2 , …, w j , …, w m T . In the node edge weight WoEfN (Weight of Edge for Nodes), each edge between nodes is associated with a weight value, and this weight value is the sum of the corresponding element weights (WoE).

[0059] As a preferred embodiment, further, in obtaining the node embedding information using the GraphSAGE graph neural network model, first, the neighbor sampling of the nodes in the graph is performed from the inside out using the sampling function, and then the nodes in the sampled graph are aggregated one by one from the outer layer to the inner layer to generate the embedding information of the nodes. Among them, the probability of each neighbor node being selected in the sampling function is set to the proportion of the corresponding node edge weight in the sum of the edge weights of all neighbor nodes.

[0060] See Figure 2 ​As shown, a set of user access records containing multivariate vectors is used as the input of the model. Each access record includes multiple types of information such as the subject (the visitor, i.e., the user), the object (the resource), the access operation (read, write, add, delete, update, etc.), and the context (time, location, the relationship between the subject and the object, etc.). Each type of information contains one or more vector information. Each access record is used as a node in the graph. The input features of the node are composed of multiple elements. The types and numbers of elements of each node are the same, but the element values may be the same or different. When the vector value of the corresponding element distance (DoE) does not exceed the element distance threshold (TfD), an element edge (EfE) is established between the corresponding elements of the two nodes. The number of element edges between the two nodes is 0 to m (m is the number of elements). To facilitate the learning of the graph neural network and without losing the edge relationship between nodes, the multi-graph is converted into a weighted graph through a normalization method. There are 0 or 1 node edges (EfN) between the nodes of the weighted graph, and the weight (WoEfN) of the node edge is the sum of the corresponding element weights (WoE). According to the Mini-Batch idea of GraphSAGE, the nodes in the graph can be sampled layer by layer from the inside out; to distinguish different association relationships between nodes, a variable of the node edge weight W’ is added to the sampling function, and the neighbor nodes corresponding to the edges with higher node edge weights are preferentially sampled. On the basis of weighted sampling, from the outer layer to the inner layer, the nodes in the graph are aggregated one by one to generate the embedding information of the nodes. The output node embedding information is clustered to obtain a set of context scenario information. Each element in the set corresponds to a context scenario for the access control policy. The context scenarios obtained by clustering are automatically generated into corresponding access control rules. Finally, when an access request is received, the access request is converted into a feature vector containing context information and matched with the access control rules to obtain the final decision result.

[0061] As a preferred embodiment, further, in the neighbor sampling of the nodes in the graph using the sampling function, first, an initial node set is set, and all first-order neighbor nodes of the initial nodes are obtained; then, the first-order neighbor nodes are sorted according to the node edge weights between the nodes and the first-order neighbor nodes, and weighted sampling is performed from the neighbor nodes; then, the sampled nodes are weighted sampled for their respective neighbor nodes, and weighted neighbor node sets are obtained by weighted sampling layer by layer according to the preset sampling layers.

[0062] The user access records exactly contain the raw data required for context acquisition, such as time, location, entity status, entity relationship, access intent, etc., which can be used as the basis for automatically learning and reasoning about the context. Here, the user access records refer to the log records after the user's successful and failed access to resources, or the information of the user's access resource request records, etc. The user access records are a collection of data units containing multiple elements, multiple dimensions, and no labels, and the amount of data is very large. Moreover, there are internal connections between the user access records and they are not independent of each other. It is necessary to model and reason to discover their internal relationships.

[0063] Taking the user access records as the original input data, each data record includes multiple types of information such as the subject, object, access operation, context, etc., and each type of information contains one or more vector information.

[0064] Assume that the access record set D contains n + 1 unlabeled samples (i.e., access records), D = {X 0 , X 1 , … X i , … X n}, (0 ≤ i ≤ n). Each sample (input data) X i includes m elements, X i = [x i1 , x i2 , …, x ij , …, x im , (0 ≤ i ≤ n, 1 ≤ j ≤ m). The element x ij is a d j -dimensional vector, x ij = [x ij1 , x ij2 , …, x ijk , … x ijdj , (0 ≤ i ≤ n, 1 ≤ j ≤ m, 1 ≤ k ≤ d j , x ij ∈ X i ). Each element corresponds to a context information. The data structure is as Figure 3 shown.

[0065] User access records are a dataset with multiple elements, multiple dimensions, multiple features, and a large quantity, and there may be various relationships among the context information in the records. In order to learn the context features, potential patterns, and internal relationships during user access from the access record dataset, the embodiments of this case depict and describe these data by constructing a graph. There are two optional methods for constructing a graph from user access records: one is a multigraph. A graph containing parallel edges is called a multigraph, that is, if the number of edges between two nodes in the graph is more than one and it is also allowed for a vertex to be associated with itself through the same edge, then it is called a multigraph. Each access record can be used as a node, and each relationship between the records can be used as an edge, thus constructing a multigraph. The other is a heterogeneous graph. A heterogeneous graph includes multiple types of nodes and edges. The elements in the access record can be used as nodes, and the relationships between the elements can be used as edges to construct a heterogeneous graph. In this way, each access record (containing multiple elements) will correspond to multiple graph nodes and be associated with multiple types of edges. When the types of elements and relationships are numerous, it will cause a significant increase in the scale and complexity of the graph.

[0066] In the embodiments of this case, the multigraph G=(V, E) is used to model the features and relationships of the context information in the access record set. Each access record is used as a node, and the relationships between the context information in the record (hereinafter simply referred to as the relationships between records) are used as edges. Whether there is an edge between two nodes is related to the corresponding element values in the nodes. When a pair of corresponding element values in two access records are similar or the same, for example, the access time Time is the same, then there is a relationship (the same time) between these two access records, and this relationship can be represented by establishing an edge between these two user records (the corresponding nodes). When multiple pairs of corresponding element values in two access records are similar or the same, then there are multiple relationships between these two access records. For example, the access time Time is the same and the access location Location is the same, then there are two relationships (the same time, the same location). These relationships can be represented by establishing multiple edges between these two access records (the corresponding nodes). As Figure 4 shown, there is only one edge between the nodes of a simple graph, while there is one or more edges between the nodes of a multigraph. V 0 and V 1 have 3 edges between them, V 1 and V 2 have two edges between them, and there is one edge between the nodes V 1 and V 3 The edges between the nodes form the edge set E.

[0067] The graph node Vi consists of m graph node elements. Then, the set of all access records corresponds to the node set of the graph Each graph node has elements of the same type and quantity, which represent context information during access, such as time, location, resources, user type, user status, etc. The input feature vector of a graph node is X i = [x i1 , x i2 , …, x ij , …, x im , (0 ≤ i ≤ n, 1 ≤ j ≤ m), and the graph node element x ij is a d j -dimensional vector, x ij = [x ij1 , x ij2 , …, x ijk , …x ijdj , (0 ≤ i ≤ n, 1 ≤ j ≤ m, 1 ≤ k ≤ d j , and x ij ∈ X i ). The dimensions of each element form a vector D x = [d 1 , d 2 , …, d j , …, d m T , and the dimension of this vector is the sum of the dimensions of each element, that is

[0068] When the element distance DoE of the corresponding elements of two nodes is less than the element distance threshold TfD, it means that the two nodes are related, that is, there are the same context elements in the two access records, and an element edge EfE is established between the corresponding elements of the two nodes. If there are multiple pairs of identical or similar elements, multiple element edges will be established, indicating that there are multiple corresponding context elements that are identical or similar in the two access records. The more element edges there are between two nodes, the closer the relationship between them. For subsequent convenient statistical counting of the number of element edges, the element edge flag s is used to indicate whether there is an element edge between the corresponding elements of two nodes, and each element edge flag forms an element edge vector SV

[0069] Based on the relationship of the corresponding elements in the nodes, the multi-graph describes the multiple relationships between the nodes. However, there are multiple element edges between two nodes, which is not convenient for the learning of the graph neural network model. Therefore, in the embodiments of this case, the weight of the edge is introduced, multiple edges are converted into weighted edges, and accordingly the multi-graph is converted into a weighted graph. A weighted graph is a graph model that associates a weight or cost with each edge, that is, on the basis of a simple graph, a weight is added to each edge

[0070] The process of converting a multi-graph into a weighted graph is as Figure 5 shown. If two nodes V p , V q ​There is one or more element edges, which are uniformly changed to a node edge EfN. The weight of this node edge WoEfN is calculated through the element edge vector SV and the element weight WoE:

[0071]

[0072] For the convenience of data processing and calculation, the edge weights are processed by min-max normalization:

[0073]

[0074] When SV pq =[s pq1 , s pq2 ,…, s pqj ,…, s pqm is a vector of all zeros, the weight W' pq =0; when SV pq =[s pq1 , s pq2 ,…, s pqj ,…, s pqm is a vector of all ones, the weight W' pq =1. By calculating the normalized weights of all node edges in G, the adjacency matrix of graph G can be obtained, denoted as an N-order square matrix A∈{W'} N×N . An example of the normalized edge weights of a weighted graph is shown as Figure 6 .

[0075] After learning by the WGraphSAGE model, node embedding information is output, which not only retains the feature information of the input data in the low-dimensional vector data, but also incorporates the association information between nodes in the graph structure (here the association includes two aspects: multiple relationships, multiple weights (different degrees of relationship closeness)). The feature vector of the node embedding information can be used for downstream calculations and applications. Clustering operations can be performed on it, and the resulting node clusters (sets) will gather access records with the same or similar context information together. Then, the feature description of this node cluster can be used as the basis for the corresponding context scenario information description. Further, the K-means clustering algorithm is used to perform clustering operations on the feature vectors of the node embedding information, and the node clusters are used to cluster access records with the same or similar context information, so that each node cluster after clustering corresponds to a context scenario, and the feature information of the node cluster represents the feature information of the corresponding context scenario.

[0076] The feature vectors of the node embedding information are clustered through the K-means algorithm, and the clustering results can be visually presented through the t-SNE algorithm. By clustering the node embedding information, the points within the same cluster (with closer distances) are aggregated more closely, and the points between different clusters (with farther distances) are more distant. The clustering results can be used as the inferred context scenario information and as the basis for formulating context-based access control policies. The schematic diagram of the clustering is as shown in Figure 7 Figure 1. After clustering the node embedding information, each node cluster corresponds to a context scenario, and the feature information of the node cluster represents the feature information of the corresponding context scenario. The system will automatically formulate context-aware access control rules based on the clustering results (context clusters). The context-aware access control rules are represented as <RuleID, UserClass, Contextscenario, ClusterCenter, ControlResult>. Among them, RuleID represents the rule number, which is the same as the scenario serial number. UserClass represents the classification identifier of the user, which can be represented by the user role and feature attributes. Contextscenario is the description of the clustering result of the corresponding context information. ClusterCenter represents the feature vector of the clustering center, and ControlResult represents the access control result.

[0077] The main difference between the weighted neighbor sampling and weighted aggregation algorithm and the existing GraphSAGE algorithm is that the node edge weights are added during the neighbor sampling and aggregation process. When sampling, the neighbor nodes with high node edge weights are preferentially selected, and when aggregating, the neighbor node vectors are multiplied by the corresponding "node edge weights" and then the aggregation operation is performed. Through the weighted neighbor sampling and aggregation operations, the different degrees of weight relationships between the node and its neighbor nodes can be characterized, and the contribution degrees of neighbor nodes with different weights to the node embedding can be distinguished. The final output node vector representation aggregates more information of the "high-weight" neighbor nodes, which is conducive to clustering the context information in the closely related adjacent nodes more closely together.

[0078] Different from the random neighbor sampling, the weighted neighbor sampling algorithm first specifies the initial node set, and then obtains all the first-order neighbor nodes N(u) of the initial nodes. There are node edges between the nodes and the first-order neighbor nodes. Then, the first-order neighbor nodes are sorted according to the node edge weights, and then S k nodes are weighted sampled from the N(u) neighbor nodes according to the reservoir algorithm. According to the above method, S k nodes are weighted sampled for their respective S k-1 neighbor nodes, and weighted sampling is performed layer by layer in turn to obtain all the qualified weighted neighbor node sets. Among them, k represents the number of sampling layers. To prevent the problem of over-smoothing, it is generally set to 2. The weighted neighbor sampling process is as shown inFigure 8 As shown. In the example, K = 2, S 1 = 3, S 0 = 2. First, for node V 0 , at the first layer (inner layer), from all neighbor nodes {V 1 , V 2 , V 3 , V 4 , V 5}, 3 nodes with higher "node-edge weights" are sampled with weights from {V 1 , V 3 , V 5}. At the 0th layer (outer layer), each first-order neighbor node is sampled with weights separately, and 2 neighbor nodes are obtained for each, resulting in 6 neighbor nodes {V 6 , V 7 , V 9 , V 10 , V 12 , V 13}.

[0079] In GraphSAGE, the neighbor sampling function uses a fixed-length random sampling method, while the neighbor sampling function NW(u) in the embodiments of this case is different. When sampling neighbors, it considers the node-edge weight information associated with them. The weighted neighbor sampling algorithm can be as shown in Algorithm 1:

[0080]

[0081] The algorithm is divided into two stages. In stage 1, all first-order neighbor nodes of node u are obtained from the graph (i.e., the nodes with node edges to u), and they are sorted according to the node-edge weights. The algorithm complexity of stage 1 is O(n). In stage 2, S neighbors are sampled with weights from the N(u) neighbor nodes of u. S is a hyperparameter determined in advance. The problem to be solved in stage 2 is essentially a weighted random sampling problem, that is, each sample in the sample set is attached with a weight w i > 0, and the probability of each sample being drawn is determined by w i . In the embodiments of this case, the reservoir sampling algorithm can be used to implement the sampling. If the number of neighbors N(u) is greater than or equal to the sampling target number S, then call A-Res(N(u), W’ u , S) for sampling without replacement, where the weight of the sample is W’ u , and the reservoir size is S. If the number of neighbors N(u) is less than the sampling target number S, then create S reservoirs of size 1. For each sample V i, the A-Res algorithm is independently run on each reservoir for sampling with replacement. The algorithm complexity of stage 2 is O(m*log(n)).

[0082] As a preferred embodiment, further, in the clustering operation, first, the representation information of the outermost neighbor nodes of each node is multiplied by the node weight, and the products are aggregated; then, the aggregated value of the neighbor nodes is concatenated with the current representation information of the second outermost layer nodes, and the second outermost layer node representation information is obtained through a non-linear activation function transformation; then, the representation information of the second outermost layer neighbor nodes is used to continue the weighted aggregation until the aggregation operation reaches the innermost layer node.

[0083] Different from the aggregation function in GraphSAGE, in the solution of this case, the weighted aggregation function not only aggregates the representation information of neighbor nodes, but also adds the edge weight information between the node and the neighbor nodes. The weighted aggregation is performed layer by layer from the outer layer to the inner layer based on the weighted neighbor sampling results. First, the representation information of the outermost neighbor nodes of each node is multiplied by the node edge weight, secondly, these products are aggregated and calculated, thirdly, the aggregated value of the neighbor nodes is concatenated with the current representation information of the second outermost layer nodes, and then the concatenated vector is transformed through the fully connected layer of the non-linear activation function to obtain the second outermost layer node representation information. Using the representation information of the second outermost layer neighbor nodes, continue the weighted aggregation operation. And so on, until the aggregation operation reaches the innermost layer node. As Figure 9 shown, the weighted sampling neighbor nodes of V 0 are {V 1 , V 3 , V 5}, while the weighted sampling neighbors of V 1 , V 3 , V 5 are {V 6 , V 7}, {V 9 , V 10}, {V 12 , V 13} respectively. When the node V 1 performs weighted aggregation, first calculate the neighbor weighted aggregation information of V 1 Then calculate the joint information of V 1 Similarly, obtain the weighted aggregation representation information of the nodes V 3 , V 5 respectively, and then calculate the neighbor weighted aggregation representation information of the node V 0 as The joint information is AGG represents the name of the aggregation function. The optional aggregation functions include Mean, LSTM, Mean-Pooling, and Max-Pooling. After aggregation, normalization processing is calculated for the representation information of the nodes.

[0084] By referring to the idea of the Mini-Batch aggregation algorithm of GraphSAGE, the Mini-Batch weighted aggregation algorithm can be as shown in Algorithm 2.

[0085]

[0086]

[0087] The meanings of the symbols in the algorithm are as follows. X v represents the input features of a certain node v. represents the initial vector representation of node v; represents the vector representation of node v after k iterations; z v represents the final output vector of a certain node v after passing through the WGraphSAGE model. k represents the number of layers, and the number of layers refers to the farthest distance that the node features can be transmitted. For example, in the first layer of WGraphSAGE, each node can only obtain information from its neighbor nodes, and the process of each node collecting information is independent and synchronous. When adding another layer on the basis of the first layer, the process of collecting information is repeated, but this time when collecting information, the neighbor nodes already have the information of their own neighbors (from the previous step). This makes the number of layers the maximum number of hops that each node can take.

[0088] Phase 1: Select a small batch subset B of the node set V as the initial node set B K . Lines 2-7 execute two nested loop bodies. The first loop performs weighted sampling processing layer by layer from the inside out, and the second loop performs weighted sampling on each neighbor of the nodes in this layer. Starting from the innermost B K , call the weighted neighbor sampling function NW(u) to sample the neighbors of all nodes u in B K , and the sampling result is merged with B K to obtain B K-1 , and so on to obtain the small batch weighted sampling set B 0 .

[0089] Stage 2: In the 8th line of the algorithm, the feature vectors of all nodes in the input graph are first initialized. Lines 9 - 15 also execute two nested loops. The first loop performs weighted aggregation processing layer by layer from the outside to the inside, and the second loop performs weighted aggregation on each node u in this layer. In the second loop, in line 11, the aggregation function is used to perform weighted aggregation on the information of the neighbor nodes of u, that is, first multiply the feature information of the neighbor nodes by the node edge weight (the edge between node u and this neighbor node), and then aggregate the products. In line 12, the neighbor aggregation information and the current embedding information of node u are concatenated, and then the embedding information of u itself is updated through a non - linear transformation. In line 13, normalization processing calculation is performed on the node embedding information. In line 16, finally, the result after two - layer loop calculation is used as the output z u , that is, the final node embedding information.

[0090] It should be noted that the node edge weight W’ here is different from the weight matrix W. The node edge weight W’ is determined when the weighted graph is established, while the weight matrix W is a parameter to be learned by the graph neural network.

[0091] In the embodiments of this case, the context scenario information is modeled and inferred from the access record dataset, and this dataset is unlabeled sample data. Therefore, the learning of the context scenario belongs to the form of unsupervised learning. The graph - based loss function hopes that adjacent nodes have similar vector representations, and at the same time makes the representations of separated nodes as different as possible. The loss function can adopt the loss function of the unsupervised learning method of GraphSAGE, and the function is as follows:

[0092]

[0093] where z u is the output vector of node u (i.e., the node embedding), v is the neighbor node that appears near u through the weighted sampling algorithm, P n (v) is the probability distribution of negative sampling, σ is the sigmoid function, z vn is the negative sampling distribution of node v, and Q is the number of negative samples. In the sampling data, points that are far from the target node or unreachable nodes are called negative samples. The loss function means that when the similarity of the embeddings of adjacent nodes is as large as possible, the expected similarity of the embeddings of non - adjacent nodes is made as small as possible. Weighted sampling is performed on the nodes in graph G to obtain the positive sampling set {z u} and the negative sampling set {z vn} respectively, and the loss function is calculated. When the function value is smaller, the output effect is better. According to the loss function value, the parameter weight matrix W k and the parameters inside the aggregation function are adjusted by the stochastic gradient descent method. As the loss function value decreases, the node embedding information better contains the node feature information, the relationship between the node and its neighbors, and the tightness of the relationship.

[0094] Further, based on the above method, an embodiment of the present invention further provides a context-aware system based on weighted GraphSAGE for big data access control, including: a data acquisition module, a problem modeling module, and an analysis and clustering module, where,

[0095] The data acquisition module is used to obtain the original context information describing the access control entity data from the user access record dataset;

[0096] The problem modeling module is used to construct a weighted graph of the context information, and model the access control context-aware problem as an inductive learning clustering problem represented by graph nodes. Among them, in the weighted graph, the context information in the user access record is used as the graph node, and the relationship between the context information is used as the edge. Each node consists of multiple elements, and each element is the same as the corresponding position element of the access record, and the edge weight is set according to the node elements;

[0097] The analysis and clustering module is used to use the GraphSAGE graph neural network model to learn the graph node representation and obtain the node embedding information, where the node embedding information at least includes node features, node relationship features, and node relationship strength; and use the clustering algorithm to perform inductive clustering on the node embedding information and generate scenario information composed of a set of context information, and automatically generate access control rules for dynamically controlling access requests from the scenario information.

[0098] Further, an embodiment of the present invention further provides a data access control method, which uses the above context-aware method to generate access control rules, and matches the user access request with the generated access control rules and makes an access control permission decision. Among them, the matching process includes: First, convert the user access request information into a context vector and obtain the context information; then, calculate the distance between the context information and each clustering center; then, sort the multiple distance values, select the clustering center with the closest distance as the matching item, and use the access control rule corresponding to the matching item to respond to the user access request.

[0099] To verify the effectiveness of the solution in this case, the following further explains with experimental data:

[0100] To verify the availability of the WGraphSAGE model and algorithm in the solution of this case, four datasets of user access information resources, namely Appusage, Carat, Frappe, and WS-Dream, are selected for experimental analysis.

[0101] Each access record in the dataset file is regarded as a node, and corresponding fields are selected from each access record as the input data for each node. The data includes a node serial number and node elements. The node serial number is a sequential number ID (i.e., node number VID), and the node elements are a <u, t, l, r> quadruple, including user category User, time Time, location Location, and resource information Resource. Taking Appusage as an example, the node input data is obtained from App_Usage_Trace.txt. Each element of the graph node is represented by a vector. If the vector length of the element is set to l bits, then the total length of the node vector is 4l bits. For example, the vector lengths of the elements in the Appusage dataset are 100 bits respectively. Element edges are established. If the values of the corresponding elements of two nodes are the same (such as the same user category User, location Location, or resource information Resource), then an element edge is established for the corresponding elements of the two nodes (such as the same location); if the times of two nodes are close, then an element edge is established for the time elements of the two nodes. All nodes are scanned in sequence to establish element edges related to the corresponding elements of all nodes. Node edges are established and the weights of the node edges are calculated. In the experiment, it is assumed that the element weights are the same, and the element weight WoE is 1 / 4 = 0.25. Based on this, the edge weights between any two nodes are calculated to obtain a weighted graph.

[0102] Since the current CAAC model does not support automatic context awareness yet, the experiment is divided into three steps: automatic context awareness, context-aware access control decision-making, and ablation experiment. Among them, for automatic context awareness, the WGraphSAGE model in the embodiments of this case is compared with graph representation models Node2vec, GCN, and GraphSAGE, and the running results of this model with different aggregation functions (according to the aggregation function selection method, the WGraphSAGE model is further divided into four cases: mean, seq, maxpool, and meanpool) are compared. After learning through the graph neural network, the clustering results and effects of the graph nodes are analyzed, and the running time, clustering evaluation metrics, etc. are compared. For context-aware access control decision-making, the access control effects of the WGraphSAGE model are compared with those of the OntCAAC and FB-CAAC models. For the ablation experiment, the impacts of weighted neighbor sampling and weighted aggregation algorithms in the WGraphSAGE model on the clustering results and access control results are analyzed.

[0103] In a computer hardware environment with an Intel Core i9-10980XE 3.00GHz CPU and 256GB DDR4-3200 SDRAM memory, the Node2vec, GCN, GraphSAGE, and WGraphSAGE models were implemented through TensorFlow. The parameter settings were as follows: the number of Mini-Batches was 512, the number of layers K for graph aggregation was 2, the number of samples S 1 = 25, S 2 = 10, the output vector Out_dim1 = 128, Out_dim2 = 128, Out_dim = 256, and the number of negative samples S neg = 20.

[0104] I. Context-Aware Experiment

[0105] Select different numbers of graph nodes, and execute models such as Node2vec, GCN, GraphSAGE, and WGraphSAGE respectively. Statistically analyze the relationship between the number of nodes and the running time, as shown in Figure 10As shown in the figure. In the figure, (A, B, C) respectively represent the time comparison of running four models, namely Node2vec, GCN, GraphSAGE (using the Mean function) and WGraphSAGE (using the weighted Mean function), on three datasets of Appusage, Carat, and Frappe (excluding the time for building the graph). (D) in the figure represents the time comparison of running four aggregation functions of GraphSAGE and WGraphSAGE respectively on the WS-Dream dataset. (A, B, C) show that the operation time of the four models shows a polynomial growth trend as the number of nodes increases. When the number of nodes is the same, the operation times of GraphSAGE and WGraphSAGE are both higher than those of the Node2vec and GCN models. This is because, without predicting the embedding information of unknown nodes, the overhead of the aggregation functions used by GraphSAGE and WGraphSAGE is higher than the skip-gram method of Node2vec and the frequency-domain graph convolution method of GCN. (D) shows that when the same aggregation function is used, the operation time of WGraphSAGE is slightly lower than that of GraphSAGE. For example, Time (WGraphSAGE-mean’) < Time(GraphSAGE-mean), Time(WGraphSAGE-maxpool’) < Time (GraphSAGE-maxpool). This is because the GraphSAGE model needs to perform random sampling during each iteration calculation, while after the weighted neighbor sampling of the WGraphSAGE model, a list of neighbor nodes closely related to each node and the node edge weights are obtained, and repeated sampling is not required during subsequent calculations, reducing the time overhead of sampling. Although the multiplication calculation of the edge weight and the aggregation vector is added in the aggregation function of this model, the multiplication calculation complexity is at the constant level and does not increase the overhead. In the WGraphSAGE model, the time overhead of the weighted aggregation functions is Time(mean’) < Time(meanpool’) < Time(maxpool’) < Time(LSTM’), which is caused by the time complexity of the four aggregation functions themselves, and the weighted operation does not affect their original time complexity.

[0106] Loss function J G (z u ) means that when the similarity of the embeddings of adjacent nodes is as large as possible, the expected similarity of the embeddings of non-adjacent nodes is ensured to be as small as possible. The loss function value Loss represents the effect of the output information of the graph nodes after learning by the graph neural network. The lower the Loss value, the better the learning effect. To determine the appropriate number of iterations to achieve a lower loss function value, the relationship between the loss function value Loss and the number of iterations epochs after the model is executed was statistically analyzed.Figure 11 Among them, (A - D) respectively represent the changes in the loss function values after several iterative calculations of four models, namely Node2vec, GCN, GraphSAGE, and WGraphSAGE, executed on four datasets, Appusage, Carat, Frappe, and WS - Dream (the learning rate is set to μ = 0.0001). In (A), the changes in the loss function values after iterative calculations of WGraphSAGE using four different weighted aggregation functions are also added. For each model, as the number of epochs increases, the change in the loss function value Loss. In the same dataset, the Loss value of WGraphSAGE is lower than that of Node2vec, GCN, and GraphSAGE. When the number of epochs = 2, after 2 iterations, the loss function value Loss of WGraphSAGE has a significant decrease, and then as epochs increase, the Loss value basically stabilizes. This is because after adding the node edge weight parameter, WGraphSAGE improves the similarity of adjacent nodes and the distinguishability between non - adjacent nodes during neighbor sampling and aggregation calculations, so the loss function value is lower than that of similar models. In addition, since WGraphSAGE only performs weighted sampling and aggregation calculations on local neighbor nodes, after several iterative calculations for all nodes, the parameters of its node embedding method will reach a relatively stable value, and correspondingly, the Loss value will also be relatively stable, and even continuing multiple iterations has little impact on it.

[0107] The K-Means++ algorithm is used to perform clustering operations on the output vectors of graph nodes, and it is necessary to determine a suitable value of the number of clusters k_clusters in advance. In unsupervised learning models, a popular method for determining the number of clusters is the Elbow method. This method is based on the within-cluster sum of squares. If adding another cluster does not result in better data modeling (i.e., the elbow point of the graph), the current k_clusters will be selected as the value of the number of clusters. However, the disadvantage of this method is that it is difficult to select the appropriate point where the true bend occurs. To overcome the problems of the Elbow method, the Silhouette (SIL), Calinski-Harabaz Index (CHI), and Davies-Bouldin Index (DBI) are selected to comprehensively calculate and determine the number of clusters, abbreviated as k_clusters. The Silhouette coefficient is calculated using the average within-cluster distance (a) and the average nearest inter-cluster distance (b) of each sample. The Silhouette coefficient of a single sample is (b - a) / max(a, b), and the average Silhouette coefficient of a group of samples is the average of the Silhouette coefficients of the samples in the group. The measurement range of the average Silhouette coefficient is [-1, 1], with the best value being 1 and the worst value being -1. Theoretically, when the average Silhouette coefficient value is the highest, the clustering result is the best. The Calinski-Harabaz Index (CHI) is obtained by the ratio of the mean between-clusters dispersion to the within-cluster dispersion (where the dispersion is defined as the sum of the squared distances). The higher the index, the better the clustering effect. DBI represents the average "similarity" between clusters, where the similarity is a measure that compares the distance between clusters with the size of the clusters themselves. The lowest score of DBI is 0, and the closer the value is to 0, the better the clustering result. The process of determining the number of clusters is as follows. First, for different values of the number of clusters k_clusters, run the clustering algorithm and calculate the values of SIL, CHI, and DBI respectively. Then, considering the SIL, CHI, and DBI metrics comprehensively, select the appropriate value of k_clusters such that relatively optimal values are obtained for all three metrics. Figure 12Among them, (A), (B), (C), and (D) respectively represent the relationship between the number of clusters and metrics of the four datasets Appusage, Carat, Frappe, and WS-Dream. For example, in (A), when k_clusters = 203, CHI = 21506.02 and DBI = 0.36, reaching the optimal values. At this time, SIL = 0.95, reaching the maximum value. From the SIL curve, there is an obvious inflection point here, and the tangent slope drops significantly. Subsequently, as k_clusters increases, SIL shows a downward trend. Similarly for (C), when k_clusters = 55, the three metrics reach relatively good values. This is because when the number of clusters k_clusters takes an appropriate value, qualitatively, the nodes within the same cluster are more concentrated, and those between different clusters are more dispersed; quantitatively, the SIL, CHI, and DBI metrics will reach the optimal or relatively good values, and at this time, the clustering of the context reaches the optimal or relatively good state. Therefore, the results of clustering can be comprehensively judged well through these 3 metrics.

[0108] The results after clustering are represented by t-SNE dimensionality reduction as follows Figure 13 as shown. Among them, A, B, C, and D respectively correspond to the sub-clustering results of the four datasets Appusage, Carat, Frappe, and WS-Dream, and the four columns in the figure respectively correspond to the clustering results of the four models Node2vec, GCN, GraphSAGE, and WGraphSAGE. Since the number of clusters corresponding to the entire access record dataset is relatively large, here, to briefly indicate the number of clusters, 5000 records are randomly selected as the analysis object. Node2vec does not depict the relationship between nodes and does not achieve the clustering effect. GCN clusters the nodes. Its advantage is that the clusters are relatively dispersed, that is, the distance between clusters is large, but the disadvantage is that the distance within the clusters is large, and the number of clusters is excessive. In the clustering results of GraphSAGE, the advantage is that the number of clusters is significantly reduced and the distance within the clusters is small, but the disadvantage is that the distance between clusters is small, which is not conducive to distinguishing different context scenarios. The aggregation effect of WGraphSAGE is significantly better than that of the Node2vec, GCN, and GraphSAGE models. The distance within the clusters is small, the distance between clusters is large, and the number of clusters is appropriate. This shows that: through the weighted neighbor sampling and weighted aggregation of WGraphSAGE, the embedding information of graph nodes can better describe the sample characteristics of access record information and depict the relationship between them. Associated nodes can be better clustered in the same cluster, and the clusters composed of these nodes correspond to relevant context scenario information, realizing the modeling of context information and the reasoning of context scenarios, thus verifying the feasibility of the WGraphSAGE model.

[0109] II. Access Control Decision Experiment

[0110] First, context-aware access control rules are formulated for four datasets respectively. For the same dataset, the number of rules generated by the OntCAAC and FB-CAAC models is more than that of the WGraphSAGE model, as shown in (A) of Figure 14 . This is because the baseline methods OntCAAC and FB-CAAC models formulate rules manually in an ontology way, resulting in more redundant rules. In contrast, the WGraphSAGE model in the embodiments of this case automatically generates context-aware access control rules from the clustering results, reducing the amount of redundant information.

[0111] Secondly, the generated rule set is used for control decisions during user access. In the embodiments of this case, the test data part in the dataset is used as a simulated access request to test the total response time from when the system receives the access request to generating the access control decision result. Ignoring the network transmission delay, the decision response time (Tr) = context-aware time (Tca) + access control rule matching time (Tacm). The OntCAAC, FB-CAAC, and WGraphSAGE methods are successively used to test the 4 datasets, and the differences in access control response times are compared. The test results are shown in (B) of Figure 14 .

[0112] The experimental results show that the response time of the access control decision of the WGraphSAGE method is significantly lower than that of the two baseline models. The reasons are as follows: on the one hand, before the WGraphSAGE model makes an access control decision, the rules have been generated by the context-aware method. Its context-aware time (Tca) is equal to the time of converting the access control request into a context information feature vector, and the time-consuming is much lower than the corresponding time of the other two models. On the other hand, the access control rule matching process of the WGraphSAGE model is essentially a process of selecting the clustering center node with the shortest distance. This time (Tacm) is the time of calculating the distance between the request vector and the feature vectors of each clustering center node and sorting them. The test results show that the access control rule matching time of the WGraphSAGE model is basically the same as the corresponding time of the two baseline methods. In addition, the number of access control rules of the WGraphSAGE model is less than that of the baseline methods, thus reducing the number of rule matching times. Therefore, the response time of the WGraphSAGE access control decision is lower than that of the baseline methods.

[0113] III. Ablation Experiment

[0114] To address the problems of automatic context information perception, automatic context relationship modeling and reasoning in big data access control, based on the baseline method GraphSAGE, a graph neural network model WGraphSAGE was constructed using weighted neighbor sampling and weighted aggregation algorithms. Ablation experiments were conducted to verify the importance of the weighted neighbor sampling and weighted aggregation steps for the WGraphSAGE model and their roles in solving access control problems.

[0115] In the ablation experiments, four methods, namely GraphSAGE, GraphSAGE+WS, GraphSAGE+WA, and WGraphSAGE, were used to conduct context awareness experiments and access control decision experiments on four datasets for comparison, as Figure 15 shown. A, B, C, and D correspond to the experimental results of the Appusage, Carat, Frappe, and WS-Dream datasets respectively. Among them, GraphSAGE+WS means that the GraphSAGE model adds weighted neighbor sampling on the basis of weighted graph transformation but does not perform the weighted aggregation step. GraphSAGE+WA means that the GraphSAGE model adds weighted aggregation of the sampled neighbor nodes without weighted sampling of neighbors on the basis of weighted graph transformation, still using the original random neighbor sampling method of GraphSAGE. SIL and DBI represent the measurement indicators of the clustering effects of each model. Numbers of Rules represents the number of access control rules automatically generated by each model from context clusters, and Response Time represents the response time of each model for access control decisions.

[0116] The results show that the WGraphSAGE model has the best performance in all indicators in the experiments on the four datasets. Adding only one of the weighted neighbor sampling or weighted aggregation steps to the original GraphSAGE model improves the clustering effect and access control results compared to the original GraphSAGE model, but is significantly weakened compared to the WGraphSAGE model. For example, for GraphSAGE+WS and GraphSAGE+WA, when training and testing on the four datasets, the experimental results of the three indicators of DBI, Numbers of Rules, and ResponseTime are all higher than the corresponding values of WGraphSAGE, and the SIL indicator result is lower than the corresponding value of WGraphSAGE. Therefore, the results of this ablation experiment show that the weighted neighbor sampling and weighted aggregation steps in the WGraphSAGE model play important roles in context information clustering and access control decision-making, directly affecting the experimental results and improving the context awareness effect on the target dataset and the efficiency of access control decision-making based on context.

[0117] Unless otherwise specifically stated, the relative steps, numerical expressions, and numerical values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0118] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.

[0119] The units and method steps of the examples described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.

[0120] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software functional module. The present invention is not limited to any specific form of the combination of hardware and software.

[0121] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, and not to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A context-aware method based on weighted GraphSAGE for big data access control, characterized in that, it includes the following content: Obtain the original context information describing the access control entity data from the user access record dataset; Construct a weighted graph of context information and model the access control context awareness problem as an inductive learning clustering problem represented by graph nodes. In the weighted graph, the context information in the user access record is used as graph nodes, and the relationships between context information are used as edges. Each node consists of multiple elements, each element being the same as the corresponding position element in the access record, and the edge weights are set according to the node elements. When modeling the access control context awareness problem, a multi-graph G = (V, E) is used to model the context information features and relationships, and each user access record X i is used as a node of the graph V i , and each graph node V i consists of m elements. An element edge is established between two nodes whose corresponding element distances in the node are less than the preset element distance threshold. Each element is associated with an element weight, and the edge weight is the sum of the corresponding element weights; Use the GraphSAGE graph neural network model to learn the graph node representation, use the sampling function to sample the neighbors of the nodes in the graph from the inside out, aggregate the sampled nodes in the graph one by one from the outer layer to the inner layer, and generate the embedding information of the nodes. The probability of each neighbor node being selected in the sampling function is set to the ratio of the edge weight of the corresponding node to the sum of the edge weights of all neighbor nodes. The node embedding information at least includes node features, node relationship features, and node relationship strength; and use the clustering algorithm to inductively cluster the node embedding information and generate the scenario information composed of the set of context information, and automatically generate the access control rules for dynamically controlling the access requests from the scenario information.

2. The context-aware method based on weighted GraphSAGE according to claim 1, characterized in that, the context information is represented as a multi-tuple composed of access control entity data elements, where the access control entity data elements at least include subject status information, object status information, environmental status information, operation request information, relationship status information, and status information of the access purpose.

3. The context-aware method based on weighted GraphSAGE according to claim 1, characterized in that, in the neighbor sampling of the nodes in the graph using the sampling function, first set the initial node set and obtain all the first-order neighbor nodes of the initial nodes; then, sort the first-order neighbor nodes according to the node edge weights between the nodes and the first-order neighbor nodes, and perform weighted sampling from the neighbor nodes; then, perform weighted sampling on the sampled nodes for their respective neighbor nodes, and sequentially perform weighted sampling layer by layer according to the preset sampling layers to obtain the weighted neighbor node set.

4. The context-aware method based on weighted GraphSAGE according to claim 1, characterized in that, in the aggregation operation, first multiply the representation information of the outermost neighbor nodes of each node by the node weight and aggregate the products; then connect the aggregated value of the neighbor nodes with the current representation information of the second outermost layer nodes and obtain the representation information of the second outermost layer nodes through the transformation of the non-linear activation function; then, continue to perform weighted aggregation using the representation information of the second outermost layer neighbor nodes until the aggregation operation reaches the innermost layer nodes.

5. The context-aware method based on weighted GraphSAGE according to claim 1, characterized in that, use the K-means clustering algorithm to perform clustering operations on the feature vectors of the node embedding information, use the node clusters to cluster the access records with the same or similar context information, so that each node cluster after clustering corresponds to a context scenario, and the feature information of the node cluster represents the feature information of the corresponding context scenario.

6. The context-aware method based on weighted GraphSAGE according to claim 1, characterized in that, The access control rules automatically generated from the scenario information are represented as <RuleID, UserClass, Contextscenario, ClusterCenter, ControlResult>. Among them, RuleID represents the rule number, and this rule number is the same as the context scenario serial number. UserClass represents the user classification identifier represented by the user role and characteristic attributes. Contextscenario represents the description of the clustering result of the corresponding context information. ClusterCenter represents the characteristic vector of the clustering center. ControlResult represents the access control result.

7. A context-aware system based on weighted GraphSAGE for big data access control, Characterized in that, It includes: a data acquisition module, a problem modeling module, and an analysis and clustering module. Among them, The data acquisition module is used to obtain the original context information describing the access control entity data from the user access record dataset; The problem modeling module is used to construct a weighted graph of context information and model the access control context awareness problem as an inductive learning clustering problem represented by graph nodes. In the weighted graph, the context information in the user access record is used as graph nodes, and the relationships between context information are used as edges. Each node consists of multiple elements, each element being the same as the corresponding position element in the access record, and the edge weights are set according to the node elements. When modeling the access control context awareness problem, a multi-graph G = (V, E) is used to model the context information features and relationships, and each user access record X i is used as a node of the graph V i , and each graph node V i consists of m elements. An element edge is established between two nodes whose corresponding element distances in the node are less than the preset element distance threshold. Each element is associated with an element weight, and the edge weight is the sum of the corresponding element weights; The analysis and clustering module is used to learn the graph node representation by using the GraphSAGE graph neural network model, sample the neighbors of the nodes in the graph from the inside out by using the sampling function, aggregate the sampled nodes in the graph one by one from the outer layer to the inner layer, and generate the embedding information of the nodes. The probability of each neighbor node being selected in the sampling function is set to the proportion of the edge weight of the corresponding node in the sum of the edge weights of all neighbor nodes. The node embedding information at least includes node features, node relationship features, and node relationship strength; and uses the clustering algorithm to inductively cluster the node embedding information and generate the scenario information composed of a set of context information, and automatically generates the access control rules for dynamically controlling the access request from the scenario information.

8. A data access control method, Characterized in that, Generate access control rules based on the method described in claim 1, match the user access request with the generated access control rules and make an access control permission decision. Among them, the matching process includes: First, convert the user access request information into a context vector and obtain the context information; then, calculate the distance between the context information and each clustering center; then, sort the multiple distance values, select the clustering center with the closest distance as the matching item, and use the access control rule corresponding to the matching item to respond to the user access request.