Knowledge Graph Access Control System
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-07-19
- Publication Date
- 2026-08-14
Smart Images

Figure 0007905227000001 
Figure 0007905227000002 
Figure 0007905227000003
Abstract
Description
Technical Field
[0001] The present invention relates to access control to a knowledge graph.
Background Art
[0002] A knowledge graph systematically connects various kinds of knowledge and represents it in a graph structure. The graph structure is represented by a set of nodes and a set of arcs, and an arc is represented by a start node and an end node. In many cases in a knowledge graph, information is linked to nodes and arcs.
[0003] Recently, digital transformation has accelerated in various industries, and it is required to cope with rapid business changes. Therefore, it is useful to systematically organize each customer case and utilize it to obtain suggestions for other customer cases. For example, this corresponds to organizing the flow of funds, information, etc. between stakeholders in each business, or organizing the relationship between the business issues of customers and the individual technical application issues. Such information is summarized as a knowledge graph.
[0004] For example, Patent Document 1 holds a knowledge graph as a hierarchical (tree) structure of subgraphs, manages the access rights to each, and displays an appropriate knowledge graph for the user. Patent Document 1 manages the subgraph structure of a knowledge graph as a hierarchical structure and controls the disclosure and non-disclosure of each node (subgraph).
Prior Art Documents
Patent Documents
[0005]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] When dealing with individual cases, the disclosure of sensitive information becomes a bottleneck, hindering the use of knowledge graphs across different organizations. Therefore, when utilizing knowledge graphs that contain sensitive case information, access control based on information granularity is necessary. However, redefining the level of knowledge graphs that can be disclosed from scratch by humans is a time-consuming and costly process. Furthermore, managing knowledge graphs with different disclosure levels also incurs significant costs.
[0007] One aspect of the present invention has been made in view of these circumstances, and aims to provide an efficient access control technology for knowledge graphs that can promote the utilization of knowledge graphs. [Means for solving the problem]
[0008] A representative example of the invention disclosed in this application is as follows: The knowledge graph access control system includes a computing device and a memory device, the memory device storing ambiguation structure information that defines the inclusion relationships between elements of the knowledge graph with different degrees of ambiguation, and access control information that manages the user's access rights to each element included in the ambiguation structure information; the computing device acquires a knowledge graph to be ambiguized, refers to the ambiguation structure information and the access control information to ambiguize the target knowledge graph for a first user to generate an ambiguized knowledge graph, and in the ambiguation of the target knowledge graph, converts the original elements included in the target knowledge graph into ambiguized elements to which the first user has access rights in the access control information and which include the original elements in the ambiguation structure information. [Effects of the Invention]
[0009] According to one aspect of the present invention, efficient access control to the knowledge graph can promote the utilization of the knowledge graph. Other issues, configurations, and effects will be clarified by the following description of the embodiments. [Brief explanation of the drawing]
[0010] [Figure 1] This is a diagram showing an example of the configuration of a computer in an embodiment of this specification. [Figure 2] An example of a knowledge graph is shown. [Figure 3A] An example of the configuration of personal management information is shown. [Figure 3B] An example of the configuration of team management information is shown. [Figure 3C] An example of the configuration of personal and team relationship information is shown. [Figure 4A] An example of the configuration of graph management information is shown. [Figure 4B] An example of the configuration of node information is shown. [Figure 4C] An example of the configuration of arc information is shown. [Figure 4D] An example of the configuration of graph and node relationship information is shown. [Figure 4E] An example of the configuration of graph and arc relationship information is shown. [Figure 5] An example of a node generalization structure is schematically shown. [Figure 6A] An example of the configuration of node information within a node generalization structure is shown. [Figure 6B] An example of the configuration of arc information within a node generalization structure is shown. [Figure 7] An example of the configuration of an arc generalization structure is schematically shown. [Figure 8A] An example of the configuration of node information within an arc generalization structure is shown. [Figure 8B] Arc information within an arc generalization structure is shown. [Figure 9A] An example of the configuration of node access management information is shown. [Figure 9B] An example of the configuration of arc access management information included in access control information is shown. [Figure 10] A flowchart of an example of an algorithm for generating a generalized knowledge graph is shown. [Figure 11] An example of a GUI screen for knowledge graph generalization is shown. [Figure 12] An example of the configuration of a computer in Example 2 is shown. [Figure 13] Shows a flowchart of an example of an information granularity evaluation algorithm executed by the information granularity evaluation unit. [Figure 14A] A diagram for explaining an example of the set s[n]. [Figure 14B] A diagram for explaining the number of unique teams having access rights to the target node. [Figure 14C] A diagram for explaining the number of unique teams having access rights to the nodes on the path between the target node and the node of the original source of fuzzification. [Figure 15] Shows a configuration example of the computer of Example 3. [Figure 16] Shows a flowchart of an example of a fuzzification DAG arc recommendation algorithm executed by the fuzzification information recommendation unit. [Figure 17] Shows a configuration example of the computer of Example 4. [Figure 18] Shows a flowchart of an example of a node access right recommendation algorithm for a fuzzification DAG executed by the access control information recommendation unit. [Figure 19] Shows a flowchart of an example of an algorithm for automatically generating a fuzzified knowledge graph and recommending access rights. [Figure 20] Shows an example of a GUI screen for predicting access rights to a knowledge graph described with reference to FIG. 19. [Figure 21] Shows an example of a GUI screen for predicting access rights to a knowledge graph described with reference to FIG. 19.
Mode for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. However, the present invention is not construed as being limited to the description of the embodiments shown below. It will be easily understood by those skilled in the art that the specific configuration can be changed without departing from the spirit or gist of the present invention.
[0012] In the configuration of the invention described below, identical or similar components or functions are denoted by the same reference numerals, and redundant descriptions may be omitted. The designations "1st," "2nd," "3rd," etc., used in this specification are for identifying components and do not necessarily limit their number or order.
[0013] The system in one embodiment of this specification may be a physical computer system (one or more physical computers) or a system built on a cloud infrastructure or other computing resource group (multiple computing resources). The computer system or computing resource group may include one or more interface devices (e.g., including communication devices and input / output devices), one or more storage devices (e.g., including memory (main memory) and auxiliary storage devices), and one or more arithmetic units.
[0014] When a function is realized by the execution of a program containing instruction codes by an arithmetic unit, the defined processing is carried out using memory and / or interface devices as appropriate, so the function may be at least a part of the arithmetic unit. The processing described with the function as the subject may be processing performed by the arithmetic unit or a system having that arithmetic unit. The program may be installed from the program source.
[0015] The program source may be, for example, a program distribution computer or a computer-readable storage medium (e.g., a computer-readable non-transient storage medium). The descriptions of each function are examples, and multiple functions may be combined into one function, or one function may be divided into multiple functions.
[0016] The positions, sizes, shapes, and ranges of each component shown in the drawings, etc., may not represent the actual positions, sizes, shapes, and ranges, etc., in order to facilitate understanding of the invention. Therefore, the present invention is not limited to the positions, sizes, shapes, and ranges, etc., disclosed in the drawings, etc.
[0017] One embodiment of this specification manages the elements of a knowledge graph based on information granularity and discloses only the information at a level appropriate to the user based on that information granularity. Furthermore, the system of this embodiment provides support functions such as recommending candidates for element ambiguation and predicting access rights. This embodiment makes it possible to display the knowledge graph at an appropriate information granularity according to the user. This allows the results of analyzing individual cases as a knowledge graph to be shared with parties to whom sensitive information cannot be shown, thereby expanding the scope of utilization of accumulated knowledge.
[0018] Higher-grained information is more ambiguous. In other words, the higher the granularity, the higher the degree of ambiguity. The process of transforming low-granularity information into higher-granularity information is called ambiguation. That is, one embodiment of this specification ambiguizes the descriptions of the elements of the knowledge graph, i.e., nodes and the arcs connecting nodes, in the control of access to the knowledge graph. The ambiguized description is a more semantically broader description that encompasses the description before ambiguation.
[0019] One embodiment of the system described herein manages the ambiguation of knowledge graph elements using a Directed Acyclic Graph (DAG). This allows for more appropriate and efficient management of the ambiguation structure of knowledge graph elements. The ambiguation structure defines the inclusion relationships between elements. The ambiguation structure may be managed by other formats. [Examples]
[0020] Figure 1 shows an example of the configuration of a computer according to the embodiments of this specification. The computer 100 is, for example, a personal computer, server, or workstation, and comprises a CPU (Central Processing Unit) 101, memory 102, auxiliary storage device 103, input device 104, output device 105, and communication device 106. Each hardware element is connected to the others via a bus 107.
[0021] The CPU 101 is an arithmetic unit that executes programs stored in memory 102. By executing processes according to the program, the CPU 101 operates as a functional unit (module) that realizes a specific function. In the following explanation, when a functional unit is the subject of a description of its processing, it indicates that the CPU 101 is executing the program that realizes that functional unit.
[0022] Memory 102 is a storage device such as DRAM (Dynamic Random Access Memory) that stores programs executed by CPU 101 and information used by CPU 101. Memory 102 also includes a work area temporarily used by CPU 101. The programs stored in memory 102 will be described later.
[0023] The program and information stored in memory 102 may also be stored in auxiliary storage device 103. In this case, the CPU 101 reads the program and information from auxiliary storage device 103, loads it into memory 102, and executes the program stored in memory 102.
[0024] The auxiliary storage device 103 is a storage device such as an HDD (Hard Disk Drive) or SSD (Solid State Drive) that permanently stores data. The information stored in the auxiliary storage device 103 will be described later. The auxiliary storage device 103 may also be a drive for storage media such as a CD-R (Compact Disc Recordable), DVD-RAM (Digital Versatile Disk-Random Access Memory), or silicon disk. In this case, the information and programs are stored on the storage media.
[0025] The input device 104 is, for example, a keyboard, mouse, scanner, and microphone, and is a device for inputting data into the computer 100. The output device 105 is, for example, a display, printer, and speaker, and is a device for outputting data from the computer 100 to the outside. The communication device 106 is, for example, a device for communicating via a network such as a LAN (Local Area Network).
[0026] Note that some of the components shown in Figure 1, such as the input device 104, output device 105, and / or communication device 106, may be omitted from the computer 100. The computer 100 may be connected to other terminals via a network, receive user input from those terminals via their input devices, transmit processing results to those terminals, and present them to the user on the output device 105 of those terminals.
[0027] The information stored in the auxiliary storage device 103 and the programs stored in the memory 102 will be described below. The auxiliary storage device 103 stores user management information 131, knowledge graph information 132, node ambiguation structure information 133, arc ambiguation structure information 134, access control information 135, and various programs 151.
[0028] User management information 131 manages individual users and the teams to which they belong. Knowledge graph information 132 includes multiple knowledge graphs. In the example described below, knowledge graph information 132 includes a case-based knowledge graph.
[0029] Node ambiguation structure information 133 is information referenced to ambiguize nodes in knowledge graph information 132. In the example described later, node ambiguation structure information 133 has a DAG structure. Arc ambiguation structure information 134 is information referenced to ambiguize arcs in knowledge graph information 132. In the example described later, arc ambiguation structure information 134 has a DAG structure.
[0030] Access control information 135 is information referenced to control a user's access to elements of the knowledge graph. Program 151 contains various programs that are executed on the CPU 101 loaded into memory 102.
[0031] Memory 102 stores programs that implement the access control information setting unit 121 and the accessibility information output unit 122. These programs are included in program 151 and are loaded into memory 102 for execution by the CPU 101. Note that, for each functional unit of the computer 100, multiple functional units may be combined into a single functional unit, or a single functional unit may be divided into multiple functional units according to its function.
[0032] Alternatively, this embodiment may be implemented as a computer system in which the various functional units of computer 100 are distributed among multiple computers. For example, a computer system consisting of a computer having an access control information setting unit 121, a computer having an accessibility information output unit 122, and a storage system for storing each piece of information can be considered.
[0033] The contents of the information stored in the auxiliary storage device 103 are described below. For ease of explanation, an example of a knowledge graph shown in Figure 2 is used. The information described below includes the information in the knowledge graph 200 shown in Figure 2. The knowledge graph 200 includes the E Hospital node 201, the C Insurance node 202, and the insured person node 203. "E Hospital," "C Insurance," and "insured person" are descriptions of these nodes.
[0034] Knowledge Graph 200 further includes arcs 204 through 209. Each arc proceeds from a source node to a target node. Each arc is given a description, which is shown next to each arc in Figure 2. For example, the description for arc 208, which goes from hospital node 201 to insured person node 203, is "cancer treatment," and the description for arc 209, which goes from insured person node 203 to hospital node 201, is "visit." Note that there may be arcs that only go in one direction between nodes, and there may be no description for an arc.
[0035] Figures 3A to 3C show examples of the structure of information included in user management information 131. Figure 3A shows an example of the structure of personal management information 300. Personal management information 300 includes a record ID field 301 and a personal name field 302 of the user who can use the system. Figure 3B shows an example of the structure of team management information 310. Each user belongs to one of the teams. Team management information 310 includes a record ID field 311 and a team name field 312.
[0036] Figure 3C shows an example of the configuration of individual and team relationship information 320. Individual and team relationship information 320 indicates individual users belonging to each team. Individual and team relationship information 320 includes an individual ID (PID) field 321 and a team ID (TID) field 322. The ID in the PID field 321 matches the ID in the record ID field 301 of the individual management information 300. The ID in the TID field 322 matches the ID in the record ID field 311 of the team management information 310.
[0037] Figures 4A to 4E show examples of the structure of the information contained in the knowledge graph information 132. Figure 4A shows an example of the structure of the graph management information 330. The graph management information 330 manages the knowledge graph contained in the knowledge graph information 132. The graph management information 330 includes a record ID field 331, a knowledge graph name field 332, and a knowledge graph owner field 333.
[0038] Figure 4B shows an example of the configuration of node information 340. Node information 340 is information about the nodes of the entire knowledge graph managed by graph management information 330. Node information 340 includes a record ID field 341, a description (DESC) field 342, and a node hierarchy ID field (NHID) 343. The record ID field 341 shows the ID of the node in the entire knowledge graph. The description field 342 shows the description assigned to each node. The node hierarchy ID field 343 shows the node ID within the node ambiguation structure described later.
[0039] Figure 4C shows an example of the structure of arc information 350. Arc information 350 is information about arcs in the entire knowledge graph managed by graph management information 330. Arc information 350 includes a record ID field 351, a description field 352, a source (SRC) field 353, a target (TGT) field 354, and an arc hierarchy ID field (AHID) 355. The record ID field 351 shows the ID of the arc in the entire knowledge graph. The description field 352 shows the description given to each arc. The source field 353 and target field 354 show the source and target of the arc. The arc hierarchy ID field 355 shows the ID of the node that represents the arc in the arc ambiguation structure described later.
[0040] Figure 4D shows an example of the configuration of the graph and node relationship information 360. The graph and node relationship information 360 shows a graph containing each node. The graph and node relationship information 360 includes a graph ID (GID) field 361 and a node ID (NID) field 362. The ID in the GID field 361 matches the ID in the record ID field 331 of the graph management information 330. The ID in the NID field 362 matches the ID in the record ID field 341 of the node information 340.
[0041] Figure 4E shows an example of the structure of graph and arc relationship information 370. Graph and arc relationship information 370 shows a graph containing each arc. Graph and arc relationship information 370 includes a graph ID (GID) field 371 and an arc ID (NID) field 372. The ID in the GID field 371 matches the ID in the record ID field 331 of the graph management information 330. The ID in the NID field 372 matches the ID in the record ID field 351 of the arc information 350.
[0042] Next, the node ambiguation structure information 133 will be described. The node ambiguation structure information 133 is pre-registered, for example, by a system designer. The node ambiguation structure information 133 manages the node ambiguation information in a predetermined structure. The node ambiguation information represents an ambiguous description of the node description in the knowledge graph. In one embodiment of this specification, the node ambiguation information is represented by a DAG. The DAG represents a hierarchy of ambiguous descriptions of the node description. Multiple ambiguation hierarchies allow for the creation of an ambiguous knowledge graph more suitable for each user.
[0043] Figure 5 schematically shows an example of a node ambiguation structure 400 as indicated by node ambiguation structure information 133. The node ambiguation structure 400 is a DAG, and the group of nodes 401 that do not have an input arc represent nodes of the knowledge graph that have not been ambiguized. An arc connects the source node from which ambiguation is performed to the target node from which the ambiguation is performed. The destination node of an arc represents a node that has had its description of the source node ambiguized. The further one follows the arc, the greater the degree of node ambiguation.
[0044] A single node can be obfuscated in one or more ways. For example, in node obfuscation structure 400, insurance node 405 is obfuscated into insurance node 406 and lineage node 407. It is also possible that no obfuscated nodes exist for a given node. For example, there are no obfuscated nodes for insured node 408.
[0045] The description of an ambiguous node (the ambiguized description) is a higher-level description that encompasses the description of the source node (the original description). For example, "insurance" encompasses "C insurance" and "D insurance." Also, "C series" encompasses "C insurance" and "C hospital."
[0046] In this embodiment, the node ambiguation structure 400 is composed of several tables. Figures 6A and 6B show the information that defines the node ambiguation structure 400. Figure 6A shows an example of the configuration of node information 420 within the node ambiguation structure. Node information 420 within the node ambiguation structure includes a record ID field 421 and a description (DESC) field 422. The record ID field 421 shows the ID of the node in the node ambiguation structure 400. The ID shown in the node hierarchy ID field 343 of the node information 340 is included in the record ID field 421. The description field 422 shows the description given to each node.
[0047] Figure 6B shows the arc information 430 within the node ambiguation structure. The arc information 430 within the node ambiguation structure includes a record ID field 431, a source (SRC) field 432, and a target (TGT) field 433. The record ID field 431 shows the ID of the arc in the node ambiguation structure 400. The source field 432 and target field 433 show the source and target of the arc.
[0048] Next, the arc ambiguation structure information 134 will be described. The arc ambiguation structure information 134 is pre-registered, for example, by a system designer. The arc ambiguation structure information 134 manages node ambiguation information in a predetermined structure. The arc ambiguation information shows an ambiguous description of the arc description in the knowledge graph. In one embodiment of this specification, the arc ambiguation information is represented by a DAG. The DAG shows a hierarchy of ambiguous descriptions of the arc description. Multiple ambiguation hierarchies allow for the creation of an ambiguous knowledge graph more suitable for each user.
[0049] Figure 7 schematically shows an example of the arc fuzzy structure 450 as indicated by the arc fuzzy structure information 134. The arc fuzzy structure 450 is a DAG, and the group of nodes 451 where no input arcs exist represent arcs in the unfuzzy knowledge graph. The nodes of the arc fuzzy structure 450 represent descriptions of arcs in the unfuzzy or fuzzy knowledge graph.
[0050] In the arc ambiguation structure 450 shown in Figure 7, the arc connects the source node from which the ambiguation originates and the target node from which the ambiguation is applied. The output node of the arc (the arc in the knowledge graph) is a node that ambiguizes the description of the source node (the arc in the knowledge graph). The further you follow the arc, the greater the degree of ambiguation of the node (the arc in the knowledge graph).
[0051] An arc in the knowledge graph, that is, a node within the arc fuzzy structure 450, can be fuzzy in one or more ways. Furthermore, it is possible that no fuzzy arc exists for a given arc in the knowledge graph.
[0052] An ambiguous arc description is a higher-level description that encompasses the original arc description. For example, "medical expenses" encompasses "medical expenses*", where "*" represents any string. Also, "medical" encompasses "cancer treatment".
[0053] In this embodiment, the arc ambiguation structure 450 is composed of several tables. Figures 8A and 8B show the information that defines the arc ambiguation structure 450. Figure 8A shows an example of the configuration of node information 470 within the arc ambiguation structure. Node information 470 within the arc ambiguation structure includes a record ID field 471 and a description (DESC) field 472. The record ID field 471 shows the ID of the node in the arc ambiguation structure 450. The ID of the arc hierarchy ID field 355 of the arc information 350 is included in the record ID field 471. The description field 422 shows the description given to each node.
[0054] Figure 8B shows the arc information 480 within the arc fuzzy structure. The arc information 480 within the arc fuzzy structure includes a record ID field 481, a source (SRC) field 482, and a target (TGT) field 483. The record ID field 431 indicates the ID of the arc in the arc fuzzy structure 450. The source field 482 and target field 483 indicate the source and target of the arc.
[0055] Next, the access control information 135 will be described. The access control information 135 manages the user's access rights to the nodes and arcs of the knowledge graph. The access control information setting unit 121 generates the access control information 135, for example, according to user input.
[0056] Figure 9A shows an example of the configuration of node access management information 500 included in access control information 135. In this example, node access management information 500 manages the teams that have access rights to each node in node ambiguation structure information 133.
[0057] The node access management information 500 includes a record ID field 501, a node hierarchy ID (NHID) field 502, and a team ID (TID) field 503. The node hierarchy ID field 502 indicates the ID of the node in the node ambiguity structure information 133 and is included in the record ID field 421 of the node information 420 within the node ambiguity structure. The team ID field 503 indicates the ID of the team that has access rights to the corresponding node in the node ambiguity structure information 133.
[0058] Figure 9B shows an example of the configuration of arc access management information 510 included in access control information 135. In this example, arc access management information 510 manages the teams that have access rights to each node (representing an arc) in arc ambiguation structure information 134.
[0059] Arc access management information 510 includes a record ID field 511, an arc hierarchy ID field (AHID) 512, and a team ID (TID) field 513. The arc hierarchy ID field 512 indicates the ID of the node (representing an arc) in the arc ambiguation structure information 134 and is included in the record ID field 471 of the node information 470 within the arc ambiguation structure. The team ID field 513 indicates the ID of the team that has access rights to the corresponding node (representing an arc) in the arc ambiguation structure information 134.
[0060] Furthermore, the node access management information 500 and arc access management information 510 may indicate a user ID instead of a team ID. Also, access rights to ambiguous nodes and arcs may be managed for each knowledge graph. In this configuration, the node access management information 500 and arc access management information 510 further include a graph ID field.
[0061] Next, the method for generating an ambiguous knowledge graph will be explained. Figure 10 shows a flowchart of an example algorithm for generating an ambiguous knowledge graph. The accessible information output unit 122 performs the processing shown in Figure 10 for the user who presents the ambiguous knowledge graph and the knowledge graph to be ambiguized, as specified via the input device 104.
[0062] The accessible information output unit 122 executes steps S10 to S15 on the node set and arc set of the knowledge graph information 132. In step S10, the accessible information output unit 122 extracts all elements of the specified knowledge graph, that is, all nodes and all arcs, from the knowledge graph information 132.
[0063] Specifically, the Accessible Information Output Unit 122 identifies the ID of the knowledge graph specified in the graph management information 330, and extracts the node ID and arc ID associated with that ID from the graph and node relationship information 360 and the graph and arc relationship information 370. The Accessible Information Output Unit 122 then extracts the node and arc information of the extracted ID from the node information 340 and the arc information 350.
[0064] The Accessible Information Output Unit 122 sequentially executes steps S11 to S15 for the extracted elements. In step S11, the Accessible Information Output Unit 122 identifies the selected element in the node ambiguation structure information 133 or arc ambiguation structure information 134 and designates it as element e. Specifically, the Accessible Information Output Unit 122 searches for the node hierarchy ID indicated by the node information 340 or the arc hierarchy ID indicated by the arc information 350 in the node ambiguation structure node information 420 or arc ambiguation structure node information 470.
[0065] Next, in step S12, the accessibility information output unit 122 determines whether the specified user has access rights to element e. Specifically, the accessibility information output unit 122 obtains the ID of the specified user by referring to the personal management information 300 in the user management information 131. Furthermore, it obtains the ID of the team to which the user belongs from the personal and team relationship information 320. The accessibility information output unit 122 checks the access rights to the node hierarchy ID or arc hierarchy ID of element e by referring to the node access management information 500 or arc access management information 510 in the access control information 135.
[0066] If the specified user has access rights to element e (S12: YES), in step S13, the accessibility information output unit 122 determines that element e is to be displayed in the ambiguous knowledge graph.
[0067] If the specified user does not have access rights to element e (S12: NO), in step S14, the accessibility information output unit 122 identifies element e as the target node (adjacent node) for which element e is the source (starting point) in the node ambiguation structure information 133 or arc ambiguation structure information 134. The adjacent node can be identified by referring to the node ambiguation structure arc information 430 or arc ambiguation structure arc information 480. The search order may be either breadth-first or depth-first.
[0068] Next, in step S15, it is determined whether element e is null. If element e is null, i.e., does not exist (S15: YES), the next element in the specified knowledge graph is selected. If element e is not null, i.e., exists (S15: NO), the flow returns to step S12.
[0069] When steps S11 to S15 are executed for all nodes and arcs in the specified knowledge graph, the display method for all nodes and arcs is determined. Each element of a node or arc in the knowledge graph is assigned either its original description or an ambiguous description, or it is excluded from display.
[0070] In step S16, the accessible information output unit 122 deletes arcs where there are no nodes at either end. Furthermore, in step S17, the accessible information output unit 122 abbreviates adjacent nodes if they are identical. This results in a more visually appealing knowledge graph. Steps S16 and S17 may be omitted. Finally, in step S18, the accessible information output unit 122 outputs the created ambiguous knowledge graph to the output device 105.
[0071] Within a DAG with an ambiguous structure, multiple ambiguation nodes can exist for a single node. In the example above, the first one found in the search order is selected. Other examples may search all nodes and then randomly select one. Still other examples may present multiple found ambiguation nodes for the user to choose from.
[0072] Figure 11 shows an example of a GUI (Graphical User Interface) screen for knowledge graph fuzzying. On the GUI screen, the user specifies the user to whom the fuzzy knowledge graph should be presented and the original knowledge graph to be fuzzymed. In the example in Figure 11, "User 2" and the "Medical Insurance Case" knowledge graph are specified.
[0073] When the user selects the "Recommend" button, the Accessible Information Output Unit 122 generates an ambiguous knowledge graph, as explained with reference to Figure 10, and displays it on the GUI screen. In the example in Figure 11, the ambiguous knowledge graph is displayed together with the original knowledge graph. [Examples]
[0074] One embodiment of this specification presents a quantitative index of the sensitivity of each node in an ambiguous structure. This allows users creating ambiguous knowledge graphs to know how sensitive each element of the knowledge graph is. The lower the degree of sensitivity of a node (information), the higher the degree of ambiguation of that node. The following mainly describes the differences from Embodiment 1.
[0075] Figure 12 shows an example configuration of a computer 100 according to one embodiment of this specification. Compared to the example configuration shown in Figure 1, an information granularity evaluation unit 123 is added. Figure 13 shows a flowchart of an example of an information granularity evaluation algorithm executed by the information granularity evaluation unit 123. The information granularity evaluation unit 123 may execute the processing shown in Figure 13 in the node ambiguation structure information 133 or the arc ambiguation structure information 134.
[0076] In step S30, the information granularity evaluation unit 123 creates a copy graph G of the ambiguous structure (ambiguous DAG). Next, in step S31, the information granularity evaluation unit 123 initializes set a with all nodes of the copy graph G that do not have an input arc. Furthermore, in step S32, the information granularity evaluation unit 123 initializes s[n] with an empty set for each node n of the copy graph G.
[0077] Next, in step S33, the information granularity evaluation unit 123 determines whether set a is empty. If set a is empty (S33: YES), this flow ends. If set a is not empty (S33: NO), in step S34, the information granularity evaluation unit 123 removes one element from set a and makes it node n. Furthermore, the information granularity evaluation unit 123 adds a team with access rights to node n to set s[n] and finalizes set s[n].
[0078] Next, the information granularity evaluation unit 123 performs the following processing for each combination of the output arc e of node n and the adjacent node n'. In step S36, the information granularity evaluation unit 123 removes arc e from the copy graph G. In step S37, the information granularity evaluation unit 123 adds set s[n] to set s[n'].
[0079] In step S38, the information granularity evaluation unit 123 determines whether the adjacent node n' has an input arc. If the adjacent node n' does not have an input arc (S38: NO), in step S39, the information granularity evaluation unit 123 adds the adjacent node n' to set a and proceeds to the next loop. If the adjacent node n' has an input arc (S38: YES), the information granularity evaluation unit 123 proceeds to the next loop without executing step S39. Once all loops from steps S36 to S39 are completed, the flow returns to step S33.
[0080] In the process described with reference to Figure 13, the set s[n] is the set of teams that have access rights to either the target node n or its descendant node, and the number of teams is the number of unique teams that have access rights to these nodes. Note that the set of teams consists of different teams, and there are no duplicate teams. The number of unique teams is the number of different teams.
[0081] Figure 14A is a diagram illustrating an example of the set s[n]. The node in the knowledge graph from which ambiguation is applied is the C insurance node 231, and the node that is ambiguated from it is the finance node 241. The set s[n] of finance nodes 241 is the set of teams that have access rights to either finance node 241 or any of its descendant nodes (nodes 231-236).
[0082] The number of teams constituting the set s[n] is a quantitative indicator of the sensitivity of node n. A larger number of teams indicates a lower degree of sensitivity for that node, that is, a greater degree of ambiguity. The information granularity evaluation unit 123 may present the calculated degree of sensitivity (degree of ambiguity) of the node to the system user creating the ambiguous knowledge graph using the output device 105.
[0083] The information granularity evaluation unit 123 may automatically set access rights based on the relationship between the degree of sensitivity (degree of ambiguity) and the threshold. For example, the information granularity evaluation unit 123 receives the specification of the user to whom access rights should be set from the system user. If the degree of sensitivity calculated based on the number of people with access rights to each node as described above is lower than the threshold (the degree of ambiguity is high), the access rights of the specified user to that node are set in the access control information 135. The threshold may be specified by the system user or set by the system design.
[0084] In other examples, the quantitative indicator of the sensitivity of node n may be the number of unique teams that have access rights to node n. In the process shown in Figure 13, by omitting the propagation of the set s[n], we can focus only on each node and count the number of unique teams that have access rights.
[0085] Figure 14B illustrates the number of unique teams with access rights to the target node. Similar to Figure 14A, the node in the knowledge graph from which ambiguation is applied is the C insurance node 231, and the node that is ambiguized is the finance node 241. The number of teams with access rights to finance node 241 is a quantitative indicator of the sensitivity of finance node 241.
[0086] In other examples, the quantitative measure of the sensitivity of node n may be the number of unique teams that have access rights to nodes on the path between that node and the original source node for ambiguation. The original source node for ambiguation is a node in the knowledge graph before ambiguation. In the process shown in Figure 13, the number of unique teams can be counted by changing the initialization of set a from all leaves to the target leaves.
[0087] Figure 14C illustrates the number of unique teams with access rights to nodes on the path between the target node and the original source node for obfuscation. Similar to Figure 14A, the source node in the knowledge graph is the C insurance node 231, and the node that obfuscated it is the finance node 241. Only insurance node 236 exists on the path between C insurance node 231 and finance node 241. The set s[n] is the set of teams that have access rights to either C insurance node 231, insurance node 236, or finance node 241. The number of teams in this set s[n] is a quantitative indicator of the sensitivity of finance node 241.
[0088] In other examples, the quantitative metric for sensitivity may be the total number of teams (the same team can be counted multiple times) instead of the number of unique teams. In the process shown in Figure 13, the total number of teams can be calculated by replacing the set with a list. Alternatively, the number of users may be used instead of the number of teams. Alternatively, a percentage of the total number of nodes may be used instead of the number of teams or users. Each of the above quantitative metrics allows users creating an ambiguous knowledge graph to know how sensitive each element of the knowledge graph is. Note that the quantitative metric for sensitivity may be calculated only for some nodes of the ambiguous DAG.
[0089] As described above, the degree of sensitivity of a node can be appropriately represented by quantifying it based on the number of access holders for that node within the ambiguous DAG. The number of access holders may be represented by the number of teams, as described above, or by the number of users that make up the team. [Examples]
[0090] One embodiment of this specification predicts the target of fuzzying and recommends it to the user. This allows the user to efficiently set the fuzzying information for each element in the fuzzy structure. The differences from Embodiment 1 will be mainly explained below. Figure 15 shows an example configuration of the computer 100 of one embodiment of this specification. Compared to the example configuration shown in Figure 1, a fuzzying information recommendation unit 124 has been added. Figure 16 shows a flowchart of an example of the fuzzy DAG arc recommendation algorithm executed by the fuzzying information recommendation unit 124. The fuzzying information recommendation unit 124 can perform this processing on the fuzzy DAG (fuzzy structure) of node fuzzy structure information 133 or arc fuzzy structure information 134.
[0091] In step S50, the fuzzy information recommendation unit 124 obtains natural language features of the nodes of the fuzzy DAG using a natural language processing model such as BERT or Word2vec.
[0092] Next, in step S51, the ambiguous information recommendation unit 124 trains a link prediction model using a machine learning model such as a GNN (Graph Neural Network). The graph structure of the ambiguous DAG and the natural language features obtained in step S50 are used for training. The link prediction model takes the graph structure of the ambiguous DAG and the natural language features of the nodes of the ambiguous DAG as input and calculates the probability (score) that a link exists between nodes.
[0093] Next, in step S52, the ambiguation information recommendation unit 124 uses the learned link prediction model to calculate the score of the links between each node in the ambiguous DAG and other nodes. Furthermore, in step S53, the ambiguation information recommendation unit 124 outputs a predetermined number of links (node pairs) with high scores. The condition may be that the score of the output links is higher than a threshold. The user sets the links that they determine to be appropriate from the presented links into the ambiguous DAG.
[0094] The ambiguation information recommendation unit 124 may accept the designation of one or more nodes to perform link prediction within the ambiguation DAG. After training the link prediction model, the ambiguation information recommendation unit 124 may add new nodes to the ambiguation DAG and, upon receiving the designation of such nodes, predict their ambiguation destinations. The ambiguation information recommendation unit 124 outputs links with high scores for the designated nodes. The ambiguation information recommendation unit 124 may automatically set links with scores exceeding a threshold in the ambiguation DAG. [Examples]
[0095] One embodiment of this specification predicts access rights to each node in an ambiguous DAG and provides recommendations to the user. This allows for efficient configuration of access rights to nodes.
[0096] Figure 17 shows an example configuration of computer 100 according to one embodiment of this specification. Compared to the configuration example shown in Figure 1, an access control information recommendation unit 125 is added. Figure 18 shows a flowchart of an example of a node access rights recommendation algorithm for an ambiguous DAG executed by the access control information recommendation unit 125. The access control information recommendation unit 125 can perform this processing on an ambiguous DAG (ambiguous structure) of node ambiguation structure information 133 or arc ambiguation structure information 134.
[0097] In step S70, the access control information recommendation unit 125 obtains natural language features of the nodes of the fuzzy DAG using a natural language processing model such as BERT or Word2vec.
[0098] Next, in step S71, the access control information recommendation unit 125 uses a machine learning model such as a GNN to perform semi-supervised learning of an access rights prediction model that binary classifies the access rights of teams. In semi-supervised learning, the graph structure of the ambiguous DAG, the natural language features obtained in step S70, and the access control information 135 indicating the teams to which each node is granted access rights are used. Alternatively, co-occurrence analysis using association rules may be used instead of a GNN.
[0099] In step S72, the access control information recommendation unit 125 predicts the access rights score for the target node using the learned access rights prediction model. The access rights prediction model uses the graph structure of the fuzzy DAG and the natural language features of the nodes of the fuzzy DAG to calculate the probability (score) that each team has access rights to each node.
[0100] Furthermore, in step S73, the access control information recommendation unit 125 outputs a predetermined number of node-team pairs that were not used by the teacher, i.e., node-team pairs that have not already been granted access rights, starting with the pairs with the highest scores. The condition may be that the scores of the output pairs must be higher than a threshold. The user sets the access rights for the pairs that they deem appropriate from the presented pairs.
[0101] The access control information recommendation unit 125 may accept the designation of teams for which access rights should be predicted. The access control information recommendation unit 125 presents the user with a predetermined number of nodes that have high access rights scores with the designated teams. In another example, the access control information recommendation unit 125 may accept the designation of nodes for which access rights should be predicted. The access control information recommendation unit 125 presents the user with a predetermined number of teams that have high access rights scores with the designated nodes. The access control information recommendation unit 125 may automatically set access rights for pairs whose scores exceed a threshold.
[0102] In addition to or instead of the team's access rights to each node of the ambiguous DAG, the Access Control Information Recommendation Unit 125 predicts the team's access rights to the ambiguous knowledge graph and makes recommendations to the user. For example, the Access Control Information Recommendation Unit 125 utilizes the ambiguous structure to generate multiple ambiguous knowledge graphs, displays several candidates that the specified team is most likely to have access to, and allows the user to select one.
[0103] Figure 19 shows a flowchart of an example of an automated generation and access rights recommendation algorithm for an ambiguous knowledge graph. This allows for efficient setting of access rights.
[0104] In step S90, the access control information recommendation unit 125 obtains natural language features of the nodes and arcs of the ambiguous DAG using natural language processing models such as BERT and Word2vec.
[0105] In step S91, the access control information recommendation unit 125 uses a machine learning model such as a GNN to train a binary classification model that predicts access rights to the knowledge graph. In the training process, natural language features obtained in step S90 and the set of ambiguous knowledge graphs that each team is accessing are used.
[0106] The trained access permission prediction model predicts the probability (score) of a team's access to a knowledge graph based on the graph structure of the ambiguous knowledge graph, the natural language features of the nodes and arcs, and the team identifier.
[0107] In step S92, the access control information recommendation unit 125 randomly generates multiple ambiguous knowledge graphs for the target knowledge graph using nodes that have a higher degree of ambiguity than each element (node or arc) in the node or arc ambiguation DAG.
[0108] In step S93, the access control information recommendation unit 125 predicts access rights to the generated knowledge graphs using a trained access rights prediction model and outputs a predetermined number of graphs from the ambiguous knowledge graphs with high scores. The condition may be that the scores of the output graphs are higher than a threshold. The user sets the team's access rights to the graphs that they determine to be appropriate from the presented graphs. The access control information recommendation unit 125 may automatically set access rights to knowledge graphs whose scores exceed a threshold.
[0109] Figures 20 and 21 show examples of GUI screens for predicting access rights to a knowledge graph, as explained with reference to Figure 19. In the GUI screen 600 shown in Figure 20, the user specifies the knowledge graph and team that will serve as the basis for the ambiguous knowledge graph whose access rights are predicted. Specifically, the original knowledge graph is specified in section 601, and the team is specified in section 602. When the user recommendation button 603 is selected, the process explained with reference to Figure 19 is executed.
[0110] GUI screen 600 displays several ambiguous knowledge graphs with high access rights prediction scores in section 604. The user selects one or more knowledge graphs from the displayed ambiguous knowledge graphs.
[0111] Figure 21 shows an example of screen 620 displaying a knowledge graph selected on GUI screen 600. On screen 620, the user can edit the selected ambiguous knowledge graph. The access control information recommendation unit 125 accepts the user's adjustment of the ambiguous knowledge graph.
[0112] For example, a user can change one node or arc to another node or arc. In the example in Figure 21, the C Insurance Node and the Insurance Node are presented as alternatives to the Financial Node. These are nodes on a single ambiguation path. In this example, the Financial Node has the highest degree of ambiguation, and the C Insurance Node has the lowest degree of ambiguation. When the OK button is selected, the access rights of the specified team to the displayed nodes and arcs are set in the access control information 135.
[0113] It should be noted that the present invention is not limited to the embodiments described above, and various modifications are included. Furthermore, for example, the embodiments described above are detailed explanations of the configuration in order to clearly illustrate the present invention, and are not necessarily limited to those having all the configurations described. In addition, some of the configurations in each embodiment can be added to, deleted from, or replaced with other configurations.
[0114] Furthermore, each of the above-mentioned configurations, functions, processing units, processing means, etc., may be implemented in hardware, in whole or in part, for example, by designing them as integrated circuits. The present invention can also be implemented by software program code that realizes the functions of the embodiment. In this case, a storage medium on which the program code is recorded is provided to a computer, and the processor of that computer reads the program code stored in the storage medium. In this case, the program code read from the storage medium itself realizes the functions of the embodiment described above, and the program code itself and the storage medium on which it is stored constitute the present invention. Examples of storage media used to supply such program code include flexible disks, CD-ROMs, DVD-ROMs, hard disks, SSDs (Solid State Drives), optical disks, magneto-optical disks, CD-Rs, magnetic tapes, non-volatile memory cards, ROMs, and the like.
[0115] Furthermore, the program code that implements the functions described in this embodiment can be implemented in a wide range of programming or scripting languages, such as assembler, C / C++, Perl, Shell, PHP, Python, and Java (registered trademark).
[0116] Furthermore, the program code for the software that implements the functions of the embodiment may be distributed via a network and stored in a storage means such as a computer's hard disk or memory, or in a storage medium such as a CD-RW or CD-R, and the computer's processor may read and execute the program code stored in the storage means or storage medium.
[0117] In the above-described embodiment, the control lines and information lines shown are those deemed necessary for explanation and do not necessarily represent all control lines and information lines in the actual product. All components may be interconnected. [Explanation of Symbols]
[0118] 100 calculator 101 CPU 102 memory 103 Auxiliary storage device 104 Input device 105 Output device 106 Communication equipment 107 Bus 121 Access control information setting unit 122 Accessible Information Output Unit 131 User Management Information 132 Knowledge Graph Information 133 Node-Fuzzy Structure Information 134 Arc ambiguity structure information 135 Access Control Information
Claims
1. A knowledge graph access control system, The computing unit and Includes a storage device, The aforementioned storage device is Ambiguity structure information defines the inclusion relationships between elements with different degrees of ambiguity in the knowledge graph, It stores access control information that manages the user's access rights to each element included in the aforementioned ambiguous structure information, The aforementioned computing device is Obtain the knowledge graph to be obfuscated, Referencing the aforementioned ambiguous structure information and the access control information, the target knowledge graph is ambiguized for the first user to generate an ambiguous knowledge graph. In the ambiguation of the target knowledge graph, the original elements included in the target knowledge graph are converted into ambiguation elements that the first user has access rights to in the access control information and that include the original elements in the ambiguation structure information. The degree of ambiguity of the first element included in the ambiguous structure information is determined based on the number of holders of access rights to the first element. A knowledge graph access control system that displays information indicating the determined degree of ambiguity on an output device.
2. A knowledge graph access control system according to Claim 1, Knowledge graph access control system, wherein the computing device determines the degree of ambiguation of the first element based on the number of holders of access rights to the elements encompassed by the first element.
3. A knowledge graph access control system according to Claim 1, The aforementioned computing device is a knowledge graph access control system that sets the access rights of a second user to the first element when the determined degree of ambiguity exceeds a threshold.
4. A knowledge graph access control system, The computing unit and Includes a storage device, The aforementioned storage device is Ambiguity structure information defines the inclusion relationships between elements with different degrees of ambiguity in the knowledge graph, It stores access control information that manages the user's access rights to each element included in the aforementioned ambiguous structure information, The aforementioned computing device is Obtain the knowledge graph to be obfuscated, Referencing the aforementioned ambiguous structure information and the access control information, the target knowledge graph is ambiguized for the first user to generate an ambiguous knowledge graph. In the ambiguation of the target knowledge graph, the original elements included in the target knowledge graph are converted into ambiguation elements that the first user has access rights to in the access control information and that include the original elements in the ambiguation structure information. A knowledge graph access control system that predicts the access rights of a third user to a third element in the ambiguous structure information based on the relationship between the elements of the ambiguous structure information indicated by the access control information and the access rights holder, and the inclusion relationship, and presents the result of the prediction on an output device.
5. A knowledge graph access control system, The computing unit and Includes a storage device, The aforementioned storage device is Ambiguity structure information defines the inclusion relationships between elements with different degrees of ambiguity in the knowledge graph, It stores access control information that manages the user's access rights to each element included in the aforementioned ambiguous structure information, The aforementioned computing device is Obtain the knowledge graph to be obfuscated, Referencing the aforementioned ambiguous structure information and the access control information, the target knowledge graph is ambiguized for the first user to generate an ambiguous knowledge graph. In the ambiguation of the target knowledge graph, the original elements included in the target knowledge graph are converted into ambiguation elements that the first user has access rights to in the access control information and that include the original elements in the ambiguation structure information. A candidate ambiguous knowledge graph for the first knowledge graph is generated based on the ambiguous structure information. A knowledge graph access control system that uses a pre-prepared prediction model to predict the access rights of a fourth user to the candidate ambiguous knowledge graph, and presents the prediction results on an output device.
6. A knowledge graph access control system according to claim 5, The aforementioned computing device is The output device presents a candidate ambiguation knowledge graph in which the prediction satisfies predetermined conditions. A knowledge graph access control system that accepts user adjustments to the presented candidate ambiguous knowledge graph.
7. A method for controlling access to a knowledge graph by a system, The aforementioned system, Ambiguity structure information defines the inclusion relationships between elements with different degrees of ambiguity in the knowledge graph, It stores access control information that manages the user's access rights to each element included in the aforementioned ambiguous structure information, The above method is performed by the system, Obtain the knowledge graph to be obfuscated, Referencing the aforementioned ambiguous structure information and the access control information, the target knowledge graph is ambiguized for the first user to generate an ambiguous knowledge graph. In the ambiguation of the target knowledge graph, the original elements included in the target knowledge graph are converted into ambiguation elements that the first user has access rights to in the access control information and that include the original elements in the ambiguation structure information. The degree of ambiguity of the first element included in the ambiguous structure information is determined based on the number of holders of access rights to the first element. A method for displaying information indicating the determined degree of ambiguity on an output device.
8. A method for controlling access to a knowledge graph by a system, The aforementioned system, Ambiguity structure information defines the inclusion relationships between elements with different degrees of ambiguity in the knowledge graph, It stores access control information that manages the user's access rights to each element included in the aforementioned ambiguous structure information, The above method is performed by the system, Obtain the knowledge graph to be obfuscated, Referencing the aforementioned ambiguous structure information and the access control information, the target knowledge graph is ambiguized for the first user to generate an ambiguous knowledge graph. In the ambiguation of the target knowledge graph, the original elements included in the target knowledge graph are converted into ambiguation elements that the first user has access rights to in the access control information and that include the original elements in the ambiguation structure information. A method for predicting a third user's access rights to a third element in ambiguous structure information based on the relationship between the elements of the ambiguous structure information indicated by the access control information and the access rights holder, and the inclusion relationship, and presenting the result of the prediction on an output device.
9. A method for controlling access to a knowledge graph by a system, The aforementioned system, Ambiguity structure information defines the inclusion relationships between elements with different degrees of ambiguity in the knowledge graph, It stores access control information that manages the user's access rights to each element included in the aforementioned ambiguous structure information, The above method is performed by the system, Obtain the knowledge graph to be obfuscated, Referencing the aforementioned ambiguous structure information and the access control information, the target knowledge graph is ambiguized for the first user to generate an ambiguous knowledge graph. In the ambiguation of the target knowledge graph, the original elements included in the target knowledge graph are converted into ambiguation elements that the first user has access rights to in the access control information and that include the original elements in the ambiguation structure information. A candidate ambiguous knowledge graph for the first knowledge graph is generated based on the ambiguous structure information. A method for predicting a fourth user's access rights to the candidate ambiguous knowledge graph using a pre-prepared prediction model, and presenting the prediction results on an output device.
Citation Information
Patent Citations
Business information masking system
JP2011227536A
Anonymization system
JP2015125646A
Dynamic Access Control for Knowledge Graphs
JP2021513138A