Related information retrieval device, related information retrieval method, and related information retrieval program
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-26
- Publication Date
- 2026-04-14
Smart Images

Figure 00000015_0000 
Figure 00000015_0001 
Figure 00000016_0000
Description
Technical Field
[0001] The present disclosure relates to a related information retrieval device, a related information retrieval method, and a related information retrieval program.
Background Art
[0002] There is a search technology based on a knowledge base. Patent Document 1 discloses a technology that realizes a function of presenting knowledge necessary as an appropriate measure for a generated failure when a failure occurs in a facility device at a manufacturing site based on the knowledge base.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] According to the technology disclosed in Patent Document 1, previously accumulated documents are converted into a knowledge graph, and related past knowledge is retrieved according to the current situation. However, in this technology, since only labels are extracted from individual documents, the accuracy of the connection between documents is low (not connected if the expressions are different, or connected only because the words match), and there is a problem that cross-search cannot be performed between different databases.
[0005] An object of the present disclosure is to make the accuracy of the connection between documents relatively high in a search technology based on a knowledge graph and to enable cross-search between different databases.
Means for Solving the Problems
[0006] The related information retrieval device according to the present disclosure is A static link unit connects the first node and the second node when the similarity between the first text corresponding to the first node in the first knowledge graph corresponding to the document and the second text corresponding to the second node in the second knowledge graph corresponding to the document is greater than or equal to the similarity threshold. It is equipped with. [Effects of the Invention]
[0007] According to this disclosure, the static linking section determines whether to connect the first node and the second node based on the similarity between the first text corresponding to the first node in the first knowledge graph and the second text corresponding to the second node in the second knowledge graph. Therefore, according to this disclosure, in a knowledge graph-based search technology, the accuracy of connections between documents can be made relatively high, and cross-database searches can be performed. [Brief explanation of the drawing]
[0008] [Figure 1] A diagram illustrating the outline of Embodiment 1. [Figure 2] A diagram illustrating the outline of Embodiment 1. [Figure 3] A diagram showing an example configuration of the related information retrieval device 100 according to Embodiment 1. [Figure 4] A diagram illustrating the knowledge graph according to Embodiment 1. [Figure 5] This is a diagram illustrating the knowledge graph according to Embodiment 1, where (a) is a table showing the meaning of node types, and (b) is a table showing triples in the knowledge graph. [Figure 6] A diagram illustrating the processing of the knowledge acquisition unit 110 according to Embodiment 1. [Figure 7] A diagram illustrating the processing of the static link section 130 according to Embodiment 1. [Figure 8] A diagram illustrating the processing of the static link section 130 according to Embodiment 1. [Figure 9] A diagram illustrating the processing of the related information inference unit 150 according to Embodiment 1. [Figure 10] FIG. for explaining the processing of the relevance calculation unit 160 according to Embodiment 1, (a) is a diagram for explaining each path, and (b) is a table for explaining the score calculation processing. [Figure 11] FIG. for explaining the processing of the display data generation unit 170 according to Embodiment 1. [Figure 12] FIG. for explaining the processing of the display data generation unit 170 according to Embodiment 1. [Figure 13] FIG. showing a hardware configuration example of the related information retrieval device 100 according to Embodiment 1. [Figure 14] Flowchart showing the operation of the related information retrieval device 100 according to Embodiment 1. [Figure 15] FIG. showing a hardware configuration example of the related information retrieval device 100 according to a modification example of Embodiment 1. [Figure 16] FIG. showing a configuration example of the related information retrieval device 100 according to Embodiment 2. [Figure 17] FIG. for explaining the processing of the connection condition specifying unit 210 according to Embodiment 2. <00000Flowchart showing the operation of the related information search device 100 according to Embodiment 4.
Mode for Carrying Out the Invention
[0009] In the description of the embodiments and the drawings, the same elements and corresponding elements are denoted by the same reference numerals. The description of the elements denoted by the same reference numerals will be omitted or simplified as appropriate. The arrows in the drawings mainly indicate the flow of data or the flow of processing. Also, "section" may be appropriately read as "circuit", "step", "procedure", "process", or "circuitry".
[0010] Embodiment 1. Hereinafter, this embodiment will be described in detail with reference to the drawings. FIG. 1 is a diagram for explaining the outline of this embodiment. In this embodiment, one of the objectives is to construct a knowledge graph based on a file comprehensively showing influences and defects. More specifically, one of the objectives is to construct a wide-ranging knowledge graph from in-house documents indicating design, manufacturing processes, market defects, etc. Each knowledge graph corresponds to an information source. The file may also be expressed as a document. At this time, by utilizing document information, hidden relationships are predicted, and based on the predicted relationships, multiple knowledge graphs are connected. Note that by using a language model specialized for a domain such as a company, the calculation cost may be reduced. Also, a knowledge graph embedding technique incorporating document information is utilized. As a specific example, the document is a business document. As a specific example, this embodiment is utilized when, for example, when changing the material of a certain member to PPS (polyphenylene sulfide), it is desired to explore phenomena that may occur in relation to PPS.
[0011] FIG. 2 is a diagram for explaining the outline of this embodiment. In FIG. 2, knowledge graphs regarding each data type are shown. Each node of each knowledge graph is a node determined according to the content of the knowledge. In this embodiment, as shown by the dashed lines in Figure 2, nodes representing similar content are linked between different knowledge graphs. This semantically connects the knowledge graphs obtained from each of the multiple databases. Therefore, according to this embodiment, cross-database searches become possible by tracing semantic relationships between multiple databases.
[0012] ***Explanation of the structure*** Figure 3 shows an example configuration of the related information retrieval device 100 according to this embodiment. As shown in Figure 3, the related information retrieval device 100 comprises a knowledge acquisition unit 110, a knowledge registration unit 120, a static link unit 130, a DB operation unit 140, a related information inference unit 150, a relatedness calculation unit 160, a display data generation unit 170, and a user interface unit 180. The related information retrieval device 100 also stores a knowledge graph DB 190. DB is an abbreviation for database.
[0013] The knowledge acquisition unit 110 receives data representing files as a set of documents. The knowledge acquisition unit 110 acquires information about the nodes and edges of the knowledge graph from the input files. Here, edges correspond to the relationships between nodes. The generation of the knowledge graph constitutes knowledge acquisition.
[0014] Figure 4 shows a concrete example of a knowledge graph. The text within each node indicates the node type corresponding to that node. Each edge is accompanied by text indicating the edge type corresponding to that edge. Note that in Figure 4, the knowledge graph is a directed graph, but the knowledge graph may also be an undirected graph. A knowledge graph is created from a single file, as shown in Figure 4. Such a knowledge graph is created for each data point, and each created knowledge graph is stored in the knowledge graph DB190.
[0015] Figure 5 is a table corresponding to Figure 4. Figure 5(a) is a table showing the meaning of node types. Figure 5(b) is a table showing triples in the knowledge graph. Figure 5(b) shows the "content" and the knowledge graph triple corresponding to the "content". A knowledge graph triple is defined for each combination of node types corresponding to two nodes. In Figure 5(b), for each combination of node types corresponding to two nodes, the edge type corresponding to the edge connecting the two nodes is shown.
[0016] As a specific example, if the input file is a tabular file, the knowledge acquisition unit 110 extracts text from each cell of the tabular file and uses it as a node. In this case, the node type corresponding to each text may be defined in advance. Alternatively, the knowledge acquisition unit 110 may estimate the node type based on the content of each text and use the estimated node type as the node type corresponding to each text. Furthermore, the knowledge acquisition unit 110 determines the relationships (edges) between multiple nodes using information such as whether the corresponding cells for multiple nodes are in the same row or column. Figure 6 illustrates the process of generating a knowledge graph from a tabular file. In Figure 6, node types corresponding to each column are predefined, and each pair of nodes is appropriately connected by an edge having a corresponding edge type. Each node corresponds to the text in each cell. The knowledge acquisition unit 110 may decide whether or not to draw an edge between two nodes based on the similarity between the texts corresponding to the two nodes. Each edge type may be determined based on the structural information of the tabular file.
[0017] The knowledge registration unit 120 registers information indicating nodes and edges acquired by the knowledge acquisition unit 110 in the knowledge graph DB 190 via the DB operation unit 140.
[0018] The static linking unit 130 connects the first node and the second node when the target similarity is equal to or greater than the similarity threshold. The target similarity is the similarity between the first text and the second text. The first text corresponds to the first node included in the first knowledge graph corresponding to the document. The second text corresponds to the second node included in the second knowledge graph corresponding to the document. The target similarity may also be the similarity between the vector corresponding to the first text and the vector corresponding to the second text. The static linking unit 130 may calculate the target similarity based on the edit distance. The similarity threshold may be defined in any way. As a specific example, the static linking unit 130 converts the text corresponding to all nodes registered in the knowledge graph DB 190 into vectors, and uses each converted vector to calculate the vector similarity between any two nodes. Subsequently, the static linking unit 130 draws an edge between two nodes with high corresponding vector similarity. The static linking unit 130 may also link semantically similar nodes by text embedding or graph embedding, and connect the linked nodes with an edge. The two nodes to which an edge is drawn may be nodes in different knowledge graphs, or they may be two nodes in the same knowledge graph.
[0019] The static link section 130 may vectorize text in any way. Methods for obtaining vector representations corresponding to text include methods for aggregating vector representations of words and methods for vectorizing the text itself. The static link section 130 may also vectorize each node using a graph. In other words, when vectorizing nodes, the static link section 130 may use text embedding (vectorizing nodes using only text information) or graph embedding (vectorizing nodes using information from a subgraph). Specific examples of methods for vectorizing text include BoW (Bag of Words), TF-IDF (term frequency-inverse document frequency), the average of Word2vec (word embeddings), and Sentence BERT (Bidirectional Encoder Representations from Transformers). These methods use the text information contained in the nodes to obtain vectors. BoW is a method for calculating the sum of local representations of words that make up a text. Using TF-IDF allows you to increase the value of words that have features not found in other documents. Word2vec averaging is a method for calculating the average of the word vectors of the words that make up a sentence. Sentence BERT is a method that fine-tunes BERT to generate relatively high-quality sentence vectors.
[0020] The static link unit 130 vectorizes each node and, as a specific example, calculates the similarity between vectors, i.e., the similarity between nodes, by determining the cosine similarity shown in [Equation 1]. Based on the calculated similarity, the static link unit 130 decides whether or not to create an edge between the nodes.
[0021]
number
[0022] Figure 7 shows a concrete example of how edges (SIMILAR_TO) are drawn between nodes whose similarity is above a threshold in knowledge graphs constructed from three separate datasets. In Figure 7, for the knowledge graphs, the similarity between nodes with SIMILAR_TO edges is above the threshold, while the similarity between nodes without SIMILAR_TO edges is below the threshold. SIMILAR_TO edges are edges drawn between similar nodes. Figure 8 shows a specific example of node type combinations corresponding to two nodes on which a SIMIALAR_TO edge may be established.
[0023] The DB operation unit 140 has the function of operating the knowledge graph DB 190.
[0024] The related information inference unit 150 selects a parent node from the first knowledge graph and the second knowledge graph based on the text indicated by the query, and searches for one or more nodes by traversing edges from the parent node in at least one of the first knowledge graph and the second knowledge graph. As a concrete example, the related information inference unit 150 infers relevant information in response to the user's query through roughly the following two steps.
[0025] (Step 1: Node acquisition) The related information inference unit 150 retrieves the node in the knowledge graph that is most similar to the query entered by the user. In this case, the related information inference unit 150 may use natural language search or vector search, etc., to retrieve one or more nodes corresponding to texts that have a relatively high similarity to the text indicated by the entered query. Figure 9 is a diagram illustrating the processing of the related information inference unit 150. In the example shown in Figure 9, the query is the text "Changed the feather material from metal to plastic." Nodes #123, #456, and #789 are each parent nodes. Here, node #123 is the node corresponding to the change "Feather material: metal ⇒ plastic". Node #456 is the node corresponding to the countermeasure "Changed the material to plastic A." Node #789 is the node corresponding to the countermeasure "The material is hot-dip galvanized steel sheet...".
[0026] (Step 2: Graph Search) The related information inference unit 150 uses the node obtained in step 1 as the starting point (parent node) and obtains other nodes (child nodes) that can be reached from the parent node in a predetermined number of hops, as well as the path from the parent node to each child node. Each child node corresponding to the parent node is an related node corresponding to the parent node. Furthermore, the related information inference unit 150 can infer grandchild nodes starting from child nodes and trace relationships by repeating step 2. In the example shown in Figure 9, each node connected to each parent node is retrieved. Here, node #062 is another node accessed from node #123 and node #456, respectively, and corresponds to the phenomenon "ultraviolet degradation". Node #554 is another node accessed from node #789, and corresponds to the phenomenon "material strength variation".
[0027] The relevance calculation unit 160 treats each of the one or more nodes found as a target node and calculates the relevance corresponding to the target node based on the path from the parent node to the target node. The relevance calculation unit 160 calculates a score for each child node obtained by the relevance information inference unit 150 based on the path from the parent node to the child node. Figure 10 is a diagram illustrating the processing of the relevance calculation unit 160. As a specific example, as shown in Figure 10, the relevance calculation unit 160 calculates a score corresponding to each node such that the closer the node is and the more paths it can be reached by, the higher the corresponding relevance. In the example shown in Figure 10, the relevance calculation unit 160 enumerates all possible paths from node S to node G, calculates the length of each enumerated path, calculates a score for each path based on the calculated path length, and aggregates the calculated scores to calculate the score corresponding to node G. Figure 10(a) is a diagram illustrating the paths from node S to node G. Figure 10(b) is a table illustrating the process of calculating the score corresponding to node G. In this example, the length of each edge is defined. The path length is the sum of the lengths of each edge on the path. The path score corresponding to each path is the reciprocal of the path length. This is to ensure that the longer the path (the greater the distance from node S to node G), the lower the path score. In Figure 10, the score of node G for node S is 1.66. The final score is also called the search score.
[0028] The display data generation unit 170 generates data to display each node acquired by the related information inference unit 150, based on the score calculated by the relatedness calculation unit 160. As a specific example, the display data generation unit 170 generates a screen that displays the obtained parent nodes and child nodes in descending order of their corresponding scores. Figures 11 and 12 show specific examples of screens generated by the display data generation unit 170. Figure 11 shows each parent node and each child node associated with that parent node. As shown in Figure 12, when a displayed child node is selected, the path from the parent node to that child node may be shown. In Figure 12, the text corresponding to each node, the file containing each text, and the date on which each text was registered are shown.
[0029] The user interface unit 180 constitutes the user interface. The user interface unit 180 has functions such as receiving user queries and outputting search results corresponding to the received queries.
[0030] Knowledge Graph DB190 is a database that stores data representing knowledge graphs.
[0031] Figure 13 shows an example of the hardware configuration of the related information retrieval device 100 according to this embodiment. The related information retrieval device 100 consists of a computer. The related information retrieval device 100 may consist of multiple computers.
[0032] As shown in this figure, the related information retrieval device 100 is a computer equipped with hardware such as a processor 11, a volatile storage device 12, a non-volatile storage device 13, and an interface 14. These hardware components are connected as appropriate via signal lines.
[0033] The processor 11 is an integrated circuit (IC) that performs arithmetic operations and controls the hardware of the computer. Specific examples of the processor 11 include a CPU (Central Processing Unit), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit). The related information retrieval device 100 may include multiple processors that replace the processor 11. The multiple processors share the role of the processor 11.
[0034] The volatile memory device 12 is also called main memory, and a specific example is RAM (Random Access Memory). Data stored in the volatile memory device 12 is saved to the non-volatile memory device 13 as needed.
[0035] The non-volatile storage device 13 is also called an auxiliary storage device, and specific examples include ROM (Read Only Memory), HDD (Hard Disk Drive), or flash memory. Data stored in the non-volatile storage device 13 is loaded into the volatile storage device 12 as needed. The volatile storage device 12 and the non-volatile storage device 13 may be configured as a single unit.
[0036] Interface 14 has the function of communicating with other devices. Interface 14 may receive voice data or transmit text data. Interface 14 may include, as a specific example, a communication chip or a NIC (Network Interface Card). Each part of the related information retrieval device 100 may use the interface 14 as appropriate when communicating with other devices.
[0037] The non-volatile memory device 13 stores the related information retrieval program. The related information retrieval program is a program that enables the computer to implement the functions of each part of the related information retrieval device 100. The related information retrieval program is loaded into the volatile memory device 12 and executed by the processor 11. The functions of each part of the related information retrieval device 100 are implemented by software.
[0038] Data used when executing the related information retrieval program, and data obtained by executing the related information retrieval program, are appropriately stored in the memory device. Each part of the related information retrieval device 100 utilizes the memory device as appropriate. The memory device consists, specifically, of at least one of a volatile memory device 12, a non-volatile memory device 13, a register in the processor 11, and a cache memory in the processor 11. Note that the terms "data" and "information" may have the same meaning. The memory device may be independent of the computer. The functions of the volatile storage device 12 and the non-volatile storage device 13 may be realized by other storage devices.
[0039] The related information retrieval program may be recorded on a computer-readable non-volatile recording medium. Specific examples of non-volatile recording media include optical discs or flash memory. The related information retrieval program may also be provided as a program product.
[0040] ***Explanation of operation*** The operating procedure of the related information retrieval device 100 corresponds to the related information retrieval method. Furthermore, the program that implements the operation of the related information retrieval device 100 corresponds to the related information retrieval program.
[0041] Figure 14 is a flowchart illustrating an example of the operation of the related information retrieval device 100. This operation will be explained using Figure 14.
[0042] (Step S101) The knowledge acquisition unit 110 receives input from a set of documents and acquires knowledge from the received set of documents, that is, it generates knowledge graphs based on the received set of documents. The knowledge registration unit 120 registers each knowledge graph generated by the knowledge acquisition unit 110 into the knowledge graph DB 190 via the DB operation unit 140.
[0043] (Step S102) The static linking unit 130 links semantically similar nodes together for each knowledge graph registered in the knowledge graph DB 190. The knowledge registration unit 120 updates the knowledge graph DB 190 via the DB operation unit 140 so that edges are drawn between nodes linked by the static link unit 130.
[0044] (Step S103) The related information inference unit 150 receives a query input from the user and searches the knowledge graph DB 190 for the node most similar to the received query.
[0045] (Step S104) The related information inference unit 150 takes each node found in step S103 as a parent node and, starting from each parent node, traverses related nodes in order to obtain each related node corresponding to each parent node.
[0046] (Step S105) The relevance calculation unit 160 calculates a score for each related node obtained by the relevance information inference unit 150 for each parent node.
[0047] (Step S106) The display data generation unit 170 displays the nodes on the screen in descending order of the scores calculated by the relevance calculation unit 160. In this case, the display data generation unit 170 may appropriately classify each node as either a parent node or a child node.
[0048] ***Explanation of the effects of Embodiment 1*** In conventional technology, when constructing a knowledge graph from individual files, multiple nodes are treated as related only if there is no variation in notation between them or if they contain common keywords. However, in all other cases, these nodes are treated as unrelated. On the other hand, in this embodiment, by providing a static link section 130, the similarity between nodes is calculated among multiple knowledge graphs constructed from each file, and semantically similar nodes are linked together based on the calculated similarity. Furthermore, this embodiment is not limited to internal company data, but can also be applied to public data. While this embodiment is explained using data indicating design and defects, it can also be applied to other types of data, as long as the information can be represented by a knowledge graph.
[0049] ***Other configurations*** <Example 1> Figure 15 shows an example of the hardware configuration of the related information retrieval device 100 according to this modified example. The related information retrieval device 100 includes a processing circuit 18 instead of a processor 11, a processor 11 and a volatile storage device 12, a processor 11 and a non-volatile storage device 13, or a processor 11, a volatile storage device 12, and a non-volatile storage device 13. The processing circuit 18 is hardware that implements at least a part of each component of the related information retrieval device 100. The processing circuit 18 may be dedicated hardware, or it may be a processor that executes a program stored in the volatile memory device 12.
[0050] When the processing circuit 18 is dedicated hardware, specific examples of the processing circuit 18 include a single circuit, a composite circuit, a programmed processor, a parallel programmed processor, an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a combination thereof. The related information retrieval device 100 may include multiple processing circuits that replace the processing circuit 18. The multiple processing circuits share the role of the processing circuit 18.
[0051] In the related information retrieval device 100, some functions may be implemented by dedicated hardware, while the remaining functions may be implemented by software or firmware.
[0052] The processing circuit 18 can be implemented, in specific examples, by hardware, software, firmware, or a combination thereof. The processor 11, the volatile memory device 12, the non-volatile memory device 13, and the processing circuit 18 are collectively referred to as the "processing circuitry." In other words, the functions of each functional component of the related information retrieval device 100 are realized by the processing circuitry. The related information retrieval device 100 according to other embodiments may also have a configuration similar to this modified example.
[0053] Embodiment 2. The following will explain the differences from the previously described embodiment, primarily with reference to the drawings. This embodiment aims to appropriately remove noise caused by the cross-sectional connectivity of multiple databases.
[0054] ***Explanation of the structure*** Figure 16 shows an example of the configuration of the related information retrieval device 100 according to this embodiment. The related information retrieval device 100 further includes a connection condition specification unit 210.
[0055] The connection condition specification unit 210 specifies the connection conditions. The connection condition specification unit 210 is also called the registration condition specification unit. The connection conditions are conditions that the two nodes to be connected must satisfy. The connection conditions may also be conditions relating to the combination of the node type of the first node and the node type of the second node. The connection condition specification unit 210 sets constraints on multiple nodes linked by the static link unit 130 according to a knowledge graph schema defined in advance by the user. These constraints constitute connection conditions. Connection conditions are conditions for linking nodes that have a semantically similar relationship. As a specific example, in this embodiment, instead of calculating the similarity between all nodes and drawing edges, a restriction is imposed so that only nodes corresponding to combinations of node types listed as "SIMILAR_TO" in the table shown in Figure 17 are considered candidates for drawing SIMILAR_TO edges. Figure 18 corresponds to Figure 17 and illustrates the processing of the static link unit 130 according to this embodiment. As a specific example, the static link unit 130 links nodes whose node type is Defect, according to the table shown in Figure 17.
[0056] The static link unit 130 according to this embodiment connects the first node and the second node when the connection conditions are met for both the first node and the second node. As a specific example, if the static link section 130 has similar or related node types, it aggregates the similarity of multiple nodes, each having a similar or related node type, to determine whether or not to create an edge for multiple nodes. Figure 19 is a diagram illustrating the processing of the static link unit 130. In Figure 19, each case corresponds to two nodes. In the example shown in Figure 19, "D (failure)" and "P (phenomenon)" are similar to or related to each other. Therefore, the static link unit 130 will create an edge only when both the failures and the phenomena are similar for each pair of nodes having "D (failure)" and "P (phenomenon)" as node types. In the cases 1 and 2 shown in Figure 19, the defects are not similar, so extending an edge between the two nodes is likely to link phenomena that are related to completely different failures. In the case of case 5, the phenomena are not similar, so extending an edge between the two nodes is likely to link failures that are related to completely different phenomena. If edges were extended between the two nodes corresponding to each of cases 1, 2, and 5, each extended edge is likely to become noise. Therefore, the static link unit 130 does not extend edges between these pairs of nodes. In the cases 3 and 4 shown in Figure 19, since the contents of both the Defect and the Phenomenon are similar, the static link section 130 creates edges between the two nodes.
[0057] ***Explanation of operation*** Figure 20 is a flowchart illustrating an example of the operation of the related information retrieval device 100. This operation will be explained using Figure 20.
[0058] (Step S201) The connection condition specification unit 210 reads the schema of the knowledge graph that the user has defined in advance. The read schema indicates the connection conditions.
[0059] (Step S202) The static linking unit 130 links together nodes that are semantically similar and satisfy the connection conditions indicated by the loaded schema.
[0060] (Step S203) The knowledge registration unit 120 updates the knowledge graph DB 190 via the DB operation unit 140 so that edges are drawn between nodes linked by the static link unit 130.
[0061] ***Explanation of the effects of Embodiment 2*** As described above, according to this embodiment, by setting constraints on the candidates for edge formation, it is possible to exclude edges that should not be connected as use cases, even if they have a semantic connection locally. Therefore, according to this embodiment, more effective search becomes possible.
[0062] Embodiment 3. The following will explain the differences from the previously described embodiment, primarily with reference to the drawings. Even when using the knowledge graph DB190, which consists of registered knowledge graphs with noise removed, depending on the user's search criteria, the search results may include a mix of information the user wants and information that is unnecessary for the user. Therefore, the objective of this embodiment is to allow the user to specify the node type or the relationship between node types that they want to search for.
[0063] ***Explanation of the structure*** Figure 21 shows an example of the configuration of the related information retrieval device 100 according to this embodiment. The related information retrieval device 100 further includes a search condition specification unit 310.
[0064] The search condition specification unit 310 specifies the search conditions. The search conditions are conditions that each node being searched must satisfy. The search conditions are conditions corresponding to the node type of an intermediate node, and may also be conditions corresponding to the node type that the related information inference unit 150 should follow after the intermediate node. An intermediate node is either a parent node or a node reached by following edges from the parent node. As a specific example, the search condition specification unit 310 receives instructions from the user regarding the node type to be searched for, or the relationship between them. Specifically, by having the search condition specification unit 310 accept the specification of at least one node type between a parent node and a child node, relationships that are not necessary for the search can be explicitly excluded. The search condition specification unit 310 specifies the received instructions as search conditions. Figure 22 is a diagram illustrating the processing of the search condition specification unit 310. The search condition specification unit 310 accepts the specification of the node type of the parent node and the specification of the node type of the child node.
[0065] The related information inference unit 150 in this embodiment searches for each node that satisfies the search conditions, that is, it acquires each node according to the search conditions specified by the search condition specification unit 310.
[0066] ***Explanation of operation*** Figure 23 is a flowchart illustrating an example of the operation of the related information retrieval device 100. This operation will be explained using Figure 23.
[0067] (Step S301) The search condition specification unit 310 reads the node type of each node defined in advance by the user and treats the read node type as a search condition.
[0068] (Step S302) This step is the same as step S103. However, the related information inference unit 150 searches for nodes according to the search conditions.
[0069] (Step S303) This step is the same as step S104. However, the related information inference unit 150 obtains each related node according to the search conditions.
[0070] ***Explanation of the effects of Embodiment 3*** As described above, according to this embodiment, the search target is limited to nodes that meet the user's requirements. Therefore, according to this embodiment, it becomes easier for the user to find the information they want.
[0071] Embodiment 4. The following will explain the differences from the previously described embodiment, primarily with reference to the drawings.
[0072] ***Explanation of the structure*** Figure 24 shows an example configuration of the related information retrieval device 100 according to this embodiment. The related information retrieval device 100 further comprises a dynamic link section 410.
[0073] The dynamic link unit 410 edits at least one of the first knowledge graph and the second knowledge graph based on at least a portion of the range traversed when searching for the target node. The dynamic linking unit 410 reads the path taken by the related information inference unit 150 and dynamically edits the knowledge graph based on the read path. Specifically, the dynamic linking unit 410 filters each node or dynamically calculates the importance of each node. As another specific example, the dynamic linking unit 410 uses the information of the nodes and edges taken by the related information inference unit 150, or the path taken by the related information inference unit 150 from parent nodes to child nodes, grandchild nodes, and so on as context to dynamically connect nodes or dynamically calculate importance, thereby thinning out some of the statically linked edges. In this case, the dynamic linking unit 410 may delete edges linked by the static linking unit 130 that do not have a similar relationship with the embedded subgraph containing the path taken by the related information inference unit 150.
[0074] Figure 25 is a diagram illustrating the processing of the dynamic link unit 410. In the example shown in Figure 25, the association information inference unit 150 is assumed to have traversed from the parent node to the child node. Furthermore, each node with a diagonal line drawn inside is assumed to be a candidate for a grandchild node. At this time, the dynamic link unit 410 extracts a subgraph that includes the area surrounding each candidate grandchild node. Next, the dynamic link unit 410 retains the grandchild node corresponding to the extracted subgraph if the vector obtained by graph embedding corresponding to the extracted subgraph is similar to the vector obtained by graph embedding corresponding to the path followed by the related information inference unit 150. Through the processing of the dynamic link unit 410, as a concrete example, only countermeasures with common components or operations can be inferred.
[0075] The related information inference unit 150 in this embodiment may search for a target destination node if the similarity between the embedded representation of the subgraph corresponding to the target destination node and the embedded representation of the subgraph corresponding to at least a portion of the traced edges is equal to or greater than the representation similarity threshold. The target destination node is each node that can be traced from the parent node. The representation similarity threshold can be determined in any way.
[0076] The relevance calculation unit 160 in this embodiment uses the knowledge graph edited by the dynamic link unit 410 when calculating the relevance. Specifically, the relevance calculation unit 160 uses the edited first knowledge graph when the first knowledge graph is edited, and uses the edited second knowledge graph when the second knowledge graph is edited.
[0077] ***Explanation of operation*** Figure 26 is a flowchart illustrating an example of the operation of the related information retrieval device 100. This operation will be explained using Figure 26.
[0078] (Step S401) The dynamic link unit 410 reads the path taken by the related information inference unit 150 and edits the knowledge graph based on the read path.
[0079] ***Explanation of the effects of Embodiment 4*** As described above, according to this embodiment, by dynamically editing the knowledge graph, it is possible to display more accurate search results.
[0080] ***Other Embodiments*** The embodiments described above can be freely combined, any component of each embodiment can be modified, or any component can be omitted in each embodiment. Furthermore, the embodiments are not limited to those shown in Embodiments 1 to 4, and various modifications can be made as needed. The procedures described using flowcharts and the like may be modified as appropriate. [Explanation of Symbols]
[0081] 11 Processor, 12 Volatile memory device, 13 Non-volatile memory device, 14 Interface, 18 Processing circuit, 100 Related information retrieval device, 110 Knowledge acquisition unit, 120 Knowledge registration unit, 130 Static link unit, 140 DB operation unit, 150 Related information inference unit, 160 Relevance calculation unit, 170 Display data generation unit, 180 User interface unit, 190 Knowledge graph DB, 210 Connection condition specification unit, 310 Search condition specification unit, 410 Dynamic link unit.
Claims
1. A static link unit connects the first node and the second node when the similarity between the first text corresponding to the first node in the first knowledge graph corresponding to the document and the second text corresponding to the second node in the second knowledge graph corresponding to the document is greater than or equal to the similarity threshold. A connection condition specification unit that specifies the connection conditions that the two connected nodes must satisfy, and A related information retrieval device comprising: The static link unit connects the first node and the second node when the first node and the second node satisfy the connection conditions. The aforementioned connection conditions are conditions relating to the combination of the node type of the first node and the node type of the second node in the related information retrieval device.
2. A static link unit connects the first node and the second node when the similarity between the first text corresponding to the first node in the first knowledge graph corresponding to the document and the second text corresponding to the second node in the second knowledge graph corresponding to the document is greater than or equal to the similarity threshold. A related information inference unit selects a parent node from the first knowledge graph and the second knowledge graph based on the text indicated by the query, and searches for one or more nodes by traversing edges from the parent node in at least one of the first knowledge graph and the second knowledge graph. A correlation calculation unit that takes each of the one or more nodes found as a target node and calculates the correlation degree corresponding to the target node based on the path from the parent node to the target node. A related information retrieval device equipped with the following features.
3. The related information retrieval device according to claim 1 or 2, wherein the target similarity is the similarity between the vector corresponding to the first text and the vector corresponding to the second text.
4. The static link portion calculates the target similarity based on the edit distance in the related information retrieval device according to claim 1 or 2.
5. The aforementioned related information retrieval device further, Connection condition specification section: Specifies the connection conditions that the two connected nodes must satisfy. Equipped with, The related information retrieval device according to claim 2, wherein the static link section connects the first node and the second node when the first node and the second node satisfy the connection conditions.
6. The related information retrieval device according to claim 5, wherein the connection condition is a condition relating to a combination of the node type of the first node and the node type of the second node.
7. The aforementioned related information retrieval device further, A related information inference unit selects a parent node from the first knowledge graph and the second knowledge graph based on the text indicated by the query, and searches for one or more nodes by traversing edges from the parent node in at least one of the first knowledge graph and the second knowledge graph. A correlation calculation unit that takes each of the one or more nodes found as a target node and calculates the correlation degree corresponding to the target node based on the path from the parent node to the target node. The related information retrieval device according to claim 1, comprising:
8. The aforementioned related information retrieval device further, Search condition specification section: Specifies the search conditions that each node being searched must satisfy. Equipped with, The related information retrieval device according to claim 2 or 7, wherein the related information inference unit searches each node that satisfies the search conditions.
9. The related information retrieval device according to claim 8, wherein the search condition is a condition corresponding to the node type of an intermediate node which is either the parent node or a node reached by traversing an edge from the parent node, and the condition corresponds to the node type that the related information inference unit should traverse to the next node after the intermediate node.
10. The aforementioned related information retrieval device further, A dynamic linking unit that edits at least one of the first knowledge graph and the second knowledge graph based on at least a portion of the range traversed when searching for the target node. Equipped with, The related information retrieval device according to claim 2 or 7, wherein the relatedness calculation unit uses the edited first knowledge graph when the first knowledge graph is edited, and uses the edited second knowledge graph when the second knowledge graph is edited.
11. The related information retrieval device according to claim 10, wherein each node that can be traced from the parent node is designated as a target destination node, and the related information inference unit searches for the target destination node if the similarity between the embedded representation of the subgraph corresponding to the target destination node and the embedded representation of the subgraph corresponding to at least a portion of the traced edge is equal to or greater than the representation similarity threshold.
12. The computer connects the first node and the second node if the object similarity, which is the similarity between the first text corresponding to the first node in the first knowledge graph corresponding to the document and the second text corresponding to the second node in the second knowledge graph corresponding to the document, is equal to or greater than the similarity threshold. The aforementioned computer specifies connection conditions, which are conditions that the two connected nodes must satisfy. A related information retrieval method in which the computer connects the first node and the second node when the connection conditions are met, The connection conditions are related to a combination of node types possessed by the first node and node types possessed by the second node, and are related information retrieval method.
13. A static linking process is performed to connect the first node and the second node when the similarity between the first text corresponding to the first node in the first knowledge graph corresponding to the document and the second text corresponding to the second node in the second knowledge graph corresponding to the document is equal to or greater than the similarity threshold. A connection condition specification process that specifies the connection conditions that the two nodes to be connected must satisfy, and A related information retrieval program that causes a related information retrieval device, which is a computer, to execute, In the static linking process, the first node and the second node are connected when the connection conditions are met. The aforementioned connection conditions are related information retrieval program conditions relating to the combination of node types possessed by the first node and node types possessed by the second node.
14. The computer connects the first node and the second node if the object similarity, which is the similarity between the first text corresponding to the first node in the first knowledge graph corresponding to the document and the second text corresponding to the second node in the second knowledge graph corresponding to the document, is equal to or greater than the similarity threshold. The computer selects a parent node from the first knowledge graph and the second knowledge graph based on the text indicated by the query, and searches for one or more nodes by traversing edges from the parent node in at least one of the first knowledge graph and the second knowledge graph. A related information retrieval method in which the computer determines each of the one or more nodes found as a target node and calculates the degree of relevance corresponding to the target node based on the path from the parent node to the target node.
15. A static linking process is performed to connect the first node and the second node when the similarity between the first text corresponding to the first node in the first knowledge graph corresponding to the document and the second text corresponding to the second node in the second knowledge graph corresponding to the document is equal to or greater than the similarity threshold. A related information inference process that selects a parent node from the first knowledge graph and the second knowledge graph based on the text indicated by the query, and searches for one or more nodes by traversing edges from the parent node in at least one of the first knowledge graph and the second knowledge graph, A correlation calculation process that takes each of the one or more nodes found as a target node and calculates the correlation to the target node based on the path from the parent node to the target node. A related information retrieval program that causes a related information retrieval device, which is a computer, to execute.