Information processing device, information processing method, and program
Patent Information
- Application Number
- PCT/JP2026/012029
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-25
- Filing Date
- 2026-03-25
- Publication Date
- 2026-10-01
Smart Images

Figure JP2026012029_01102026_PF_FP_ABST
Abstract
Description
Information processing device, information processing method, and program
[0001] This invention relates to an information processing apparatus, an information processing method, and a program.
[0002] Knowledge graphs are known, which structure data obtained from databases and information based on entities and their relationships. For example, Non-Patent Document 1 discloses a technique in which an explanatory model trained using supervised learning, reinforcement learning, or unsupervised learning outputs an explanation corresponding to input text.
[0003] Patent No. 7599622
[0004] Conventional reinforcement learning methods optimize the search path by using labels that indicate the correct path in order to find inference paths in an incomplete knowledge graph. However, while conventional knowledge graphs are optimized through learning the search path, they do not take into account the diversity that reflects the characteristics of the paths. Therefore, conventional knowledge graphs have the problem that software agents are less likely to find undiscovered paths or relationships.
[0005] The present invention aims to provide an information processing device, an information processing method, and a program that can improve the diversity of searches in a knowledge graph.
[0006] To solve the above-mentioned problems and achieve the objective, the information processing device according to the present invention includes: a storage unit that stores data of a knowledge graph structured by associating multiple entities with feature vectors; a calculation unit that calculates the heterogeneity of features between the entities in the knowledge graph and the next entity associated with the entity by the feature vector; and an assignment unit that assigns reward data that can identify the calculated heterogeneity to the path between the entity and the next entity.
[0007] To solve the above-mentioned problems and achieve the objective, the present invention provides an information processing method for an information processing device that stores data of a knowledge graph structured by associating a plurality of entities with feature vectors in a storage unit, the method comprising: a calculation step of calculating the heterogeneity of features between an entity in the knowledge graph and the next entity associated with that entity by the feature vector; and a grant step of granting reward data that can identify the calculated heterogeneity to the path between the entity and the next entity.
[0008] To solve the above-mentioned problems and achieve the objective, the program according to the present invention causes an information processing device, which stores data of a knowledge graph structured by associating multiple entities with feature vectors in a storage unit, to perform a calculation step of calculating the heterogeneity of features between the entity in the knowledge graph and the next entity associated with the entity by the feature vector, and a grant step of granting reward data that can identify the calculated heterogeneity to the path between the entity and the next entity.
[0009] According to the present invention, the diversity of searches in a knowledge graph can be improved.
[0010] Figure 1 shows an example of the system configuration of a policy network. Figure 2 shows the label generation process for a self-supervised knowledge graph. Figure 3 is a diagram illustrating an example configuration of the information processing device according to the first embodiment. Figure 4 is a flowchart showing an example of the processing procedure of the information processing device according to the first embodiment. Figure 5 shows an example of assigning reward data to a knowledge graph. Figure 6 is a diagram illustrating an example configuration of the information processing device according to the second embodiment. Figure 7 is a flowchart showing an example of the processing procedure of the information processing device according to the second embodiment. Figure 8 shows an example of assigning exploration progress reward data to a knowledge graph. Figure 9 is a diagram illustrating another example of calculating heterogeneity reward data.
[0011] Hereinafter, embodiments according to the present invention will be described in detail with reference to the accompanying drawings. The present invention is not limited by this embodiment, and when there are a plurality of embodiments, the present invention also includes configurations formed by combining the respective embodiments. In the following embodiments, the same parts are denoted by the same reference numerals, and repeated descriptions are omitted.
[0012] (Policy Network) FIG. 1 is a diagram showing an example of the system configuration of a policy network. As shown in FIG. 1, the policy network 100 is used to predict the probability π=P(S_t, A_t) of possible actions at time t. The training data has a source entity: e_s, a relation: r_q, and a target entity: e_q, and is used to learn a path from the source entity e_s to the target entity e_q.
[0013] The policy network 100 includes an LSTM 110, an embedding unit 120, an encoder 130, and a softmax function 140.
[0014] The LSTM 110 is Long Short-Term Memory, which encodes state-dependent information into a vector h t . The LSTM 110 receives, as input, the vector representation of the edge r t-1 selected at the previous time t-1 and the visited entity e t . The LSTM 110 outputs the encoded vector h t to the embedding unit 120.
[0015] The embedding unit 120 receives a query r q and a vector h t as input, and concatenates the query r q and the vector h t . The embedding unit 120 inputs the query r q and the vector h t into a feedforward network having ReLU nonlinearity. The feedforward network encodes information about the query r q and outputs a feature quantity z tThe embedded part 120 generates the generated feature quantity z. t This is output to encoder 130.
[0016] The encoder 130 encodes all possible actions and sets a policy π for action selection. t Calculate the policy π. t This is the feature z of the action selection vector. t and Action embedding matrix A t This is determined by calculating the inner product of the action embedding matrix A. t This is, for example, a matrix of combinations of visitable entities and edges. The encoder 130 calculates the policy π t Output this to the softmax function 140.
[0017] The softmax function 140 is based on policy π. t The probability of each action is calculated based on this. This allows the policy network 100 to predict what action should be taken next. The softmax function 140 has a first stage 141 and a second stage 142.
[0018] Stage 141 involves self-supervised learning (SL). Self-supervised learning (SL) is a part of the training process in which the agent (software agent) learns using pre-generated labels. In the SL phase of self-supervised learning, the agent trains the policy network 100 and learns the correct actions (labels) in specific states.
[0019] The SL (Simulation Level) stage includes, for example, the processes of label generation, training of the policy network 100, and minimizing cross-entropy loss.
[0020] The labeling process involves the agent creating labels based on the correct path as it explores the environment (knowledge graph). t This generates a guideline for the agent, showing them the correct course of action at each step.
[0021] The training process for policy network 100 involves generating labels y tThe policy network 100 is trained using this method. The agent learns to select the correct action at this stage, and will be able to make better choices in subsequent reinforcement learning (RL) stages.
[0022] The process of minimizing cross-entropy loss is the policy π of the output of policy network 100. t The generated label y t Minimize the cross-entropy loss between (d(π) t , y t By doing this, the agent learns to make the optimal action choices.
[0023] Stage 2, 142, involves reinforcement learning (RL). Reinforcement learning (RL) is the stage where the agent further improves its actions to maximize rewards using a reward function. In the RL stage of reinforcement learning, the label is set to 1 when the agent reaches the target entity, i.e., when it reaches the correct answer, and to 0 otherwise.
[0024] (Knowledge Graph) Figure 2 shows the label generation process of a self-supervised knowledge graph. Figure 2 visually illustrates how the paths explored by the agent and their labels are generated. In Figure 2, the knowledge graph 200 structures data (information) into entities and their relationships, with the relationships between entities represented by nodes (points) and edges (lines). Entities are components of the knowledge graph and hold specific things or concepts such as people, places, and organizations in image or text format. The knowledge graph can intuitively show search results based on the relationships between things and concepts represented in images or text.
[0025] In scenario C1 of Figure 2, the knowledge graph 200 shows entities e0, e1, e3, e4, e5, e6, e7, and e8, which are adjacent nodes within three hops of the starting entity e2. In Figure 2, entities are represented by circles in the node, and only numbers are shown inside the circles. For example, in the case of entity e2, only "2" is shown inside the circle in the node. The knowledge graph 200 has entities e5 and e8 as the end entities, and all paths connecting entity e2 to entities e5 and e8 are correct paths. The dashed line B represents a missing link that needs to be inferred.
[0026] In scenario C2 of Figure 2, the knowledge graph 200 has had all self-loops removed except for the target entities e5 and e8. The self-loops are retained to allow the agent to remain on the target node after reaching it.
[0027] In scenario C3 of Figure 2, a breadth-first search (BFS) is performed on the knowledge graph 200 to find the correct path to reach the target entity. Breadth-first search involves visiting the starting node, adding its elements to a queue, repeatedly visiting adjacent unvisited nodes, removing nodes from the queue, and adding newly visited nodes to the queue for further exploration. Explored nodes are added to set M to avoid being processed again. Visited nodes are marked with an empty circle (white circle) to indicate that they will not be explored further.
[0028] In scenes C4 and C5 of Figure 2, all nodes in the correct path are added to set C, and parent nodes are added backward until the start entity is reached. This process is repeated recursively. In scene C4, set C has nodes 1, 4, and 3 added, which are in the correct path. In scene C5, set C has nodes 1, 4, 3, and 0 added, which are in the correct path.
[0029] In scene C6 of FIG. 2, finally, labels 300 are generated for all nodes included in set C=(1,4,3,0). For the label 300 of a node, the leftmost bit is marked as "0" if the destination has not been reached. Edges included in the correct path are labeled "1", and other edges are marked as "0". This label generation process can assist the agent in searching for a correct inference path on the knowledge graph 200 and efficiently determining the next action to be taken.
[0030] As shown in FIG. 2, in the conventional knowledge graph 200, the label 300 indicating a correct path is represented by 0 or 1 for the purpose of improving learning efficiency, so diversity that can reflect the characteristics of paths is not considered. For this reason, in the conventional knowledge graph 200, the probability that the agent finds undiscovered paths and relationships, that is, the probability of obtaining new knowledge, is low. In the present embodiment, efficient and diverse search is realized by providing a new reward that considers the characteristics between a starting entity and a next entity in the search of the knowledge graph 200.
[0031] [First Embodiment] (Information Processing Apparatus) A configuration example of an information processing apparatus according to a first embodiment will be described with reference to FIG. 3. FIG. 3 is a diagram for explaining the configuration example of the information processing apparatus according to the first embodiment.
[0032] As shown in FIG. 3, the information processing apparatus 10 includes a communication unit 20, a storage unit 22, and a control unit 24. The information processing apparatus 10 can be implemented by, for example, a server apparatus that manages and provides the knowledge graph 200, an agent apparatus that uses the knowledge graph 200, or the like.
[0033] The communication unit 20 is a communication interface that executes communication between the information processing apparatus 10 and an external apparatus. For example, the communication unit 20 executes communication between the information processing apparatus 10 and the external apparatus.
[0034] The storage unit 22 stores various types of information. The storage unit 22 stores information such as calculation contents of the control unit 24 and programs. The storage unit 22 includes at least one of, for example, a main storage device such as a RAM (Random Access Memory) and a ROM (Read Only Memory), and an external storage device such as an HDD (Hard Disk Drive).
[0035] The storage unit 22 stores a program D1. The program D1 includes a program that causes the control unit 24 of the information processing apparatus 10 to execute an information processing method and the like. The storage unit 22 stores knowledge graph data D2 capable of representing a knowledge graph 200. Note that the storage unit 22 may be configured to store the knowledge graph data D2 in a storage device or the like external to the information processing apparatus 10.
[0036] The knowledge graph data D2 includes data of the knowledge graph 200 that indicates a plurality of entities obtained from databases and information, relationships between the entities, and the like. For example, the knowledge graph data D2 is structured by associating an entity with other entities via feature vectors (attributes). The feature vectors indicate, for example, attributes, features, relationships and the like of the entity. For example, when the entity is the word "Tokyo Skytree", it is associated with other entities such as "634 m" and "Sumida Ward, Tokyo" via feature vectors such as "height" and "location". The storage unit 22 may store the knowledge graph data D2 permanently or temporarily. The storage unit 22 can store reward data D3 to be assigned to the knowledge graph 200. The reward data D3 includes information that enables identification of heterogeneity in paths between a plurality of entities in the knowledge graph 200.
[0037] The control unit 24 controls each part of the information processing device 10. The control unit 24 includes, for example, an information processing device such as a CPU (Central Processing Unit) or an MPU (Micro Processing Unit), and a storage device such as a RAM (Random Access Memory) or a ROM (Read Only Memory). The control unit 24 executes a program that controls the operation of the information processing device 10 according to the present invention. The control unit 24 may be implemented by an integrated circuit such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array). The control unit 24 may be implemented by a combination of hardware and software.
[0038] The control unit 24 comprises a calculation unit 30 and an assignment unit 32.
[0039] The calculation unit 30 calculates heterogeneity between an entity in the knowledge graph 200 and the next entity associated with that entity by its feature vector. Heterogeneity means that there are diverse features, elements, properties, relationships, etc., between entities in the knowledge graph 200. The calculation unit 30 calculates heterogeneity based on, for example, norms, cosine similarity, etc. Norms are the distance between the feature vector of an entity and the feature vector of the next entity. In particular, the L2 norm corresponds to the Euclidean distance. Cosine similarity is the cosine value of the angle between two vectors. In order to realize efficient and diverse searches, the calculation unit 30 calculates a high heterogeneity reward for paths with high heterogeneity between entities in the knowledge graph 200 and the next entity. In this embodiment, the calculation unit 30 calculates a reward of, for example, "2" when heterogeneity is high, "0" when heterogeneity is low, and "1" when heterogeneity is in between.
[0040] The assignment unit 32 assigns the heterogeneity-indicating reward data D3 calculated by the calculation unit 30 to the path between an entity in the knowledge graph 200 and the next entity. The assignment unit 32 associates the reward data D3 with the corresponding path in the knowledge graph data D2 and stores it in the storage unit 22. The assignment unit 32 assigns the heterogeneity-indicating reward data D3 to the path if the heterogeneity satisfies the determination condition. For example, if the determination condition is for determining a path where the distance between entities is large and the similarity is low, the assignment unit 32 associates the reward data D3 with the corresponding path in the knowledge graph data D2 if the heterogeneity satisfies the determination condition, and does not associate the reward data D3 with the path if the heterogeneity does not satisfy the determination condition. Here, the assignment unit 32 may determine heterogeneity based on either the distance between entities or the similarity, or both conditions alone.
[0041] The knowledge graph 200 has label data attached to the paths between entities and entities associated with feature vectors, allowing for the identification of the path used for exploration. In this case, the assignment unit 32 assigns reward data D3 to the paths in the knowledge graph 200 that have label data attached. This allows the information processing device 10 to assign new reward data D3 according to the heterogeneity between entities in the knowledge graph 200, thereby achieving both diversity and efficiency in the agent's exploration.
[0042] (Processing Procedure of the Information Processing Device) Figure 4 is a flowchart showing an example of the processing procedure of the information processing device according to the first embodiment. The processing procedure shown in Figure 4 is realized when the control unit 24 executes program D1.
[0043] As shown in Figure 4, the control unit 24 of the information processing device 10 acquires the feature vector of the entity to be calculated (step S101). For example, the control unit 24 acquires data representing the entity to be calculated and its feature vector from the knowledge graph data D2. When the processing in step S101 is completed, the control unit 24 proceeds to step S102.
[0044] The control unit 24 obtains the feature vector of the next entity (step S102). For example, the control unit 24 obtains data from the knowledge graph data D2 showing all of the next entities associated with the target vector of the entity to be calculated, and their feature vectors. When the processing in step S102 is completed, the control unit 24 proceeds to step S103.
[0045] The control unit 24 calculates heterogeneity based on the feature vectors of the entity and the next entity (step S103). For example, the control unit 24 compares the feature vectors of the entity and the next entity and calculates heterogeneity for each next entity based on norm, cosine similarity, etc. When the processing in step S103 is completed, the control unit 24 proceeds to step S104.
[0046] The control unit 24 assigns heterogeneous-identifiable reward data D3 to the paths in the knowledge graph 200 (step S104). For example, the control unit 24 creates calculated heterogeneous-identifiable reward data D3 and assigns it to the paths in the knowledge graph 200 between entities and the next entity. When the processing in step S104 is completed, the control unit 24 terminates the processing procedure shown in Figure 4.
[0047] In this way, the control unit 24 can assign reward data D3 to the paths of the knowledge graph 200 by executing the processing procedure shown in Figure 4 for each of the multiple entities.
[0048] (Processing details of the information processing device) Figure 5 shows an example of assigning reward data to a knowledge graph. In the example shown in Figure 5, the knowledge graph 200 shows entities e0, e1, e3, e4, e5, e6, e7, and e8, which are adjacent nodes within 3 hops from the starting entity e2, as described above. The knowledge graph 200 performs breadth-first search (BFS) and shows the correct path to reach the target entity in a tree structure. For example, the starting entity e2 shows only the number "2" inside a circle in the node, and the feature vector is shown as an arrow pointing to entities e0, e1, and e3, respectively. In the knowledge graph 200, the target entities e5 and e8 exist as the end entities, and all paths connecting entity e2 to entities e5 and e8 are correct paths. The knowledge graph 200 uses the reinforcement learning method already explained to find inference paths in the incomplete knowledge graph 200, and the search path is made more efficient by using labels 300 that indicate the correct path.
[0049] The information processing device 10 calculates the heterogeneity between the feature vector of the starting entity e2 and the feature vectors of the next entities e0, e1, and e3. The information processing device 10 assigns reward data D3 to paths with high heterogeneity (long distance and low similarity). For example, the information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to paths with high heterogeneity (long distance and low similarity). The information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to each path from entity e2 to the next entities e0, e1, and e3, thereby creating paths that combine labels 300 and reward data D3.
[0050] Next, the information processing device 10 calculates the heterogeneity between the feature vector of entity e0 and the feature vector of entity e1, which is the correct path among the next entities e1 and e2. If the heterogeneity is high, the information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to the path from entity e0 to the next entity e1, thereby creating a path that combines the label 300 and the reward data D3.
[0051] Next, the information processing device 10 calculates the heterogeneity between the feature vector of entity e1 and the feature vectors of entities e0, e4, and e5 that represent the correct path among the following entities e0, e2, e4, and e5. If the heterogeneity is high, the information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to the path from entity e1 to the following entities e0, e4, and e5, thereby creating a path that combines the label 300 and the reward data D3.
[0052] Next, the information processing device 10 calculates the heterogeneity between the feature vector of entity e3 and the feature vector of entity e8, which is the correct path among the following entities e2, e7, and e8. If the heterogeneity is high, the information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to the path from entity e3 to the next entity e8, thereby creating a path that combines the label 300 and the reward data D3.
[0053] Furthermore, the information processing device 10 calculates the heterogeneity between the feature vector of the next entity e4 following entity e1 and the feature vector of entity e5 of the next entities e1, e5, e6, and e7 that represent the correct path. If the heterogeneity is high, the information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to the path from entity e4 to the next entity e5, thereby creating a path that combines the label 300 and the reward data D3.
[0054] As described above, when the agent searches for an answer to a question from the knowledge graph 200, the information processing device 10 can increase the likelihood of finding undiscovered paths and relationships by prioritizing the reward data D3 when searching for an inference path. As a result, the information processing device 10 can achieve greater diversity in the paths to the target entity (answer) than by simply referring to the labels 300 in the knowledge graph 200 when searching for a path. Path diversity means reaching the same objective from different processes, including both basic and leap processes.
[0055] The knowledge graph 200 assigns labels 300 to paths between entities associated with the next entity by a feature vector, indicating whether or not it is a search path. The labeling unit 32 of the information processing device 10 can assign reward data D3 to paths with labels 300. This allows the information processing device 10 to achieve both diversity and efficiency in the search by assigning reward data D3 to paths with labels 300 that correspond to the heterogeneity between entities in the knowledge graph 200.
[0056] In this embodiment, the case in which the information processing device 10 has labels 300 attached to the knowledge graph 200 has been described, but it is not limited to this. For example, the information processing device 10 may be configured to attach only reward data D3 to the knowledge graph 200 which does not have labels 300 attached.
[0057] [Second Embodiment] A second embodiment will now be described. In the second embodiment, the information processing device 10 performs processing to add adaptive search capabilities to the knowledge graph 200.
[0058] (Information Processing Device) An example of the configuration of the information processing device according to the second embodiment will be explained using Figure 6. Figure 6 is a diagram illustrating an example of the configuration of the information processing device according to the second embodiment.
[0059] As shown in Figure 6, the information processing device 10 comprises a communication unit 20, a storage unit 22, and a control unit 24.
[0060] The memory unit 22 can store the exploration progress reward data D4. The exploration progress reward data D4 contains information that allows for the identification of heterogeneity, emphasizing diversity in the initial stages of exploration, focusing on paths with large feature vector distances, and gradually converging the focus of the exploration as the exploration progresses.
[0061] The control unit 24 comprises a calculation unit 30, an assignment unit 32, and a creation unit 34.
[0062] The creation unit 34 creates exploration progress reward data D4, in which heterogeneity converges as time elapses while exploring paths in the knowledge graph 200. The creation unit 34 determines, for example, whether the stage of the exploration process is the initial, middle, or late stage, based on the elapsed time since the reward data D3 was assigned. The initial, middle, and late stages of exploration are set, for example, by thresholds. In this embodiment, the case where the exploration process is divided into the initial, middle, and late stages of exploration is described, but for example, there may be two or four or more stages. For example, the creation unit 34 sets the reward for the initial stage of exploration to be the same as the heterogeneity reward described above, the reward for the middle stage of exploration to be "1", and the calculation result of 1 / heterogeneity reward for the late stage of exploration. The creation unit 34 creates the exploration progress reward data D4 based on the calculation result. As a result, the information processing device 10 can ensure diversity in the initial stage of exploring the knowledge graph 200 and ultimately obtain highly accurate results.
[0063] The assignment unit 32 assigns the exploration progress reward data D4 created by the creation unit 34 to the paths in the knowledge graph 200. If reward data D3 is already assigned to the paths in the knowledge graph 200, the assignment unit 32 assigns the exploration progress reward data D4 together with the reward data D3 by associating it with the corresponding paths in the knowledge graph data D2.
[0064] (Processing Procedure of Information Processing Device) Figure 7 is a flowchart showing an example of the processing procedure of the information processing device according to the second embodiment. The processing procedure shown in Figure 7 is realized by the control unit 24 executing program D1. Note that the processing from step S101 to step S103 is the same as the processing from step S101 to step S103 shown in Figure 4, so the explanation is omitted.
[0065] As shown in Figure 7, when the processing in step S103 is completed, the control unit 24 of the information processing device 10 creates search progress reward data D4 in which heterogeneity converges as the time elapsed in the path search (step S111). For example, based on the elapsed time since the reward data D3 was assigned, the control unit 24 determines whether the stage of the search is the initial, middle, or late stage of the search, and creates search progress reward data D4 that can identify heterogeneity according to that period.
[0066] The control unit 24 adds heterogeneity-identifiable reward data D3 and exploration progress reward data D4 to the paths in the knowledge graph 200 (step S112). For example, the control unit 24 creates the calculated heterogeneity-identifiable reward data D3 and adds it, along with the exploration progress reward data D4 created in step S111, to the paths in the knowledge graph 200 between entities and the next entity. When the processing in step S111 is completed, the control unit 24 terminates the processing procedure shown in Figure 7.
[0067] In this way, the control unit 24 can assign reward data D3 and exploration progress reward data D4 to the paths of the knowledge graph 200 by executing the processing procedure shown in Figure 7 for each of the multiple entities. As a result, the information processing device 10 can ensure diversity of heterogeneity between entities in the knowledge graph 200 at the initial stage of exploration, and then converge the heterogeneity as the exploration progresses, thereby ultimately obtaining highly accurate results.
[0068] (Processing details of the information processing device) Figure 8 shows an example of assigning search progress reward data to a knowledge graph. In the example shown in Figure 8, the knowledge graph 200 shows entities e0, e1, e3, e4, e5, e6, e7, and e8, which are adjacent nodes within 3 hops from the starting entity e2, as described above. The knowledge graph 200 performs breadth-first search (BFS) and shows the correct path to reach the target entity in a tree structure. For example, the starting entity e2 shows only the number "2" inside a circle in the node, and the feature vector is shown as an arrow pointing to entities e0, e1, and e3, respectively. In the knowledge graph 200, the target entities e5 and e8 exist as the end entities, and all paths connecting entity e2 to entities e5 and e8 are correct paths. The knowledge graph 200 uses the reinforcement learning method already described to find inference paths in the incomplete knowledge graph 200, and the search path is made more efficient by using labels 300 that indicate the correct path.
[0069] The information processing device 10 calculates the heterogeneity between the feature vector of the starting entity e2 and the feature vectors of the next entities e0, e1, and e3. The information processing device 10 assigns reward data D3 to paths with high heterogeneity (long distance and low similarity) and creates exploration progress reward data D4 based on that heterogeneity and the exploration progress. For example, the information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to paths with high heterogeneity (long distance and low similarity). The information processing device 10 creates exploration progress reward data D4 that identifies heterogeneity according to the exploration progress. The information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 and exploration progress reward data D4 to each path from entity e2 to the next entities e0, e1, and e3, thereby creating paths that combine the label 300, reward data D3, and exploration progress reward data D4.
[0070] Next, the information processing device 10 calculates the heterogeneity between the feature vector of entity e0 and the feature vector of entity e1, which is the correct path among the next entities e1 and e2. If the heterogeneity is high, the information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to the path from entity e0 to the next entity e1, thereby creating a path that combines the label 300 and the reward data D3. The information processing device 10 creates search progress reward data D4 that shows the heterogeneity according to the search progress, and assigns it to the path from entity e0 to the next entity e1, thereby creating a path that combines the label 300, the reward data D3, and the search progress reward data D4.
[0071] Next, the information processing device 10 calculates the heterogeneity between the feature vector of entity e1 and the feature vectors of entities e0, e4, and e5 that represent the correct path among the following entities e0, e2, e4, and e5. If the heterogeneity is high, the information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to the path from entity e1 to the next entities e0, e4, and e5, thereby creating a path that combines the label 300 and the reward data D3. The information processing device 10 creates search progress reward data D4 that shows the heterogeneity according to the search progress and assigns it to the path from entity e1 to the next entities e0, e4, and e5, thereby creating a path that combines the label 300, the reward data D3, and the search progress reward data D4.
[0072] Next, the information processing device 10 calculates the heterogeneity between the feature vector of entity e3 and the feature vector of entity e8, which is the correct path among the next entities e2, e7, and e8. If the heterogeneity is high, the information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to the path from entity e3 to the next entity e8, thereby creating a path that combines the label 300 and the reward data D3. The information processing device 10 creates search progress reward data D4 that shows the heterogeneity according to the search progress, and assigns it to the path from entity e3 to the next entity e8, thereby creating a path that combines the label 300, the reward data D3, and the search progress reward data D4.
[0073] Furthermore, the information processing device 10 calculates the heterogeneity between the feature vector of the next entity e4 following entity e1 and the feature vector of entity e5 of the next entities e1, e5, e6, and e7 that represent the correct path. If the heterogeneity is high, the information processing device 10 assigns the calculated heterogeneity-identifiable reward data D3 to the path from entity e4 to the next entity e5, thereby creating a path that combines the label 300 and the reward data D3. The information processing device 10 creates search progress reward data D4 that shows the heterogeneity according to the search progress and assigns it to the path from entity e4 to the next entity e5, thereby creating a path that combines the label 300, the reward data D3, and the search progress reward data D4.
[0074] As described above, the information processing device 10 can assign reward data D3 and exploration progress reward data D4 to the path from one entity to the next in the knowledge graph 200. As a result, when an agent searches for an answer to a question from the knowledge graph 200, the information processing device 10 prioritizes the exploration progress reward data D4 when searching for an inference path, thereby ensuring diversity in the early stages of the search and ultimately obtaining highly accurate results. Furthermore, since the information processing device 10 can change the process of reaching the search result between the early and late stages of the search, the diversity of the search can be further improved.
[0075] [Other Embodiments] Next, other embodiments will be described. The case in which the information processing device 10 calculates heterogeneity from the feature vector of one entity and the feature vector of the next entity has been described, but it is not limited to this.
[0076] Figure 9 illustrates another example of calculating heterogeneity reward data. In the example shown in Figure 9, the starting entity e0 of the knowledge graph 200 is associated with the four subsequent entities e1, e2, e3, and e4 by feature vectors. In this case, the information processing device 10 can be configured to select the entity e4 whose feature vectors are most different from the four subsequent entities e1, e2, e3, and e4, and to calculate heterogeneity from the feature vectors of entity e0 and the feature vectors of the subsequent entity e4. This allows the information processing device 10 to show the path with the highest heterogeneity in the reward data D3, thereby improving the possibility of finding undiscovered paths and relationships. For example, this configuration can calculate the average vector of each feature vector of the subsequent entities e1, e2, e3, and e4, and the path to the entity furthest from the calculated average vector can be considered the path with the highest heterogeneity from the starting entity e0.
[0077] Although embodiments of the present invention have been described above, the present invention is not limited by the content of these embodiments. Furthermore, the aforementioned components include those that can be easily conceived by those skilled in the art, those that are substantially the same, and those that fall within the so-called equivalent range. Moreover, the aforementioned components can be combined as appropriate. Furthermore, various omissions, substitutions, or modifications of the components can be made without departing from the gist of the embodiments described above.
[0078] 10 Information Processing Unit 20 Communication Unit 22 Storage Unit 24 Control Unit 30 Calculation Unit 32 Assignment Unit 34 Creation Unit 100 Policy Network 110 LSTM 120 Embedding Unit 130 Encoder 141 First Stage 142 Second Stage 200 Knowledge Graph 300 Label D1 Program D2 Knowledge Graph Data D3 Reward Data D4 Search Progress Reward Data
Claims
1. An information processing device comprising: a storage unit for storing data of a knowledge graph structured by associating multiple entities with feature vectors; a calculation unit for calculating the heterogeneity of features between an entity in the knowledge graph and the next entity associated with that entity by its feature vector; and an assignment unit for assigning reward data that can identify the calculated heterogeneity to the path between the entity and the next entity.
2. The information processing apparatus according to claim 1, wherein the knowledge graph is provided with labels indicating whether or not the paths between the entities and the entities associated with the feature vectors are search paths, and the assignment unit assigns the reward data to the paths to which the labels are assigned.
3. The information processing apparatus according to claim 2, further comprising a creation unit that creates exploration progress reward data in which the heterogeneity converges as time elapses while exploring the path, and the assignment unit assigns the exploration progress reward data to the path.
4. An information processing method for an information processing device that stores data of a knowledge graph structured by associating multiple entities with feature vectors in a storage unit, the method comprising: a calculation step of calculating feature heterogeneity between an entity in the knowledge graph and a next entity associated with that entity by its feature vector; and a grant step of granting reward data that can identify the calculated heterogeneity to a path between the entity and the next entity.
5. A program that causes an information processing device, which stores data of a knowledge graph structured by associating multiple entities with feature vectors in a storage unit, to execute: a calculation step of calculating the heterogeneity of features between an entity in the knowledge graph and the next entity associated with that entity by its feature vector; and a grant step of granting reward data that can identify the calculated heterogeneity to the path between the entity and the next entity.