Knowledge graph evaluation method and device, electronic equipment and computer storage medium

By building a training environment and evaluating the knowledge graph using deep Q networks, the problem of a large amount of manual supervision in the existing technology is solved, and the automatic evaluation and update of the knowledge graph is realized, and efficiency and accuracy are improved.

CN120105042AInactive Publication Date: 2025-06-06THREE GORGES GROUP IND DEVELOPMENT (BEIJING) CO LTD +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510590183.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-08
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the construction and update of the knowledge graph requires a lot of manual supervision and expert annotations, which are inefficient and costly.

Method used

By building a training environment for training neural network models, using deep Q networks to evaluate and update the knowledge graph, automatically adjust the confidence of triples, and reduce dependence on manual supervision.

Benefits of technology

Automatic evaluation and update of knowledge graphs is realized, evaluation efficiency and accuracy are improved, labor costs are reduced, and reasoning errors caused by errors or low-quality knowledge are avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105042A_ABST
    Figure CN120105042A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a knowledge graph evaluation method and device, electronic equipment and a computer readable storage medium, and relates to the technical field of knowledge graphs, and the method comprises the steps: an intelligent agent carries out the initialization of an original neural network model, obtains an initial neural network model, and constructs a training environment for training the initial neural network model; performing data preprocessing on the sample knowledge graph for training to obtain to-be-processed data; training the initial neural network model by using the to-be-processed data in the training environment to obtain a trained target neural network model; and adopting the target neural network model to perform confidence measurement on a to-be-measured knowledge graph to obtain a measurement result. According to the embodiment of the invention, the efficiency of confidence evaluation of the intelligent agent by adopting the trained neural network model is improved, and the problem that a large amount of manual supervision and manual annotation are needed is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of knowledge graph technology, and in particular to a knowledge graph evaluation method, a knowledge graph evaluation device, an electronic device and a computer-readable storage medium. Background Art

[0002] With the continuous enrichment of data from hydropower stations and renewable energy power stations, knowledge graphs have been widely used in production operations, fault diagnosis, expert database decision support and other fields. Therefore, timely updating and correction of knowledge graphs is necessary. Most traditional knowledge graph construction methods require a lot of manual supervision or expert annotations, but with the expansion of the scale of the power generation industry, manual supervision or expert annotation methods are not only inefficient, but also cost a lot of labor costs. Summary of the invention

[0003] In view of the above problems, embodiments of the present invention are proposed to provide a knowledge graph evaluation method, a knowledge graph evaluation device, an electronic device and a computer-readable storage medium that overcome the above problems or at least partially solve the above problems.

[0004] In order to solve the above problems, an embodiment of the present invention discloses a knowledge graph evaluation method, the method comprising: initializing an original neural network model to obtain an initial neural network model, and constructing a training environment for training the initial neural network model; Perform data preprocessing on the sample knowledge graph used for training to obtain data to be processed; Using the data to be processed to train the initial neural network model in the training environment to obtain a trained target neural network model; The target neural network model is used to perform confidence measurement on the knowledge graph to be measured to obtain a measurement result.

[0005] In one or more embodiments, the initializing the original neural network model to obtain the initial neural network model includes: Construct a neural network structure based on the deep Q network to obtain the original neural network model; The weight coefficients and model parameters of the original neural network model are initialized to obtain an initial neural network model.

[0006] In one or more embodiments, constructing a training environment for training the initial neural network model includes: The subject, relationship and subject in each triple in the sample knowledge graph form a state, and the triple states corresponding to each triple are obtained; Set the action to be performed based on the confidence of the triple; Set the reward for the accuracy of the triplet confidence estimate for the action performed; Set the termination conditions for training; An experience replay buffer is set; the experience replay buffer is used to store the experience generated during the training process.

[0007] In one or more embodiments, the data preprocessing of the sample knowledge graph for training to obtain the data to be processed includes: Acquire at least one triple data from the sample knowledge graph; Divide each triple data into training set, validation set and test set; Each triplet data in the training set, the validation set and the test set is encoded to obtain an encoded sample training set, a sample validation set and a sample test set.

[0008] In one or more embodiments, dividing each triple data to obtain a training set, a validation set, and a test set includes: The random stratified sampling method is used to divide each triple data into training set, validation set and test set.

[0009] In one or more embodiments, the using the to-be-processed data to train the initial neural network model in the training environment to obtain a trained target neural network model includes: S1, obtaining a current triple to be processed from a sample training set in the data to be processed; S2. Obtain the current triplet state corresponding to the current triplet; S3, determining a target action corresponding to the current triple state from the training environment according to the current triple state; S4, executing the target action; S5, obtaining corresponding experience feedback when the target action is completed; S6, storing the experience feedback in an experience playback buffer; S7, obtaining target experience feedback from the experience playback buffer; S8, using the target experience feedback to update the initial neural network model to obtain an updated candidate neural network model; S9, determining whether the termination condition in the training environment is met, if so, obtaining the trained target neural network model; if not, executing S10; S10, obtaining the next triplet from the sample training set in the data to be processed, taking the next triplet as the current triplet, and executing S2.

[0010] In one or more embodiments, determining a target action corresponding to the current triple state from the training environment according to the current triple state includes: According to the current triplet state, a target action corresponding to the current triplet state is selected from the training environment using a greedy strategy.

[0011] In one or more embodiments, the step of selecting a target action corresponding to the current triple state from the training environment using a greedy strategy according to the current triple state includes: According to the current triplet state, randomly select a first action from the training environment with probability e; Using the initial neural network model to calculate the Q value corresponding to each action in the training environment, and selecting the second action with the largest Q value with probability 1-e; The action with a higher probability between the first action and the second action is determined as the target action.

[0012] In one or more embodiments, obtaining experience feedback corresponding to the completion of executing the target action includes: Obtaining the next triplet state and reward of the current triplet from the training environment, and obtaining a termination flag at the current moment; the reward is calculated by the training environment using a preset reward function, and the termination flag is used to indicate whether the training is terminated; The current triplet state, the target action, the reward, the next triplet state and the termination flag are combined to obtain experience feedback corresponding to the completion of the target action.

[0013] In one or more embodiments, obtaining target experience feedback from the experience playback buffer includes: Determining the number of samples used to update the initial neural network model; A target number of experience feedback samples is obtained from the experience playback buffer.

[0014] In one or more embodiments, the updating of the initial neural network model using the target experience feedback to obtain an updated candidate neural network model includes: The initial neural network model is used to calculate the Q value estimation of executing the target action in the current triple state in each target experience feedback; Calculate the target Q value corresponding to each target experience feedback; The loss function corresponding to each target experience feedback is calculated using the Q value estimate and the target Q value corresponding to each target experience feedback; The initial neural network model is updated using the loss function corresponding to each target experience feedback to obtain an updated candidate neural network model.

[0015] In one or more embodiments, calculating the target Q value corresponding to each target experience feedback includes: For any target experience feedback, when the termination flag is reaching the termination state, the corresponding target Q value is the reward; When the termination flag is not the reaching termination state, the target Q value is calculated using the reward, the next triplet state and a preset discount factor.

[0016] In one or more embodiments, the using the Q value estimate and the target Q value corresponding to each target experience feedback to calculate the loss function corresponding to each target experience feedback includes: The mean square error is used as the loss function to calculate the error between the Q value estimate corresponding to each target experience feedback and the target Q value.

[0017] In one or more embodiments, the determining whether a termination condition in the training environment is satisfied includes: Determine whether at least one of the following conditions is met: a maximum number of training steps is reached and a performance indicator of the validation set reaches a threshold; If so, the termination condition in the training environment is met.

[0018] In one or more embodiments, further comprising: If the confidence in the measurement result is different from the confidence in the knowledge graph to be measured, the confidence in the measurement result is used to replace the confidence in the knowledge graph to be measured.

[0019] Accordingly, an embodiment of the present invention discloses a knowledge graph evaluation device, the device comprising: An initialization module is used to initialize the original neural network model to obtain an initial neural network model; A construction module, used to construct a training environment for training the initial neural network model; A preprocessing module is used to preprocess the sample knowledge graph used for training to obtain data to be processed; A training module, used to train the initial neural network model using the data to be processed in the training environment to obtain a trained target neural network model; The measurement module is used to use the target neural network model to perform confidence measurement on the knowledge graph to be measured to obtain a measurement result.

[0020] In one or more embodiments, the initialization module is specifically used to: Construct a neural network structure based on the deep Q network to obtain the original neural network model; The weight coefficients and model parameters of the original neural network model are initialized to obtain an initial neural network model.

[0021] In one or more embodiments, the building block is specifically used to: The subject, relationship and subject in each triple in the sample knowledge graph form a state, and the triple states corresponding to each triple are obtained; Set the action to be performed based on the confidence of the triple; Set the reward for the accuracy of the triplet confidence estimate for the action performed; Set the termination conditions for training; An experience replay buffer is set; the experience replay buffer is used to store the experience generated during the training process.

[0022] In one or more embodiments, the preprocessing module comprises: A first acquisition submodule, used to acquire at least one triple data from the sample knowledge graph; The partitioning submodule is used to partition each triplet data into a training set, a validation set, and a test set; The encoding submodule is used to encode each triplet data in the training set, the verification set and the test set to obtain the encoded sample training set, sample verification set and sample test set.

[0023] In one or more embodiments, the division submodule is specifically used to: The random stratified sampling method is used to divide each triple data into training set, validation set and test set.

[0024] In one or more embodiments, the training module includes: A second acquisition submodule is used to acquire the current triple to be processed from the sample training set in the data to be processed; A third acquisition submodule is used to acquire a current triplet state corresponding to the current triplet; A determination submodule, configured to determine a target action corresponding to the current triple state from the training environment according to the current triple state; An execution submodule, used for executing the target action; A fourth acquisition submodule is used to obtain experience feedback corresponding to the completion of the target action; A storage submodule, used for storing the experience feedback in an experience playback buffer; A fifth acquisition submodule, used to acquire target experience feedback from the experience playback buffer; An updating submodule, used to update the initial neural network model using the target experience feedback to obtain an updated candidate neural network model; A judgment submodule is used to determine whether the termination condition in the training environment is met, and if so, obtain the trained target neural network model; if not, call the sixth acquisition submodule; The sixth acquisition submodule is used to acquire the next triplet from the sample training set in the data to be processed, take the next triplet as the current triplet, and call the second acquisition submodule.

[0025] In one or more embodiments, the determining submodule is specifically used to: According to the current triplet state, a target action corresponding to the current triplet state is selected from the training environment using a greedy strategy.

[0026] In one or more embodiments, the determining submodule is further used to: According to the current triplet state, randomly select a first action from the training environment with probability e; Using the initial neural network model to calculate the Q value corresponding to each action in the training environment, and selecting the second action with the largest Q value with probability 1-e; The action with a higher probability between the first action and the second action is determined as the target action.

[0027] In one or more embodiments, the fourth acquisition submodule is specifically used to: Obtaining the next triplet state and reward of the current triplet from the training environment, and obtaining a termination flag at the current moment; the reward is calculated by the training environment using a preset reward function, and the termination flag is used to indicate whether the training is terminated; The current triplet state, the target action, the reward, the next triplet state and the termination flag are combined to obtain experience feedback corresponding to the completion of the target action.

[0028] In one or more embodiments, the fifth acquisition submodule is specifically used to: Determining the number of samples used to update the initial neural network model; A target number of experience feedback samples is obtained from the experience playback buffer.

[0029] In one or more embodiments, the updating submodule includes: A first calculation unit, configured to calculate, using the initial neural network model, a Q value estimate of executing a target action in a current triplet state in each target experience feedback; The second calculation unit is used to calculate the target Q value corresponding to each target experience feedback; A third calculation unit is used to calculate the loss function corresponding to each target experience feedback by using the Q value estimate and the target Q value corresponding to each target experience feedback; An updating unit is used to update the initial neural network model using a loss function corresponding to each target experience feedback to obtain an updated candidate neural network model.

[0030] In one or more embodiments, the second computing unit is specifically configured to: For any target experience feedback, when the termination flag is reaching the termination state, the corresponding target Q value is the reward; When the termination flag is not the reaching termination state, the target Q value is calculated using the reward, the next triplet state and a preset discount factor.

[0031] In one or more embodiments, the third computing unit is specifically configured to: The mean square error is used as the loss function to calculate the error between the Q value estimate corresponding to each target experience feedback and the target Q value.

[0032] In one or more embodiments, the judgment submodule is specifically used to: Determine whether at least one of the following conditions is met: a maximum number of training steps is reached and a performance indicator of the validation set reaches a threshold; If so, the termination condition in the training environment is met.

[0033] In one or more embodiments, further comprising: An updating module is used to replace the confidence in the knowledge graph to be measured with the confidence in the measurement result if the confidence in the measurement result is different from the confidence in the knowledge graph to be measured.

[0034] Correspondingly, an embodiment of the present invention discloses an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program implements the various steps of the above-mentioned knowledge graph evaluation method embodiment when executed by the processor.

[0035] Correspondingly, an embodiment of the present invention discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various steps of the above-mentioned knowledge graph evaluation method embodiment are implemented.

[0036] The embodiments of the present invention include the following advantages: The agent initializes the original neural network model to obtain the initial neural network model, and constructs a training environment for training the initial neural network model; performs data preprocessing on the sample knowledge graph used for training to obtain the data to be processed; uses the data to be processed to train the initial neural network model in the training environment to obtain the trained target neural network model; uses the target neural network model to measure the confidence of the knowledge graph to be measured to obtain the measurement result. In this way, the agent uses the sample knowledge graph in the constructed training environment to train the neural network model so that the neural network model can quantify its semantic correctness and the truth of the facts expressed. During the training process, the confidence, as the weight or credibility indicator of the reasoning path, can guide the reasoning in a more reliable direction. The reasonable use of confidence can avoid reasoning errors caused by erroneous or low-quality knowledge, and improve the accuracy and interpretability of the reasoning results. This enables the agent to automatically evaluate the confidence of the triples through reinforcement learning, and learn the optimal strategy through the rewards fed back by the training environment, without the need for manual judgment one by one. Moreover, the experience replay mechanism enables the agent to learn from historical experience, thereby further optimizing the confidence assessment and improving the efficiency of the agent in confidence assessment using the trained neural network model, solving the problem of requiring a large amount of manual supervision and manual annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 It is a flowchart of the steps of Embodiment 1 of a knowledge graph evaluation method of the present invention; Figure 2 This is a flowchart of the steps of Embodiment 2 of a knowledge graph evaluation method of the present invention; Figure 3 It is a structural block diagram of an embodiment of a knowledge graph evaluation device of the present invention. DETAILED DESCRIPTION

[0038] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0039] One of the core concepts of the embodiment of the present invention is that the intelligent agent uses a sample knowledge graph in a constructed training environment to train the neural network model so that the neural network model can quantify its semantic correctness and the truth of the facts expressed. During the training process, confidence, as the weight or credibility indicator of the reasoning path, can guide the reasoning in a more reliable direction. Reasonable use of confidence can avoid reasoning errors caused by erroneous or low-quality knowledge, and improve the accuracy and interpretability of the reasoning results. This enables the intelligent agent to automatically evaluate the confidence of triples through reinforcement learning, and learn the optimal strategy through rewards fed back by the training environment, without the need for manual judgment one by one. Moreover, the experience replay mechanism allows the intelligent agent to learn from historical experience, thereby further optimizing the confidence assessment, improving the efficiency of the intelligent agent using the trained neural network model for confidence assessment, and solving the problem of requiring a large amount of manual supervision and manual annotation.

[0040] Reference Figure 1 , showing a flow chart of the steps of the first embodiment of a knowledge graph evaluation method of the present invention, which can be applied to an agent, which refers to an agent that can perceive the environment and take actions to achieve specific goals. It can be software, hardware or a system with autonomy, adaptability and interaction capabilities. The agent perceives changes in the environment (such as through sensors or data input), makes judgments and decisions based on the knowledge and algorithms it has learned, and then performs actions to affect the environment or achieve predetermined goals. Agents are widely used in the field of artificial intelligence, and are commonly found in automated systems, robots, virtual assistants, and game characters. The core of agents is the ability to learn autonomously and continuously evolve to better complete tasks and adapt to complex environments.

[0041] Based on this, in an embodiment of the present invention, the intelligent agent obtains information by receiving the state of triples in the knowledge graph (including the triples themselves and contextual information); it has autonomy and can make decisions independently according to its own neural network structure and learning algorithm (such as the strategy of DQN (Deep Q-Network)), and choose actions to adjust the confidence of triples; it has adaptability and gradually adapts to the characteristics of the knowledge graph by continuously receiving environmental feedback and updating neural model parameters during the training process, and optimizes the judgment of triple confidence; it also has interactive capabilities and can interact with the knowledge graph environment to obtain the new state, reward and termination flag returned by the training environment according to the action.

[0042] That is to say, in the whole method, starting from initializing the neural network structure, the agent selects actions according to the triplet state in the training environment, performs operations to adjust the confidence, receives environmental feedback, and then continuously optimizes its decision-making ability through experience playback and parameter updates, and finally realizes the measurement of the confidence of the knowledge graph triplet. The whole process revolves around the agent, and the construction of the training environment and data preprocessing all provide support and conditions for the learning and decision-making of the agent.

[0043] The method may specifically include the following steps: Step 101, initialize the original neural network model to obtain an initial neural network model, and construct a training environment for training the initial neural network model.

[0044] The intelligent agent can initialize the constructed neural network model (referred to as the "original neural network model" for easy distinction) to obtain an initialized neural network model (referred to as the "initial neural network model"), so as to train the initial neural network model. At the same time, a training environment for training the initial neural network model can also be constructed.

[0045] In the embodiment of the present invention, the initializing the original neural network model to obtain the initial neural network model includes: Construct a neural network structure based on the deep Q network to obtain the original neural network model; The weight coefficients and model parameters of the original neural network model are initialized to obtain an initial neural network model.

[0046] Specifically, the original neural network model can be constructed based on the DQN network. The original neural network model can include an input layer, several hidden layers, and an output layer. The input layer is used to receive the encoding representation of the triple, the hidden layer is used to extract features, and the output layer is used to output the Q value of the action. The Q value can be calculated using the following formula:

[0047] in, The triple state at time t is , take action , the model parameters are The Q value at the time is the value evaluation of the triple state-action pair, which is used to measure the Take action The expected cumulative reward that can be obtained after the summation. n is the upper limit of the summation, indicating that there are n items involved in the accumulation operation. is the weight coefficient used to measure different functions The contribution to the Q value calculation is different The value can highlight or weaken the effect of the corresponding function. is the characteristic function, based on the triple state and actions Extracted relevant features.

[0048] After obtaining the original neural network model, the weight coefficients and model parameters can be initialized to obtain the initial neural network model.

[0049] Among them, when initializing the weight coefficients and model parameters, a random initialization method can be used, such as a normal distribution method or a uniform distribution method, so as to ensure that the parameters are within a reasonable range and have a certain degree of randomness, and avoid falling into a local optimal solution.

[0050] Of course, in addition to the random initialization method, other initialization methods may also be used. In practical applications, the initialization method of the weight coefficient and the model parameter may be set according to actual needs, and the embodiment of the present invention does not limit this.

[0051] In an embodiment of the present invention, the step of constructing a training environment for training the initial neural network model includes: The subject, relationship and subject in each triple in the sample knowledge graph form a state, and the triple states corresponding to each triple are obtained; Set the action to be performed based on the confidence of the triple; Set the reward for the accuracy of the triplet confidence estimate for the action performed; Set the termination conditions for training; An experience replay buffer is set; the experience replay buffer is used to store the experience generated during the training process.

[0052] Specifically, the triple (S, P, O) in the knowledge graph, where S is the subject, P is the relationship, and O is the subject, directly constitutes the basic unit of the interactive environment of the reinforcement learning agent, and each triple becomes a state in the state space. When the agent makes a decision, it is based on specific triples. For example, for the triple "generator (S)-operating state (P)-normal (O)", it represents a specific state, and the agent needs to judge how to adjust its confidence based on this state, that is, to take corresponding actions.

[0053] The state space is not limited to the triple itself, but also extends to include the contextual information of the triple, such as adjacent nodes, node attributes, etc. Taking the power station knowledge graph as an example, the triple of an equipment entity may have "transformer (S) - connection (P) - bus (O)", and its adjacent nodes may have other equipment connected to the transformer. The node attributes include the rated capacity of the transformer, the voltage level of the bus, etc. These contextual information combined with the triple can enable the intelligent agent to have a more comprehensive understanding of the current state, so as to more accurately evaluate the confidence of the triple and make more reasonable decisions.

[0054] The triplet state may include the entity's own state, relationship state, and context state. The entity's own state includes attribute state and operating state, such as the rated parameters and operating conditions of the equipment; the relationship state includes connection and causal relationship state, reflecting the entity association and causal relationship; the context state involves adjacent nodes and historical states. The adjacent node state is affected by the connected equipment, and the historical state is based on past records. These states together constitute the complete state information of the triplet. Of course, in addition to the above-mentioned states, the triplet state may also include other states. In actual applications, it can be set according to actual needs, and the embodiments of the present invention are not limited to this.

[0055] In the confidence measurement scenario, the action can be an adjustment operation on the triple confidence, such as increasing the confidence, decreasing the confidence, or keeping it unchanged. The number of nodes in the output layer corresponds to the number of actions.

[0056] Furthermore, the actions performed on the confidence of the triples may include: confidence adjustment class, such as increasing confidence, decreasing confidence or keeping it unchanged; related information query class, such as querying entity attributes and related triples; verification class, that is, verifying single or batch similar triples with the help of external sources or rules to assist in judging and adjusting confidence. Of course, in addition to the above actions, the actions performed may also include other actions. In actual applications, they can be set according to actual needs, and the embodiments of the present invention do not limit this.

[0057] In the embodiment of the present invention, the reward can be the feedback of the training environment to the action of the agent, and rewards or penalties are given according to the accuracy of the triple confidence assessment of the action. The function can be customized according to the scenario. For example, if the agent assigns high confidence to the correct triple (based on some external verification or prior knowledge), a positive reward is given; if it is assigned to the wrong triple, a negative reward is given. The reward function can be customized according to the specific application scenario, and the factors considered include but are not limited to the importance of the triple, the consistency with other triples, etc.

[0058] The termination flag is used to determine whether the training is finished. When the maximum number of training steps is reached or the performance index of the validation set no longer improves or reaches the preset threshold, it is set to True, otherwise it is False. When the termination state is reached, the agent stops the current training and then uses the trained neural network model to measure the triple confidence.

[0059] The experience replay buffer is a queue of limited size that is used to store the historical experience generated by the agent during the training process. Experience can be a series of data accumulated by the agent during its interaction with the training environment, which records the behavior of the agent and the feedback of the training environment.

[0060] Step 102, preprocessing the sample knowledge graph used for training to obtain data to be processed.

[0061] After obtaining the initial neural network model, since the initial neural network model cannot process the knowledge graph, the knowledge graph used for training (referred to as the "sample knowledge graph") is preprocessed to obtain data that can be processed by the initial neural network model (referred to as the "data to be processed").

[0062] In the embodiment of the present invention, the data preprocessing of the sample knowledge graph used for training to obtain the data to be processed includes: Acquire at least one triple data from the sample knowledge graph; Divide each triple data into training set, validation set and test set; Each triplet data in the training set, the validation set and the test set is encoded to obtain an encoded sample training set, a sample validation set and a sample test set.

[0063] Specifically, for a sample knowledge graph, at least one triple data can be obtained from it, and then all the triple data can be divided into a training set, a validation set, and a test set.

[0064] The training set is used to train the model, that is, to determine the model's parameters such as weights and biases. The data in the training set directly participates in the model's learning process. Through repeated iterations and optimization, the model can fit the training data as accurately as possible.

[0065] The validation set is used for model selection and hyperparameter adjustment. The validation set does not participate in the learning process of model parameters, but is used to evaluate the performance of the model under different hyperparameter settings, so as to select the optimal hyperparameter combination.

[0066] The test set is used to evaluate the generalization ability of the final model. The test set is used after the model training is completed to measure the performance of the model on unseen data. The test set does not participate in the selection process of model parameters and hyperparameters and is used for the final evaluation of the model.

[0067] By reasonably dividing the training set, validation set, and test set, the generalization ability of the model can be effectively improved, overfitting can be avoided, and better results can be achieved in practical applications.

[0068] After obtaining the training set, validation set, and test set, each triplet of data in the training set, validation set, and test set can be encoded, and the entities and relationships can be converted into vector form using the word vector method so that the intelligent agent can process them, thereby obtaining the encoded training set (referred to as the "sample training set"), validation set (referred to as the "sample validation set"), and test set (referred to as the "sample test set").

[0069] In the embodiment of the present invention, the division of each triple data to obtain a training set, a validation set and a test set includes: The random stratified sampling method is used to divide each triple data into training set, validation set and test set.

[0070] Specifically, when dividing the training set, validation set, and test set, a random stratified sampling method can be used. First, the triples in the sample knowledge graph are classified by relationship type, such as device attribute relationship, fault association relationship, etc. Then, for the triples of each type of relationship, they are randomly selected according to a certain ratio (for example, 70% as training set, 15% as validation set, and 15% as test set). This ensures that each set contains triples of various types of relationships, and the ratio is similar to the original knowledge graph, so that the model can fully learn the characteristics of different relationships during the training process.

[0071] Of course, in addition to the above-mentioned ratios, the random stratified sampling method may also adopt other ratios. In practical applications, the specific values ​​of the ratios may be set according to actual needs, and the embodiments of the present invention do not limit this.

[0072] Moreover, in addition to the random stratified sampling method, other methods can also be used to divide the training set, validation set and test set. In practical applications, the specific division method can also be set according to actual needs, and the embodiment of the present invention does not limit this.

[0073] Step 103, training the initial neural network model using the data to be processed in the training environment to obtain a trained target neural network model.

[0074] After obtaining the data to be processed, the initial neural network model can be trained using the data to be processed in the training environment, so that the intelligent agent can select actions according to the triplet states in the training environment, perform operations to adjust the confidence, receive environmental feedback, and then continuously optimize its decision-making ability through experience playback and parameter updates, and finally achieve the measurement of the confidence of the knowledge graph triples, thereby obtaining the trained neural network model (referred to as the "target neural network model").

[0075] In an embodiment of the present invention, the step of training the initial neural network model using the data to be processed in the training environment to obtain a trained target neural network model includes: S1, obtaining a current triple to be processed from a sample training set in the data to be processed; S2. Obtain the current triplet state corresponding to the current triplet; S3, determining a target action corresponding to the current triple state from the training environment according to the current triple state; S4, executing the target action; S5, obtaining corresponding experience feedback when the target action is completed; S6, storing the experience feedback in an experience playback buffer; S7, obtaining target experience feedback from the experience playback buffer; S8, using the target experience feedback to update the initial neural network model to obtain an updated candidate neural network model; S9, determining whether the termination condition in the training environment is met, if so, obtaining the trained target neural network model; if not, executing S10; S10, obtaining the next triplet from the sample training set in the data to be processed, taking the next triplet as the current triplet, and executing S2.

[0076] Specifically, the triple to be processed (referred to as "current triple") can be obtained from the sample training set in the data to be processed. The current triple can be one or more. The embodiment of the present invention takes one current triple as an example for explanation. In practical applications, the number of current triples can be set according to actual needs, and the embodiment of the present invention does not limit this.

[0077] After obtaining the current triplet, the triplet state corresponding to the current triplet can be further obtained (referred to as "current triplet state"). Since multiple actions have been defined in the training environment, the action corresponding to the current triplet state can be determined from the training environment according to the current triplet state (referred to as "target action") and the target action can be executed.

[0078] The training environment can return the corresponding experience (referred to as "experience feedback") when the target action is completed. After obtaining the experience feedback, the agent can store the experience feedback in the experience playback buffer, and then obtain a certain amount of experience feedback (referred to as "target experience feedback") from the experience playback buffer. In other words, the target experience feedback may include. Among them, the target experience feedback can be one or more. In actual applications, the number of target experience feedback can be set according to actual needs, and the embodiment of the present invention does not limit this.

[0079] After obtaining the target experience feedback, the target experience feedback can be used to update the initial neural network model to obtain an updated neural network model (referred to as a "candidate neural network model").

[0080] After the update is completed, it can be detected whether the termination condition defined in the training environment is met. If it is met, it means that the training is completed and the training can be terminated to obtain the trained neural network model (recorded as the "target neural network model"); if it is not met, the next triplet to be processed can be obtained from the sample training set, and for the next triplet, jump to the step of "obtaining the current triplet state corresponding to the current triplet" and continue training until the termination condition is met after the training of a certain triplet is completed.

[0081] In an embodiment of the present invention, determining a target action corresponding to the current triple state from the training environment according to the current triple state includes: According to the current triplet state, a target action corresponding to the current triplet state is selected from the training environment using a greedy strategy.

[0082] Specifically, after obtaining the current triplet state, the agent can use a greedy strategy to select the target action corresponding to the current triplet state from the training environment.

[0083] The core idea of ​​the greedy strategy is to take the best option at each step. The advantages of the greedy strategy are that it is simple, intuitive, easy to implement, and can get results that are very close to the optimal solution or the exact optimal solution for some problems.

[0084] Of course, in addition to the greedy strategy, other methods can also be used to select the target action from the training environment. In practical applications, the specific selection method can be set according to actual needs, and the embodiment of the present invention does not limit this.

[0085] Wherein, the step of selecting a target action corresponding to the current triplet state from the training environment using a greedy strategy according to the current triplet state includes: According to the current triplet state, randomly select a first action from the training environment with probability e; Using the initial neural network model to calculate the Q value corresponding to each action in the training environment, and selecting the second action with the largest Q value with probability 1-e; The action with a higher probability between the first action and the second action is determined as the target action.

[0086] Specifically, the greedy strategy can randomly select an action (referred to as the "first action") from the training environment with probability e according to the current triple state, and use the initial neural network model to calculate the Q value corresponding to each action in the training environment (see the calculation formula in step 101 for details), and select the action with the largest Q value (referred to as the "second action") with probability 1-e.

[0087] Then compare the probability of the first action with the probability of the second action, and take the action with a higher probability as the target action.

[0088] It should be noted that e is a parameter that gradually decays during the training process. It is mainly used to control the balance between exploration and utilization of action selection. The initial value is high to encourage exploration, and it gradually decreases as the training progresses to increase utilization. In other words, in the action selection strategy of reinforcement learning, although only one action is executed in the end, the selection method is based on probability. The purpose of randomly selecting an action with probability e is to explore different action effects in the environment and prevent the agent from falling into a local optimal solution. This is because in the early stages of training, the agent has limited knowledge of the environment, and completely relying on the current estimated Q value to select actions (that is, selecting the action with the largest current Q value with a probability of 1-e) may miss some potential better strategies. As training progresses, e gradually decreases, and the agent makes more use of the knowledge it has learned and selects the action with the largest current Q value, that is, using existing experience to make decisions. Therefore, it is actually selecting from two selection methods with different probabilities to ultimately determine the only action to be executed.

[0089] Furthermore, the action performed by the agent is not entirely the action with the largest calculated Q value. During the reinforcement learning training process, the agent adopts a greedy strategy to select actions. An action is randomly selected with probability e to explore the environment and discover new possible strategies; the action with the largest current Q value is selected with probability 1-e, and the action currently considered to be the best is selected using the learned knowledge. As mentioned earlier, e is a parameter that gradually decays during the training process. The initial value is high to encourage exploration, and it gradually decreases as the training progresses to increase utilization.

[0090] After the training is completed, when the trained reinforcement learning model is used to measure the confidence of the knowledge graph triple, the confidence corresponding to the action with the largest calculated Q value will be selected as the confidence estimate of the triple. At this time, the probability e is no longer used for random exploration, and the optimal strategy learned by the model can be used.

[0091] In the embodiment of the present invention, the obtaining of the experience feedback corresponding to the completion of the target action includes: Obtaining the next triplet state and reward of the current triplet from the training environment, and obtaining the termination flag at the current moment; the reward is calculated by the training environment using a preset reward function; The current triplet state, the target action, the reward, the next triplet state and the termination flag are combined to obtain experience feedback corresponding to the completion of the target action.

[0092] Specifically, after the agent completes the target action, the training environment can return the next triplet state of the current triplet, the reward, and the termination flag of the current moment.

[0093] The reward is calculated by the training environment using a preset reward function, and is used as feedback for the agent to perform the target action. The termination flag is used to indicate whether the training is finished, that is, when the termination flag is True, it means that the training is finished, and the training is terminated. When the termination flag is False, it means that the training needs to continue.

[0094] Then, the current triplet state, target action, reward, next triplet state, and termination flag are combined to obtain the experience feedback corresponding to the completion of the target action. For example, the experience feedback can be expressed as . is the current triplet state, For the target action, For reward, is the next triple state, It is the termination sign.

[0095] It should be noted that in reinforcement learning, although the current triplet state can be obtained, the agent can select the target action based on the current triplet state to act on the training environment. The training environment feeds back the new triplet state according to its own rules and target action. At the same time, it gives rewards and a sign of whether the terminal state has been reached. This series of information constitutes experience feedback and is recorded. The agent uses this to review learning and optimize decision-making strategies to complete the triplet confidence measurement task.

[0096] Furthermore, when determining rewards, the Q value is used to guide the agent to make more accurate triplet confidence assessment decisions. At the same time, the accuracy of triplet confidence assessment can be judged by external verification or prior knowledge, which in turn affects the update of the Q value. The Q value represents the cumulative reward that the agent expects to obtain after taking a certain action in a specific state. In the triplet confidence assessment scenario, the agent's action is to adjust the triplet confidence (such as increasing confidence, decreasing confidence, or keeping it unchanged). If an action can make the triplet confidence assessment more accurate, then the Q value based on this action will be higher. For example, if correctly assigning high confidence to a real triplet can obtain a positive reward, the agent will find that the Q value corresponding to this action is higher in subsequent learning; conversely, the action that incorrectly adjusts the confidence and leads to a negative reward will have a lower Q value. By constantly trying different actions, the agent gradually learns to choose actions that can improve the accuracy of triplet confidence assessment based on the feedback of the Q value.

[0097] In the embodiment of the present invention, the step of obtaining target experience feedback from the experience playback buffer includes: Determining the number of samples used to update the initial neural network model; A target number of experience feedback samples is obtained from the experience playback buffer.

[0098] Specifically, when obtaining the target number of experience feedback, the number of samples used to update the model parameters in the initial neural network model can be determined first, and then the target number of experience feedback samples is obtained from the experience playback buffer. , is the number of samples used to update the model parameters in the initial neural network model.

[0099] Among them, when obtaining the target experience feedback of the sample quantity, a random sampling method can be used, which can break the correlation between the target experience feedbacks, thereby improving the stability and efficiency of the training. Of course, in addition to the random sampling method, other methods can also be used. In practical applications, the specific method of obtaining can be set according to actual needs, and the embodiment of the present invention does not limit this.

[0100] It should be noted that a larger number of samples can improve training efficiency, but may require more computing resources; a smaller number of samples may lead to unstable updates. Therefore, a suitable sample size can be selected according to the size of the data set and computing resources, such as 32, 64, 128, etc. Of course, in practical applications, the number of target experience feedbacks can be the same as the number of samples used to update the model parameters, or other numbers can be used. The specific number of target experience feedbacks can be set according to actual needs, and the embodiment of the present invention does not limit this.

[0101] In an embodiment of the present invention, the step of updating the initial neural network model using the target experience feedback to obtain an updated candidate neural network model includes: The initial neural network model is used to calculate the Q value estimation of executing the target action in the current triple state in each target experience feedback; Calculate the target Q value corresponding to each target experience feedback; The loss function corresponding to each target experience feedback is calculated using the Q value estimate and the target Q value corresponding to each target experience feedback; The initial neural network model is updated using the loss function corresponding to each target experience feedback to obtain an updated candidate neural network model.

[0102] Specifically, when updating the initial neural network model, for any target experience feedback among all target experience feedbacks, the initial neural network model can be used to calculate the Q value of executing the target action under the current triplet state (referred to as "Q value estimate"), and calculate the actual Q value corresponding to the target experience feedback (referred to as "target Q value"), and then calculate the loss function of the Q value estimate of the target experience feedback and the target Q value, and use the loss function to update the initial neural network model.

[0103] The Q value is estimated as ,in, is the current triplet state, is the target action to be performed in the current triple state.

[0104] When each target experience feedback completes the above process, an updated candidate neural network model can be obtained.

[0105] The step of calculating the target Q value corresponding to each target experience feedback includes: For any target experience feedback, when the termination flag is reaching the termination state, the corresponding target Q value is the reward; When the termination flag is not the reaching termination state, the target Q value is calculated using the reward, the next triplet state and a preset discount factor.

[0106] Specifically, when calculating the target Q value corresponding to any target experience feedback, it is necessary to refer to the termination flag of the target experience feedback. If the termination flag is True, that is, the termination state is reached, then the target Q value ,in, The reward in the experience feedback for any target.

[0107] If the termination flag is False, that is, the termination state has not been reached, then the target Q value ,in, The reward in the experience feedback of any target, is a discount factor used to weigh the importance of the reward, is the next triple state, a' For the next triple state An action from the set of possible actions that the agent can take.

[0108] It should be noted that The value range of can be set according to actual needs, and the embodiment of the present invention does not limit this.

[0109] The method of using the Q value estimation and the target Q value corresponding to each target experience feedback to calculate the loss function corresponding to each target experience feedback includes: The mean square error is used as the loss function to calculate the error between the Q value estimate corresponding to each target experience feedback and the target Q value.

[0110] Specifically, when calculating the loss function, the mean square error (MSE) function can be used as the loss function, and the error between the Q value estimate and the target Q value can be calculated using the following formula:

[0111] in, are the model parameters of the neural network. The gradient of the loss function with respect to the model parameters is calculated through the back-propagation algorithm, and the model parameters are updated using an optimizer (such as SGD (stochastic gradient descent), Adam (Adaptive Moment Estimation), etc.) to minimize the loss function. The update formula of the optimizer can be ,in is the learning rate, which can range from 0.0001 to 0.1. It can be dynamically adjusted according to the performance changes during training using a learning rate adjustment strategy (such as learning rate decay) to control the step size of the model parameter update.

[0112] Of course, in addition to using the mean square error function as the loss, other functions can also be used as the loss function. In practical applications, the specific loss function can be set according to actual needs, and the embodiment of the present invention does not limit this.

[0113] In an embodiment of the present invention, the determining whether a termination condition in the training environment is satisfied includes: Determine whether at least one of the following conditions is met: a maximum number of training steps is reached and a performance indicator of the validation set reaches a threshold; If so, the termination condition in the training environment is met.

[0114] Specifically, when determining whether the termination condition is met, it can be determined whether the training has reached the maximum number of training steps, and / or whether the performance indicator of the validation set is no longer improved, and / or whether the performance indicator of the validation set reaches a threshold. If at least one of these is met, then it can be determined that the termination condition in the training environment is met.

[0115] Of course, in addition to determining whether the termination condition is met in the above manner, other methods may be used to determine whether the termination condition is met. In practical applications, the specific determination method may be set according to actual needs, and the embodiment of the present invention does not limit this.

[0116] Step 104: Use the target neural network model to perform confidence measurement on the knowledge graph to be measured to obtain a measurement result.

[0117] After the training is completed and the target neural network model is obtained, the triples in the knowledge graph to be measured can be encoded, and then the encoded triples (and context information, if any) are input into the target neural network model as triple states.

[0118] The target neural network model calculates the Q value of each possible action (in the confidence measurement scenario, actions correspond to different confidence evaluation results, such as high confidence, medium confidence, low confidence, etc., and the action space can be defined according to actual needs) based on its learned strategy and network structure. The confidence of the triplet is determined based on the calculated Q value, and the confidence corresponding to the action with the largest Q value is selected as the confidence estimate of the triplet, and the final triplet confidence is output as the measurement result (probability between 0 and 1).

[0119] In an embodiment of the present invention, the agent initializes the original neural network model to obtain the initial neural network model, and constructs a training environment for training the initial neural network model; performs data preprocessing on the sample knowledge graph used for training to obtain the data to be processed; uses the data to be processed to train the initial neural network model in the training environment to obtain the trained target neural network model; uses the target neural network model to measure the confidence of the knowledge graph to be measured to obtain the measurement result. In this way, the agent uses the sample knowledge graph in the constructed training environment to train the neural network model so that the neural network model can quantify its semantic correctness and the truth of the facts expressed. During the training process, the confidence, as the weight or credibility indicator of the reasoning path, can guide the reasoning in a more reliable direction. Reasonable use of confidence can avoid reasoning errors caused by erroneous or low-quality knowledge, and improve the accuracy and interpretability of the reasoning results. Thus, the agent can automatically evaluate the confidence of the triples through reinforcement learning, and learn the optimal strategy through the rewards fed back by the training environment, without manual judgment one by one. Moreover, the experience replay mechanism enables the agent to learn from historical experience, thereby further optimizing the confidence assessment and improving the efficiency of the agent in confidence assessment using the trained neural network model, solving the problem of requiring a large amount of manual supervision and manual annotation.

[0120] Reference Figure 2 , shows a flow chart of the steps of Embodiment 2 of a knowledge graph evaluation method of the present invention, which may specifically include the following steps: Step 201, initialize the original neural network model to obtain an initial neural network model, and construct a training environment for training the initial neural network model.

[0121] Step 202: preprocess the sample knowledge graph used for training to obtain data to be processed.

[0122] Step 203 uses the data to be processed to train the initial neural network model in the training environment to obtain a trained target neural network model.

[0123] Step 204: Use the target neural network model to perform confidence measurement on the knowledge graph to be measured to obtain a measurement result.

[0124] Among them, step 201 to step 204 are substantially the same as step 101 to step 104, and will not be described here to avoid repetition.

[0125] Step 205: If the confidence in the measurement result is different from the confidence in the knowledge graph to be measured, the confidence in the measurement result is used to replace the confidence in the knowledge graph to be measured.

[0126] Specifically, after obtaining the measurement results, if the confidence of a triple in the measurement results is different from the confidence of the triple in the knowledge graph to be measured, then the confidence in the measurement results can be used to replace the confidence in the knowledge graph to be measured, thereby realizing the update and correction of the confidence of the triple in the knowledge graph to be measured.

[0127] In an embodiment of the present invention, the intelligent agent uses a sample knowledge graph in a constructed training environment to train the neural network model so that the neural network model can quantify its semantic correctness and the truth of the facts expressed. During the training process, confidence, as the weight or credibility indicator of the reasoning path, can guide the reasoning in a more reliable direction. Reasonable use of confidence can avoid reasoning errors caused by erroneous or low-quality knowledge, and improve the accuracy and interpretability of the reasoning results. As a result, the intelligent agent can automatically evaluate the confidence of the triples through reinforcement learning, and learn the optimal strategy through the rewards fed back by the training environment, without the need for manual judgment one by one. Moreover, the experience replay mechanism allows the intelligent agent to learn from historical experience, thereby further optimizing the confidence assessment and improving the efficiency of the intelligent agent using the trained neural network model for confidence assessment, thereby improving the updating and correction efficiency of the knowledge graph, reducing labor costs, and solving the problem of requiring a large amount of manual supervision and manual annotation.

[0128] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.

[0129] Reference Figure 3 , shows a structural block diagram of an embodiment of a knowledge graph evaluation device of the present invention, which may specifically include the following modules: Initialization module 301, used to initialize the original neural network model to obtain an initial neural network model; A construction module 302 is used to construct a training environment for training the initial neural network model; A preprocessing module 303 is used to perform data preprocessing on the sample knowledge graph used for training to obtain data to be processed; A training module 304 is used to train the initial neural network model using the data to be processed in the training environment to obtain a trained target neural network model; The measurement module 305 is used to use the target neural network model to perform confidence measurement on the knowledge graph to be measured to obtain a measurement result.

[0130] In the embodiment of the present invention, the initialization module is specifically used to: Construct a neural network structure based on the deep Q network to obtain the original neural network model; The weight coefficients and model parameters of the original neural network model are initialized to obtain an initial neural network model.

[0131] In the embodiment of the present invention, the building module is specifically used for: The subject, relationship and subject in each triple in the sample knowledge graph form a state, and the triple states corresponding to each triple are obtained; Set the action to be performed based on the confidence of the triple; Set the reward for the accuracy of the triplet confidence estimate for the action performed; Set the termination conditions for training; An experience replay buffer is set; the experience replay buffer is used to store the experience generated during the training process.

[0132] In an embodiment of the present invention, the preprocessing module includes: A first acquisition submodule, used to acquire at least one triple data from the sample knowledge graph; The partitioning submodule is used to partition each triplet data into a training set, a validation set, and a test set; The encoding submodule is used to encode each triplet data in the training set, the verification set and the test set to obtain the encoded sample training set, sample verification set and sample test set.

[0133] In the embodiment of the present invention, the division submodule is specifically used for: The random stratified sampling method is used to divide each triple data into training set, validation set and test set.

[0134] In an embodiment of the present invention, the training module includes: A second acquisition submodule is used to acquire the current triple to be processed from the sample training set in the data to be processed; A third acquisition submodule is used to acquire a current triplet state corresponding to the current triplet; A determination submodule, configured to determine a target action corresponding to the current triple state from the training environment according to the current triple state; An execution submodule, used for executing the target action; A fourth acquisition submodule is used to obtain experience feedback corresponding to the completion of the target action; A storage submodule, used for storing the experience feedback in an experience playback buffer; A fifth acquisition submodule, used to acquire target experience feedback from the experience playback buffer; An updating submodule, used to update the initial neural network model using the target experience feedback to obtain an updated candidate neural network model; A judgment submodule is used to determine whether the termination condition in the training environment is met, and if so, obtain the trained target neural network model; if not, call the sixth acquisition submodule; The sixth acquisition submodule is used to acquire the next triplet from the sample training set in the data to be processed, take the next triplet as the current triplet, and call the second acquisition submodule.

[0135] In the embodiment of the present invention, the determining submodule is specifically used for: According to the current triplet state, a target action corresponding to the current triplet state is selected from the training environment using a greedy strategy.

[0136] In the embodiment of the present invention, the determining submodule is further used to: According to the current triplet state, randomly select a first action from the training environment with probability e; Using the initial neural network model to calculate the Q value corresponding to each action in the training environment, and selecting the second action with the largest Q value with probability 1-e; The action with a higher probability between the first action and the second action is determined as the target action.

[0137] In the embodiment of the present invention, the fourth acquisition submodule is specifically used for: Obtaining the next triplet state and reward of the current triplet from the training environment, and obtaining a termination flag at the current moment; the reward is calculated by the training environment using a preset reward function, and the termination flag is used to indicate whether the training is terminated; The current triplet state, the target action, the reward, the next triplet state and the termination flag are combined to obtain experience feedback corresponding to the completion of the target action.

[0138] In the embodiment of the present invention, the fifth acquisition submodule is specifically used for: Determining the number of samples used to update the initial neural network model; A target number of experience feedback samples is obtained from the experience playback buffer.

[0139] In the embodiment of the present invention, the updating submodule includes: A first calculation unit, configured to calculate, using the initial neural network model, a Q value estimate of executing a target action in a current triplet state in each target experience feedback; The second calculation unit is used to calculate the target Q value corresponding to each target experience feedback; A third calculation unit is used to calculate the loss function corresponding to each target experience feedback by using the Q value estimate and the target Q value corresponding to each target experience feedback; An updating unit is used to update the initial neural network model using a loss function corresponding to each target experience feedback to obtain an updated candidate neural network model.

[0140] In the embodiment of the present invention, the second computing unit is specifically configured to: For any target experience feedback, when the termination flag is reaching the termination state, the corresponding target Q value is the reward; When the termination flag is not the reaching termination state, the target Q value is calculated using the reward, the next triplet state and a preset discount factor.

[0141] In this embodiment of the present invention, the third computing unit is specifically configured to: The mean square error is used as the loss function to calculate the error between the Q value estimate corresponding to each target experience feedback and the target Q value.

[0142] In the embodiment of the present invention, the judgment submodule is specifically used to: Determine whether at least one of the following conditions is met: a maximum number of training steps is reached and a performance indicator of the validation set reaches a threshold; If so, the termination condition in the training environment is met.

[0143] In an embodiment of the present invention, it also includes: An updating module is used to replace the confidence in the knowledge graph to be measured with the confidence in the measurement result if the confidence in the measurement result is different from the confidence in the knowledge graph to be measured.

[0144] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0145] An embodiment of the present invention further provides an electronic device, including: It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, the various processes of the above-mentioned knowledge graph evaluation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0146] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned knowledge graph evaluation method embodiment are implemented, and the same technical effect can be achieved. To avoid repetition, it will not be repeated here.

[0147] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0148] It should be understood by those skilled in the art that the embodiments of the present invention can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0149] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0150] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0151] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0152] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.

[0153] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0154] The above is a detailed introduction to a knowledge graph evaluation method and a knowledge graph evaluation device provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A knowledge graph evaluation method, characterized in that: The method comprises: Initializing the original neural network model to obtain an initial neural network model, and constructing a training environment for training the initial neural network model; Perform data preprocessing on the sample knowledge graph used for training to obtain data to be processed; Using the data to be processed to train the initial neural network model in the training environment to obtain a trained target neural network model; The target neural network model is used to perform confidence measurement on the knowledge graph to be measured to obtain a measurement result.

2. The knowledge graph evaluation method according to claim 1, characterized in that: The initializing of the original neural network model to obtain the initial neural network model includes: Construct a neural network structure based on the deep Q network to obtain the original neural network model; The weight coefficients and model parameters of the original neural network model are initialized to obtain an initial neural network model.

3. The knowledge graph evaluation method according to claim 1, characterized in that: The constructing of a training environment for training the initial neural network model includes: The subject, relationship and subject in each triple in the sample knowledge graph form a state, and the triple states corresponding to each triple are obtained; Set the action to be performed based on the confidence of the triple; Set the reward for the accuracy of the triplet confidence estimate for the action performed; Set the termination conditions for training; An experience replay buffer is set; the experience replay buffer is used to store the experience generated during the training process.

4. The knowledge graph evaluation method according to claim 1, characterized in that: The data preprocessing of the sample knowledge graph for training to obtain the data to be processed includes: Acquire at least one triple data from the sample knowledge graph; Divide each triple data into training set, validation set and test set; Each triplet data in the training set, the validation set and the test set is encoded to obtain an encoded sample training set, a sample validation set and a sample test set.

5. The knowledge graph evaluation method according to claim 4, characterized in that: The triple data are divided into a training set, a validation set and a test set, including: The random stratified sampling method is used to divide each triple data into training set, validation set and test set.

6. The knowledge graph evaluation method according to claim 1, characterized in that: The step of training the initial neural network model using the data to be processed in the training environment to obtain a trained target neural network model includes: S1, obtaining a current triple to be processed from a sample training set in the data to be processed; S2. Obtain the current triplet state corresponding to the current triplet; S3, determining a target action corresponding to the current triple state from the training environment according to the current triple state; S4, executing the target action; S5, obtaining corresponding experience feedback when the target action is completed; S6, storing the experience feedback in an experience playback buffer; S7, obtaining target experience feedback from the experience playback buffer; S8, using the target experience feedback to update the initial neural network model to obtain an updated candidate neural network model; S9, determining whether the termination condition in the training environment is met, if so, obtaining the trained target neural network model; if not, executing S10; S10, obtaining the next triplet from the sample training set in the data to be processed, taking the next triplet as the current triplet, and executing S2.

7. The knowledge graph evaluation method according to claim 6, characterized in that: The determining, from the training environment according to the current triplet state, a target action corresponding to the current triplet state comprises: According to the current triplet state, a target action corresponding to the current triplet state is selected from the training environment using a greedy strategy.

8. The knowledge graph evaluation method according to claim 7, characterized in that: The step of selecting a target action corresponding to the current triplet state from the training environment using a greedy strategy according to the current triplet state includes: According to the current triplet state, randomly select a first action from the training environment with probability e; Using the initial neural network model to calculate the Q value corresponding to each action in the training environment, and selecting the second action with the largest Q value with probability 1-e; The action with a higher probability between the first action and the second action is determined as the target action.

9. The knowledge graph evaluation method according to claim 6, characterized in that: The obtaining of experience feedback corresponding to the completion of the target action includes: Obtaining the next triplet state and reward of the current triplet from the training environment, and obtaining a termination flag at the current moment; the reward is calculated by the training environment using a preset reward function, and the termination flag is used to indicate whether the training is terminated; The current triplet state, the target action, the reward, the next triplet state and the termination flag are combined to obtain experience feedback corresponding to the completion of the target action.

10. The knowledge graph evaluation method according to claim 6, characterized in that: The obtaining target experience feedback from the experience playback buffer includes: Determining the number of samples used to update the initial neural network model; A target number of experience feedback samples is obtained from the experience playback buffer.

11. The knowledge graph evaluation method according to claim 6, characterized in that: The step of updating the initial neural network model by using the target experience feedback to obtain an updated candidate neural network model includes: The initial neural network model is used to calculate the Q value estimation of executing the target action in the current triple state in each target experience feedback; Calculate the target Q value corresponding to each target experience feedback; The loss function corresponding to each target experience feedback is calculated using the Q value estimate and the target Q value corresponding to each target experience feedback; The initial neural network model is updated using the loss function corresponding to each target experience feedback to obtain an updated candidate neural network model.

12. The knowledge graph evaluation method according to claim 11, characterized in that: The calculating of the target Q value corresponding to each target experience feedback includes: For any target experience feedback, when the termination flag is reaching the termination state, the corresponding target Q value is the reward; When the termination flag is not the reaching termination state, the target Q value is calculated using the reward, the next triplet state and a preset discount factor.

13. The knowledge graph evaluation method according to claim 11, characterized in that: The Q value estimation and the target Q value corresponding to each target experience feedback are used to calculate the loss function corresponding to each target experience feedback, including: The mean square error is used as the loss function to calculate the error between the Q value estimate corresponding to each target experience feedback and the target Q value.

14. The knowledge graph evaluation method according to claim 6, characterized in that: The determining whether a termination condition in the training environment is satisfied includes: Determine whether at least one of the following conditions is met: a maximum number of training steps is reached and a performance indicator of the validation set reaches a threshold; If so, the termination condition in the training environment is met.

15. The knowledge graph evaluation method according to claim 1, characterized in that: Also includes: If the confidence in the measurement result is different from the confidence in the knowledge graph to be measured, the confidence in the measurement result is used to replace the confidence in the knowledge graph to be measured.

16. A knowledge graph evaluation device, characterized in that: The device comprises: An initialization module is used to initialize the original neural network model to obtain an initial neural network model; A construction module, used to construct a training environment for training the initial neural network model; A preprocessing module is used to preprocess the sample knowledge graph used for training to obtain data to be processed; A training module, used to train the initial neural network model using the data to be processed in the training environment to obtain a trained target neural network model; The measurement module is used to use the target neural network model to perform confidence measurement on the knowledge graph to be measured to obtain a measurement result.

17. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of the knowledge graph evaluation method as described in any one of claims 1 to 15 are implemented.

18. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the knowledge graph evaluation method as described in any one of claims 1 to 15 are implemented.

Citation Information

Patent Citations

  • Entity relation path reasoning method and system for regulation and control knowledge graph

    CN113128689A

  • Knowledge graph triple reliability evaluation method, system and device and medium

    CN115238582A

  • Knowledge reasoning method and system based on agent dynamic path completion strategy

    CN115526321A

  • Knowledge graph reasoning method based on logic rules and reinforcement learning

    CN115660086A

  • Knowledge graph multi-hop reasoning method based on reinforcement learning

    CN117217305A