Decision-making method, device, equipment and medium based on knowledge embedded reinforcement learning
By introducing knowledge embedding technology into the reinforcement learning model, the original image is integrated with prior knowledge and the graph vector is generated for decision-making, the problem of low learning efficiency and low effect of existing reinforcement learning models is solved, and more accurate decision-making is achieved.
Patent Information
- Application Number
- CN202311086572.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-28
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-08-28
AI Technical Summary
The existing Actor-Critic reinforcement learning model learns based on the original information of the environment, resulting in low learning efficiency and effectiveness, and a large difference between decisions and expected decisions.
Using a reinforcement learning method based on knowledge embedding, the original image is fused with prior knowledge through the knowledge fusion module, a graph vector containing prior knowledge is generated, and a strategy network is used to make decisions based on this graph vector.
The learning efficiency and effectiveness of the reinforcement learning model are improved and more accurate decisions can be made.
Smart Images

Figure CN117115608B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a decision-making method, device, equipment and medium based on knowledge embedding reinforcement learning. Background Art
[0002] Reinforcement learning is an important learning method for intelligent agents. It explores the world by actively interacting with the environment and adjusting its own strategy based on the feedback from the environment to achieve the goal of making environmental changes in line with its own expectations. At present, Actor-Critic is the mainstream model of reinforcement learning, which mainly includes a policy network (Actor) and a critic network (Critic). Among them, the policy network makes corresponding decisions based on the given environmental state, and the critic network is used to evaluate the quality of the decision.
[0003] At present, the Actor-Critic reinforcement learning model is based on the original information of the environment and uses a blind trial-and-error method to learn. Since the original information of the environment is low-level information, the amount of information represented is small, resulting in low learning efficiency and learning effect of the reinforcement learning model, and the decision made based on the model is greatly different from the expected decision. Summary of the invention
[0004] Based on the technical problem that the decision made by using the existing reinforcement learning model is significantly different from the expected decision, the embodiments of the present invention provide a decision-making method, device, equipment and medium based on knowledge embedded reinforcement learning.
[0005] In a first aspect, an embodiment of the present invention provides a decision-making method based on knowledge embedding reinforcement learning, comprising:
[0006] Obtain the original image of the target environment to be decided;
[0007] The original image to be decided is input into a pre-trained reinforcement learning model, and a decision corresponding to the original image to be decided is output; the pre-trained reinforcement learning model includes a policy network, an evaluation network, a reward function and a knowledge fusion module, the knowledge fusion module is used to fuse the input original image with prior knowledge to obtain a graph vector containing prior knowledge, and the policy network is used to output a decision to the target environment based on the graph vector.
[0008] In one possible design, the pre-trained reinforcement learning model is trained in the following manner:
[0009] S1, obtaining the current original image of the target environment;
[0010] S2, inputting the current original image into the knowledge fusion module, so that the knowledge fusion module fuses the current original image with the prior knowledge to obtain a graph vector containing the prior knowledge;
[0011] S3, inputting the graph vector containing prior knowledge into the policy network and the evaluation network, so that the policy network outputs a decision and the evaluation network outputs an evaluation value, and recording the evaluation value;
[0012] S4, applying the decision output by the policy network to the target environment to obtain an original image of the target environment after the decision is applied, calculating the original image using the reward function to obtain a reward value of the decision, and recording the reward value;
[0013] S5, determining whether the number of currently recorded return values reaches a preset number;
[0014] If not, the original image is used as the current original image and the process returns to S2;
[0015] If so, the currently recorded reward value is used to determine whether the reinforcement learning model has been trained. If not, the currently recorded reward value is used to update the parameters of the evaluation network and the currently recorded evaluation value is used to update the network parameters of the policy network. The original image is used as the current original image, the recorded evaluation value and reward value are cleared, and the process returns to execute S2. If the training is completed, the current policy network and evaluation network are used as the final policy network and evaluation network to obtain a trained reinforcement learning model.
[0016] In a possible design, the knowledge fusion module includes a scene understanding module and a domain knowledge graph; S2 includes:
[0017] Inputting the current original image into the scene understanding module, using the scene understanding module to identify at least one preset target from the current original image, and outputting the type and location information of each preset target;
[0018] Based on the type and position information of each of the preset targets, a foreground image of the current original image is generated using a semantic relationship graph network;
[0019] Generate a background image corresponding to the foreground image based on the ontology relationship of the foreground image and the prior knowledge corresponding to the foreground image provided by the current domain knowledge graph;
[0020] The foreground image and the background image are merged to obtain a scene image containing prior knowledge;
[0021] The scene graph is compressed based on graph embedding technology to obtain a graph vector containing prior knowledge.
[0022] In a possible design, the knowledge fusion module further includes a data buffer module and a knowledge induction module, and training the reinforcement learning model further includes:
[0023] The scene graph and the graph vector obtained in S2, and the decision corresponding to the scene graph and the graph vector obtained in S3 are stored in the data buffer module;
[0024] In response to the number of currently recorded reward values reaching a preset number, based on the stability relationship between the scene graph, the graph vector and the decision, extracting target knowledge from the data buffer module using the knowledge induction module;
[0025] The domain knowledge graph is updated based on the target knowledge, and the updated domain knowledge graph is used as the current domain knowledge graph, and the process returns to execute S2.
[0026] In a possible design, determining whether the reinforcement learning model is trained by using the currently recorded reward value includes:
[0027] Determine whether the number of reward values greater than the reward threshold in the currently recorded reward values is greater than the set number; if so, determine that the reinforcement learning model training is completed; if not, determine that the reinforcement learning model training is not completed.
[0028] In a possible design, the updating of the parameters of the evaluation network using the currently recorded reward value includes:
[0029] Perform weighted processing on the return value of the current record to obtain the weighted return value;
[0030] Based on the weighted returns, using proximal strategy optimization, and / or
[0031] The trust region strategy optimization algorithm updates the parameters of the evaluation network.
[0032] In a possible design, the updating of network parameters of the strategy network using the currently recorded evaluation values includes:
[0033] Based on the currently recorded evaluation value, use the proximal strategy optimization, and / or
[0034] The trust region policy optimization algorithm updates the parameters of the policy network.
[0035] In a second aspect, an embodiment of the present invention further provides a decision-making device based on knowledge embedding reinforcement learning, comprising:
[0036] An acquisition module is used to acquire the original image of the target environment to be decided;
[0037] An input module is used to input the original image to be decided into a pre-trained reinforcement learning model and output a decision that meets expectations; the pre-trained reinforcement learning model includes a policy network, an evaluation network, a reward function and a knowledge fusion module, the knowledge fusion module is used to fuse the input original image with prior knowledge to obtain a graph vector containing prior knowledge, and the policy network is used to output a decision to the target environment based on the graph vector.
[0038] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method described in any embodiment of this specification is implemented.
[0039] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, enables the computer to execute the method described in any embodiment of this specification.
[0040] An embodiment of the present invention provides a decision-making method based on knowledge-embedded reinforcement learning. In the method, a pre-trained reinforcement learning model includes a policy network, an evaluation network, a reward function and a knowledge fusion module. The model first uses the knowledge fusion module to fuse the input original image with the prior knowledge to obtain a graph vector containing the prior knowledge, and then uses the policy network to output a decision to the target environment based on the graph vector. Since the graph vector not only contains the low-level information of the original image, but also incorporates the higher-level abstract knowledge of humans, the graph vector contains more image information. It can be seen that the pre-trained reinforcement learning model has higher learning efficiency and learning effect, and can make more accurate decisions. In the present invention, by inputting the original image of the target environment to be decided into the pre-trained reinforcement learning model, since the model has a good learning effect, it can output a decision that is more in line with expectations. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0042] Figure 1 It is a structural diagram of a decision-making method based on knowledge embedding reinforcement learning provided by an embodiment of the present invention;
[0043] Figure 2 is a hardware architecture diagram of an electronic device provided by an embodiment of the present invention;
[0044] Figure 3 is a structural diagram of a decision-making device based on knowledge embedding reinforcement learning provided by an embodiment of the present invention;
[0045] Figure 4 It is a schematic diagram of a process of performing target detection and recognition on an original image of a target environment provided by an embodiment of the present invention;
[0046] Figure 5 An embodiment of the present invention provides Figure 4 Schematic diagram of extracting semantic relationship from the background image shown;
[0047] Figure 6 An embodiment of the present invention provides Figure 4 and Figure 5 Schematic diagram of knowledge fusion of background image and foreground image shown;
[0048] Figure 7 An embodiment of the present invention provides Figure 6 Schematic diagram of graph network compression of scene graph shown;
[0049] Figure 8 is a schematic diagram of a pre-trained reinforcement learning model provided in one embodiment of the present invention. DETAILED DESCRIPTION
[0050] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0051] Please refer to Figure 1 , an embodiment of the present invention provides a decision-making method based on knowledge embedding reinforcement learning, the method comprising:
[0052] Step 100, obtaining an original image of the target environment to be decided;
[0053] Step 102, input the original image to be decided into a pre-trained reinforcement learning model, and output a decision corresponding to the original image to be decided; the pre-trained reinforcement learning model includes a policy network, an evaluation network, a reward function and a knowledge fusion module, the knowledge fusion module is used to fuse the input original image with the prior knowledge to obtain a graph vector containing the prior knowledge, and the policy network is used to output a decision to the target environment based on the graph vector.
[0054] In an embodiment of the present invention, a pre-trained reinforcement learning model first uses a knowledge fusion module to fuse the input original image with prior knowledge to obtain a graph vector containing prior knowledge, and then uses a policy network to output a decision to the target environment based on the graph vector. Since the graph vector not only contains the low-level information of the original image, but also incorporates higher-level abstract knowledge of humans, the graph vector contains more image information. It can be seen that the pre-trained reinforcement learning model has higher learning efficiency and learning effect, and can make more accurate decisions. In the present invention, by inputting the original image of the target environment to be decided into the pre-trained reinforcement learning model, since the model has a good learning effect, it can output a decision that is more in line with expectations.
[0055] Described below Figure 1 How the various steps are performed.
[0056] First, with respect to step 100, an original image of the target environment to be decided is obtained.
[0057] In this step, the pre-trained reinforcement learning model is applied to an intelligent agent, such as a robot. The target environment is any scene outside the intelligent agent. The original image can be an image directly observed by the detector on the intelligent agent, or it can be obtained through an external device and then input into the reinforcement learning model. The original image can represent the current state of the target environment. The reinforcement learning model outputs a decision that is adapted to the current state based on the current state to guide the intelligent agent to make corresponding actions towards the target environment, so that the change of the target environment meets expectations.
[0058] Then, for step 102, the original image to be decided is input into a pre-trained reinforcement learning model, and a decision corresponding to the original image to be decided is output; the pre-trained reinforcement learning model includes a policy network, an evaluation network, a reward function and a knowledge fusion module, the knowledge fusion module is used to fuse the input original image with the prior knowledge to obtain a graph vector containing the prior knowledge, and the policy network is used to output a decision to the target environment based on the graph vector.
[0059] In this step, the decision can be used to guide the agent to make corresponding actions. The action acts on the target environment and will have an impact on the target environment, causing the target environment to change accordingly. Therefore, the quality of the reinforcement learning model training directly affects the quality of the decision. In addition, the prior knowledge is the same knowledge as the input original image domain, such as the domain knowledge graph, which is obtained based on historical data and historical experience, and is a semantic network that can reveal the relationship between entities in the original image.
[0060] It should also be noted that the network architecture of the policy network and the evaluation network can adopt a structure similar to the existing Actor-Critic architecture, but the inputs of both are graph vectors. This application does not specifically limit the specific network structure of the policy network and the evaluation network.
[0061] In some embodiments, the pre-trained reinforcement learning model is trained in the following manner:
[0062] S1, obtain the current original image of the target environment;
[0063] S2, inputting the current original image into the knowledge fusion module, so that the knowledge fusion module fuses the current original image with the prior knowledge to obtain a graph vector containing the prior knowledge;
[0064] S3, inputting the graph vector containing prior knowledge into the policy network and the evaluation network, so that the policy network outputs a decision and the evaluation network outputs an evaluation value, and recording the evaluation value;
[0065] S4, applying the decision output by the policy network to the target environment to obtain an original image of the target environment after being affected by the decision, calculating the original image using the reward function to obtain a reward value of the decision, and recording the reward value;
[0066] S5, determining whether the number of currently recorded return values reaches a preset number;
[0067] If not, the original image is used as the current original image and the process returns to S2;
[0068] If so, the currently recorded reward value is used to determine whether the reinforcement learning model has been trained. If not, the currently recorded reward value is used to update the parameters of the evaluation network and the currently recorded evaluation value is used to update the network parameters of the policy network. The original image is used as the current original image, the recorded evaluation value and reward value are cleared, and the process returns to execute S2. If the training is completed, the current policy network and evaluation network are used as the final policy network and evaluation network to obtain a trained reinforcement learning model.
[0069] In this embodiment, through step S2, human higher-level abstract knowledge can be integrated into the original image, so that the obtained graph vector contains more information. In step S3, the graph vector is used as the input of the policy network and the evaluation network for training, which can speed up the training speed and training effect of the policy network and the evaluation network. In addition, when training the reinforcement learning model, multiple rounds of calculations are required, and each round requires a preset number of calculations, such as 100 times. In each round of calculation, the parameters of the policy network and the evaluation network are kept unchanged. After the calculation of this round is completed, a preset number of reward values and evaluation values can be obtained. Based on the preset number of reward values, it is judged whether the conditions for the end of training are met. If they are met, the training is stopped, and the parameters of the policy network and the evaluation network in this round are used as the final parameters to obtain a trained reinforcement learning model. If not, the reward values of the preset number of this round are used to train and update the evaluation network of this round, and the evaluation values of the preset number of this round are used to train and update the policy network of this round to obtain new policy network parameters and evaluation network parameters, and used for the next round of training, and the above operations are repeated until the training is completed.
[0070] In some implementations, the knowledge fusion module includes a scene understanding module and a domain knowledge graph; step S2 includes:
[0071] A1, input the current original image into the scene understanding module, use the scene understanding module to identify at least one preset target from the current original image, and output the type and location information of each preset target.
[0072] In this step, the preset target is the target that the user is interested in. It can be set by the user according to the needs. The preset target can be one or more. The scene understanding module uses target detection and recognition technology, such as YOLO, Faster RCNN or Mask RCNN, to identify the target and output the type and position coordinate frame of each target, such as Figure 4 shown.
[0073] A2, based on the type and location information of each preset target, uses the semantic relationship graph network to generate the foreground image of the current original image, such as Figure 5 In this step, the semantic relationship graph network can be a GNN network.
[0074] A3, based on the ontological relationship of the foreground image and the prior knowledge corresponding to the foreground image provided by the current domain knowledge graph, a background image corresponding to the foreground image is generated.
[0075] In this step, the ontology relationship is a higher level of knowledge, such as Figure 4In the example, the chef is an individual and the person is the ontology, that is, the chef can be seen as a person's attribute. The ontology can give the individual more relevant information, so that more information related to the individual can be expanded based on the foreground graph. The domain knowledge graph is used to express abstract knowledge (generally a triple: subject, predicate, object), which is represented in the form of a graph network. The nodes represent the subject and object, and the edges represent the relationship or predicate. The domain knowledge graph provides background knowledge for scene understanding. Based on the foreground graph, this step further infers the ontology relationship of the foreground graph according to the pre-given domain knowledge graph corresponding to the input image domain, and obtains the corresponding background graph.
[0076] A4, fuses the foreground image and the background image to obtain a scene image containing prior knowledge.
[0077] In this step, the foreground image and the background image are fused through logical relationships to form a larger knowledge graph network, so as to achieve the purpose of supplementing the original image information with human abstract knowledge and obtain a scene graph containing prior knowledge, such as Figure 6 shown.
[0078] A5, compresses the scene graph based on graph embedding technology to obtain a graph vector containing prior knowledge.
[0079] In step A4, the semantic relationship graph network combines the original image with prior knowledge to represent it as a scene graph, which achieves information expansion. However, the scene graph is high-dimensional information. In order to use the reinforcement learning model for processing, it needs to be compressed without losing information. Therefore, this step is based on the graph embedding technology of the graph network, which can compress the scene graph into a low-dimensional vector, namely the graph vector, such as Figure 7 shown.
[0080] like Figure 8 As shown, in some implementations, the knowledge fusion module further includes a data buffer module and a knowledge induction module, and training the reinforcement learning model further includes:
[0081] B1, store the scene graph and graph vector obtained in S2, and the decision corresponding to the scene graph and graph vector obtained in S3 into the data buffer module.
[0082] The data buffer module is used to store process data. The basic structure is:<Vs(k),Gs(k),A(k),Vs(k+1),Gs(k+1)> , where Vs(k) represents the graph vector at the kth moment, Gs(k) represents the scene graph at the kth moment, A(k) represents the decision at the kth moment, Vs(k+1) represents the graph vector at the k+1th moment, and Gs(k+1) represents the scene graph at the k+1th moment. Each moment corresponds to a state of the target environment, and each state is presented in the form of an original image. The decision at the current moment affects the state of the target environment at the next moment.
[0083] B2, in response to the number of currently recorded reward values reaching a preset number, based on the stability relationship between the scene graph, the graph vector and the decision, the knowledge induction module is used to extract the target knowledge from the data buffer module.
[0084] Stability relationship means that the relationship between variables is regular and certain. If the relationship between variables changes randomly, it is an unstable relationship. In this step, the knowledge induction module is responsible for abstracting the stable relationship between variables in the reinforcement learning model during the learning process into knowledge and expressing it in the form of a knowledge graph. For example, extract data from the data buffer module, count<Gs(k),A(k),Gs(k+1)> The stable corresponding relationship in the domain knowledge graph is established, and the new relationship is added to the domain knowledge graph by using the fusion technology of the knowledge graph to realize the generation of knowledge.
[0085] B3, updates the domain knowledge graph based on the target knowledge, uses the updated domain knowledge graph as the current domain knowledge graph, and returns to execute S2.
[0086] In this step, by continuously updating the domain knowledge graph during the training process, the domain knowledge graph can be made more accurate, accelerating the learning effect and learning effect of the reinforcement learning model.
[0087] In some implementations, determining whether the reinforcement learning model is trained using the currently recorded reward value includes:
[0088] Determine whether the number of reward values greater than the reward threshold in the currently recorded reward values is greater than the set number; if so, determine that the reinforcement learning model training is completed; if not, determine that the reinforcement learning model training is not completed.
[0089] In this embodiment, the reward value ranges from 0 to 1. The larger the value, the more accurate the decision output by the strategy network is and the better the training effect is. In addition, the reward threshold and the set number are determined according to the user's training accuracy of the reinforcement learning model. The larger the reward threshold and the set number, the better the training effect, but the calculation time also increases accordingly. For example, the reward threshold can be 0.9, and the set number is greater than 90% of the preset number. Taking 100 training rounds per time as an example, 100 reward values are obtained. Then, when the number of reward values greater than 0.9 exceeds 90, the training is considered to be completed.
[0090] In some implementations, updating the parameters of the evaluation network using the currently recorded reward value includes:
[0091] Perform weighted processing on the return value of the current record to obtain the weighted return value;
[0092] Based on the weighted returns, using proximal strategy optimization, and / or
[0093] The trust region strategy optimization algorithm updates the parameters of the evaluation network.
[0094] In some implementations, updating the network parameters of the policy network using the currently recorded evaluation values includes:
[0095] Based on the currently recorded evaluation value, use the proximal strategy optimization, and / or
[0096] The trust region policy optimization algorithm updates the parameters of the policy network.
[0097] In the above steps, the updated policy network and evaluation network are used as network parameters for the next round of training.
[0098] like Figure 2 , Figure 3 As shown, the embodiment of the present invention provides a decision-making device based on knowledge embedding reinforcement learning. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. From the hardware level, Figure 2 As shown, it is a hardware architecture diagram of an electronic device where a decision-making device based on knowledge embedding reinforcement learning provided by an embodiment of the present invention is located, except Figure 2 In addition to the processor, memory, network interface, and non-volatile memory shown in the figure, the electronic device in which the device is located in the embodiment may also generally include other hardware, such as a forwarding chip responsible for processing messages, etc. Taking software implementation as an example, Figure 3 As shown, as a device in a logical sense, the CPU of the electronic device in which it is located reads the corresponding computer program in the non-volatile memory into the internal memory and runs it.
[0099] This embodiment provides a decision-making device based on knowledge embedding reinforcement learning, including:
[0100] The acquisition module 300 is used to acquire the original image of the target environment to be decided.
[0101] The input module 302 is used to input the original image to be decided into a pre-trained reinforcement learning model and output an expected decision; the pre-trained reinforcement learning model includes a policy network, an evaluation network, a reward function and a knowledge fusion module. The knowledge fusion module is used to fuse the input original image with the prior knowledge to obtain a graph vector containing the prior knowledge. The policy network is used to output a decision to the target environment based on the graph vector.
[0102] In the embodiment of the present invention, the acquisition module 300 may be used to execute step 102 in the above method embodiment, and the input module 302 may be used to execute step 102 in the above method embodiment.
[0103] In some embodiments, the pre-trained reinforcement learning model is trained in the following manner:
[0104] S1, obtain the current original image of the target environment;
[0105] S2, inputting the current original image into the knowledge fusion module, so that the knowledge fusion module fuses the current original image with the prior knowledge to obtain a graph vector containing the prior knowledge;
[0106] S3, inputting the graph vector containing prior knowledge into the policy network and the evaluation network, so that the policy network outputs a decision and the evaluation network outputs an evaluation value, and recording the evaluation value;
[0107] S4, applying the decision output by the policy network to the target environment to obtain an original image of the target environment after being affected by the decision, calculating the original image using the reward function to obtain a reward value of the decision, and recording the reward value;
[0108] S5, determining whether the number of currently recorded return values reaches a preset number;
[0109] If not, the original image is used as the current original image and the process returns to S2;
[0110] If so, the currently recorded reward value is used to determine whether the reinforcement learning model has been trained. If not, the currently recorded reward value is used to update the parameters of the evaluation network and the currently recorded evaluation value is used to update the network parameters of the policy network. The original image is used as the current original image, the recorded evaluation value and reward value are cleared, and the process returns to execute S2. If the training is completed, the current policy network and evaluation network are used as the final policy network and evaluation network to obtain a trained reinforcement learning model.
[0111] In some implementations, the knowledge fusion module includes a scene understanding module and a domain knowledge graph; S2 includes:
[0112] Input the current original image into the scene understanding module, use the scene understanding module to identify at least one preset target from the current original image, and output the type and location information of each preset target;
[0113] Based on the type and location information of each preset target, a foreground image of the current original image is generated using a semantic relationship graph network;
[0114] Based on the ontological relationship of the foreground image and the prior knowledge corresponding to the foreground image provided by the current domain knowledge graph, a background image corresponding to the foreground image is generated;
[0115] The foreground image and the background image are fused to obtain a scene image containing prior knowledge;
[0116] The scene graph is compressed based on graph embedding technology to obtain a graph vector containing prior knowledge.
[0117] In some implementations, the knowledge fusion module further includes a data buffer module and a knowledge induction module, and training the reinforcement learning model further includes:
[0118] The scene graph and graph vector obtained in S2, and the decision corresponding to the scene graph and graph vector obtained in S3 are stored in a data buffer module;
[0119] In response to the number of currently recorded reward values reaching a preset number, based on a stability relationship between the scene graph, the graph vector, and the decision, using the knowledge induction module to extract target knowledge from the data buffer module;
[0120] Update the domain knowledge graph based on the target knowledge, use the updated domain knowledge graph as the current domain knowledge graph, and return to execute S2.
[0121] In some implementations, determining whether the reinforcement learning model is trained using the currently recorded reward value includes:
[0122] Determine whether the number of reward values greater than the reward threshold in the currently recorded reward values is greater than the set number; if so, determine that the reinforcement learning model training is completed; if not, determine that the reinforcement learning model training is not completed.
[0123] In some implementations, updating the parameters of the evaluation network using the currently recorded reward value includes:
[0124] Perform weighted processing on the return value of the current record to obtain the weighted return value;
[0125] Based on the weighted returns, using proximal strategy optimization, and / or
[0126] The trust region strategy optimization algorithm updates the parameters of the evaluation network.
[0127] In some implementations, updating the network parameters of the policy network using the currently recorded evaluation values includes:
[0128] Based on the currently recorded evaluation value, use the proximal strategy optimization, and / or
[0129] The trust region policy optimization algorithm updates the parameters of the policy network.
[0130] It is understandable that the structure illustrated in the embodiment of the present invention does not constitute a specific limitation on a decision-making device based on knowledge-embedded reinforcement learning. In other embodiments of the present invention, a decision-making device based on knowledge-embedded reinforcement learning may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or arrange the components differently. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0131] The information interaction, execution process and other contents between the modules in the above-mentioned device are based on the same concept as the embodiment of the method of the present invention. For the specific contents, please refer to the description in the embodiment of the method of the present invention, and no further description is given here.
[0132] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, a decision-making method based on knowledge-embedded reinforcement learning in any embodiment of the present invention is implemented.
[0133] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor executes a decision-making method based on knowledge-embedded reinforcement learning in any embodiment of the present invention.
[0134] Specifically, a system or device equipped with a storage medium can be provided, on which software program code that implements the functions of any of the above-mentioned embodiments is stored, and a computer (or CPU or MPU) of the system or device can be enabled to read and execute the program code stored in the storage medium.
[0135] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.
[0136] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer by a communication network.
[0137] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.
[0138] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or to a memory provided in an expansion module connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or expansion module is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.
[0139] It should be noted that, in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical factors in the process, method, article or device including the elements.
[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A decision-making method based on knowledge embedding reinforcement learning, It is characterized in that include: Obtain the original image of the target environment to be decided; Input the original image to be decided into a pre-trained reinforcement learning model, and output a decision corresponding to the original image to be decided; the pre-trained reinforcement learning model includes a policy network, an evaluation network, a reward function and a knowledge fusion module, the knowledge fusion module is used to fuse the input original image with prior knowledge to obtain a graph vector containing prior knowledge, and the policy network is used to output a decision to the target environment based on the graph vector; The knowledge fusion module includes a scene understanding module and a domain knowledge graph; The process of obtaining the graph vector by the knowledge fusion module includes: Inputting the current original image into the scene understanding module, using the scene understanding module to identify at least one preset target from the current original image, and outputting the type and location information of each preset target; Based on the type and position information of each of the preset targets, a foreground image of the current original image is generated using a semantic relationship graph network; Generate a background image corresponding to the foreground image based on the ontology relationship of the foreground image and the prior knowledge corresponding to the foreground image provided by the current domain knowledge graph; The foreground image and the background image are merged to obtain a scene image containing prior knowledge; The scene graph is compressed based on graph embedding technology to obtain a graph vector containing prior knowledge.
2. The method according to claim 1, It is characterized in that The pre-trained reinforcement learning model is trained in the following way: S1, obtaining the current original image of the target environment; S2, inputting the current original image into the knowledge fusion module, so that the knowledge fusion module fuses the current original image with the prior knowledge to obtain a graph vector containing the prior knowledge; S3, inputting the graph vector containing prior knowledge into the policy network and the evaluation network, so that the policy network outputs a decision and the evaluation network outputs an evaluation value, and recording the evaluation value; S4, applying the decision output by the policy network to the target environment to obtain an original image of the target environment after the decision is applied, calculating the original image using the reward function to obtain a reward value of the decision, and recording the reward value; S5, determining whether the number of currently recorded return values reaches a preset number; If not, the original image is used as the current original image and the process returns to S2; If so, the currently recorded reward value is used to determine whether the reinforcement learning model has been trained. If not, the currently recorded reward value is used to update the parameters of the evaluation network and the currently recorded evaluation value is used to update the network parameters of the policy network. The original image is used as the current original image, the recorded evaluation value and reward value are cleared, and the process returns to execute S2. If the training is completed, the current policy network and evaluation network are used as the final policy network and evaluation network to obtain a trained reinforcement learning model.
3. The method according to claim 2, It is characterized in that The knowledge fusion module further includes a data buffer module and a knowledge summarization module, and training the reinforcement learning model further includes: storing the scene graph and the graph vector, and the decision corresponding to the scene graph and the graph vector obtained in S3 in the data buffer module; In response to the number of currently recorded reward values reaching a preset number, based on the stability relationship between the scene graph, the graph vector and the decision, extracting target knowledge from the data buffer module using the knowledge induction module; The domain knowledge graph is updated based on the target knowledge, and the updated domain knowledge graph is used as the current domain knowledge graph, and the process returns to execute S2.
4. The method according to claim 2, It is characterized in that The method of using the currently recorded reward value to determine whether the reinforcement learning model has been trained includes: Determine whether the number of reward values greater than the reward threshold in the currently recorded reward values is greater than the set number; if so, determine that the reinforcement learning model training is completed; if not, determine that the reinforcement learning model training is not completed.
5. The method according to claim 4, It is characterized in that The method of updating the parameters of the evaluation network using the currently recorded reward value includes: Perform weighted processing on the return value of the current record to obtain the weighted return value; Based on the weighted returns, using proximal strategy optimization, and / or The trust region strategy optimization algorithm updates the parameters of the evaluation network.
6. The method according to claim 5, It is characterized in that The method of updating the network parameters of the strategy network using the currently recorded evaluation values includes: Based on the currently recorded evaluation value, use the proximal strategy optimization, and / or The trust region policy optimization algorithm updates the parameters of the policy network.
7. A decision-making device based on knowledge embedded reinforcement learning, It is characterized in that include: An acquisition module is used to acquire the original image of the target environment to be decided; An input module, used to input the original image to be decided into a pre-trained reinforcement learning model, and output a decision that meets expectations; the pre-trained reinforcement learning model includes a policy network, an evaluation network, a reward function and a knowledge fusion module, the knowledge fusion module is used to fuse the input original image with prior knowledge to obtain a graph vector containing prior knowledge, and the policy network is used to output a decision to the target environment based on the graph vector; The knowledge fusion module includes a scene understanding module and a domain knowledge graph; The process of obtaining the graph vector by the knowledge fusion module includes: Inputting the current original image into the scene understanding module, using the scene understanding module to identify at least one preset target from the current original image, and outputting the type and location information of each preset target; Based on the type and position information of each of the preset targets, a foreground image of the current original image is generated using a semantic relationship graph network; Generate a background image corresponding to the foreground image based on the ontology relationship of the foreground image and the prior knowledge corresponding to the foreground image provided by the current domain knowledge graph; The foreground image and the background image are merged to obtain a scene image containing prior knowledge; The scene graph is compressed based on graph embedding technology to obtain a graph vector containing prior knowledge.
8. A computing device, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Medical image segmentation method for introducing priori knowledge based on reward function
CN115187571A
Background modeling method and device and computer storage medium
CN115937241A