Reinforcement Learning Decision Method and Device for Complex Scenarios

By using event generation network model to determine the event state set in complex scenarios and inputting the reinforcement learning network model in the current state, the problems of low reinforcement learning efficiency and poor effect in complex scenarios are solved, and more accurate and efficient decision-making is achieved.

CN117493884BActive Publication Date: 2025-06-20BEIJING INST OF CONTROL ENG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311533174.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-11-16
Publication Date
2025-06-20
Estimated Expiration
2043-11-16

AI Technical Summary

Technical Problem

In complex scenarios, when an agent performs reinforcement learning through original information, the learning efficiency is low and the effect is poor, which affects the accuracy of decision-making.

Method used

By obtaining the current state of the target environment, using the pre-trained event generation network model to determine the event state set corresponding to the current state, and input the current state and event state sets into the pre-trained reinforcement learning network model to output more accurate decisions.

Benefits of technology

Improve the efficiency of reinforcement learning training and the accuracy of results in complex scenarios, ensuring that the agent can make accurate decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117493884B_ABST
    Figure CN117493884B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and particularly relates to a reinforcement learning decision-making method and device for complex scenarios. Obtain the current state of the target environment and the set of event states corresponding to the current state, where the set of event states is determined by a pre-trained event generation network model based on the current state; the event generation network model is trained based on a sample set including multiple sample pairs, and each sample pair includes the environmental state of the target environment and the probabilities of occurrence of each event in the event set corresponding to the environmental state; input the current state and the set of event states into a pre-trained reinforcement learning network model, and output a decision corresponding to the current state, where the reinforcement learning network model is trained with the environmental state of the target environment and the set of event states output by the event generation network model as inputs. The method of the present invention can make accurate decisions for complex scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a reinforcement learning decision-making method and device for complex scenarios. Background Art

[0002] With the development of technology, intelligent agents such as robots are widely used in complex scenarios such as space operations. In practical applications, the intelligent agent first obtains the original information of the environment according to a camera or a sensor, etc., and performs reinforcement learning based on the original information. After the learning is completed, it can be used for decision-making in the actual scenario.

[0003] For simple scenarios, the intelligent agent can perform reinforcement learning based on the original information of the scenario and obtain good learning results, so as to make accurate decisions in practical applications. For complex scenarios, since the deep information contained in the original information is less, if the original information is still used for reinforcement learning, not only is the learning efficiency low, but also the learning effect is poor, which affects the accuracy of the intelligent agent's decision-making.

[0004] Therefore, there is an urgent need for a reinforcement learning decision-making method and device for complex scenarios to solve the above technical problems. Summary of the Invention

[0005] Embodiments of the present invention provide a reinforcement learning decision-making method and device for complex scenarios, which can make accurate decisions on complex scenarios.

[0006] In a first aspect, embodiments of the present invention provide a reinforcement learning decision-making method for complex scenarios, including:

[0007] Obtain the current state of the target environment and the set of event states corresponding to the current state, where the set of event states is determined by a pre-trained event generation network model based on the current state; the event generation network model is trained based on a sample set including multiple sample pairs, and each sample pair includes the environmental state of the target environment and the probabilities of occurrence of each event in the event set corresponding to the environmental state;

[0008] Input the current state and the set of event states into a pre-trained reinforcement learning network model, and output a decision corresponding to the current state. The reinforcement learning network model is trained with the environmental state of the target environment and the set of event states output by the event generation network model as inputs.

[0009] In a second aspect, embodiments of the present invention further provide a reinforcement learning decision-making device for complex scenarios, including:

[0010] An acquisition module, configured to acquire the current state of a target environment and an event state set corresponding to the current state, where the event state set is determined by a pre-trained event generation network model based on the current state; the event generation network model is trained based on a sample set including a plurality of sample pairs, and each sample pair includes an environment state of the target environment and the probabilities of occurrence of each event in the event set corresponding to the environment state.

[0011] An input module, configured to input the current state and the event state set into a pre-trained reinforcement learning network model, and output a decision corresponding to the current state, where the reinforcement learning network model is trained with the environment state of the target environment and the event state set output by the event generation network model as inputs.

[0012] In a third aspect, an embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the method described in any embodiment of this specification is implemented.

[0013] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the method described in any embodiment of this specification.

[0014] An embodiment of the present invention provides a reinforcement learning decision-making method and device for complex scenarios. In this method, when facing a complex scenario, only the current environment state needs to be known to obtain the corresponding event state set. In addition, the reinforcement learning network model is trained based on the environment state of the target environment and the event state set corresponding to the environment state. Since the environment state can represent the original information of the target environment, and the event state set can represent the higher-level information of the target environment, the training efficiency of the reinforcement learning network model is higher and the training result is more accurate through this method. As can be seen from the above, when using the method of the present invention to process complex scenarios, as long as the current state of the scenario is known, the corresponding event state set can be determined based on the event generation network model, and then the current state and the event state set are input into the reinforcement learning network model at the same time, and a more accurate decision can be obtained. Description of the Drawings

[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 It is a schematic structural diagram of a reinforcement learning decision-making method for complex scenarios provided by an embodiment of the present invention;

[0017] Figure 2 It is a hardware architecture diagram of an electronic device provided by an embodiment of the present invention;

[0018] Figure 3 It is a structural diagram of a reinforcement learning decision-making device for complex scenarios provided by an embodiment of the present invention;

[0019] Figure 4 It is a schematic structural diagram of an event generation network model and a reinforcement learning network model provided by an embodiment of the present invention;

[0020] Figure 5 It is a schematic structural diagram of an event generation network model provided by an embodiment of the present invention;

[0021] Figure 6 It is a hierarchical reinforcement learning strategy diagram based on events provided by an embodiment of the present invention. Detailed implementation manners

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0023] To better understand the solution, some terms related to the present invention will be explained below first.

[0024] Objects, relationships, and events are important factors constituting complex scenarios. How to identify these factors is the main difficulty that needs to be overcome for an agent to autonomously understand and then make decisions about complex scenarios. As a special form of information, an event refers to a specific event involving one or more participants that occurs at a specific time and specific location, and can usually be described as a change in state, which is a highly abstract concept.

[0025] Currently, learning and reasoning using reinforcement learning are completely based on raw data information (such as images, sensor information, etc.), without rising to the abstract level of events, so the efficiency is extremely low. However, many human decision-making processes are based on events, and decisions made based on events are more accurate. Therefore, extracting events from raw information is particularly important for constructing a higher-level policy learning.

[0026] Based on the above concept, the inventor proposes that reinforcement learning can be performed using the environmental state and the corresponding set of events, and the environmental state and the corresponding set of events are used as inputs to output more accurate decisions.

[0027] Please refer to Figure 1 , an embodiment of the present invention provides a reinforcement learning decision-making method for complex scenarios, and the method includes:

[0028] Step 100, obtaining the current state of the target environment and the set of event states corresponding to the current state, where the set of event states is determined by a pre-trained event generation network model based on the current state; the event generation network model is trained based on a sample set including multiple sample pairs, and each sample pair includes the environmental state of the target environment and the probabilities of occurrence of each event in the set of events corresponding to the environmental state;

[0029] Step 102, inputting the current state and the set of event states into a pre-trained reinforcement learning network model, and outputting a decision corresponding to the current state, where the reinforcement learning network model is trained with the environmental state of the target environment and the set of event states output by the event generation network model as inputs.

[0030] This embodiment provides a reinforcement learning decision-making method for complex scenarios. When facing complex scenarios, only the current environmental state needs to be known to obtain the corresponding set of event states. In addition, the reinforcement learning network model is trained based on the environmental state of the target environment and the set of event states corresponding to the environmental state. Since the environmental state can represent the original information of the target environment, and the set of event states can represent higher-level information of the target environment, the training efficiency of the reinforcement learning network model is higher and the training result is more accurate through this method. As can be seen from the above, when using the method of the present invention to process complex scenarios, as long as the current state of the scenario is known, the corresponding set of event states can be determined based on the event generation network model, and then the current state and the set of event states are input into the reinforcement learning network model at the same time to obtain a more accurate decision.

[0031] It should also be noted that the pre-trained reinforcement learning network model is applied to an intelligent agent, such as a robot, etc. The target environment is any scenario outside the intelligent agent, and the environmental state of the target environment can be the information directly observed by a detector on the intelligent agent or the information obtained through an external device. The reinforcement learning network model outputs a decision adapted to the current state based on the current state and the set of event states corresponding to the current state, and the decision can be used to guide the intelligent agent to perform corresponding actions, and the actions act on the target environment and will have an impact on the target environment to make the change of the target environment meet the expectations.

[0032] The following description Figure 1The execution manner of each step shown.

[0033] First, for step 100, obtain the current state of the target environment and the set of event states corresponding to this current state, where the set of event states is determined by a pre-trained event generation network model based on this current state; the event generation network model is trained based on a sample set containing multiple sample pairs, and each sample pair includes the environmental state of the target environment and the probabilities of occurrence of each event in the event set corresponding to this environmental state.

[0034] In some embodiments, the sample set is determined in the following manner:

[0035] Construct a simulation model of the target task, where the target task corresponds to an event set composed of multiple known events, and each known event corresponds to a probability function, and the probability function is used to characterize the probability of occurrence of the known event;

[0036] For each environmental state of the target task, perform the following: use the simulation model to calculate the value of the probability function corresponding to each known event in this environmental state, and obtain the probabilities of occurrence of each known event in the event set corresponding to this environmental state; use this environmental state as the input and the probabilities of occurrence of each known event in the event set corresponding to this environmental state as the output to obtain a sample pair of this sample set.

[0037] In this embodiment, first define a mathematical simulation model related to the target task, and this model includes a simulation environment and multiple known events related to this target task. Taking a robot's operation of grasping as an example, events can be defined as collision, grasping success, failure, object tipping, etc. Each known event is defined as a scalar function f i (x) ∈ [0, 1], representing the probability e i , where x is the internal state of the environment. During the operation of the entire system, f i (x) is calculated in real time as an indicator of whether the event occurs. For example, the collision event can be defined using a collision detection function, where collision is 1 and no collision is 0. It should be noted that the probability function of the known event can be designed manually according to needs, and this application does not make specific limitations.

[0038] When performing mathematical simulation on the entire target task, calculate the value of the probability function f i (x) of each event according to the real-time change of the environmental state S, and obtain the event state set Establish an event sample set with S and E as input-output pairs.

[0039] In addition, the environmental state S is the information obtained through sensors (vision, touch), which is called observable information.

[0040] In some embodiments, the specific training process of the event generation network model includes:

[0041] Based on a supervised learning model, use a sample set to train the event generation network model;

[0042] Determine the loss function of the supervised learning model;

[0043] For each round of training, correct the parameters of the event generation network model based on the loss function until the model converges, and obtain the trained event generation network model.

[0044] In some embodiments, the loss function is

[0045] In the formula, e i is the probability that event i occurs in the sample set, is the estimated probability that event i occurs output by the event generation network model, E represents taking the mean, i = 1, 2,..., n, and n is the number of known events.

[0046] In this embodiment, using an artificially designed abstract event function as the supervision information, learning the probability of event occurrence from a large number of scenarios generated by random exploration during the learning process can achieve autonomous learning without annotation effects and avoid manual annotation.

[0047] In some embodiments, as Figure 5 shown, the network structure of the event generation network model includes a convolutional network CNN, a long short-term memory network LSTM, and an attention network Attention.

[0048] Then, for step 102, input the current state and the event state set into a pre-trained reinforcement learning network model, and output a decision corresponding to the current state. The reinforcement learning network model is trained based on a sample set.

[0049] In some embodiments, as Figure 5 shown, the pre-trained reinforcement learning network model includes a policy network, a value network, and a reward function;

[0050] The policy network includes a first policy and a second policy; the input of the first policy is the event state set output by the event generation network model, and the output is an operation target corresponding to the event state set. This operation target serves as the target of the second policy; the input of the second policy is the environmental state corresponding to the event state set, and the output is a decision made based on this operation target and this environmental state. This decision acts on the target environment and generates a new environmental state;

[0051] The reward function takes this new environmental state as the input and calculates the reward value of this decision based on this new environmental state;

[0052] The evaluation network takes the environmental state before decision-making and the reward value output by the reward function as inputs, and is used to evaluate the pros and cons of the evaluation policy network.

[0053] In this embodiment, the first policy takes the event e i as an input, takes g i as an output, and uses it as the target of the second policy; the second policy takes the state S as an input, takes the output of the first policy as the target, and outputs the specific action a i , where i = 1, 2,... n, and n is the number of known events. In Figure 5 , π h represents the first policy, and π l represents the second policy. The first policy responds to the environment by perceiving changes in events and provides a target for the second policy. The second policy completes local sub-goals, and in this way, the decomposition of complex tasks is achieved. Among them, the network structure of the first policy can adopt structures such as LSTM structure, RNN structure, or MLP structure, and the network structure of the second policy can adopt structures such as LSTM structure, RNN structure, or MLP structure. The overall reinforcement learning network model adopts the actor-critic architecture, proposes a hierarchical reinforcement learning strategy based on events, and realizes the purpose that the first policy makes decisions and plans based on abstract information and the low-level policy makes decisions and plans based on raw information. Compared with traditional reinforcement learning methods, it can solve complex problems more efficiently.

[0054] It should be noted that for the environmental state at a certain moment, multiple events can correspond to it. Multiple events are input into the first policy at the same time, and multiple targets are output. The second policy takes the environmental state and multiple targets as inputs and outputs corresponding actions. Taking Figure 5 as an example, the environmental state S at a certain moment corresponds to two events, and two targets g0 and g1 are output. The second network takes the environmental state S as an input and g0 and g1 as targets, and outputs actions a0 to a5. Through this method, the decision made is more accurate.

[0055] In some embodiments, the pre-trained reinforcement learning network model is trained by the following method:

[0056] Obtain the current environmental state of the target environment; input the current environmental state into the event generation network model to obtain the probabilities of occurrence of each event in the current event set corresponding to the current environmental state; input the probabilities of occurrence of each event in the current event set into the first policy to output the current operation target; use the current operation target as the target of the second policy, and input the current environmental state into the second policy to obtain the current decision corresponding to the current environmental state and the current event set. The current decision acts on the target environment to generate a new environmental state;

[0057] Based on a preset reinforcement learning algorithm, the parameters of the reinforcement learning network model are updated using the new environmental state until the model converges, and a trained reinforcement learning network model is obtained.

[0058] In this embodiment, the reinforcement learning network model is trained with the environmental state and the event set corresponding to the environmental state. Since the environmental state can represent the original information of the target environment, and the event set can represent the higher-level information of the target environment, therefore, by training the reinforcement learning network model through this method, the training efficiency is higher and the training result is more accurate.

[0059] Such as Figure 2 、 Figure 3 As shown, an embodiment of the present invention provides a reinforcement learning decision-making device for complex scenarios. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. From the hardware level, as Figure 2 shown, it is a hardware architecture diagram of an electronic device where a reinforcement learning decision-making device for complex scenarios provided by an embodiment of the present invention is located. In addition to Figure 2 the processor, memory, network interface, and non-volatile memory shown, the electronic device where the device is located in the embodiment usually may also include other hardware, such as a forwarding chip responsible for processing packets, etc. Taking software implementation as an example, as Figure 3 shown, as a logically meaningful device, it is formed by the CPU of its corresponding electronic device reading the computer program in the non-volatile memory into the memory and running it.

[0060] A reinforcement learning decision-making device for complex scenarios provided in this embodiment includes:

[0061] An acquisition module 300, configured to acquire the current state of the target environment and the event state set corresponding to the current state, where the event state set is determined by a pre-trained event generation network model based on the current state; the event generation network model is trained based on a sample set including multiple sample pairs, and each sample pair includes the environmental state of the target environment and the probability of occurrence of each event in the event set corresponding to the environmental state;

[0062] An input module 302, configured to input the current state and the event state set into a pre-trained reinforcement learning network model, and output a decision corresponding to the current state, where the reinforcement learning network model is trained with the environmental state of the target environment and the event state set output by the event generation network model as inputs.

[0063] In an embodiment of the present invention, the acquisition module 300 can be used to execute step 102 in the above method embodiment, and the input module 302 can be used to execute step 102 in the above method embodiment.

[0064] In some embodiments, the sample set is determined as follows:

[0065] Construct a simulation model for the target task, where the target task corresponds to an event set composed of multiple known events, and each known event corresponds to a probability function, and the probability function is used to characterize the probability of the known event occurring;

[0066] For each environmental state of the target task, the following operations are performed: use the simulation model to calculate the values of the probability functions corresponding to each known event in this environmental state, and obtain the probabilities of occurrence of each known event in the event set corresponding to this environmental state; take this environmental state as the input and the probabilities of occurrence of each known event in the event set corresponding to this environmental state as the output, and obtain a sample pair of this sample set.

[0067] In some embodiments, the specific training process of the event generation network model includes:

[0068] Determine the loss function of the supervised learning model;

[0069] Based on the supervised learning model, use the sample set to train the event generation network model;

[0070] For each round of training, correct the parameters of the event generation network model based on the loss function until the model converges, and obtain the trained event generation network model.

[0071] In some embodiments, the loss function is:

[0072] where e i is the probability of event i occurring in the sample set, is the estimated probability of event i occurring output by the event generation network model, E represents taking the mean, i = 1, 2,..., n, and n is the number of known events.

[0073] In some embodiments, the network structure of the event generation network model includes a convolutional network, a long short-term memory network, and an attention network.

[0074] In some embodiments, the pre-trained reinforcement learning network model includes a policy network, a value network, and a reward function;

[0075] The policy network includes a first policy and a second policy; the input of the first policy is the event state set output by the event generation network model, and the output is the operation target corresponding to this event state set, and this operation target serves as the target of the second policy; the input of the second policy is the environmental state corresponding to this event state set, and the output is the decision made based on this operation target and this environmental state, and this decision acts on the target environment and generates a new environmental state;

[0076] The reward function takes this new environmental state as input and calculates the reward value of the decision based on this new environmental state;

[0077] The evaluation network takes the environmental state before the decision and the reward value output by the reward function as input, and is used to evaluate the pros and cons of the evaluation policy network.

[0078] In some embodiments, the pre-trained reinforcement learning network model is trained by the following method:

[0079] Obtain the current environmental state of the target environment; input the current environmental state into the event generation network model to obtain the probabilities of occurrence of each event in the current event set corresponding to the current environmental state; input the probabilities of occurrence of each event in the current event set into the first policy to output the current operation target; use the current operation target as the target of the second policy, and input the current environmental state into the second policy to obtain the current decision corresponding to the current environmental state and the current event set, and the current decision acts on the target environment to generate a new environmental state;

[0080] Based on a preset reinforcement learning algorithm, use the new environmental state to update the parameters of the reinforcement learning network model until the model converges to obtain the trained reinforcement learning network model.

[0081] It can be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on a reinforcement learning decision-making device for complex scenarios. In other embodiments of the present invention, a reinforcement learning decision-making device for complex scenarios may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components can be implemented in hardware, software, or a combination of software and hardware.

[0082] Regarding the information interaction, execution process, etc. between the various modules in the above device, since it is based on the same concept as the method embodiments of the present invention, the specific content can be referred to the description in the method embodiments of the present invention and will not be elaborated here.

[0083] The embodiments of the present invention also provide an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements a reinforcement learning decision-making method for complex scenarios in any embodiment of the present invention.

[0084] The embodiments of the present invention also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the processor is enabled to execute a reinforcement learning decision-making method for complex scenarios in any embodiment of the present invention.

[0085] Specifically, a system or device equipped with a storage medium can be provided, on which software program code for implementing the functions of any one of the above embodiments is stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program code stored in the storage medium.

[0086] In this case, the program code read from the storage medium itself can implement the functions of any one of the above embodiments, so the program code and the storage medium storing the program code constitute a part of the present invention.

[0087] Examples of the storage medium for providing the program code include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0088] In addition, it should be clear that not only can the functions of any one of the above embodiments be realized by executing the program code read by the computer, but also by causing an operating system or the like operating on the computer based on the instructions of the program code to complete part or all of the actual operations.

[0089] In addition, it can be understood that the program code read from the storage medium is written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion module connected to the computer, and then based on the instructions of the program code, the CPU or the like installed on the expansion board or the expansion module is caused to execute part or all of the actual operations, so as to realize the functions of any one of the above embodiments.

[0090] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0091] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A reinforcement learning decision-making method for complex scenarios, characterized in that, Including: Obtain the current state of the target environment and the set of event states corresponding to this current state. The environmental state is observable information obtained through sensors, and the set of event states is determined by a pre-trained event generation network model based on this current state; the event generation network model is trained based on a sample set containing multiple sample pairs, and each sample pair includes the environmental state of the target environment and the probabilities of occurrence of each event in the event set corresponding to this environmental state; Input the current state and the set of event states into a pre-trained reinforcement learning network model, and output a decision corresponding to this current state. The reinforcement learning network model is trained with the environmental state of the target environment and the set of event states output by the event generation network model as inputs; The sample set is determined by the following method: Construct a simulation model of the target task. The target task corresponds to an event set composed of multiple known events, and each known event corresponds to a probability function, which is used to characterize the probability of occurrence of the known event; For each environmental state of the target task, perform: use the simulation model to calculate the values of the probability functions corresponding to each known event in this environmental state, and obtain the probabilities of occurrence of each known event in the event set corresponding to this environmental state; Use this environmental state as the input and the probabilities of occurrence of each known event in the event set corresponding to this environmental state as the output to obtain a sample pair of the sample set; The specific training process of the event generation network model includes: Determine the loss function of the supervised learning model; Based on the supervised learning model, use the sample set to train the event generation network model; For each round of training, correct the parameters of the event generation network model based on the loss function until the model converges to obtain a trained event generation network model; The pre-trained reinforcement learning network model includes a policy network, a value network, and a reward function; The policy network includes a first policy and a second policy; the input of the first policy is the set of event states output by the event generation network model, and the output is an operation target corresponding to this set of event states, and this operation target serves as the target of the second policy; the input of the second policy is the environmental state corresponding to this set of event states, and the output is a decision made based on this operation target and this environmental state, and this decision acts on the target environment and generates a new environmental state; The reward function takes this new environmental state as the input and calculates the reward value of this decision based on this new environmental state; The value network takes the environmental state before the decision and the reward value output by the reward function as inputs, and is used to evaluate the quality of the policy network; The pre-trained reinforcement learning network model is trained by the following method: Obtain the current environmental state of the target environment; input the current environmental state into the event generation network model to obtain the probabilities of occurrence of each event in the current event set corresponding to the current environmental state; input the probabilities of occurrence of each event in the current event set into the first policy to output the current operation target; use the current operation target as the target of the second policy and input the current environmental state into the second policy to obtain the current decision corresponding to the current environmental state and the current event set, and the current decision acts on the target environment to generate a new environmental state; Based on a preset reinforcement learning algorithm, use the new environmental state to update the parameters of the reinforcement learning network model until the model converges to obtain a trained reinforcement learning network model.

2. The method according to claim 1, characterized in that, The loss function is as follows: ; wherein, is the probability of the occurrence of event i in the sample set, is the estimated probability of the occurrence of event i output by the event generation network model, E represents taking the mean, i = 1, 2, …… n, and n is the number of known events.

3. The method according to claim 1, characterized in that, The network structure of the event generation network model includes a convolutional network, a long short-term memory network, and an attention network.

4. A reinforcement learning decision-making device for complex scenarios, characterized in that, Applied to the method according to any one of claims 1-3, the device includes: An acquisition module, configured to acquire the current state of the target environment and an event state set corresponding to the current state, where the event state set is determined by a pre-trained event generation network model based on the current state; the event generation network model is trained based on a sample set including a plurality of sample pairs, and each sample pair includes the environmental state of the target environment and the probabilities of occurrence of each event in the event set corresponding to the environmental state; An input module, configured to input the current state and the event state set into a pre-trained reinforcement learning network model, and output a decision corresponding to the current state, where the reinforcement learning network model is trained with the environmental state of the target environment and the event state set output by the event generation network model as inputs.

5. A computing device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method described in any one of claims 1 - 3 is implemented.

6. A computer-readable storage medium, on which a computer program is stored, and when the computer program is executed on a computer, the computer is made to execute the method described in any one of claims 1 - 3.

Citation Information

Patent Citations

  • Multi-agent strategy prediction method and device

    CN112329948A

  • Reinforcement learning method and device based on environment dynamic model

    CN116362349A