An Accelerated Construction Method for a Troop Behavior Decision Model Based on the Combination of Offline and Online Training

By combining offline training and online training methods, using expert sample reuse and enhancement mechanisms, the problem of time-consuming reinforcement learning training is solved, and the rapid construction of military behavior decision-making models is achieved and efficient data utilization is improved, and training efficiency and decision-making level are improved.

CN115062761BActive Publication Date: 2025-08-05BEIHANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210642647.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-08
Publication Date
2025-08-05
Estimated Expiration
2042-06-08

AI Technical Summary

Technical Problem

In the prior art, strengthening learning and training military behavior decision-making models take a long time, making it difficult to quickly reach the expected level under limited computing resources, and the efficiency of historical interactive data utilization is low, resulting in a long training cycle.

Method used

Combining the methods of offline training and online training, high-quality data sets are constructed through the expert sample reuse mechanism, offline pre-training is used to use the behavioral cloning algorithm, and online training is carried out in combination with the expert sample enhancement mechanism to alleviate cascade errors and improve training efficiency.

Benefits of technology

The training process of the military behavior decision-making model has been accelerated, the interaction time with the simulation environment has been reduced, data utilization efficiency has been improved, and intelligent confrontation decisions have been achieved that quickly reach the expected level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115062761B_ABST
    Figure CN115062761B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for accelerating the construction of a force behavior decision model based on a combination of offline and online training, and belongs to the technical field of computer-generated force confrontation decision-making. A method for constructing an offline data set based on an expert sample reuse mechanism is proposed to support subsequent offline behavior cloning and online reinforcement learning processes; an offline pre-training mechanism is proposed, which utilizes an expert interaction data set and combines a behavior cloning algorithm to avoid interaction with the underlying simulation environment and obtain an initial strategy with relatively good performance; an online training method based on an expert example sample enhancement mechanism is proposed, which regularly conducts strategy evaluation, and online reinforcement learning completes strategy improvement based on the knowledge of the initial strategy connotation. The technical solution of the present invention can effectively accelerate the model tuning process, quickly obtain a force behavior decision model of the expected level, and at the same time correct the cascade error problem that may exist in the behavior cloning algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer-generated force confrontation decision-making, and in particular to the technical field of accelerated construction of a reinforcement learning decision-making model in intelligent confrontation force behavior decision-making. Background Art

[0002] Computer Generated Forces (CGF) technology has become a crucial element in the field of military simulation. It builds intelligent force agents based on deep reinforcement learning algorithms and continuously interacts with the battlefield environment, constantly learning from experience, updating deep neural networks, and assisting them in making continuous behavioral decisions. It is a key technology in the current field of military intelligent confrontation behavioral decision-making.

[0003] Typically, building and training a reinforcement learning-based force behavior decision-making model relies on a large amount of interaction data generated by the online interaction between the force agent and the adversarial simulation environment. Reaching the desired level of performance for a specific task requires a long period of policy iteration. The following three reasons make the entire reinforcement learning policy iteration and training process very time-consuming:

[0004] (1) Reinforcement learning training itself requires a lot of exploration. Reinforcement learning is essentially a trial-and-error learning method. The core idea is to efficiently and stably optimize strategies from existing interaction experience to approach the mission objectives. When faced with complex force decision-making tasks, it is generally necessary for the force agent to conduct thousands of rounds of complete interaction with the simulation environment in order to explore and obtain a sufficient number of valid samples and train a better strategy.

[0005] (2) The efficiency of reinforcement learning training depends on the advancement rate of the simulation environment. Reinforcement learning requires that the military agent gradually completes the exploration and reinforcement of strategies during the process of online interaction with the simulation environment. At each time step, the agent must complete a two-way interaction with the simulation environment. Therefore, the advancement rate of the simulation environment largely determines the efficiency of reinforcement learning training. In addition, complex joint force decision-making tasks are difficult to run at high rates under limited hardware capabilities and computing resources.

[0006] (3) Reinforcement learning data cannot be reused efficiently. The mechanism of reinforcement learning is that the military agent alternately completes the two key processes of "collecting interaction data" and "optimizing iterative strategies" during the interaction with the simulation environment. In conventional same-strategy reinforcement learning algorithms, the agent uses the interaction data between its latest strategy and the environment to complete strategy evaluation and self-evolution. The interaction data between the old version of the iterative strategy and the environment can only be used in the current iteration. Even for different-strategy reinforcement learning algorithms that can use the old version of the strategy, the sampling efficiency problem of reinforcement learning itself has not been well solved. Therefore, in the multi-round, long-term interactive training and learning process in large-scale military simulation scenarios, the collection of effective interaction data is time-consuming and the utilization efficiency is not high.

[0007] In most cases, designers need to experiment with different reinforcement learning algorithms, search under various hyperparameter configurations, and train the network under limited computing resources to find the most efficient reinforcement learning method for force behavior decision modeling. Therefore, it is necessary to accelerate the entire reinforcement learning training process so that designers can obtain feedback on algorithm training results in a short period of time, speeding up the tuning process and quickly achieving the desired force behavior decision model.

[0008] In summary, accelerating the construction of reinforcement learning-based intelligent adversarial force behavior decision-making models is a key issue that requires breakthroughs. Designing a rapid learning and training mechanism within a data-driven paradigm that efficiently utilizes historical interaction data between the intelligent agent and the adversarial simulation environment, reduces the cold start time of the reinforcement learning force behavior decision-making model training system, and accelerates the optimization of the agent to the ideal intelligent adversarial level through iterative reinforcement learning strategy training is a key technical challenge that urgently needs to be overcome in the application of reinforcement learning methods to force behavior decision-making modeling in joint force adversarial simulation scenarios. Summary of the Invention

[0009] This invention addresses the problem of accelerating the construction of reinforcement learning decision models for force confrontation and proposes a method for accelerating the construction of decision models that combines offline and online training. This method uses a behavioral cloning algorithm to supervise the learning of expert samples, making full use of expert example data. However, because offline behavioral cloning does not consider the long-term impact of the current state, subtle errors will be gradually amplified in the sequential decision-making process, resulting in cascading errors. The method proposed in this invention combines offline and online training to alleviate the cascading error problem existing in offline behavioral cloning. The specific technical solutions of this invention are as follows:

[0010] A method for accelerating the construction of a force behavior decision model based on a combination of offline and online training includes the following steps:

[0011] S1: Offline datasets are constructed based on the expert sample reuse mechanism, and the interaction data between different types of strategies and simulation environments are integrated to form a high-quality dataset that supports subsequent offline training. Specifically, when facing specific force decision-making tasks, different types of expert strategies based on rule reasoning, flowcharts, and finite state machines interact with the simulation environment, and interaction data with reward information is generated after the interaction. Based on subsequent offline and online learning of different paradigms, the interaction data is specifically reconstructed and processed to form a "behavior-action" expert dataset for offline imitation learning. At the same time, the rewarded expert dataset is used as a permanent subset of the sample pool of the online deep Q network DQN. While performing online reinforcement learning, sampling is performed from the interaction data of the DQN strategy and the expert dataset to achieve continuous retention of expert interaction data.

[0012] S2: Offline pre-training step, specifically including: using the behavior cloning algorithm (BC) to perform offline supervised training based on existing expert example data. During the offline pre-training phase, interaction with the underlying simulation environment is avoided. After offline pre-training, an initial policy that meets preset conditions is obtained, where the preset conditions are related to the performance of the policy.

[0013] S3: Online training based on the expert example sample enhancement mechanism, taking advantage of the fact that the different-policy DQN can fully utilize the interaction data of any behavior strategy, combined with the expert example data reuse mechanism, the expert data is always used as a subset of the experience sample pool. At the same time, an expert dataset enhancement mechanism is proposed. During the DQN online training process, strategy evaluation is performed regularly. According to the different improvement thresholds achieved by the strategy, different proportions of the online DQN interaction dataset are stored in the expert data.

[0014] Furthermore, the specific process of step S1 is as follows:

[0015] S1-1: Define the reward function r t (s t ,a t );

[0016] S1-2: Make the military agent follow the expert strategy π E After several rounds of interaction with the adversarial environment, a series of rewarded expert strategy interaction sequences {τ1,τ2,…,τ m}, each expert strategy interaction sequence contains state, action and corresponding reward The expert strategy interaction sequence is called expert example data, which reflects what kind of behavior decision the expert strategy will make when facing a certain confrontation situation. All m sequence data form a data set.

[0017] S1-3: Initialize two empty data sets

[0018] S1-4: For the dataset D E The following operations are performed on the sequence in the data set: all the "state-behavior" pairs in the sequence are extracted in turn as the data set Each "state-behavior" pair (s, a) is used as a training example, with the state vector s as the feature vector feature and the parameterized behavior a as the label label, supporting subsequent supervised behavior cloning imitation learning;

[0019] S1-5: For the dataset D E Operate the sequence in: extract all the short sequences of "state-behavior-reward-new state" in the sequence and put them into the data sample pool In , each short sequence of “state-behavior-reward-next state” (s, a, r, s′) is used as a training sample, where The sequence looks like this:

[0020]

[0021] So far, an expert dataset for behavior cloning imitation learning and an expert dataset for online reinforcement learning have been constructed based on the expert sample reuse mechanism.

[0022] Furthermore, the specific process of step S2 is as follows:

[0023] S2-1: In the discrete action space, the strategy for discrete actions is In the discrete action space One-Hot encoding is performed on it to make it a reasonable multi-classification label, and a training algorithm is trained to classify the output into the target action when the input state s is given. classifier;

[0024] S2-2: Construct a fully connected neural network as a discrete Actor network. The input layer width is the dimension of the state vector, and the output layer dimension is k, where k is the number of optional discrete actions. The k is also the discrete action space. The Actor network is used to represent the strategy obtained by imitating the expert strategy, wherein the Actor network is for the discrete action space Perform supervised imitation learning to obtain strategy π d ;

[0025] S2-3: The Actor network takes the state vector as input, and each output f of the network a1 ,f a2 ,…,f akCorresponding to k discrete actions, the softmax(f) operation is applied to the output layer to map the k network output values to the probability value of each discrete action, and then random sampling is performed according to the probability distribution determined by this probability value to obtain the predicted discrete action a;

[0026] S2-4: Based on the classification characteristics of the discrete Actor network, the cross entropy CrossEntropy shown in the following formula is calculated as the Loss function, the loss function formula L d as follows:

[0027]

[0028] in, Indicates the action in the sample One-Hot encoding vector of d (s) represents the output vector of the discrete Actor network; softmax(·) represents the softmax operation applied to the output layer; N batch represents the number of samples in a supervised learning training; H(·,·) represents the cross entropy, and the cross entropy formula is as follows:

[0029] H(p,q)=-p·log(q)

[0030] Among them, p represents the classification label vector in One-Hot form; q represents the probability value vector of each classification predicted by the network.

[0031] S2-5: Back propagation updates parameters based on the loss function L d Perform gradient descent to complete the network parameter update, where α is the parameter update learning rate, and the calculation formula is as follows:

[0032]

[0033] S2-6: Complete the offline training of the Actor network and obtain the initial strong policy model π through behavior cloning imitation learning start (a|s;θ) discrete action a selected by the Actor network k Sure.

[0034] Furthermore, the specific process of step S3 is as follows:

[0035] S3-1: Define algorithm parameters: total number of training rounds M, experience pool size N D , the network training frequency is N m step, each batch of samples has a sample size of N batch , set the reward discount factor γ and the strategy evaluation frequency to N M Step, the N M Nm An integer multiple of , the number of strategy evaluation interaction rounds k;

[0036] S3-2: Store the DQN sample pool;

[0037] S3-3: Initialize the DQN value network based on the initial strong policy model π start (a|s;θ) interacts with the environment and transforms the sample (s i ,a i ,r i ,s i+1 ) Deposit into experience pool D dynamic middle;

[0038] S3-4: Judgment Experience Pool D dynamic Is the number of samples in less than N? D If it is less than, return to S3-3, otherwise, execute S3-5;

[0039] S3-5: Start DQN online training and set the interaction count N i = 0, each interaction with the environment executes N i =N i +1;

[0040] S3-6: Interact with the environment and judge N i Whether N m If the result is not divisible, then repeat S3-6. If the result is divisible, then execute steps S3-7 to S3-8 in sequence.

[0041] S3-7: Uniformly sample N from expert datasets and preset excellent DQN datasets batch The size of the data sample;

[0042] S3-8: Calculate the network loss function according to the DQN algorithm update method, and update the parameters of the DQN value network based on the gradient descent method;

[0043] S3-9: Behavioral decision-making and environmental interaction based on DQN value network m Step 1, dynamically update the DQN sample pool, and put the first N m The first step is to delete the sample and store the newly generated sample in the sample pool;

[0044] S3-10: Determine N i Whether N M If it is not divisible, then go back to S3-7 and continue executing. If it is, then go to S3-11.

[0045] S3-11: Perform strategy evaluation and continue to interact with the environment for k rounds based on the behavior decision generated by the current DQN value network. At the same time, store the sample in D tempIn the process, determine whether the average reward value of round k is greater than the average reward in the expert data set. If not, go back to S3-6 and continue executing. If satisfied, go to S3-12.

[0046] S3-12: Determine whether the average reward value of round k is greater than 150% of the average reward in the expert dataset. If it satisfies, randomly sample D temp 50% of the samples are stored in Jump to step S3-15 to execute. If not satisfied, execute S3-13;

[0047] S3-13: Determine whether the average reward value of round k is greater than 125% of the average reward in the expert dataset. If it satisfies, randomly sample D temp 25% of the samples are stored in Jump to step S3-15 to execute. If not satisfied, execute S3-14;

[0048] S3-14: Random Sampling D temp 10% of the samples are stored in

[0049] S3-15: D dynamic Sample pool update, respectively from Sample pool and Sample evenly from the sample pool and store in D dynamic In the sample pool;

[0050] S3-16: Determine whether the total number of training rounds is greater than M. If so, return to S3-7 for execution; otherwise, execute S3-17.

[0051] S3-17: End.

[0052] The beneficial effects of the present invention are:

[0053] (1) A mechanism for reusing expert example samples is proposed to form an expert dataset for offline and online learning, providing effective support for the construction of expert example datasets in model acceleration methods.

[0054] (2) Aiming at the problem of slow and long training cycle of reinforcement learning in military confrontation, a set of offline and online reinforcement learning decision model acceleration construction solutions that can fully utilize expert example data is provided. The offline pre-training stage is introduced to accelerate the model construction process, avoid interaction with the underlying simulation system, and alleviate the cascade error problem existing in offline behavior cloning.

[0055] (3) An expert sample enhancement mechanism is proposed. During the DQN online training, strategy evaluation is performed regularly to retain the better DQN strategies, thereby enhancing the expert dataset. Based on the expert sample enhancement mechanism proposed in this invention, online reinforcement learning completes strategy improvement on the basis of retaining the knowledge of the initial strategy connotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. By referring to the drawings, the features and advantages of the present invention will be more clearly understood. The drawings are schematic and should not be understood as limiting the present invention in any way. Those skilled in the art can derive other drawings based on these drawings without inventive effort. Among them:

[0057] Figure 1 It is the overall framework diagram

[0058] Figure 2 This is a schematic diagram of offline pre-training

[0059] Figure 3 This is the online DQN flow chart DETAILED DESCRIPTION

[0060] In order to make the technical solution of the present invention clearer, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. Figure 1 The overall framework diagram illustrates many specific details to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited to the specific embodiments disclosed below.

[0061] Specifically, a method for accelerating the construction of a force behavior decision model based on a combination of offline and online training includes the following steps:

[0062] S1: Offline datasets are constructed based on the expert sample reuse mechanism, and the interaction data between different types of strategies and simulation environments are integrated to form a high-quality dataset that supports subsequent offline training. Specifically, when facing specific force decision-making tasks, different types of expert strategies based on rule reasoning, flowcharts, and finite state machines interact with the simulation environment, and interaction data with reward information is generated after the interaction. Based on subsequent offline and online learning of different paradigms, the interaction data is specifically reconstructed and processed to form a "behavior-action" expert dataset for offline imitation learning. At the same time, the rewarded expert dataset is used as a permanent subset of the sample pool of the online deep Q network DQN. While performing online reinforcement learning, sampling is performed from the interaction data of the DQN strategy and the expert dataset to achieve continuous retention of expert interaction data.

[0063] S2: Offline pre-training step, specifically including: using the behavior cloning algorithm (BC) to perform offline supervised training based on existing expert example data. During the offline pre-training phase, interaction with the underlying simulation environment is avoided. After offline pre-training, an initial policy that meets preset conditions is obtained, where the preset conditions are related to the performance of the policy.

[0064] S3: Online training based on the expert example sample enhancement mechanism, taking advantage of the fact that the different-policy DQN can fully utilize the interaction data of any behavior strategy, combined with the expert example data reuse mechanism, the expert data is always used as a subset of the experience sample pool. At the same time, an expert dataset enhancement mechanism is proposed. During the DQN online training process, strategy evaluation is performed regularly. According to the different improvement thresholds achieved by the strategy, different proportions of the online DQN interaction dataset are stored in the expert data.

[0065] Furthermore, the algorithm flow for constructing the offline dataset based on the expert sample reuse mechanism in step S1 is as follows:

[0066] S1-1: Define the reward function r t (s t ,a t ), the reward function is based on the input state s t and action a t , giving a reward value r t ;

[0067] S1-2: Make the military agent follow the expert strategy π E After several rounds of interaction with the adversarial environment, a series of rewarded expert strategy interaction sequences {τ1,τ2,…,τ m}, each expert strategy interaction sequence contains state, action and corresponding reward The τ i The superscript i in the sequence represents the i-th sequence, and the subscript represents the time step in the sequence. The expert strategy interaction sequence is called expert example data, which reflects what kind of behavior decision the expert strategy will make when facing a certain confrontation situation. All m sequence data form a data set

[0068] S1-3: Initialize two empty data sets The dataset The dataset is used to store data that supports behavioral cloning imitation learning. Used to store data that supports online reinforcement learning;

[0069] S1-4: For the dataset D EThe following operations are performed on the sequence in the data set: all the "state-behavior" pairs in the sequence are extracted in turn as the data set Each "state-behavior" pair (s, a) is used as a training example, with the state vector s as the feature vector feature and the parameterized behavior a as the label label, supporting subsequent supervised behavior cloning imitation learning;

[0070] S1-5: For the dataset D E Operate the sequence in: extract all the short sequences of "state-behavior-reward-new state" in the sequence and put them into the data sample pool In , each short sequence of “state-behavior-reward-next state” (s, a, r, s′) is used as a training sample, where The sequence looks like this:

[0071]

[0072] So far, an expert dataset for behavior cloning imitation learning and an expert dataset for online reinforcement learning have been constructed based on the expert sample reuse mechanism.

[0073] Furthermore, the algorithm flow of offline pre-training using the behavior cloning algorithm in step S2 is as follows, and its offline pre-training schematic diagram is as follows: Figure 2 As shown:

[0074] S2-1: In the discrete action space, the strategy for discrete actions is In the discrete action space One-Hot encoding is performed on the above. A simple example in the field of tank force behavior decision-making is used to illustrate One-Hot encoding. The tank agent has two discrete actions to choose from, firing a1 and retreating a2. The discrete action space is The discrete action a1 is mapped to a One-Hot encoded vector label (1,0), and the discretized action a2 is mapped to a One-Hot encoded vector label (0,1). The One-Hot encoding makes it a reasonable multi-classification label to train a classifier that can classify the output into the target action when the input state s is s. classifier;

[0075] S2-2: Construct a fully connected neural network as a discrete Actor network. The input layer width is the dimension of the state vector, and the output layer dimension is k, where k is the number of optional discrete actions. The k is also the discrete action space. The Actor network is used to represent the strategy obtained by imitating the expert strategy, wherein the Actor network is for the discrete action space Perform supervised imitation learning to obtain strategy π d ;

[0076] S2-3: The Actor network takes the state vector as input, and each output f of the network a1 ,f a2 ,…,f ak Corresponding to k discrete actions, a softmax(f) operation is applied to the output layer. The softmax function normalizes the network output and maps the k network output values to the probability value of each discrete action. Then, random sampling is performed according to the probability distribution determined by the probability value to obtain the predicted discrete action a.

[0077] S2-4: Based on the classification characteristics of the discrete Actor network, the cross entropy CrossEntropy shown in the following formula is calculated as the Loss function, the loss function formula L d as follows:

[0078]

[0079] in, Indicates the action in the sample One-Hot encoding vector of d (s) represents the output vector of the discrete Actor network; softmax(·) represents the softmax operation applied to the output layer; N batch represents the number of samples in a supervised learning training; H(·,·) represents the cross entropy, and the cross entropy formula is as follows:

[0080] H(p,q)=-p·log(q)

[0081] Among them, p represents the classification label vector in One-Hot form; q represents the probability value vector of each classification predicted by the network.

[0082] S2-5: Back propagation updates parameters based on the loss function L d Perform gradient descent to complete the network parameter update, where α is the parameter update learning rate, and the calculation formula is as follows:

[0083]

[0084] S2-6: Complete the offline training of the Actor network and obtain the initial strong policy model π through behavior cloning imitation learning start (a|s;θ) discrete action a selected by the Actor network k Sure.

[0085] At this point, offline pre-training based on the behavior cloning algorithm has been completed. Since it does not require online interaction with the environment, does not involve the estimation of the value function, and can use the graphics processing unit (GPU) to complete accelerated calculations, the supervised imitation learning training process is more efficient and faster than reinforcement learning training.

[0086] Furthermore, in step S3, the initial strong strategy model π determined by the Actor network is obtained by using supervised imitation learning after sufficient rounds of iterative training. start (a|s;θ). Based on this model, if Figure 3 As shown in the flowchart, online reinforcement learning training is performed based on DQN to further optimize the cascading error problem that may exist in behavior cloning imitation learning:

[0087] S3-1: Define algorithm parameters: total number of training rounds M, experience pool size N D , the network training frequency is N m Step, the network training frequency is N m Step by step to update the network parameters of the reinforcement learning algorithm, and the sample size of each batch is N batch , set the reward discount factor γ and the strategy evaluation frequency to N M Step, the N M N m The number of strategy evaluation interaction rounds k is an integer multiple of , where the number of strategy evaluation interaction rounds is k steps of interaction between a single strategy evaluation and the environment;

[0088] S3-2: Store the DQN sample pool;

[0089] S3-3: Initialize the DQN value network based on the initial strong policy model π start (a|s;θ) interacts with the environment and transforms the sample (s i ,a i ,r i ,s i+1 ) Deposit into experience pool D dynamic middle;

[0090] S3-4: Judgment Experience Pool D dynamic Is the number of samples in less than N? D If it is less than, return to S3-3, otherwise, execute S3-5;

[0091] S3-5: Start DQN online training and set the interaction count N i = 0, each interaction with the environment executes N i =N i +1;

[0092] S3-6: Interact with the environment and judge N i Whether N m If the result is not divisible, then repeat S3-6. If the result is divisible, then execute steps S3-7 to S3-8 in sequence.

[0093] S3-7: Uniformly sample N from expert datasets and preset excellent DQN datasets batch The size of the data sample;

[0094] S3-8: Calculate the network loss function according to the update method of the DQN algorithm, and update the parameters of the DQN value network based on the gradient descent method. The DQN algorithm network loss function L i (θ i ) is as follows:

[0095] L i (θ i )=E[(y i -Q(s,a;θ i )) 2 ]

[0096] Wherein, the subscript i represents the i-th iteration, θ is the DQN value network parameter, E[·] is the mathematical expectation function, Q(s,a) is the DQN algorithm value function when taking action a in the current state s, and y i is the target value of the current i-th iteration, and its expression is as follows:

[0097] y i =E[r+γ′max a ′Q(s′,a′;θ i-1 )|s,a]

[0098] Where E[·|s,a] is the mathematical expectation of taking action a in the current state s, s′ is the state of the next time step, a′ is the action of the next time step, r is the current reward function, and γ′ is the discount factor;

[0099] S3-9: Behavioral decision-making and environmental interaction based on DQN value network m Step 1, dynamically update the DQN sample pool, and put the first N m The first step is to delete the sample and store the newly generated sample in the sample pool;

[0100] S3-10: Determine N i Whether N M If it is not divisible, then go back to S3-7 and continue executing. If it is, then go to S3-11.

[0101] S3-11: Perform strategy evaluation and continue to interact with the environment for k rounds based on the behavior decision generated by the current DQN value network. At the same time, store the sample in Dtemp In the D temp A temporary storage pool for sample sequences that interact with the environment for k rounds, to determine whether the average reward value for k rounds is greater than the average reward in the expert dataset Where γ is the reward discount factor. If it is not satisfied, the process returns to S3-6 to continue the execution. If it is satisfied, the process goes to S3-12.

[0102] S3-12: Determine whether the average reward value of round k is greater than 150% of the average reward in the expert dataset. If it satisfies, randomly sample D temp 50% of the samples are stored in Jump to step S3-15 to execute. If not satisfied, execute S3-13;

[0103] S3-13: Determine whether the average reward value of round k is greater than 125% of the average reward in the expert dataset. If it satisfies, randomly sample D temp 25% of the samples are stored in Jump to step S3-15 to execute. If not satisfied, execute S3-14;

[0104] S3-14: Random Sampling D temp 10% of the samples are stored in

[0105] S3-15: D dynamic Sample pool update, respectively from Sample pool and Sample evenly from the sample pool and store in D dynamic In the sample pool, the D dynamic The sample pool has been enriched with better DQN strategy samples, and the strategy has been improved;

[0106] S3-16: Determine whether the total number of training rounds is greater than M. If so, return to S3-7 for execution; otherwise, execute S3-17.

[0107] S3-17: End.

[0108] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for accelerating the construction of a force behavior decision model based on a combination of offline and online training, characterized by: The method specifically comprises the following steps: S1: Offline datasets are constructed based on the expert sample reuse mechanism, integrating the interaction data between different types of strategies and simulation environments to form a high-quality dataset that supports subsequent offline training. Specifically, when facing specific force decision-making tasks, different types of expert strategies based on rule reasoning, flowcharts, and finite state machines interact with the simulation environment, generating interaction data with reward information. Based on subsequent offline and online learning of different paradigms, the interaction data is specifically reconstructed and processed to form a "behavior-action" expert dataset for offline imitation learning. At the same time, the rewarded expert dataset is used as a permanent subset of the sample pool of the online deep Q network DQN. Samples are taken from the interaction data of the DQN strategy and the expert dataset during online reinforcement learning to achieve continuous retention of expert interaction data. S2: Offline pre-training step, specifically including: using the behavior cloning algorithm (BC) to perform offline supervised training based on existing expert example data. During the offline pre-training stage, interaction with the underlying simulation environment is avoided. After offline pre-training, an initial policy that meets preset conditions is obtained, where the preset conditions are related to the performance of the policy. S3: Online training based on the expert example sample enhancement mechanism, taking advantage of the fact that the different-policy DQN can fully utilize the interaction data of any behavior strategy, combined with the expert example data reuse mechanism, the expert data is always used as a subset of the experience sample pool. At the same time, an expert dataset enhancement mechanism is proposed. During the DQN online training process, strategy evaluation is performed regularly. According to the different improvement thresholds achieved by the strategy, different proportions of the online DQN interaction dataset are stored in the expert data.

2. The method according to claim 1, characterized in that The specific process of S1 is as follows: S1-1: Define the reward function r t (s t , a t ); S1-2: Make the military agent follow the expert strategy π E After several rounds of interaction with the adversarial environment, a series of rewarded expert strategy interaction sequences {τ1, τ2, ..., τ m }, each expert strategy interaction sequence contains state, action and corresponding reward The expert strategy interaction sequence is called expert example data, which reflects what kind of behavior decision the expert strategy will make when facing a certain confrontation situation. All m sequence data form a data set. S1-3: Initialize two empty data sets S1-4: For the dataset D E The following operations are performed on the sequence in the data set: all the "state-behavior" pairs in the sequence are extracted in sequence as the data set Each "state-behavior" pair (s, a) is used as a training example, with the state vector s as the feature vector and the parameterized behavior a as the label, supporting subsequent supervised behavior cloning imitation learning; S1-5: For the dataset D E The sequence in the sequence is operated: all the short sequences of "state-behavior-reward-new state" in the sequence are extracted and put into the data sample pool. In the example, each short sequence (s, a, r, s′) of “state-behavior-reward-next-state” is used as a training sample, where The sequence looks like this: So far, an expert dataset for behavior cloning imitation learning and an expert dataset for online reinforcement learning have been constructed based on the expert sample reuse mechanism.

3. The method according to claim 1, characterized in that The specific process of S2 is as follows: S2-1: In the discrete action space, the strategy for discrete actions is In the discrete action space One-Hot encoding is performed on it to make it a reasonable multi-classification label, and a training algorithm is trained to classify the output into the target action when the input state s is given. classifier; S2-2: Construct a fully connected neural network as a discrete Actor network. The input layer width is the dimension of the state vector, and the output layer dimension is k, where k is the number of optional discrete actions. The k is also the discrete action space. The Actor network is used to represent the strategy obtained by imitating the expert strategy, wherein the Actor network is for the discrete action space Perform supervised imitation learning to obtain strategy π d ; S2-3: The Actor network takes the state vector as input, and each output f of the network a1 , f a2 ,...,f ak Corresponding to k discrete actions, the softmax(f) operation is applied to the output layer to map the k network output values to the probability value of each discrete action, and then random sampling is performed according to the probability distribution determined by this probability value to obtain the predicted discrete action a; S2-4: Based on the classification characteristics of the discrete Actor network, the cross entropy shown in the following formula is calculated as the Loss function. The loss function formula L d as follows: in, Indicates the target action One-Hot encoding vector of d (s) represents the output vector of the discrete Actor network; softmax(·) represents the softmax operation applied to the output layer; N batch represents the number of samples in a supervised learning training; H(·,·) represents the cross entropy, and the cross entropy formula is as follows: H(p,q)=-p·log(q) Where p represents the classification label vector in the form of One-Hot; q represents the probability value vector of each classification predicted by the network; S2-5: Back propagation updates parameters based on the loss function L d Perform gradient descent to complete the network parameter update, where α is the parameter update learning rate, and the calculation formula is as follows: S2-6: Complete the offline training of the Actor network and obtain the initial strong policy model π through behavior cloning imitation learning start (a|s;θ) discrete action a selected by the Actor network k Sure.

4. The method according to claim 1, wherein The specific process of S3 is as follows: S3-1: Define algorithm parameters: total number of training rounds M, experience pool size N D , the network training frequency is N m step, each batch of samples has a sample size of N batch , set the reward discount factor γ and the strategy evaluation frequency to N M Step, the N M N m An integer multiple of , the number of strategy evaluation interaction rounds k; S3-2: Store the DQN sample pool; S3-3: Initialize the DQN value network based on the initial strong policy model π start (a|s;θ) interacts with the environment and transforms the sample (s i , a i , r i , s i+1 ) Deposit into experience pool D dynamic middle; S3-4: Judgment Experience Pool D dynamic Is the number of samples in less than N? D If it is less than, return to S3-3, otherwise, execute S3-5; S3-5: Start DQN online training and set the interaction count N i = 0, each interaction with the environment executes N i =N i +1; S3-6: Interact with the environment and judge N i Whether N m If the result is not divisible, then repeat S3-6. If the result is divisible, then execute steps S3-7 to S3-8 in sequence. S3-7: Uniformly sample N from expert datasets and preset excellent DQN datasets batch The size of the data sample; S3-8: Calculate the network loss function according to the DQN algorithm update method, and update the parameters of the DQN value network based on the gradient descent method; S3-9: Behavioral decision-making and environmental interaction based on DQN value network m Step 1, dynamically update the DQN sample pool, and put the first N m The first step is to delete the sample and store the newly generated sample in the sample pool; S3-10: Determine N i Whether N M If it is not divisible, then go back to S3-7 and continue. If it is, then go to S3-11. S3-11: Perform strategy evaluation, continue interacting with the environment for k rounds based on the behavioral decisions generated by the current DQN value network, and store the samples in Dtemp to determine whether the average reward value of k rounds is greater than the average reward in the expert dataset. If not, go back to S3-6 and continue executing. If satisfied, go to S3-12. S3-12: Determine whether the average reward value of round k is greater than 150% of the average reward in the expert dataset. If it satisfies, randomly sample D temp 50% of the samples are stored in Jump to step S3-15 to execute. If not satisfied, execute S3-13; S3-13: Determine whether the average reward value of round k is greater than 125% of the average reward in the expert dataset. If it satisfies, randomly sample D temp 25% of the samples are stored in Jump to step S3-15 to execute. If not satisfied, execute S3-14; S3-14: Random Sampling D temp 10% of the samples are stored in S3-15: D dynamic Sample pool update, respectively from Sample pool and Sample evenly from the sample pool and store in D dynamic In the sample pool; S3-16: Determine whether the total number of current training rounds is greater than M. If it is less, return to S3-7 to execute, otherwise execute S3-17; S3-17: End.

Citation Information

Patent Citations

  • Model privacy protection method and system for deep reinforcement learning

    CN113420326A

  • Unmanned aerial vehicle intelligent decision-making system based on imitation learning and reinforcement learning

    CN113741533A