An intelligent agent game confrontation method based on key-value pair attention mechanism

The AT-Double-DQN-OAP algorithm addresses the challenges of complex opponent modeling by using a key-value attention mechanism to enhance feature extraction and Q-value estimation, resulting in faster and more accurate opponent strategy learning in reinforcement learning environments.

CN116029377BActive Publication Date: 2025-07-15SHENYANG AEROSPACE UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310073929.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-07
Publication Date
2025-07-15
Estimated Expiration
2043-02-07

AI Technical Summary

Technical Problem

There are non-stationarity problems in existing agent games. The traditional implicit modeling method has low calculation accuracy, poor prediction effect, and insufficient extraction of opponent behavior characteristics, and incomplete feature information, requiring manual prior knowledge assistance.

Method used

The AT-Double-DQN-OAP algorithm based on the key-value pair attention mechanism is adopted, and the environmental information extraction module, the opponent's behavior prediction module and our behavior learning module are combined with the key-value pair attention mechanism and the Double-DQN network to achieve efficient and accurate prediction of opponent's strategy characteristics.

Benefits of technology

It improves the ability of the agent to extract opponent's behavioral characteristics in a non-stationary environment, the algorithm converges quickly, and the parameters are updated accurately, which improves the evaluation indicators and winning rate of the game strategy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116029377B_ABST
    Figure CN116029377B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent agent game confrontation method based on a key-value pair attention mechanism, and proposes an AT-Double-DQN-OAP algorithm network, which specifically includes an environmental information extraction module, an opponent behavior prediction module, and a self-behavior learning module; first, the current confrontation environment state features are extracted, and the extracted data are respectively input into the OAP behavior prediction network and the Double-Q value learning network; a key-value pair attention mechanism is added to the OAP behavior prediction network. In the behavior prediction module, the environmental state quantity S(L,V) (key) is input through a scoring function, and the query vector r is used to score the input environmental state quantity at this time. The different scores of the environmental state quantity are normalized through the softmax function to obtain the attention weights of each part. When information aggregation is performed by combining the input (value) with the attention distribution, the opponent's policy function of this part is learned emphatically, which improves the ability of the self-intelligent agent to capture the action behavior characteristics of the opponent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent agent game confrontation technology, and in particular to an intelligent agent game confrontation method based on a key-value pair attention mechanism. Background Art

[0002] In recent years, intelligent agent games have broad application prospects in competitive games, combat simulation and deduction, and its game confrontation strategy has always been a hot topic in artificial intelligence research. In the game confrontation environment, intelligent agent games have the problem of "non-stationarity". That is, the agent is often affected not only by the fixed environment in game learning, but also by the actions of other agents. In the game environment, the decision model of each agent changes over time. Therefore, the agent learning model no longer satisfies Markov decision. Therefore, how to solve the "non-stationarity" problem in intelligent agent games has become the research focus of intelligent games.

[0003] Among them, opponent behavior modeling is a means to solve the "non-stationarity" of reinforcement learning. It models and predicts the opponent's behavior information in the environment, assists our intelligent agent in reinforcement learning, predicts its strategy and intention based on the interaction information between the opponent and the environment, and adjusts the agent's own strategy accordingly. Early opponent behavior modeling mainly adopted the explicit modeling method, which has the disadvantages of complex model, large amount of calculation, and the need for human participation in model parameter setting. With the popularization of neural networks, implicit opponent behavior modeling based on neural networks has become the mainstream modeling method. However, the current implicit modeling method is based on the Q-value learning network (DQN), with low calculation accuracy and poor prediction effect. The parameter α in some search strategies needs to be selected with the help of artificial prior knowledge. The algorithm is not intelligent enough, the strategy has certain errors, and the algorithm accuracy is not ideal. Many of its improved algorithms also have problems such as slow information extraction, incomplete feature information, and insufficient concentration of behavior feature extraction combined with environmental information in the opponent behavior feature extraction part. Summary of the invention

[0004] In view of the shortcomings of the above-mentioned existing technologies, an intelligent agent game confrontation method based on a key-value pair attention mechanism is provided and applied to an intelligent agent game confrontation environment. In the process of the game between our reinforcement learning agent and the opponent's agent, our agent can efficiently and accurately learn the strategic characteristics of the opponent's agent.

[0005] An intelligent agent game confrontation method based on key-value pair attention mechanism, first defines an AT-Double-DQN-OAP algorithm; the AT-Double-DQN-OAP algorithm is divided into three modules, namely the environmental information extraction module, the opponent behavior prediction module, and the self-behavior learning module; the environmental state feature extraction module performs feature encoding on the input environmental state S, which is used as the shared input of the following two modules for targeted in-depth extraction; the opponent behavior prediction module takes the environmental state feature information s as the input, and predicts the strategy of the opponent's action through the OAP network to obtain the opponent strategy feature; the self-learning module is used to fit the Q-value function of the intelligent agent, so that the self-intelligent agent can select the optimal action according to the local action and execute it;

[0006] The method for intelligent agent game confrontation based on the AT-Double-DQN-OAP algorithm is specifically as follows:

[0007] Step 1: Use the AT-Double-DQN-OAP algorithm to encode three different types of time, space, and statistical data information to obtain the current environmental state S;

[0008] Use a recurrent neural network to collect time information to obtain a time series, use a convolutional neural network to collect spatial information to obtain convolutional image features, and use a fully connected neural network to extract data statistical information; use the three types of information extracted by the three networks to generate the feature information s after encoding the current environmental state feature extraction; and at the AT-Double-DQN-OAP algorithm level: initialize the environmental state S, initialize the value network parameters, initialize the OAP feature function, initialize the target network parameters, and initialize the training pool parameters;

[0009] Step 2: The environmental quantity input to the self-behavior learning module is directly represented by the fully connected hidden layer of the feature information s. When input to the opponent behavior prediction module, due to the introduction of key-value pair attention, the environmental state quantity needs to be expressed as a vector expression of S(K, V);

[0010] Step 3: Input the environmental state quantity S(K, V) into the opponent behavior prediction module. The opponent behavior prediction module encodes the information with greater influence in the current vectorized environmental feature information S(K, V) through the key-value pair attention mechanism, and uses the encoded environmental feature S′(K, V) as the input to extract feature information through the key-value pair network. The feature information extracted by the key-value pair satisfies:

[0011]

[0012] where q is the task query vector, N is the number of task groups, k n is the key vector of the nth group of input information, k j is the key vector of the jth group of input information; vn is the value vector of the nth group of information;

[0013] At the level of the AT-Double-DQN-OAP algorithm: Input the initial environmental state to both agents, and the opponent agent starts to take corresponding actions according to the environmental characteristics;

[0014] Step 4: Output the opponent's policy probability distribution through the softmax function for the feature information extracted in Step 3. The policy distribution π(a|att(s),θ) satisfies:

[0015]

[0016] where a′ is the next action, a is the current action, θ is the network parameter, att(s) is the feature information after key-value pair attention extraction; π is the opponent's policy distribution;

[0017] Thus, output the probability distribution of each action of the opponent agent;

[0018] Step 5: Input the environmental feature information s into the Double-DQN learning network. Compared with the traditional DQN network, Double-DQN introduces a target network Q′ to solve the problem of overestimation of Q values in the agent learning process. The target network Q′ generates the maximum Q value of the current action, and input the maximum value Q into the value network y * Generate the optimal Q * , specifically as follows:

[0019] θ′ = θ + a(y + Q(sinθ)Q(s,a,θ))

[0020] y * = E (s,a,r,s′) [r + yQ(s,argmaxQ′(s′,a′,θ′),θ)]

[0021] Q * = Q(s,a) + a(r + ymaxQ(s′,a) - Q(s,a))

[0022] where s is the current environmental state, s′ is the environmental state at the next moment, r is the transfer factor, y is the discount factor, and y * is the discount factor after eliminating the overestimated value, θ is the network parameter, θ′ is the network parameter at the next moment, Q is the action value of the current state, and Q * is the action value at the next moment;

[0023] Step 6: Calculate the loss function for the AT-Double-DQN-OAP algorithm;

[0024] Calculate the loss functions for the Double-DQN network and the AT-OAP algorithm respectively. Among them, the calculation of the Double-DQN loss function is as follows:

[0025] a l = argmaxQ′(s′, a′, θ′)

[0026] L(θ) = E (s,a,r,s′) [r + yQ(s′, a l , θ) - Q(s, a, θ) 2

[0027] The AT-OAP loss function needs to be obtained by performing a cross-entropy operation on the policy distribution of the opponent's action predicted by our agent during the game process and the opponent's behavior strategy in the real experiment. The calculation of the AT-OAP loss function is as follows:

[0028]

[0029] Our agent extracts the environmental characteristics, combines the actions of the opponent agent, and makes corresponding predictions; stores the prediction results and the actual action results of the opponent agent in the training pool, and then conducts the training at the next moment. By continuously interacting with the environment, stores the experience data; calculates the loss function for each step and performs gradient descent, and updates and iterates the value network parameters θ of our agent according to the loss function for each step; continuously repeats the above processes of agent interaction, learning, and iteration until the AT-Double-DQN-OAP algorithm converges, saves the value network parameters, and the agent learning ends.

[0030] Advantageous technical effects of the present invention:

[0031] By introducing the key-value pair attention mechanism in the present invention, it assists the agent to complete the prediction of the opponent's behavior faster and more pertinently. The AT-OAP opponent behavior feature prediction module is proposed, and the key-value pair attention mechanism is added to extract the prominent feature points of the environmental information and the opponent's behavior, and complete the behavior prediction faster and more pertinently. By fusing the AT-OAP algorithm with the Double DQN algorithm, since the opponent's behavior characteristics are effectively predicted, the algorithm converges faster, and better results are also obtained in the evaluation indexes compared in the random policy, fixed policy, and rule policy. By designing the loss function of the AT-OAP algorithm, on the basis of the traditional Double DQN loss function, a cross-entropy operation is performed on the opponent's behavior policy distribution based on the attention mechanism and the real action distribution, so that the algorithm converges faster and the network parameters are updated more accurately. During the game process between our reinforcement learning agent and the opponent agent, our agent can efficiently and accurately learn the strategy characteristics of the opponent agent, and improve the ability of our agent to extract the action behavior characteristics of the opponent in the non-stationary environment. Description of the Drawings​

[0032] Figure 1 This is a schematic diagram of the network structure of the AT-Double-DQN-OAP algorithm in the embodiment of the present invention;

[0033] Figure 2 This is a flowchart of an agent game confrontation method based on a key-value pair attention mechanism in the embodiment of the present invention;

[0034] Figure 3 This is a comparison of the convergence curves of the algorithm of the present invention and two existing algorithms in the training stage in the embodiment of the present invention;

[0035] Figure 4 This is a comparison of the confrontation step lengths of the algorithm of the present invention and two existing algorithms in the embodiment of the present invention;

[0036] Figure 5 This is a comparison of the confrontation winning rates of the algorithm of the present invention and two existing algorithms in the embodiment of the present invention. Detailed implementation manners

[0037] The following combines the drawings and embodiments to further describe in detail the specific implementation manners of the present invention.

[0038] An agent game confrontation method based on a key-value pair attention mechanism first defines an AT-Double-DQN-OAP algorithm; the AT-Double-DQN-OAP algorithm is divided into three modules, as shown in the attached Figure 1 figure, which are an environmental information extraction module, an opponent behavior prediction module, and our behavior learning module; the environmental state feature extraction module performs feature encoding on the input environmental state S, which is used as the shared input of the latter two modules for targeted in-depth extraction; the opponent behavior prediction module takes the environmental state feature information s as the input, and predicts the strategy of the opponent's action through the OAP network to obtain the opponent strategy feature; our learning module is used to fit the Q-value function of the agent so that our agent can select the optimal action to execute according to the local action;

[0039] A method for agent game confrontation based on the AT-Double-DQN-OAP algorithm, as shown in the attached Figure 2 figure, specifically:

[0040] Step 1: Use the AT-Double-DQN-OAP algorithm to encode three different types of time, space, and statistical data information to obtain the current environmental state S;

[0041] The time information is collected by a recurrent neural network to obtain a time series, the spatial information is collected by a convolutional neural network to obtain convolutional image features, and the data statistical information is extracted by a fully connected neural network; the three types of information extracted by the three networks are used to generate the current environmental state feature extraction encoded feature information s; and at the level of the AT-Double-DQN-OAP algorithm: initialize the environmental state S, initialize the value network parameters, initialize the OAP feature function, initialize the target network parameters, and initialize the training pool parameters;

[0042] Step 2: The environmental quantity input to our behavior learning module is directly represented by the fully connected hidden layer of the feature information s. When input to the opponent behavior prediction module, due to the introduction of the key-value pair attention, the environmental state quantity needs to be represented as a vector expression of S(K, V);

[0043] Step 3: Input the environmental state quantity S(K, V) into the opponent behavior prediction module. The opponent behavior prediction module encodes the information with greater influence in the current vectorized environmental feature information S(K, V) through the key-value pair attention mechanism, and uses the encoded environmental feature S′(K, V) as the input to extract feature information through the key-value pair network. The feature information extracted by the key-value pair satisfies:

[0044]

[0045] where q is the task query vector, N is the number of task groups, k n is the key vector of the nth group of input information, k j is the key vector of the jth group of input information; v n is the value vector of the nth group of information;

[0046] Based on the AT-Double-DQN-OAP algorithm level: input the initial environmental state to both agents, and the opponent agent starts to take corresponding actions according to the environmental features;

[0047] Step 4: Output the opponent probability distribution through the softmax function for the feature information extracted in Step 3. The policy distribution π(a|att(s), θ) satisfies:

[0048]

[0049] where a′ is the next action, a is the current action, θ is the network parameter, and att(s) is the feature information extracted by the key-value pair attention; π is the opponent policy distribution;

[0050] Thus, the probability distribution of each action of the opponent agent is output;

[0051] Step 5: Input the environmental feature information s into the Double-DQN learning network. Compared with the traditional DQN network, Double-DQN introduces a target network Q′ to solve the problem of overestimation of Q values during the learning process of the agent. The target network Q′ generates the maximum Q value of the current action, and inputs the maximum value Q into the value network y * Generate the optimal Q * , specifically as follows:

[0052] θ′ = θ + a(y + Q(sinθ)Q(s, a, θ))

[0053] y * = E (s,a,r,s′) [r + yQ(s, argmaxQ′(s′, a′, θ′), θ)]

[0054] Q * = Q(s, a) + a(r + ymaxQ(s′, a) - Q(s, a))

[0055] Among them, s is the current environmental state, s′ is the environmental state at the next moment, r is the transfer factor, y is the discount factor, and y * is the discount factor after eliminating the overestimated value, θ is the network parameter, θ′ is the network parameter at the next moment, Q is the action value of the current state, and Q * is the action value at the next moment;

[0056] Step 6: Calculate the loss function for the AT-Double-DQN-OAP algorithm;

[0057] Calculate the loss function for the Double-DQN network and the AT-OAP algorithm respectively. Among them, the calculation of the Double-DQN loss function is as follows:

[0058] a l = argmaxQ′(s′, a′, θ′)

[0059] L(θ) = E (s,a,r,s′) [r + yQ(s′, a l , θ) - Q(s, a, θ) 2

[0060] The AT-OAP loss function needs to be obtained by performing cross-entropy operation on the strategy distribution of the opponent's action predicted by our agent during the game process and the opponent's behavior strategy in the real experiment. The calculation of the AT-OAP loss function is as follows:

[0061]

[0062] ​Our agent extracts environmental characteristics, combines the actions of the opponent agent, and makes corresponding predictions; stores the prediction results and the actual action results of the opponent agent in the training pool, and then proceeds with the training at the next moment. By continuously interacting with the environment, it stores experience data; calculates the loss function for each step and performs gradient descent, and updates and iterates the value network parameters θ of our agent according to the loss function for each step; continuously repeats the above processes of agent interaction, learning, and iteration until the AT-Double-DQN-OAP algorithm converges, saves the value network parameters, and the agent learning ends.

[0063] Figure 3 This is the comparison of the convergence curves of this algorithm and two existing algorithms in the training stage; Figure 4 This is the comparison of the adversarial step lengths of this algorithm and two existing algorithms; Figure 5 This is the comparison of the adversarial win rates of this algorithm and two existing algorithms. The method of this embodiment is described as follows. 1000 rounds of adversarial training are designed, in which the opponent's strategy adopts two strategies of fixed actions and random actions, and two currently popular algorithms are compared. It can be seen from Figure 3 that the convergence speed of AT-Double-DQN is much faster than that of DOUBLE-DQN-OAP and DQN-OAP, especially in the middle and early stages of training (500 - 1500 steps). The rewards obtained by AT-Double-DQN-OAP are significantly dominant, indicating that compared with the other two algorithms, the D3QN-OAP algorithm using opponent action prediction can learn countermeasures against the opponent and the offensive strategy of its own side in the environment earlier. Figure 4 and Figure 5 For the algorithm proposed by the present invention in terms of adversarial step length and adversarial win rate, the adversarial step length is significantly reduced compared with the other two algorithms, that is, the game can end at a faster speed, and at the same time, the win rate is also significantly improved.

Claims

1. An intelligent agent game confrontation method based on key-value pair attention mechanism, characterized in that, First, an AT-Double-DQN-OAP algorithm is defined. The AT-Double-DQN-OAP algorithm is divided into three modules, namely the environmental information extraction module, the opponent behavior prediction module, and our behavior learning module. The environmental state feature extraction module performs feature encoding on the input environmental state S, which serves as the shared input for the following two modules for targeted in-depth extraction. The opponent behavior prediction module takes the environmental state feature information s as the input and predicts the opponent's action strategy through the OAP network to obtain the opponent strategy feature. Our learning module is used to fit the Q-value function of the agent so that our agent can select the optimal action to execute according to the opponent's action. Step 1: Use the AT-Double-DQN-OAP algorithm to encode three different types of time, space, and statistical data information to obtain the current environmental state S. Use a recurrent neural network to collect time information to obtain a time series, use a convolutional neural network to collect spatial information to obtain convolutional image features, and use a fully connected neural network to extract data statistical information. The three types of information extracted by the three networks will be used to generate the encoded feature information s of the current environmental state feature extraction. And at the AT-Double-DQN-OAP algorithm level: Initialize the environmental state S, initialize the value network parameters, initialize the OAP feature function, initialize the target network parameters, and initialize the training pool parameters. Step 2: The environmental quantity input to our behavior learning module is directly represented by the fully connected hidden layer of the feature information s. When input to the opponent behavior prediction module, due to the introduction of key-value pair attention, the environmental state quantity needs to be expressed as a vector expression of S(K, V). Step 3: Input the environmental state quantity S(K, V) into the opponent behavior prediction module. The opponent behavior prediction module encodes the information with greater influence in the current vectorized environmental feature information S(K, V) through the key-value pair attention mechanism, and takes the encoded environmental feature S′(K, V) as the input to extract feature information through the key-value pair network. Step 4: Output the opponent strategy probability distribution of the feature information extracted in Step 3 through the softmax function. Step 5: Input the environmental feature information s into the Double-DQN learning network. Compared with the traditional DQN network, Double-DQN introduces a target network Q′ to solve the problem of overestimation of Q values during the learning process of the agent. The target network Q′ generates the maximum Q value of the current action, and inputs the maximum value Q into the value network y * Generate the optimal Q * ; Step 6: Calculate the loss function for the AT-Double-DQN-OAP algorithm.

2. The intelligent agent game confrontation method based on key-value pair attention mechanism according to claim 1, wherein The feature information extraction in Step 3 satisfies: where q is the task query vector, N is the number of task groups, k n is the key vector of the n-th group of input information, k j is the key vector of the j-th group of input information; ν n is the value vector of the n-th group of information; Based on the AT-Double-DQN-OAP algorithm level: Input the initial environmental state to both agents, and the opponent agent starts to take corresponding actions according to the environmental features.

3. The intelligent agent game confrontation method based on the key-value pair attention mechanism according to claim 1, wherein The opponent strategy distribution π(a|att(s), θ) in Step 4 satisfies: where a′ is the next action, a is the current action, θ is the network parameter, att(s) is the feature information extracted by the key-value pair attention; π is the opponent strategy distribution. Thus, the probability distribution of each action of the opponent agent is output.

4. An intelligent agent game confrontation method based on a key-value pair attention mechanism according to claim 1, characterized in that Step 5 is specifically: θ′ = θ + a(y + Q(sinθ)Q(s, a, θ)) y * = E (s,a,r,s′) [r + yQ(s, argmaxQ′(s′, a′, θ′), θ)] Q * = Q(s,a) + a(r + y max Q(s′,a) - Q(s,a)) Among them, s is the current environmental state, s′ is the environmental state at the next moment, r is the transfer factor, y is the discount factor, and y * is the discount factor after eliminating the overestimated value, θ is the network parameter, θ′ is the network parameter at the next moment, Q is the action value of the current state, and Q * is the action value at the next moment.

5. An intelligent agent game confrontation method based on a key-value pair attention mechanism according to claim 1, characterized in that Step 6 is specifically: Calculate the loss function for the Double-DQN network and the AT-OAP algorithm respectively. Among them, the calculation of the Double-DQN loss function is as follows: a l = arg max Q′(s′, a′, θ′) L(θ) = E (s,a,r,s′) [r + yQ(s′, a l , θ) - Q(s, a, θ) 2 ​ The AT-OAP loss function needs to be obtained by performing cross-entropy operation on the strategy distribution of the opponent's action predicted by our agent during the game process and the behavior strategy of the opponent in the real experiment. The calculation of the AT-OAP loss function is as follows: Our agent extracts the environmental characteristics, combines the actions of the opponent agent, and makes corresponding predictions; stores the prediction results and the actual action results of the opponent agent in the training pool, and then proceeds to the training at the next moment. By continuously interacting with the environment, experience data is stored; calculates the loss function for each step and performs gradient descent, and updates and iterates the value network parameters θ of our agent according to the loss function for each step; continuously repeats the above processes of agent interaction, learning, and iteration until the AT-Double-DQN-OAP algorithm converges, saves the value network parameters, and the agent learning ends.

Citation Information

Patent Citations

  • Behavior imitation training method for air intelligent game

    CN113221444A

  • Intelligent agent task allocation method based on deep reinforcement learning

    CN114638339A