A training method, device, equipment and storage medium for a neural network model

By training neural network models in multiple agent collaboration environments, using behavior prediction networks and value networks, the problems of value shift and lack of team awareness during AI model training are solved, and a higher level of agent simulation and team collaboration are achieved.

CN114272611BActive Publication Date: 2025-05-30TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111242582.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-25
Publication Date
2025-05-30
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

The optimization goal of existing AI models during training is to maximize the benefits of each agent, resulting in a shift in the agent's values, lack of team awareness, and a low simulation level.

Method used

Train neural network models in an environment where multiple agents collaborate, and through behavior prediction networks and value networks, the agents are encouraged to perform actions that can maximize their own value and team value, thereby improving the team's joint action quality.

Benefits of technology

A higher level of intelligent body simulation has been achieved, the effectiveness of team collaboration has been improved, and the agent can perform actions that maximize the team's value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114272611B_ABST
    Figure CN114272611B_ABST
Patent Text Reader

Abstract

The present application discloses a neural network model training method implemented based on artificial intelligence technology, including: obtaining first global information and first unit information at a first moment; based on the first global information and the first unit information, obtaining the action distribution of each agent at the first moment through an action prediction network; obtaining second global information and second unit information according to the first action distribution; determining reward information according to the first global information, the first unit information, the second global information and the second unit information; based on the global information and the unit information, obtaining the target's own value and the target team value through a value network; training the network according to the action distribution, the reward information, the target's own value and the target team value. The present application also provides a device. The present application encourages agents in the same camp to perform actions that can maximize their own value and the team value, so as to make the joint actions of the team better and achieve a higher level of agent simulation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment, and storage medium for training a neural network model. Background Art

[0002] With the development of Internet technology, there are more and more types of online games, and Multiplayer Online Battle Arena (MOBA) is one of them. In such games, players are usually divided into two or more opposing camps and compete with each other in a scattered game map. Each player controls the selected game character through an interface to fight against the other side.

[0003] The game character can not only be controlled by the player, but also be controlled by an Artificial Intelligence (AI) model to fight. At present, the reinforcement learning algorithm is mainly used in the prior art, so that the agent can learn a strategy to maximize the reward or achieve a specific goal during the interaction with the environment.

[0004] The inventor found that there are at least the following problems in the prior art. The optimization goal during the training of the AI model is to maximize the income of each agent, while the essence of a team game is to maximize the income of the whole team. Therefore, the model optimization method will cause the deviation of the agent's values and lack of team awareness, resulting in a low simulation level of the agent. Summary of the Invention

[0005] Embodiments of the present application provide a method, device, equipment, and storage medium for training a neural network model, which trains the network model used in reinforcement learning in an environment where multiple agents cooperate, encourages the agents in the same camp to execute actions that can maximize their own value and the team value, so that the joint actions of the team are better and a higher agent simulation level is achieved.

[0006] In view of this, on the one hand, the present application provides a method for training a neural network model, including:

[0007] Obtain the first global information and the first unit information corresponding to the first moment, where the first global information is used to describe the game situation of the camp at the first moment, and the first unit information is used to describe the state of the interaction unit at the first moment, and the interaction unit represents an object having an interaction relationship with the game character in the game;

[0008] Based on the first global information and the first unit information, obtain the action distribution of each of the K agents at the first moment through the behavior prediction network, where the K agents belong to the same camp, the agents are game characters controlled by artificial intelligence, and K is an integer greater than 1;

[0009] According to the action distribution of each agent, obtain the second global information and the second unit information at the second moment, where the second moment is the next moment of the first moment, the second global information is used to describe the game situation of the camp at the second moment, and the second unit information is used to describe the state of the interaction unit at the second moment;

[0010] According to the first global information, the first unit information, the second global information, and the second unit information, determine the reward information of each agent, where the reward information is the incentive obtained after triggering the action;

[0011] Based on the first global information, the first unit information, the second global information, and the second unit information, obtain the target self-value and the target team value of each agent through the value network, where the target team value represents the value of the agent in the camp;

[0012] Train the behavior prediction network and the value network according to the action distribution, the reward information, the target self-value, and the target team value of each agent.

[0013] Another aspect of this application provides a training method for a neural network model, including:

[0014] Call the behavior prediction network to conduct self-play battles to obtain training data;

[0015] Obtain the global information and the unit information at different moments from the training data, where the global information is used to describe the game situation of the camp, the unit information is used to describe the real-time state of the interaction unit, and the interaction unit represents an object that has an interaction relationship with the game character in the game;

[0016] According to the global information and the unit information at different moments, determine the reward information of each agent in the same camp, where the reward information is the incentive obtained after triggering the action, and the agent is a game character controlled by artificial intelligence;

[0017] Based on the global information and the unit information, obtain the target self-value and the target team value of each agent through the value network, where the target team value represents the value of the agent in the camp;

[0018] Train the behavior prediction network and the value network according to the reward information, the target self-value, and the target team value of each agent.

[0019] On the other hand, this application provides a model training device, including:

[0020] An acquisition module, configured to acquire first global information and first unit information corresponding to a first moment, where the first global information is used to describe the game situation of the camp at the first moment, and the first unit information is used to describe the state of the interaction unit at the first moment, and the interaction unit represents an object having an interaction relationship with the game character in the game;

[0021] The acquisition module is further configured to, based on the first global information and the first unit information, obtain the action distribution of each of the K agents at the first moment through a behavior prediction network, where the K agents belong to the same camp, the agent is a game character controlled by artificial intelligence, and K is an integer greater than 1;

[0022] The acquisition module is further configured to, according to the action distribution of each agent, obtain second global information and second unit information at a second moment, where the second moment is the next moment of the first moment, the second global information is used to describe the game situation of the camp at the second moment, and the second unit information is used to describe the state of the interaction unit at the second moment;

[0023] A determination module, configured to determine the reward information of each agent according to the first global information, the first unit information, the second global information, and the second unit information, where the reward information is the incentive obtained after triggering the action;

[0024] The acquisition module is further configured to, based on the first global information, the first unit information, the second global information, and the second unit information, obtain the target self-value and the target team value of each agent through a value network, where the target team value represents the value of the agent in the camp;

[0025] A training module, configured to train the behavior prediction network and the value network according to the action distribution, the reward information, the target self-value, and the target team value of each agent.

[0026] In a possible design, in another implementation manner of another aspect of the embodiments of this application,

[0027] The acquisition module is specifically configured to acquire the first global information corresponding to the first moment from the screen at the first moment, where the first global information includes at least one of the global map information and the camp confrontation information corresponding to the first moment;

[0028] Acquire the first unit information corresponding to the first moment from the screen at the first moment, where the first unit information includes at least one of the game character information, the non-player character (NPC) information, and the defense tower information corresponding to the first moment.

[0029] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0030] An acquisition module, specifically configured to obtain, based on the first global information and the first unit information, the first action distribution of each of the K agents at the first moment through the first action prediction network;

[0031] Based on the first global information and the first unit information, obtain the second action distribution of each of the K agents at the first moment through the second action prediction network, where the second action prediction network is generated after the first action prediction network.

[0032] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0033] An acquisition module, specifically configured to extract the first global feature based on the first global information through the global feature network included in the first action prediction network;

[0034] Based on the first unit information, extract the first unit feature through the local feature network included in the first action prediction network;

[0035] Based on the first global feature and the first unit feature, obtain the first action distribution of each of the K agents at the first moment through the fully connected layer included in the first action prediction network;

[0036] An acquisition module, specifically configured to extract the second global feature based on the first global information through the global feature network included in the second action prediction network;

[0037] Based on the first unit information, extract the second unit feature through the local feature network included in the second action prediction network;

[0038] Based on the second global feature and the second unit feature, obtain the second action distribution of each of the K agents at the first moment through the fully connected layer included in the second action prediction network.

[0039] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0040] An acquisition module, specifically configured to determine the predicted action of each agent according to the first action distribution of each agent at the first moment, where the predicted action is the action corresponding to the maximum probability in the first action distribution;

[0041] Interact the predicted action of each agent with the game environment to obtain the second global information and the second unit information at the second moment.

[0042] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0043] A determination module, specifically configured to obtain the dense incentive factor of each agent according to the first unit information and the second unit information, where the dense incentive factor includes at least one of an asset incentive factor, an experience incentive factor, and a blue amount incentive factor;

[0044] Obtain the sparse incentive factor of each agent according to the first unit information and the second unit information, where the sparse incentive factor includes at least one of a kill incentive factor, a death incentive factor, and an assist incentive factor;

[0045] Obtain the survival incentive factor of each agent according to the first unit information and the second unit information, where the survival incentive factor includes at least one of a self-survival incentive factor and an enemy survival incentive factor;

[0046] Obtain the turret incentive factor of each agent according to the first unit information and the second unit information, where the turret incentive factor includes at least one of a quantity incentive factor and a firmness incentive factor;

[0047] Obtain the game incentive factor of each agent according to the first global information and the second global information.

[0048] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0049] An acquisition module, specifically configured to obtain the first self-value of each agent at the first moment through the first self-value network based on the first global information and the first unit information;

[0050] Obtain the second self-value of each agent at the second moment through the first self-value network based on the second global information and the second unit information;

[0051] Obtain the third self-value of each agent at the first moment through the second self-value network based on the first global information and the first unit information, where the second self-value network is generated after the first self-value network;

[0052] Obtain the first team value of each agent at the first moment through the first team value network based on the first global information and the first unit information;

[0053] Obtain the second team value of each agent at the second moment through the first team value network based on the second global information and the second unit information;

[0054] Based on the first global information and the first unit information, obtain the third team value of each agent at the first moment through the second team value network, where the second team value network is generated after the first team value network.

[0055] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0056] The acquisition module is specifically configured to obtain the dense value factor of each agent at the first moment through the dense evaluation network included in the first self-value network based on the first global information and the first unit information;

[0057] Based on the first global information and the first unit information, obtain the sparse value factor of each agent at the first moment through the sparse evaluation network included in the first self-value network;

[0058] Based on the first global information and the first unit information, obtain the survival value factor of each agent at the first moment through the survival evaluation network included in the first self-value network;

[0059] Based on the first global information and the first unit information, obtain the defense tower value factor of each agent at the first moment through the defense tower evaluation network included in the first self-value network;

[0060] Based on the first global information and the first unit information, obtain the game value factor of each agent at the first moment through the game evaluation network included in the first self-value network.

[0061] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0062] The acquisition module is specifically configured to obtain the role value vector of each agent among the K agents, where the role value vector includes L feature elements, and L is an integer greater than or equal to 1;

[0063] Recombine the role value vectors of each agent among the K agents to obtain L hierarchical value vectors, where each hierarchical value vector includes the feature elements of each agent;

[0064] Based on the L hierarchical value vectors, the first global information, and the first unit information, obtain L tap value vectors through the value attention network included in the first team value network, where each tap value vector includes T tap values, and T is an integer greater than or equal to 1;

[0065] Based on the first global information, the first unit information, and the L tap value vectors, obtain the first team value of each agent at the first moment through the value fusion network included in the first team value network.

[0066] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0067] An acquisition module, specifically configured to obtain a first tap value vector through a dense value attention network included in a value attention network based on first global information, first unit information, and a dense layer value vector among L hierarchical value vectors, where the value attention network is included in a first team value network;

[0068] Based on the first global information, the first unit information, and a sparse layer value vector among the L hierarchical value vectors, obtain a second tap value vector through a sparse value attention network included in the value attention network;

[0069] Based on the first global information, the first unit information, and a survival layer value vector among the L hierarchical value vectors, obtain a third tap value vector through a survival value attention network included in the value attention network;

[0070] Based on the first global information, the first unit information, and a turret layer value vector among the L hierarchical value vectors, obtain a fourth tap value vector through a turret value attention network included in the value attention network;

[0071] Based on the first global information, the first unit information, and a game layer value vector among the L hierarchical value vectors, obtain a fifth tap value vector through a game value attention network included in the value attention network, where the first tap value vector, the second tap value vector, the third tap value vector, the fourth tap value vector, and the fifth tap value vector are included in the L tap value vectors.

[0072] In a possible design, in another implementation of another aspect of the embodiments of the present application,

[0073] An acquisition module, specifically configured to obtain a first fusion value through a dense value fusion network included in a value fusion network based on the first global information, the first unit information, and a first tap value vector among the L tap value vectors, where the value fusion network is included in the first team value network;

[0074] Based on the first global information, the first unit information, and a second tap value vector among the L tap value vectors, obtain a second fusion value through a sparse value fusion network included in the value fusion network;

[0075] Based on the first global information, the first unit information, and a third tap value vector among the L tap value vectors, obtain a third fusion value through a survival value fusion network included in the value fusion network;

[0076] Based on the first global information, the first unit information, and the fourth tap value vector among the L tap value vectors, obtain the fourth fusion value through the defense tower value fusion network included in the value fusion network;

[0077] Based on the first global information, the first unit information, and the fifth tap value vector among the L tap value vectors, obtain the fifth fusion value through the game value fusion network included in the value fusion network, where the first fusion value, the second fusion value, the third fusion value, the fourth fusion value, and the fifth fusion value are included in the first team value of each agent in the same camp at the first moment.

[0078] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,

[0079] The training module is specifically configured to determine the target loss value according to the target self-value, the target team value, the action distribution, and the reward information of each agent;

[0080] Update the network parameters of the second behavior prediction network by using the target loss value, where the second behavior prediction network is included in the behavior prediction network;

[0081] Determine the first loss value according to the third self-value, the reward information, and the second self-value of each agent at the second moment, where the third self-value is obtained through the second self-value network, the second self-value is obtained through the first self-value network, and the second self-value network is generated after the first self-value network;

[0082] Update the network parameters of the second self-value network by using the first loss value, where the second self-value network is included in the second value network in the value network;

[0083] Determine the second loss value according to the third team value, the reward information, and the second team value of each agent at the second moment, where the third team value is obtained through the second team value network, the second team value is obtained through the first team value network, and the second team value network is generated after the first team value network;

[0084] Update the network parameters of the second team value network by using the second loss value, where the second team value network is included in the second value network in the value network.

[0085] In a possible design, in another implementation of another aspect of the embodiments of the present application, the target self-value includes the first self-value at the first moment and the second self-value at the second moment obtained through the first self-value network, and the target team value includes the first team value at the first moment and the second team value at the second moment obtained through the first team value network;

[0086] The training module is specifically configured to calculate the self-advantage estimation value according to the first self-value of each agent at the first moment, the second self-value at the second moment, and the reward information;

[0087] Calculate the team advantage estimation value according to the first team value of each agent at the first moment, the second team value at the second moment, and the reward information;

[0088] Determine the self-loss value according to the action distribution of each agent and the self-advantage estimation value;

[0089] Determine the team loss value according to the action distribution of each agent and the team advantage estimation value;

[0090] Sum up the self-loss value and the team loss value to calculate the target loss value.

[0091] Another aspect of the present application provides a model training device, including:

[0092] The calling module is used to call the behavior prediction network to conduct self-play battles to obtain training data;

[0093] The acquisition module is used to obtain the global information and unit information at different moments from the training data, where the global information is used to describe the game situation of the camp, and the unit information is used to describe the real-time state of the interaction unit, and the interaction unit represents an object having an interaction relationship with the game character in the game;

[0094] The determination module is used to determine the reward information of each agent in the same camp according to the global information and unit information at different moments, where the reward information is the incentive obtained after triggering the action, and the agent is a game character controlled by artificial intelligence;

[0095] The acquisition module is further used to obtain the target self-value and target team value of each agent through the value network based on the global information and unit information, where the target team value represents the value of the agent in the camp;

[0096] The training module is used to train the behavior prediction network and the value network according to the reward information, target self-value, and target team value of each agent.

[0097] On the other hand, this application provides a computer device, including: a memory, a processor, and a bus system;

[0098] Wherein, the memory is used for storing programs;

[0099] The processor is used for executing the programs in the memory, and the processor is used for preparing to execute the methods provided in the above aspects according to the instructions in the program code;

[0100] The bus system is used for connecting the memory and the processor to enable the memory and the processor to communicate with each other.

[0101] On the other hand, this application provides a computer-readable storage medium, in which instructions are stored, and when the instructions are run on a computer, the computer is enabled to execute the methods in the above aspects.

[0102] In another aspect of this application, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, enabling the computer device to execute the methods provided in the above aspects.

[0103] As can be seen from the above technical solutions, the embodiments of this application have the following advantages:

[0104] In the embodiments of this application, a method for training a neural network model is provided. The first global information and the first unit information corresponding to the first moment are obtained. Thus, based on the first global information and the first unit information, the action distribution of each of the K agents at the first moment can be obtained through the action prediction network. Then, according to the action distribution of each agent, the second global information and the second unit information at the second moment are obtained. Based on this, according to the first global information, the first unit information, the second global information, and the second unit information, the reward information of each agent is determined. Then, based on the first global information, the first unit information, the second global information, and the second unit information of each agent, the target self-value and the target team value of each agent are obtained through the value network. Finally, according to the action distribution, the reward information, the target self-value, and the target team value of each agent, the action prediction network and the value network can be trained. Through the above method, in an environment where multiple agents (game characters controlled by an AI model) cooperate, the network model used in reinforcement learning is trained. Based on the constructed value network, value collaborative learning of multiple agents is realized, that is, agents in the same camp are encouraged to execute actions that can maximize their own value and the team value, so that the joint actions of the team are better and a higher level of agent simulation is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0105] Figure 1 It is a schematic diagram of an architecture for reinforcement learning in an embodiment of the present application;

[0106] Figure 2 It is a schematic diagram of an environment of a network model training system in an embodiment of the present application;

[0107] Figure 3 It is a schematic flowchart of a neural network model training method in an embodiment of the present application;

[0108] Figure 4 It is a schematic diagram of a decision model training process in an embodiment of the present application;

[0109] Figure 5 It is a schematic diagram of a value model training process in an embodiment of the present application;

[0110] Figure 6 It is a schematic diagram of global information and unit information in an embodiment of the present application;

[0111] Figure 7 It is a schematic diagram of a structure of a behavior prediction network in an embodiment of the present application;

[0112] Figure 8 It is a schematic diagram of a structure of a self - value network in an embodiment of the present application;

[0113] Figure 9 It is a schematic diagram of a structure of a value attention network in an embodiment of the present application;

[0114] Figure 10 It is a schematic diagram of a structure of a dense value attention network in an embodiment of the present application;

[0115] Figure 11 It is a schematic diagram of an interface of corresponding weights of an agent at different levels in an embodiment of the present application;

[0116] Figure 12 It is a schematic diagram of obtaining a tap value vector based on a value attention network in an embodiment of the present application;

[0117] Figure 13 It is a schematic diagram of a structure of a value fusion network in an embodiment of the present application;

[0118] Figure 14 It is a schematic diagram of a structure of a dense value fusion network in an embodiment of the present application;

[0119] Figure 15 It is a schematic diagram of obtaining a first team value based on a value fusion network in an embodiment of the present application;

[0120] Figure 16Another schematic diagram of the decision model training process in the embodiments of this application;

[0121] Figure 17 Another schematic diagram of the value model training process in the embodiments of this application;

[0122] Figure 18 Another flowchart of the neural network model training method in the embodiments of this application;

[0123] Figure 19 A schematic diagram of the model training device in the embodiments of this application;

[0124] Figure 20 A schematic structural diagram of the computer device in the embodiments of this application. Detailed implementation manners

[0125] The embodiments of this application provide a method, device, device, and storage medium for training a neural network model, which trains a network model used in reinforcement learning in an environment where multiple agents cooperate, encourages agents in the same camp to execute actions that can maximize their own value and the team value, so that the joint actions of the team are better and a higher level of agent simulation is achieved.

[0126] The terms "first", "second", "third", "fourth", etc. (if any) in the description and claims of this application and the above drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "corresponding to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0127] The solution provided by this application can be applied to training agents in a multi-player cooperative game, where an agent refers to a game character controlled by an AI model, and the multi-player cooperative game includes but is not limited to MOBA games, role-playing games (RPG), shooting games, and action games, etc.

[0128] Taking the MOBA game as an example, in a MOBA game, multiple players (e.g., 10 players) are divided into two opposing camps to conduct online confrontations, teamwork, and resource contention on the same map. The game is won by ultimately destroying the "crystal (i.e., the base)" of the enemy through competition. The behaviors performed by players in the game can be divided into two categories. One is macro scheduling, that is, friendly heroes in the team cooperate for long-distance scheduling or cooperate for rapid resource contention, such as long-distance support, cooperative contention for in-game resources, etc. The other is micro operation, that is, the immediate behaviors adopted by each hero in various scenarios, such as last-hitting in confrontation, skill combination release, avoiding enemy skills, etc. The MOBA game has a complex map environment, numerous optional characters, rich equipment combinations, and diverse player tactics, resulting in a relatively high complexity and strong exploratory nature of the MOBA game, making it difficult to rely on artificial rules to implement an intelligent agent. To build an intelligent agent for the MOBA game, a reinforcement learning algorithm can be used for self-play training and learning of AI. By designing reasonable incentives, the game intelligent agent can explore various action behaviors, learn the optimal behavior strategy, and obtain the maximum expected reward to achieve the ultimate victory of the game.

[0129] The technical solution provided by this application can train the intelligent agent in the game based on the (Reinforcement Learning, RL) algorithm. For ease of understanding, please refer to Figure 1 , Figure 1 which is a schematic diagram of the architecture of reinforcement learning in the embodiments of this application. As shown in the figure, the intelligent agent (agent) receives the state of the t-th frame (i.e., S t ) from the environment (environment), and then takes corresponding actions (i.e., A t ) based on the state of the t-th frame. As a result, the environment transitions to the next frame state (i.e., S t+1 ). The environment gives the corresponding reward (i.e., R t+1 ) according to the state of the (t + 1)-th frame (i.e., S t+1 ), thereby guiding subsequent action decisions. Based on this, the entire environment cycles continuously.

[0130] To improve the simulation level of the intelligent agent and achieve a better teamwork effect, this application proposes a training method for a neural network model, which is applied to Figure 2The network model training system shown in the figure, as shown in the figure, the network model training system includes a server and a terminal device, and the client is deployed on the terminal device. Among them, the client can run on the terminal device in the form of a browser, or can run on the terminal device in the form of an independent application (APP), etc. For the specific display form of the client, it is not limited here. The server involved in this application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a notebook computer, a handheld computer, a personal computer, a smart TV, a smart watch, a vehicle-mounted device, a wearable device, etc., but is not limited thereto. The terminal device and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this here. The number of the terminal device and the server is also not limited. The solution provided by this application can be completed independently by the terminal device, or can be completed independently by the server, or can be completed by the cooperation of the terminal device and the server. For this, this application does not make specific limitations.

[0131] Exemplarily, in a MOBA game, in addition to human players vs. human players, there can also be AI players vs. AI players. Taking Figure 2 as an example, player A controls game character A in the red camp through terminal device A, and player B controls game character B in the red camp through terminal device B. Therefore, game character A and game character B belong to the same camp. The one they are fighting against is the blue camp, and the blue camp includes game character C controlled by agent A and game character D controlled by agent B. Among them, the agents can be deployed on the server, and the server calls the AI model to control each agent in the same camp to execute actions.

[0132] This application provides a method for training a neural network model, which is implemented through a Machine Learning (ML) model and specifically involves ML technology and Computer Vision (CV) technology in the field of Artificial Intelligence (AI). AI is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, AI is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0133] AI technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. AI basic technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics, etc. AI software technologies mainly include several major directions such as CV technology, speech processing technology, natural language processing technology, and ML / deep learning, autonomous driving, and intelligent transportation.

[0134] CV is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes for machine vision such as target recognition, trace tracing, and measurement, and further performing graphic processing to make the computer process into images that are more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, CV studies related theories and technologies and attempts to build an AI system that can obtain information from images or multi-dimensional data. CV technology usually includes technologies such as image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, autonomous driving, intelligent transportation, etc., and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0135] ML is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. ML is the core of AI and the fundamental way to make computers intelligent, and its applications cover all fields of AI. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.

[0136] Combined with the above introduction, the training method of the neural network model in this application will be introduced below. Please refer to Figure 3 One embodiment of the training method of the neural network model in the embodiment of this application includes:

[0137] 110. Obtain the first global information and the first unit information corresponding to the first moment. Among them, the first global information is used to describe the game situation of the camp at the first moment, and the first unit information is used to describe the state of the interaction unit at the first moment. The interaction unit represents an object that has an interaction relationship with the game character in the game;

[0138] In one or more embodiments, the model training device obtains the first global information and the first unit information at the first moment. The first moment can be the t-th moment or the t-th frame. In this application, the first moment is taken as the t-th moment for illustration. Thus, after extracting the game screen corresponding to the t-th moment, the first global information and the first unit information can be extracted from the game screen. Among them, the first global information can describe the overall game situation of multiple camps at the first moment, and the first unit information can describe the survival state, skill state, standing position state, etc. of each interaction unit (for example, game unit) at the first moment.

[0139] It should be noted that the model training device can be deployed on the server, or the terminal device, or a system composed of the server and the terminal device, which is not limited here.

[0140] 120. Based on the first global information and the first unit information, obtain the action distribution of each of the K agents at the first moment through the behavior prediction network. Among them, the K agents belong to the same camp. The agent is a game character controlled by artificial intelligence, and K is an integer greater than 1;

[0141] In one or more embodiments, the model training device takes the first global information and the first unit information of the agent as the state information of the agent at the t-th moment, and inputs it into the behavior prediction network, and outputs the action distribution of the agent at the first moment through the first behavior prediction network. By analogy, the action distribution of each of the K agents belonging to the same camp at the first moment can be obtained. Among them, the behavior prediction network can specifically be a decision network.

[0142] Exemplarily, in the game, there are two camps, and each camp includes 5 agents. Then K is 5 at this time. And the action distribution at each moment can be represented as an M-dimensional probability distribution vector. Each dimension in the probability distribution vector corresponds to an action type, and M is an integer greater than 1.

[0143] 130. Obtain the second global information and the second unit information at the second moment according to the action distribution of each agent, where the second moment is the next moment of the first moment, the second global information is used to describe the game situation of the camp at the second moment, and the second unit information is used to describe the state of the interaction unit at the second moment;

[0144] In one or more embodiments, the model training device can respectively determine the actions of each of the K agents at the first moment according to the action distribution of each agent among the K agents at the first moment. Based on this, these agents respectively interact with the environment using the actions at the first moment, and thus, obtain the state information at the next moment (i.e., the second global information and the second unit information at the second moment). Among them, the second global information can describe the overall game situation of multiple camps at the second moment, while the second unit information can describe the survival state, skill state, standing position state, etc. of each interaction unit at the second moment.

[0145] 140. Determine the reward information of each agent according to the first global information, the first unit information, the second global information, and the second unit information, where the reward information is the incentive obtained after triggering the action;

[0146] In one or more embodiments, the model training device compares the state information of two adjacent moments to determine the reward information of each of the K agents.

[0147] Specifically, compare the first global information and the first unit information at the first moment with the second global information and the second unit information at the second moment, and thus, obtain the reward information of each agent. Among them, the reward information is the reward obtained after interacting with the environment using the action selected at the first moment.

[0148] 150. Based on the first global information, the first unit information, the second global information, and the second unit information, obtain the target self - value and the target team - value of each agent through the value network, where the target team - value represents the value of the agent in the camp;

[0149] In one or more embodiments, the model training device inputs the first global information and the first unit information of each of the K agents at the first moment into the value network, and outputs the self - value and the team - value of each agent at the first moment through the value network. The model training device inputs the second global information and the second unit information of each of the K agents at the second moment into the value network, and outputs the self - value and the team - value of each agent at the second moment through the value network.

[0150] It can be understood that the self - value at the first moment and the self - value at the second moment both belong to the target self - value, and the team value at the first moment and the team value at the second moment both belong to the target team value. Among them, the self - value represents the expected benefit that the agent can obtain for itself after taking actions starting from a state. The team value represents the expected benefit that the agent can obtain for the camp it belongs to after taking actions starting from a state.

[0151] 160. Train the behavior prediction network and the value network according to the action distribution, reward information, target self - value, and target team value of each agent.

[0152] In one or more embodiments, the model training device uses the action distribution, reward information, target self - value, and target team value of each of the K agents, and adopts a corresponding loss function to train the behavior prediction network and the value network.

[0153] Specifically, according to the action distribution, target self - value, and target team value of each agent, a reinforcement learning algorithm can be used to update the network parameters of the behavior prediction network. It can be understood that the reinforcement learning algorithm includes but is not limited to Proximal Policy Optimization (PPO), Deep Deterministic Policy Gradient (DDPG), and Asynchronous Advantage Actor - Critic (A3C). According to the reward information, target self - value, and target team value of each agent, algorithms such as Mean Squared Error (MSE), Root Mean Squared Error (RMSE), or Mean Absolute Error (MAE) can be used to update the network parameters of the value network.

[0154] For ease of understanding, please refer to Figure 4 , Figure 4 which is a schematic diagram of the decision model training process in the embodiments of this application. As shown in the figure, the state information (s_t) at the t - th moment is respectively input into the behavior prediction network and the value network. The action distribution (p_t) is output through the behavior prediction network, and the target value (v_t) at the t - th moment is output through the value network. Based on the action distribution (p t ) select the corresponding action (a t ). After the action (a t ) interacts with the environment, the state information (s t+1), and thus the reward information (r t ) can be obtained. In addition, the state information (s t+1 ) at the (t + 1)-th moment is input into the value network, and the target value (v t+1 ) at the (t + 1)-th moment is output through the value network. Using the action distribution (p t ), the target value (v t ) at the t-th moment, the target value (v t+1 ) at the (t + 1)-th moment, and the reward information (r t ), the behavior prediction network can be trained. Among them, the target value includes the target's own value and the target team value.

[0155] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the value model training process in the embodiment of the present application. As shown in the figure, similarly, after obtaining the reward information (r t ), the target value (v t ) at the t-th moment, and the target value (v t+1 ) at the (t + 1)-th moment, the value network can be trained.

[0156] In the embodiment of the present application, a training method for a neural network model is provided. Through the above method, the network model used in reinforcement learning is trained in an environment where multiple agents (game characters controlled by an AI model) cooperate, and value collaborative learning of multiple agents is realized based on the constructed value network, that is, agents in the same camp are encouraged to perform actions that can maximize their own value and team value, so as to make the joint actions of the team better and achieve a higher level of agent simulation.

[0157] Optionally, based on the above Figure 3 corresponding respective embodiments, in another optional embodiment provided by the embodiment of the present application, obtaining the first global information and the first unit information corresponding to the first moment may include:

[0158] Obtaining the first global information corresponding to the first moment from the picture at the first moment, where the first global information includes at least one of the global map information and the camp confrontation information corresponding to the first moment;

[0159] Obtaining the first unit information corresponding to the first moment from the picture at the first moment, where the first unit information includes at least one of the game character information, non-player character (NPC) information, and defense tower information corresponding to the first moment.

[0160] In one or more embodiments, a method for obtaining global information and unit information is introduced. As can be seen from the foregoing embodiments, for each picture corresponding to a moment, the global information and unit information therein can be extracted.

[0161] Specifically, effective game screens can be extracted from game stream data, and such screens include global information and unit information. Taking the game screen corresponding to a certain moment as an example for illustration below, please refer to Figure 6 , Figure 6 FIG. is a schematic diagram of global information and unit information in an embodiment of the present application. As shown in the figure, A1 is used to indicate global map information, that is, the "global mini-map". A2 is used to indicate faction confrontation information, that is, relevant information on faction confrontation in the current situation. The faction confrontation information includes but is not limited to information such as kill difference, assist difference, economy difference, and defense tower quantity difference. The global map information and the faction confrontation information belong to global information. A3 is used to indicate game character information of one faction. The game character information includes but is not limited to information such as the blood volume, blue volume, experience, skill status, and coordinate information of the game character. A4 is used to indicate game character information of another faction. A5 is used to indicate non-player character (NPC) information. The NPC information includes but is not limited to information such as level, blood volume, and coordinate information.

[0162] It should be noted that game characters, NPCs (such as minions and jungle monsters), and defense towers all belong to interactive units. The global information extracted at the first moment is the first global information, and the unit information extracted at the first moment is the first unit information, and so on, which will not be elaborated here. The unit information and global information involved in the present application are only for illustration. For example, the unit information can also include the relative distance between agents and other agents.

[0163] Secondly, in an embodiment of the present application, a method for obtaining global information and unit information is provided. Through the above method, more comprehensive features can be extracted from the game screen content, which is beneficial for the model to learn better, thereby improving the accuracy of model training.

[0164] Optionally, on the basis of the above Figure 3 corresponding various embodiments, in another optional embodiment provided by the embodiment of the present application, based on the first global information and the first unit information, obtaining the action distribution of each of the K agents at the first moment through the behavior prediction network may include:

[0165] Based on the first global information and the first unit information, obtaining the first action distribution of each of the K agents at the first moment through the first behavior prediction network;

[0166] Based on the first global information and the first unit information, obtaining the second action distribution of each of the K agents at the first moment through the second behavior prediction network, where the second behavior prediction network is generated after the first behavior prediction network.

[0167] In one or more embodiments, a training method based on two behavior prediction networks, an old one and a new one, is introduced. As can be seen from the foregoing embodiments, the behavior prediction network may include a first behavior prediction network (i.e., the old model behavior prediction network) and a second behavior prediction network (i.e., the new model behavior prediction network). Based on this, the first global information and the first unit information are input into the first behavior prediction network, and the first action distribution of each of the K agents belonging to the same camp at the first moment is output through the first behavior prediction network. The first global information and the first unit information are input into the second behavior prediction network, and the second action distribution of each of the K agents belonging to the same camp at the first moment is output through the second behavior prediction network.

[0168] Specifically, network models with different training times (e.g., the network model trained 30 minutes ago) can be randomly selected from the historical model pool as the old model, and training data is generated through self-play battles using the old model. The old model includes a first behavior prediction network (i.e., the old model behavior prediction network) and a first value network (i.e., the old model value network). The new model is the network model to be trained. After the new model is trained for a period of time, it is placed in the historical model pool for subsequent use as the old model, that is, the second behavior prediction network is generated after the first behavior prediction network.

[0169] It can be understood that after the new model is placed in the historical model pool, training data can be sampled. The policy distribution of these training data is closest to the currently updated model, which can provide more accurate samples and thus help accelerate training.

[0170] Secondly, in the embodiments of the present application, a training method based on two behavior prediction networks, an old one and a new one, is provided. Through the above method, the behavior prediction network in the old model can be used to generate self-play samples. Therefore, it is not necessary to synchronize the network model updated at each moment to each processing device, thereby improving the model training efficiency.

[0171] Optionally, on the basis of the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, based on the first global information and the first unit information, obtaining the first action distribution of each of the K agents at the first moment through the first behavior prediction network specifically includes:

[0172] Based on the first global information, the first global feature is extracted through the global feature network included in the first behavior prediction network;

[0173] Based on the first unit information, the first unit feature is extracted through the local feature network included in the first behavior prediction network;

[0174] Based on the first global feature and the first unit feature, obtain the first action distribution of each of the K agents at the first moment through the fully connected layer included in the first behavior prediction network;

[0175] Based on the first global information and the first unit information, obtain the second action distribution of each of the K agents at the first moment through the second behavior prediction network, specifically including:

[0176] Based on the first global information, extract the second global feature through the global feature network included in the second behavior prediction network;

[0177] Based on the first unit information, extract the second unit feature through the local feature network included in the second behavior prediction network;

[0178] Based on the second global feature and the second unit feature, obtain the second action distribution of each of the K agents at the first moment through the fully connected layer included in the second behavior prediction network.

[0179] In one or more embodiments, a method for outputting an action distribution based on a behavior prediction network is introduced. As can be seen from the foregoing embodiments, the global information includes but is not limited to global map information and camp confrontation information. The unit information includes but is not limited to game character information, NPC information, and defense tower information, and the NPC information includes but is not limited to minion information (i.e., lane information) and jungle monster information.

[0180] It can be understood that the processing methods of the first behavior prediction network and the second behavior prediction network are similar. Therefore, below, any one of the behavior prediction networks will be taken as an example for illustration.

[0181] Specifically, for ease of understanding, please refer to Figure 7 , Figure 7 is a schematic structural diagram of the behavior prediction network in an embodiment of the present application. As shown in the figure, obtain the global information (e.g., the first global information) and unit information (e.g., the first unit information) corresponding to each agent. Among them, the unit information includes minion information, game character information, defense tower information, and jungle monster information. Based on this, after inputting the global information into the corresponding global feature network, the global feature (e.g., the first global feature) can be obtained. After inputting the minion information into the corresponding local feature network, the minion unit feature can be obtained. After inputting the game character information into the corresponding local feature network, the game character unit feature can be obtained. After inputting the defense tower information into the corresponding local feature network, the defense tower unit feature can be obtained. After inputting the jungle monster information into the corresponding local feature network, the jungle monster unit feature can be obtained. Among them, the parameters of the local feature network are independent, and the global feature network and the local feature network include several convolutional layers.

[0182] It can be understood that the minion unit feature, the game character unit feature, the defense tower unit feature, and the wild monster unit feature all belong to the unit feature (for example, the first unit feature). It should be noted that the global information (for example, the first global information and the second global information) and the unit information (for example, the first unit information and the second unit information) of each agent are similar. In some cases, there may be certain differences in the content of the unit information. For example, the unit information includes the relative distance between the agent and other agents.

[0183] Input the global feature (for example, the first global feature) and the unit feature (for example, the first unit feature) into the fully connected layer. The global feature and the unit feature obtained after comprehensive abstraction by the fully connected layer are used to predict the action distribution of each agent at the first moment. Based on the probabilities of each action in the action distribution (for example, the first action distribution), select the action with the highest probability as the predicted action to be executed, and this predicted action is the optimal action.

[0184] Again, in the embodiments of the present application, a method for outputting an action distribution based on a behavior prediction network is provided. Through the above method, the behavior prediction network can respectively extract the features in the global information and the features in the unit information, and the diversified features are beneficial to improving the accuracy of action distribution prediction.

[0185] Optionally, based on the corresponding various embodiments above, in another optional embodiment provided by the embodiments of the present application, according to the action distribution of each agent, obtain the second global information and the second unit information at the second moment, specifically including: Figure 3 According to the first action distribution of each agent at the first moment, determine the predicted action of each agent, where the predicted action is the action corresponding to the maximum probability in the first action distribution;

[0186] Interact the predicted action of each agent with the game environment to obtain the second global information and the second unit information at the second moment.

[0187]

[0188] In one or more embodiments, a method for obtaining the second global information and the second unit information is introduced. As can be seen from the foregoing embodiments, after being processed by the behavior prediction network, the action distribution of the agent at a certain moment can be obtained. Hereinafter, the first action distribution of a certain agent at the first moment will be used as an example for illustration.

[0189] Specifically, assume that the first action distribution is an M-dimensional probability distribution vector, and each dimension corresponds to an action type. For example, the first action distribution is (0, 0.8, 0.1, 0.1), where it is assumed that the action type corresponding to the highest probability value "0.8" is "attack", then the predicted action of this agent is "attack". ​

[0190] Based on this, after obtaining the predicted actions corresponding to the K agents at the first moment respectively, the predicted action of each agent is interacted with the game environment. Reinforcement learning regards learning as a trial and evaluation process. The agent selects a predicted action for the environment. After the environment accepts the predicted action, the state changes, and the second global information and the second unit information at the second moment (i.e., the next moment) are obtained.

[0191] Again, in the embodiments of the present application, a method for obtaining the second global information and the second unit information is provided. Through the above method, based on the interaction between the predicted action and the environment, the state at the next moment can be obtained, thereby improving the feasibility and operability of the solution.

[0192] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application based on the corresponding various embodiments, according to the first global information, the first unit information, the second global information, and the second unit information, the reward information of each agent is determined, specifically including:

[0193] According to the first unit information and the second unit information, obtain the dense incentive factor of each agent, where the dense incentive factor includes at least one of the asset incentive factor, the experience incentive factor, and the blue amount incentive factor;

[0194] According to the first unit information and the second unit information, obtain the sparse incentive factor of each agent, where the sparse incentive factor includes at least one of the kill incentive factor, the death incentive factor, and the assist incentive factor;

[0195] According to the first unit information and the second unit information, obtain the survival incentive factor of each agent, where the survival incentive factor includes at least one of the self-survival incentive factor and the enemy survival incentive factor;

[0196] According to the first unit information and the second unit information, obtain the defense tower incentive factor of each agent, where the defense tower incentive factor includes at least one of the quantity incentive factor and the firmness incentive factor;

[0197] According to the first global information and the second global information, obtain the game incentive factor of each agent.

[0198] In one or more embodiments, a method for determining the reward information is introduced. As can be seen from the foregoing embodiments, by combining the change situation of the front and back two state information, the reward information of each of the K agents can be constructed. The present application includes five types of incentive factors in the reward information, but this should not be construed as a limitation on the classification of the reward information.

[0199] Specifically, the five types of incentive factors include dense incentive factors, sparse incentive factors, survival incentive factors, turret incentive factors, and game incentive factors. These incentive factors are used to describe the reward information of the agent and will be explained separately below.

[0200] I. Dense incentive factors;

[0201] Exemplarily, determine the agent's assets (e.g., in-game money), experience points, and mana (i.e., magic value) at the first moment according to the first unit information, and determine the agent's assets, experience points, and mana at the second moment according to the second unit information. Calculate the dense incentive factor based on the changes between these values. For example, if an agent has 200 assets at the first moment and 300 assets at the second moment, the asset incentive factor can be "+1.5".

[0202] Based on this, the asset incentive factor, experience incentive factor, and mana incentive factor can be summed up to obtain the dense incentive factor.

[0203] II. Sparse incentive factors;

[0204] Exemplarily, determine the agent's kill count, death count, and assist count at the first moment according to the first unit information, and determine the agent's kill count, death count, and assist count at the second moment according to the second unit information. Calculate the sparse incentive factor based on the changes between these values. For example, if an agent has 0 kills at the first moment and 2 kills at the second moment, the kill incentive factor can be "+2".

[0205] Based on this, the kill incentive factor, death incentive factor, and assist incentive factor can be summed up to obtain the sparse incentive factor.

[0206] III. Survival incentive factors;

[0207] Exemplarily, determine the agent's own health and damage to the enemy's health at the first moment according to the first unit information, and determine the agent's own health and damage to the enemy's health at the second moment according to the second unit information. Calculate the survival incentive factor based on the changes between these values. For example, if an agent has 1500 health at the first moment and 1300 health at the second moment, the self-survival incentive factor can be "-2".

[0208] Based on this, the self-survival incentive factor and the enemy-survival incentive factor can be summed up to obtain the survival incentive factor.

[0209] IV. Turret incentive factors;

[0210] Exemplarily, based on the first unit information, determine the number of defense towers and the strength of the defense towers of the agent's camp at the first moment, etc. Based on the second unit information, determine the number of defense towers and the strength of the defense towers of the agent's camp at the second moment, etc. Calculate the defense tower incentive factor based on the changes between these values. For example, if the number of defense towers at the first moment is 9 and the number of defense towers at the second moment is 8, the quantity incentive factor can be "-1".

[0211] Based on this, the quantity incentive factor and the durability incentive factor can be added together to obtain the defense tower incentive factor.

[0212] V. Game Incentive Factor;

[0213] Exemplarily, based on the first global information, determine the win-loss situation of the agent's camp at the first moment (for example, using the number of kills as the basis for win-loss, or using the remaining number of defense towers as the basis for win-loss, or using the existence of the base as the basis for win-loss). Based on the second global information, determine the win-loss situation of the agent's camp at the second moment. The game incentive factor can be calculated based on the win-loss situation. For example, if the win-loss is unknown at the first moment and the agent wins at the second moment, the game incentive factor can be "+2.5". Another example is that if the win-loss is unknown at the first moment and the agent loses in the game at the second moment, the game incentive factor can be "-2.5".

[0214] Secondly, in the embodiments of the present application, a method for determining reward information is provided. Through the above method, multiple types of incentive factors are determined according to a preset reward strategy. These incentive factors can describe the reward situation from different perspectives, improving the rationality and diversity of the reward information.

[0215] Optionally, based on the above Figure 3 On the basis of the corresponding respective embodiments, in another optional embodiment provided by the embodiments of the present application, based on the first global information, the first unit information, the second global information, and the second unit information, the target self-value and the target team value of each agent are obtained through a value network, specifically including:

[0216] Based on the first global information and the first unit information, obtain the first self-value of each agent at the first moment through the first self-value network;

[0217] Based on the second global information and the second unit information, obtain the second self-value of each agent at the second moment through the first self-value network;

[0218] Based on the first global information and the first unit information, obtain the third self-value of each agent at the first moment through the second self-value network, where the second self-value network is generated after the first self-value network;

[0219] Based on the first global information and the first unit information, obtain the first team value of each agent at the first moment through the first team value network;

[0220] Based on the second global information and the second unit information, obtain the second team value of each agent at the second moment through the first team value network;

[0221] Based on the first global information and the first unit information, obtain the third team value of each agent at the first moment through the second team value network, where the second team value network is generated after the first team value network.

[0222] In one or more embodiments, a method for obtaining its own value and team value based on a value network is introduced. As can be seen from the foregoing embodiments, when the behavior prediction network includes a first behavior prediction network and a second behavior prediction network, the value network includes a first value network (i.e., the old model value network) and a second value network (i.e., the new model value network). When evaluating the future benefits that an agent can obtain, it is necessary to consider both the target's own value and the target's team value. Therefore, the first value network (i.e., the old model value network) specifically includes a first own value network (i.e., the own value network in the old model) and a first team value network (i.e., the team value network in the old model), and the second value network (i.e., the new model value network) specifically includes a second own value network (i.e., the own value network in the new model) and a second team value network (i.e., the team value network in the new model).

[0223] Exemplarily, input the first global information and the first unit information corresponding to the agent into the first own value network, and output the first own value of the agent at the first moment through the first own value network. Similarly, input the first global information and the first unit information corresponding to the agent into the first team value network, and output the first team value of the agent at the first moment through the first team value network.

[0224] Exemplarily, input the second global information and the second unit information corresponding to the agent into the first own value network, and output the second own value of the agent at the second moment through the first own value network. Similarly, input the second global information and the second unit information corresponding to the agent into the second team value network, and output the second team value of the agent at the second moment through the first team value network.

[0225] Exemplarily, the first global information and the first unit information corresponding to the agent are input into the second self-value network, and the third self-value of the agent at the first moment is output through the second self-value network. Similarly, the first global information and the first unit information corresponding to the agent are input into the second team value network, and the third team value of the agent at the first moment is output through the second team value network.

[0226] Secondly, in the embodiments of the present application, a method for obtaining the self-value and the team value based on the value network is provided. Through the above method, two parts of self-value evaluation and team value evaluation are introduced. In the process of self-value evaluation, the benefits that each agent can obtain in the future are evaluated according to the state at the current moment. In the process of team value evaluation, the weight of each agent to the team value is dynamically measured, and the total value of the team joint action is evaluated. This is conducive to optimizing the behavior strategy of each agent and achieving the effect of team linkage optimization.

[0227] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application on the basis of the corresponding various embodiments, based on the first global information and the first unit information, the first self-value of each agent at the first moment is obtained through the first self-value network, which specifically includes:

[0228] Based on the first global information and the first unit information, the dense value factor of each agent at the first moment is obtained through the dense evaluation network included in the first self-value network;

[0229] Based on the first global information and the first unit information, the sparse value factor of each agent at the first moment is obtained through the sparse evaluation network included in the first self-value network;

[0230] Based on the first global information and the first unit information, the survival value factor of each agent at the first moment is obtained through the survival evaluation network included in the first self-value network;

[0231] Based on the first global information and the first unit information, the defense tower value factor of each agent at the first moment is obtained through the defense tower evaluation network included in the first self-value network;

[0232] Based on the first global information and the first unit information, the game value factor of each agent at the first moment is obtained through the game evaluation network included in the first self-value network.

[0233] In one or more embodiments, a method for evaluating its own value based on its own value network is introduced. As can be seen from the foregoing embodiments, during the reinforcement learning training process, the predicted actions determined by the action prediction network need to construct an action value evaluation system to measure the possible future rewards obtained by each agent when performing the predicted actions, and this reward can indicate the quality of the predicted actions. Exemplarily, for a game with simple operations, the learning of the agent can be guided only by the win-lose incentives of the game. Exemplarily, for a game with complex operations, due to the diversity of game characters, equipment, and tactics, and the overly long game duration, it is difficult to influence the behavior of the agent only by the game win-lose, so it is necessary to design multiple types of value factors to guide the learning of the agent. To accelerate the learning of the model, different types of value factors are classified, and the following will be introduced in conjunction with the diagrams.

[0234] It can be understood that the processing methods of the first self-value network and the second self-value network are similar. Therefore, the following will take any one of the self-value networks as an example for illustration.

[0235] Specifically, for ease of understanding, please refer to Figure 8 , Figure 8 is a schematic structural diagram of the self-value network in the embodiments of the present application. As shown in the figure, the first self-value network includes five hierarchical networks, namely a dense evaluation network, a sparse evaluation network, a survival evaluation network, a turret evaluation network, and a game evaluation network. After obtaining the global information (e.g., the first global information) and unit information (e.g., the first unit information) corresponding to the agent, the features of the global information and unit information are extracted, and then these features are used as the inputs of each hierarchical network. Thus, through the dense evaluation network, the dense value factor of the agent at the first moment is output. Through the sparse evaluation network, the sparse value factor of the agent at the first moment is output. Through the survival evaluation network, the survival value factor of the agent at the first moment is output. Through the turret evaluation network, the game value factor of the agent at the first moment is output. Through the game evaluation network, the game value factor of the agent at the first moment is output. Among them, the parameters of each hierarchical network in the first self-value network are independent, and each hierarchical network includes several fully connected layers.

[0236] It can be understood that the dense value factor at the first moment, the sparse value factor at the first moment, the survival value factor at the first moment, the turret value factor at the first moment, and the game value factor at the first moment all belong to the first self-value at the first moment. It should be noted that the value factors involved in the present application include but are not limited to this.

[0237] Again, in the embodiments of the present application, a method for evaluating its own value based on its own value network is provided. Through the above method, the self-value of the agent is evaluated in multiple aspects according to a preset reward strategy, thereby improving the rationality and diversity of self-value evaluation.

[0238] Optionally, based on the corresponding respective embodiments above, in another optional embodiment provided by the embodiments of the present application, based on the first global information and the first unit information, the first team value of each agent at the first moment is obtained through the first team value network, specifically including: Figure 3 Obtain the role value vector of each of the K agents, where the role value vector includes L feature elements, and L is an integer greater than or equal to 1;

[0239] Recombine the role value vectors of each of the K agents to obtain L hierarchical value vectors, where each hierarchical value vector includes the feature elements of each agent;

[0240] Based on the L hierarchical value vectors, the first global information, and the first unit information, obtain L tap value vectors through the value attention network included in the first team value network, where each tap value vector includes T tap values, and T is an integer greater than or equal to 1;

[0241] Based on the first global information, the first unit information, and the L tap value vectors, obtain the first team value of each agent at the first moment through the value fusion network included in the first team value network.

[0242]

[0243] In one or more embodiments, a method for evaluating the team value based on the team value network is introduced. As can be seen from the foregoing embodiments, the self-value network (i.e., the first self-value network and the second self-value network) can independently evaluate the value of each agent without considering the value of each agent to the team. This will cause auxiliary-type agents to preempt the economic resources of other agents and abandon output-type agents and run away alone during team battles, lacking the spirit of team dedication. Therefore, on the basis of self-value evaluation, team value evaluation is further introduced, and the value of the entire team at a certain moment is estimated according to the action value of each game character at that moment. Similar to self-value evaluation, team value evaluation also adopts the same hierarchical value system, and each hierarchical value represents the team value brought by the behaviors currently adopted by the K agents. When evaluating the team value, it consists of two networks, one is the value attention network, and the other is the value fusion network.

[0244] It is understandable that the processing methods of the first team value network and the second team value network are similar. Therefore, any one of the team value networks will be taken as an example for illustration below.

[0245] Specifically, obtain the role value vectors of each of the K agents. Among them, the role value vector includes L feature elements, and each feature element corresponds to one level. It should be noted that this application takes L as 5 and K as 5 as examples for introduction. Among them, these five levels include the dense level, the sparse level, the survival level, the turret level, and the game level. For ease of understanding, please refer to Table 1, which is a schematic diagram of the role value vectors corresponding to 5 agents.

[0246] Table 1

[0247] Game character Character value vector Agent 1 [a1, b1, c1, d1, e1] Agent 2 [a2, b2, c2, d2, e2] Agent 3 [a3, b3, c3, d3, e3] Agent 4 [a4, b4, c4, d4, e4] Agent 5 [a5, b5, c5, d5, e5]

[0248] Among them, the feature elements in the role value vector are represented as specific numerical values. Here, "a", "b", "c", "d", and "e" are only used as examples for illustration.

[0249] Based on this, the role value vectors of each of the K agents can be recombined to obtain L level value vectors. Each level value vector includes the feature elements of each agent in the same camp, that is, the level value vector is composed of K feature elements. For ease of understanding, please refer to Table 2, which is a schematic diagram of the level value vectors corresponding to 5 levels respectively.

[0250] Table 2

[0251] Hierarchy type Hierarchy value vector Dense hierarchy [a1, a2, a3, a4, a5] Sparse hierarchy [b1, b2, b3, b4, b5] Survival hierarchy [c1, c2, c3, c4, c5] Defense tower hierarchy [d1, d2, c4, d4, e4] Match hierarchy [e1, e2, e3, e4, e5]

[0252] Based on this, take the L level value vectors (for example, 5 level value vectors), the global information (for example, the first global information), and the unit information (for example, the first unit information) as the inputs of the value attention network, and output L tap value vectors through the value attention network, that is, each level corresponds to a group of tap value vectors. Among them, each tap value vector includes T tap values. For example, T can be set to 4, that is, 4 taps are set, and each tap outputs a tap value. Take the first global information, the first unit information, and the L tap value vectors as the inputs of the value fusion network, and output the first team value of each agent at the first moment through the value fusion network.

[0253] It should be noted that the team values of each agent in the same camp are the same. For example, the first team values of each agent at the first moment are the same.

[0254] Again, in the embodiments of the present application, a method for evaluating team value based on a team value network is provided. Through the above method, on the basis of evaluating its own value, it is also necessary to further evaluate the team value, so that multiple agents can simultaneously consider the maximization of their own value and the team value, and then select a better prediction action.

[0255] Optionally, on the basis of the corresponding respective embodiments above, in another optional embodiment provided by the embodiments of the present application, based on L hierarchical value vectors and T sets of tap coefficients, L tap value vectors are obtained through the value attention network included in the first team value network, specifically including: Figure 3 Based on the first global information, the first unit information, and the dense layer value vector in the L hierarchical value vectors, the first tap value vector is obtained through the dense value attention network included in the value attention network, where the value attention network is included in the first team value network;

[0256] Based on the first global information, the first unit information, and the sparse layer value vector in the L hierarchical value vectors, the second tap value vector is obtained through the sparse value attention network included in the value attention network;

[0257] Based on the first global information, the first unit information, and the survival layer value vector in the L hierarchical value vectors, the third tap value vector is obtained through the survival value attention network included in the value attention network;

[0258] Based on the first global information, the first unit information, and the turret layer value vector in the L hierarchical value vectors, the fourth tap value vector is obtained through the turret value attention network included in the value attention network;

[0259] Based on the first global information, the first unit information, and the game layer value vector in the L hierarchical value vectors, the fifth tap value vector is obtained through the game value attention network included in the value attention network, where the first tap value vector, the second tap value vector, the third tap value vector, the fourth tap value vector, and the fifth tap value vector are included in the L tap value vectors.

[0260] In one or more embodiments, a method for predicting tap value vectors based on a value attention network is introduced. As can be seen from the foregoing embodiments, the value attention network can extract the hierarchical value vectors corresponding to L levels. The process of outputting the hierarchical value vectors will be described below with reference to the drawings.

[0261] It can be understood that the processing methods of the first team value network and the second team value network are similar. Therefore, below, any one of the team value networks will be taken as an example for illustration.

[0262] It can be understood that the processing methods of the first team value network and the second team value network are similar. Therefore, below, any one of the team value networks will be taken as an example for illustration.

[0263] Specifically, for the sake of easy understanding, please refer to Figure 9 , Figure 9 which is a schematic structural diagram of the value attention network in the embodiment of the present application. As shown in the figure, corresponding features are extracted from the global information and unit information of each agent (for example, the first global information and the first unit information). For example, features of 4096 dimensions are obtained. In order to avoid excessive computational complexity caused by too high feature dimensions, therefore, the features corresponding to each agent can be respectively input into the corresponding fully connected layer, thereby realizing dimensionality reduction of the features. For example, features of 64 dimensions are obtained after dimensionality reduction. Based on this, these dimensionality-reduced features are concatenated to obtain target features. Then, the target features and the corresponding hierarchical value vectors in the L hierarchical value vectors are input into the attention network of the corresponding level, and the corresponding tap value vectors are output through the attention network of this level.

[0264] It should be noted that the structure of the value attention network can have different ways. For example, without reducing the input feature dimension, that is, without using a fully connected layer, etc. Figure 9 The structure shown is only for illustration and should not be construed as a limitation of the present application.

[0265] Exemplarily, the process of outputting the first tap value vector based on the dense value attention network will be introduced below in combination with Figure 10 . It can be understood that the processing methods of other types of attention networks (such as sparse value attention network, survival value attention network, turret value attention network, and game value attention network) are not described in detail here for the process of generating the second tap value vector, the third tap value vector, the fourth tap value vector, and the fifth tap value vector.

[0266] Please refer to Figure 10 , Figure 10 which is a schematic structural diagram of the dense value attention network in the embodiment of the present application. As shown in the figure, assume that the dense value attention network includes 4 taps, that is, T is equal to 4. Thus, the calculation of multi-tap self-attention can be performed to assign weights to the values of different agents at the same level, that is, to adaptively measure the total value of the current operations of different agents to the team. The higher the attention weight, the greater the importance. Based on this, the target features obtained after combining the features are respectively input into different fully connected layers, and keys and queries can be obtained, that is, key = FC1(target feature), query = FC2(target feature), where FC1 and FC2 are two different fully connected layers. Finally, based on the key and the query, the attention weights corresponding to each tap can be calculated using the self-attention algorithm. The weighted sum of the dense layer value vector and the attention weights of each tap is used to obtain the first tap value vector.

[0267] It can be understood that in practical applications, the average attention weights of T taps can also be visually displayed. For ease of understanding, please refer to Figure 11 , Figure 11 which is a schematic diagram of the weights corresponding to the agents at different levels in the embodiments of this application. As shown in the figure, the left side is the game scene, and the right side is the degree of contribution of each agent to the team value at this moment. Different rows represent the team values at different levels. At this time, assume that agent B enters the field to attack the enemy's game character. Since benefits such as game gold coins, experience, and damage values can be obtained after defeating the enemy, agent B has the greatest contribution to the team value of the four levels, that is, the highest attention weight. Taking the dense level as an example, the average attention weight of character A is 0.19. Suppose there are 4 taps. Among them, in tap 1, the attention weight of agent A is w1, in tap 2, the attention weight of agent A is p1, in tap 3, the attention weight of agent A is q1, and in tap 4, the attention weight of agent A is e1. Then, [(w1 + p1 + q1 + e1)] / 4 = 0.19 (that is, the average attention weight of character A is 0.19).

[0268] Combined with the above introduction, please refer to Figure 12 , Figure 12 which is a schematic diagram of obtaining the tap value vector based on the value attention network in the embodiments of this application. As shown in the figure, after reorganizing the K character value vectors, L hierarchical value vectors can be obtained. Then, the corresponding hierarchical value vectors are respectively used as the inputs of the corresponding hierarchical attention networks. Suppose there are 4 taps, and each tap calculates a tap value. The 4 tap values form a tap value vector corresponding to a level.

[0269] Exemplarily, the tap value v1 = w1 * a1 + w2 * a2 + w3 * a3 + w4 * a4 + w5 * a5, where w1 to w5 are the attention weights corresponding to tap 1.

[0270] Exemplarily, the tap value v2 = p1 * a1 + p2 * a2 + p3 * a3 + p4 * a4 + p5 * a5, where p1 to p5 are the attention weights corresponding to tap 2.

[0271] Exemplarily, the tap value v3 = q1 * a1 + q2 * a2 + q3 * a3 + q4 * a4 + q5 * a5, where q1 to q5 are the attention weights corresponding to tap 3.

[0272] Exemplarily, the tap value v4 = e1 * a1 + e2 * a2 + e3 * a3 + e4 * a4 + e5 * a5, where e1 to e5 are the attention weights corresponding to tap 4.

[0273] Furthermore, in the embodiments of the present application, a method for predicting the tap value vector based on the value attention network is provided. Through the above method, an attention mechanism is introduced, which can adaptively evaluate the contribution degree of the behaviors of different agents to the team value and guide the agents to learn in the direction of maximizing the team value. In addition, the value attention network can accurately show the contribution degree of different agents to the team, which is convenient for data analysis and further optimization, and improves the interpretability of the model.

[0274] Optionally, on the basis of the corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, based on the first global information, the first unit information, and L tap value vectors, the first team value of each agent at the first moment is obtained through the value fusion network included in the first team value network, specifically including: Figure 3 Based on the first global information, the first unit information, and the first tap value vector among the L tap value vectors, a first fusion value is obtained through the dense value fusion network included in the value fusion network, where the value fusion network is included in the first team value network;

[0275] Based on the first global information, the first unit information, and the second tap value vector among the L tap value vectors, a second fusion value is obtained through the sparse value fusion network included in the value fusion network;

[0276] Based on the first global information, the first unit information, and the third tap value vector among the L tap value vectors, a third fusion value is obtained through the survival value fusion network included in the value fusion network;

[0277] Based on the first global information, the first unit information, and the fourth tap value vector among the L tap value vectors, a fourth fusion value is obtained through the turret value fusion network included in the value fusion network;

[0278] Based on the first global information, the first unit information, and the fifth tap value vector among the L tap value vectors, a fifth fusion value is obtained through the game value fusion network included in the value fusion network, where the first fusion value, the second fusion value, the third fusion value, the fourth fusion value, and the fifth fusion value are included in the first team value of each agent in the same camp at the first moment.

[0279] In one or more embodiments, a method for predicting the team value based on the value fusion network is introduced. As can be seen from the foregoing embodiments, the value fusion network can extract the fusion values corresponding to L levels. The process of outputting the fusion values will be described below with reference to the drawings.

[0280] In one or more embodiments, a method for predicting the team value based on the value fusion network is introduced. As can be seen from the foregoing embodiments, the value fusion network can extract the fusion values corresponding to L levels. The process of outputting the fusion values will be described below with reference to the drawings.

[0281] It can be understood that the processing methods of the first team value network and the second team value network are similar. Therefore, below, any one of the team value networks will be taken as an example for illustration.

[0282] Specifically, for the convenience of understanding, please refer to Figure 13 , Figure 13 which is a schematic structural diagram of the value fusion network in the embodiment of the present application. As shown in the figure, corresponding features are extracted from the global information and unit information of each agent (for example, the first global information and the first unit information). For example, features of 4096 dimensions are obtained. In order to avoid excessive computational complexity caused by too high feature dimensions, therefore, the features corresponding to each agent can be respectively input into the corresponding fully connected layer, thereby realizing dimensionality reduction of the features. For example, features of 64 dimensions are obtained after dimensionality reduction. Based on this, these dimension-reduced features are concatenated to obtain target features. Then, the target features and the tap value vectors of the corresponding levels in the L tap value vectors are input into the fusion network of the corresponding level, and the corresponding fusion value is output through the fusion network of this level.

[0283] It should be noted that the structure of the value fusion network can have different forms. For example, without reducing the input feature dimension, that is, without using a fully connected layer, etc. Figure 13 The structure shown is only for illustration and should not be construed as a limitation of the present application.

[0284] Exemplarily, below, in combination with Figure 14 , the process of outputting the first fusion value based on the dense value fusion network will be introduced. It can be understood that the processing method types of other types of fusion networks (such as sparse value fusion network, survival value fusion network, turret value fusion network, and game value fusion network) are not described in detail here for the processes of generating the second fusion value, the third fusion value, the fourth fusion value, and the fifth fusion value.

[0285] Please refer to Figure 14 , Figure 14 which is a schematic structural diagram of the dense value fusion network in the embodiment of the present application. As shown in the figure, the target features obtained by concatenation can be used to generate the weights and bias values of the monotonically increasing fully connected layer. Among them, monotonic increase can ensure optimizing the team value while improving the capabilities of each agent. Acting the weights and bias values on the first tap value vector output by the value attention network can obtain the integrated first fusion value. It can be understood that, in order to enhance the non-linear connection information between values, Figure 14 the result shown uses an increasing fully connected layer with two layers superimposed.

[0286] Combined with the above introduction, please refer to Figure 15 , Figure 15This is a schematic diagram for obtaining the first team value based on the value fusion network in the embodiments of the present application. As shown in the figure, assume that the first tap value vector v is [v1, v2, v3, v4], and the matrix dimension of v can be expressed as [1, 4]. The weight of the increasing fully connected layer 1 is w1, and the matrix dimension of w1 can be expressed as [4, 32]. The bias of the increasing fully connected layer 1 is b1, and the matrix dimension of b1 can be expressed as [1, 32]. The weight of the increasing fully connected layer 2 is w2, and the matrix dimension of w2 can be expressed as [32, 1]. The bias of the increasing fully connected layer 2 is b2, and the matrix dimension of b2 can be expressed as [1, 1]. The matrix dimension of the intermediate vector h can be expressed as [1, 32]. The matrix dimension of the first fusion value V can be expressed as [1, 1].

[0287] The calculation method of the intermediate vector is as follows:

[0288] h[1, 32] = v[1, 4] * w1[4, 32] + b1[1, 32];

[0289] The calculation method of the first fusion value is as follows:

[0290] V[1, 1] = h[1, 32] * w2[32, 1] + b2[1, 1];

[0291] It should be noted that the calculation methods of other fusion values are similar and will not be elaborated here.

[0292] Furthermore, in the embodiments of the present application, a method for predicting the team value based on the value fusion network is provided. Through the above method, on the basis of introducing the attention mechanism, the values at each level are further fused, thereby improving the feasibility and operability of the solution.

[0293] Optionally, on the basis of the above Figure 3 In another optional embodiment provided by the embodiments of the present application based on the corresponding respective embodiments, according to the action distribution, reward information, target self-value, and target team value of each agent, the behavior prediction network and the value network are trained, specifically including:

[0294] Determine the target loss value according to the target self-value, target team value, action distribution, and reward information of each agent;

[0295] Update the network parameters of the second behavior prediction network by using the target loss value, where the second behavior prediction network is included in the behavior prediction network;

[0296] Determine a first loss value based on the third self-value, reward information, and second self-value of each agent at the first moment, where the third self-value is obtained through the second self-value network, the second self-value is obtained through the first self-value network, and the second self-value network is generated after the first self-value network;

[0297] Update the network parameters of the second self-value network using the first loss value, where the second self-value network is the second value network included in the value network;

[0298] Determine a second loss value based on the third team value, reward information, and second team value of each agent at the first moment, where the third team value is obtained through the second team value network, the second team value is obtained through the first team value network, and the second team value network is generated after the first team value network;

[0299] Update the network parameters of the second team value network using the second loss value, where the second team value network is the second value network included in the value network.

[0300] In one or more embodiments, a method for updating the behavior prediction network and the value network is introduced. As can be seen from the foregoing embodiments, the behavior prediction network may include a first behavior prediction network (i.e., the old model behavior prediction network) and a second behavior prediction network (i.e., the new model behavior prediction network), and the value network includes a first value network (i.e., the old model value network) and a second value network (i.e., the new model value network). The first value network includes a first self-value network and a first team value network, and the second value network includes a second self-value network and a second team value network. The target self-value includes the first self-value at the first moment and the second self-value at the second moment obtained through the first self-value network, and the target team value includes the first team value at the first moment and the second team value at the second moment obtained through the first team value network.

[0301] Specifically, please refer to Figure 16 , Figure 16 which is another schematic diagram of the decision model training process in the embodiments of the present application. As shown in the figure, according to the target value Vt (i.e., the first self-value and the first team value) of each agent, the target value V t+1 (i.e., the second self-value and the second team value) and the reward information r t , calculate the generalized advantage estimator (GAE). Then, combine the action distribution p output by the first behavior prediction network t(i.e., the first action distribution) and the action distribution p output by the second behavior prediction network k (i.e., the second action distribution), calculate the PPO result as the target loss value. Based on this, the stochastic gradient descent method can be used to update the network parameters of the second behavior prediction network using the target loss value.

[0302] Please refer to Figure 17 , Figure 17 , which is another schematic diagram of the value model training process in the embodiments of the present application. As shown in the figure, the reward information r of each agent t is summed with the target value V t+1 (i.e., the second self-value), and then the summation result is fitted with the target value V of each agent tk (i.e., the third self-value), and the MSE result is calculated as the first loss value. Based on this, the stochastic gradient descent method can be used to update the network parameters of the second self-value network using the first loss value. Similarly, the reward information r of each agent t is summed with the target value V t+1 (i.e., the second team value), and then the summation result is fitted with the target value V of each agent tk (i.e., the third team value), and the MSE result is calculated as the second loss value. Based on this, the stochastic gradient descent method can be used to update the network parameters of the second team value network using the second loss value.

[0303] In practical applications, the model can be pitted against the starting model every two hours, and the change in the winning rate is recorded. If the winning rate remains unchanged at a certain value, it is considered that the model's ability has converged and the training is stopped. Otherwise, the training of the model continues.

[0304] Secondly, in the embodiments of the present application, a method for updating the behavior prediction network and the value network is provided. Through the above method, the network parameters of the new model can be updated according to the output results of the new model and the old model, which not only improves the training efficiency but also saves the training amount, and only the second behavior prediction network and the second value network in the new model need to be updated.

[0305] Optionally, based on the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, the target self-value includes the first self-value at the first moment and the second self-value at the second moment obtained through the first self-value network, and the target team value includes the first team value at the first moment and the second team value at the second moment obtained through the first team value network;

[0306] Determine the target loss value according to the target self-value, target team value, action distribution, and reward information of each agent, specifically including:

[0307] Calculate the self-advantage estimate value according to the first self-value of each agent at the first moment, the second self-value at the second moment, and the reward information;

[0308] Calculate the team-advantage estimate value according to the first team value of each agent at the first moment, the second team value at the second moment, and the reward information;

[0309] Determine the self-loss value according to the action distribution of each agent and the self-advantage estimate value;

[0310] Determine the team-loss value according to the action distribution of each agent and the team-advantage estimate value;

[0311] Sum up the self-loss value and the team-loss value to calculate the target loss value.

[0312] In one or more embodiments, a method for obtaining the target loss value is introduced. As can be seen from the foregoing embodiments, the PPO result can be used as the target loss value. Since this application involves two parts, self-value and team value, when calculating the target loss value, it is necessary to calculate the PPO result for self-value and the PPO result for team value respectively.

[0313] Specifically, when calculating the target loss value, it is necessary to calculate the advantages related to self-value and team value respectively. The following formula can be used to calculate the self-advantage estimate value:

[0314]

[0315] Among them, represents the self-advantage estimate value. γ represents the first deviation. λ represents the second deviation. l represents the l-th moment (or, the l-th frame), usually, l can be set to 960. t represents the t-th moment in the l-th moment (i.e., the first moment). r t represents the reward information. V 自身 (s t+l ) represents the first self-value at the first moment. V 自身 (s t+l+1 ) represents the second self-value at the second moment.

[0316] The following formula can be used to calculate the team-advantage estimate value:

[0317]

[0318] Among them, Represents the estimated value of the team's advantage. γ represents the first deviation. λ represents the second deviation. l represents the l-th moment (or, the l-th frame). Usually, l can be set to 960. t represents the t-th moment in the l-th moment. r t Represents the reward information. V 团队 (s t+l ) represents the first team value at the first moment. V 团队 (s t+l+1 ) represents the second team value at the second moment.

[0319] Based on this, the following formula can be used to calculate the self-loss value:

[0320]

[0321] Among them, Represents the self-loss value. θ represents the network parameter. p θ (a t |s t ) represents the second action distribution at the t-th moment (i.e., the first moment) output by the second decision model. a t Represents the predicted action at the t-th moment (i.e., the first moment). s t Represents the global information and unit information at the t-th moment (i.e., the first global information and the first unit information). clip(·) represents the truncation calculation. Represents the first action distribution at the t-th moment (i.e., the first moment) output by the first decision model. ε represents a preset parameter, for example, 0.2. Represents the estimated value of the self-advantage.

[0322] Based on this, the following formula can be used to calculate the team loss value:

[0323]

[0324] Among them, Represents the team loss value. Represents the estimated value of the team's advantage.

[0325] Finally, the following formula is used to calculate the target loss value:

[0326]

[0327] Among them, J represents the target loss value.

[0328] Again, in the embodiments of the present application, a method for obtaining the target loss value is provided. Through the above method, the estimated value of the advantage is calculated based on the self-value and the team value respectively, and the corresponding loss values are calculated based on the PPO algorithm respectively and used to update the network, thereby improving the feasibility of the solution.

[0329] Combined with the above introduction, the training method of the neural network model in this application will be introduced below. Please refer to Figure 18 Another embodiment of the neural network model training method in the embodiments of this application includes:

[0330] 210. Invoke the behavior prediction network to conduct self-play battles to obtain training data;

[0331] In one or more embodiments, randomly select a model for a training time from the model pool. This model includes a behavior prediction network and a value network, and use this model to conduct self-play battles to generate training data.

[0332] 220. Obtain the global information and unit information at different times from the training data. Among them, the global information is used to describe the game situation of the camp, and the unit information is used to describe the real-time state of the interactive unit. The interactive unit represents an object that has an interactive relationship with the game character in the game;

[0333] In one or more embodiments, extract the global information and unit information required for training from the training data. The global information is used to describe the game situation of the camp, and the unit information is used to describe the real-time state of the interactive unit. The interactive unit represents an object that has an interactive relationship with the game character in the game.

[0334] 230. Determine the reward information of each agent in the same camp according to the global information and unit information at different times. Among them, the reward information is the incentive obtained after triggering the action, and the agent is a game character controlled by artificial intelligence;

[0335] In one or more embodiments, based on the global information and unit information at the current moment, the action distribution at the next moment can be predicted, and then the action of the agent at the next moment can be selected. Therefore, compare the global information and unit information at two adjacent times to determine the reward information of each agent in the same camp.

[0336] 240. Based on the global information and unit information, obtain the target self-value and target team value of each agent through the value network. Among them, the target team value represents the value of the agent in the camp;

[0337] In one or more embodiments, input the global information and unit information into the value network, and the value network outputs the target self-value and target team value of each agent.

[0338] 250. Train the behavior prediction network and the value network according to the reward information, target self-value and target team value of each agent.

[0339] In one or more embodiments, the neural network (i.e., the behavior prediction network and the value network) is trained by the PPO algorithm using the reward information of each agent, the target's own value, and the target team value, etc., to update the parameters of the model. After the model undergoes a certain number of iterative updates, the latest model is placed in the model pool to provide training data for the latest model and improve the speed of model update.

[0340] Finally, evaluate the ability of the model. For example, have the model battle against the starting model every two hours and record the change in the winning rate. If the winning rate remains unchanged at a certain value, it is considered that the model's ability has converged and training stops. Otherwise, continue training the model.

[0341] The model training device in the present application will be described in detail below. Please refer to Figure 19 , Figure 19 which is a schematic diagram of an embodiment of the model training device in the embodiment of the present application. The model training device 30 includes:

[0342] An acquisition module 310, configured to acquire the first global information and the first unit information corresponding to the first moment, where the first global information is used to describe the game situation of the camp at the first moment, and the first unit information is used to describe the state of the interaction unit at the first moment. The interaction unit represents an object having an interaction relationship with the game character in the game;

[0343] The acquisition module 310 is further configured to, based on the first global information and the first unit information, obtain the action distribution of each of the K agents at the first moment through the behavior prediction network, where the K agents belong to the same camp, the agent is a game character controlled by artificial intelligence, and K is an integer greater than 1;

[0344] The acquisition module 310 is further configured to, according to the action distribution of each agent, obtain the second global information and the second unit information at the second moment, where the second moment is the next moment of the first moment, the second global information is used to describe the game situation of the camp at the second moment, and the second unit information is used to describe the state of the interaction unit at the second moment;

[0345] A determination module 320, configured to determine the reward information of each agent according to the first global information, the first unit information, the second global information, and the second unit information, where the reward information is the incentive obtained after triggering the action;

[0346] The acquisition module 310 is further configured to, based on the first global information, the first unit information, the second global information, and the second unit information, obtain the target own value and the target team value of each agent through the value network, where the target team value represents the value of the agent in the camp;

[0347] A training module 330, configured to train a behavior prediction network and a value network according to the action distribution, reward information, target self-value, and target team value of each agent.

[0348] In an embodiment of the present application, a model training device is provided. By using the above device, in an environment where multiple agents (game characters controlled by an AI model) cooperate, a network model used for reinforcement learning is trained, and value collaborative learning of multiple agents is realized based on the constructed value network, that is, agents in the same camp are encouraged to execute actions that can maximize their own value and team value, so that the joint actions of the team are better and a higher level of agent simulation is achieved.

[0349] Optionally, based on the corresponding embodiment above, Figure 19 in another embodiment of the model training device 30 provided in the embodiment of the present application,

[0350] An acquisition module 310 is specifically configured to acquire first global information corresponding to the first moment from the screen of the first moment, where the first global information includes at least one of global map information and camp confrontation information corresponding to the first moment;

[0351] Acquire first unit information corresponding to the first moment from the screen of the first moment, where the first unit information includes at least one of game character information, non-player character (NPC) information, and defense tower information corresponding to the first moment.

[0352] In an embodiment of the present application, a model training device is provided. By using the above device, more comprehensive features can be extracted from the game screen content, which is beneficial for the model to learn better, thereby improving the accuracy of model training.

[0353] Optionally, based on the corresponding embodiment above, Figure 19 in another embodiment of the model training device 30 provided in the embodiment of the present application,

[0354] The acquisition module 310 is specifically configured to, based on the first global information and the first unit information, obtain the first action distribution of each of the K agents at the first moment through a first behavior prediction network;

[0355] Based on the first global information and the first unit information, obtain the second action distribution of each of the K agents at the first moment through a second behavior prediction network, where the second behavior prediction network is generated after the first behavior prediction network.

[0356] In an embodiment of the present application, a model training device is provided. With the above device, the behavior prediction network in the old model can be used to generate samples from the game, so there is no need to synchronize the network model updated at each moment to each processing device, thereby improving the model training efficiency.

[0357] Optionally, in the above Figure 19 On the basis of the corresponding embodiment, in another embodiment of the model training device 30 provided in the embodiment of the present application,

[0358] The acquisition module 310 is specifically configured to extract a first global feature through a global feature network included in the first behavior prediction network based on the first global information;

[0359] Based on the first unit information, extracting first unit features through a local feature network included in the first behavior prediction network;

[0360] Based on the first global feature and the first unit feature, obtaining a first action distribution of each of the K agents at the first moment through a fully connected layer included in the first behavior prediction network;

[0361] The acquisition module 310 is specifically configured to extract a second global feature through a global feature network included in the second behavior prediction network based on the first global information;

[0362] Based on the first unit information, extracting second unit features through a local feature network included in the second behavior prediction network;

[0363] Based on the second global feature and the second unit feature, the second action distribution of each of the K agents at the first moment is obtained through the fully connected layer included in the second behavior prediction network.

[0364] In an embodiment of the present application, a model training device is provided. By using the above device, the behavior prediction network can extract features from global information and features from unit information respectively, and the diversified features are conducive to improving the accuracy of action distribution prediction.

[0365] Optionally, in the above Figure 19 On the basis of the corresponding embodiment, in another embodiment of the model training device 30 provided in the embodiment of the present application,

[0366] An acquisition module 310 is specifically configured to determine a predicted action of each agent according to a first action distribution of each agent at a first moment, wherein the predicted action is an action corresponding to a maximum probability in the first action distribution;

[0367] The predicted actions of each agent are interacted with the game environment to obtain the second global information and the second unit information at the second moment.

[0368] In an embodiment of the present application, a model training device is provided. By using the above device, based on the interaction between the predicted action environments, the state at the next moment can be obtained, thereby improving the feasibility and operability of the solution.

[0369] Optionally, based on the corresponding embodiment above, in another embodiment of the model training device 30 provided in the embodiment of the present application, Figure 19 the determination module 320 is specifically configured to obtain the dense incentive factor of each agent according to the first unit information and the second unit information, where the dense incentive factor includes at least one of an asset incentive factor, an experience incentive factor, and a blue amount incentive factor;

[0370] Obtain the sparse incentive factor of each agent according to the first unit information and the second unit information, where the sparse incentive factor includes at least one of a kill incentive factor, a death incentive factor, and an assist incentive factor;

[0371] Obtain the survival incentive factor of each agent according to the first unit information and the second unit information, where the survival incentive factor includes at least one of a self-survival incentive factor and an enemy survival incentive factor;

[0372] Obtain the defense tower incentive factor of each agent according to the first unit information and the second unit information, where the defense tower incentive factor includes at least one of a quantity incentive factor and a firmness incentive factor;

[0373] Obtain the game incentive factor of each agent according to the first global information and the second global information.

[0374] In an embodiment of the present application, a model training device is provided. By using the above device, multiple types of incentive factors are determined according to a preset reward strategy, and these incentive factors can describe the reward situation from different perspectives, improving the rationality and diversity of the reward information.

[0375] Optionally, based on the corresponding embodiment above, in another embodiment of the model training device 30 provided in the embodiment of the present application,

[0376] the acquisition module 310 is specifically configured to obtain the first self-value of each agent at the first moment through the first self-value network based on the first global information and the first unit information; Figure 19 Obtain the second self-value of each agent at the second moment through the first self-value network based on the second global information and the second unit information;

[0377] Based on the second global information and the second unit information, obtain the second self-value of each agent at the second moment through the first self-value network;

[0378] Based on the second global information and the second unit information, obtain the second self-value of each agent at the second moment through the first self-value network;

[0379] Based on the first global information and the first unit information, obtain the third self-value of each agent at the first moment through the second self-value network, where the second self-value network is generated after the first self-value network;

[0380] Based on the first global information and the first unit information, obtain the first team value of each agent at the first moment through the first team value network;

[0381] Based on the second global information and the second unit information, obtain the second team value of each agent at the second moment through the first team value network;

[0382] Based on the first global information and the first unit information, obtain the third team value of each agent at the first moment through the second team value network, where the second team value network is generated after the first team value network.

[0383] In the embodiments of the present application, a model training device is provided. By using the above device, two parts of self-value evaluation and team-value evaluation are introduced. In the process of self-value evaluation, the benefits that each agent can obtain in the future are evaluated according to the state at the current moment. In the process of team-value evaluation, the weight of each agent on the team value is dynamically measured, and the total value of the team joint action is evaluated. This is beneficial to optimizing the behavior strategy of each agent and achieving the effect of team linkage optimization.

[0384] Optionally, on the basis of the corresponding embodiments above Figure 19 In another embodiment of the model training device 30 provided in the embodiments of the present application,

[0385] The acquisition module 310 is specifically configured to obtain the dense value factor of each agent at the first moment through the dense evaluation network included in the first self-value network based on the first global information and the first unit information;

[0386] Based on the first global information and the first unit information, obtain the sparse value factor of each agent at the first moment through the sparse evaluation network included in the first self-value network;

[0387] Based on the first global information and the first unit information, obtain the survival value factor of each agent at the first moment through the survival evaluation network included in the first self-value network;

[0388] Based on the first global information and the first unit information, obtain the defense tower value factor of each agent at the first moment through the defense tower evaluation network included in the first self-value network;

[0389] Based on the first global information and the first unit information, obtain the game value factor of each agent at the first moment through the game evaluation network included in the first self-value network.

[0390] In the embodiments of the present application, a model training device is provided. By using the above device, the self-value of the agent is evaluated in multiple aspects according to a preset reward strategy, so as to improve the rationality and diversity of self-value evaluation.

[0391] Optionally, on the basis of the corresponding embodiment above, in another embodiment of the model training device 30 provided by the embodiments of the present application, Figure 19

[0392] The acquisition module 310 is specifically configured to acquire the role value vector of each agent among the K agents, where the role value vector includes L feature elements, and L is an integer greater than or equal to 1;

[0393] Recombine the role value vectors of each agent among the K agents to obtain L hierarchical value vectors, where each hierarchical value vector includes the feature elements of each agent;

[0394] Based on the L hierarchical value vectors, the first global information, and the first unit information, obtain L tap value vectors through the value attention network included in the first team value network, where each tap value vector includes T tap values, and T is an integer greater than or equal to 1;

[0395] Based on the first global information, the first unit information, and the L tap value vectors, obtain the first team value of each agent at the first moment through the value fusion network included in the first team value network.

[0396] In the embodiments of the present application, a model training device is provided. By using the above device, on the basis of evaluating the self-value, it is also necessary to further evaluate the team value, so that multiple agents can simultaneously consider the maximization of the self-value and the team value, and then select a better prediction action.

[0397] Optionally, on the basis of the corresponding embodiment above, in another embodiment of the model training device 30 provided by the embodiments of the present application, Figure 19

[0398] The acquisition module 310 is specifically configured to obtain the first tap value vector through the dense value attention network included in the value attention network based on the first global information, the first unit information, and the dense layer value vector among the L hierarchical value vectors, where the value attention network is included in the first team value network;

[0399] ​​Based on the first global information, the first unit information, and the sparse layer value vector among the L hierarchical value vectors, obtain the second tap value vector through the sparse value attention network included in the value attention network;

[0400] Based on the first global information, the first unit information, and the survival layer value vector among the L hierarchical value vectors, obtain the third tap value vector through the survival value attention network included in the value attention network;

[0401] Based on the first global information, the first unit information, and the defense tower layer value vector among the L hierarchical value vectors, obtain the fourth tap value vector through the defense tower value attention network included in the value attention network;

[0402] Based on the first global information, the first unit information, and the game layer value vector among the L hierarchical value vectors, obtain the fifth tap value vector through the game value attention network included in the value attention network, where the first tap value vector, the second tap value vector, the third tap value vector, the fourth tap value vector, and the fifth tap value vector are included in the L tap value vectors.

[0403] In the embodiments of the present application, a model training device is provided. By using the above device, an attention mechanism is introduced, which can adaptively evaluate the contribution degree of the behaviors of different agents to the team value, and guide the agents to learn in the direction of maximizing the team value. In addition, the value attention network can accurately show the contribution degree of different agents to the team, which is convenient for data analysis and further optimization, and improves the interpretability of the model.

[0404] Optionally, on the basis of the corresponding embodiments above, Figure 19 In another embodiment of the model training device 30 provided by the embodiments of the present application,

[0405] The acquisition module 310 is specifically configured to obtain the first fusion value through the dense value fusion network included in the value fusion network based on the first global information, the first unit information, and the first tap value vector among the L tap value vectors, where the value fusion network is included in the first team value network;

[0406] Based on the first global information, the first unit information, and the second tap value vector among the L tap value vectors, obtain the second fusion value through the sparse value fusion network included in the value fusion network;

[0407] Based on the first global information, the first unit information, and the third tap value vector among the L tap value vectors, obtain the third fusion value through the survival value fusion network included in the value fusion network;

[0408] Based on the first global information, the first unit information, and the fourth tap value vector among the L tap value vectors, obtain the fourth fusion value through the defense tower value fusion network included in the value fusion network;

[0409] Based on the first global information, the first unit information, and the fifth tap value vector among the L tap value vectors, obtain the fifth fusion value through the game value fusion network included in the value fusion network, where the first fusion value, the second fusion value, the third fusion value, the fourth fusion value, and the fifth fusion value are included in the first team value of each agent in the same camp at the first moment.

[0410] In the embodiments of the present application, a model training device is provided. By using the above device, on the basis of introducing the attention mechanism, the values at each level are further fused, thereby improving the feasibility and operability of the solution.

[0411] Optionally, on the basis of the corresponding embodiments above Figure 19 In another embodiment of the model training device 30 provided in the embodiments of the present application,

[0412] The training module 330 is specifically configured to determine the target loss value according to the target self-value, the target team value, the action distribution, and the reward information of each agent;

[0413] Update the network parameters of the second behavior prediction network by using the target loss value, where the second behavior prediction network is included in the behavior prediction network;

[0414] Determine the first loss value according to the third self-value, the reward information, and the second self-value of each agent at the second moment, where the third self-value is obtained through the second self-value network, the second self-value is obtained through the first self-value network, and the second self-value network is generated after the first self-value network;

[0415] Update the network parameters of the second self-value network by using the first loss value, where the second self-value network is included in the second value network in the value network;

[0416] Determine the second loss value according to the third team value, the reward information, and the second team value of each agent at the second moment, where the third team value is obtained through the second team value network, the second team value is obtained through the first team value network, and the second team value network is generated after the first team value network;

[0417] Update the network parameters of the second team value network by using the second loss value, where the second team value network is included in the second value network in the value network.

[0418] In the embodiments of the present application, a model training device is provided. By using the above device, the network parameters of the new model can be updated according to the output results of the new model and the old model, which not only improves the training efficiency, but also can save the training amount, and only the second behavior prediction network and the second value network in the new model need to be updated.

[0419] Optionally, based on the corresponding embodiments above, in another embodiment of the model training device 30 provided in the embodiments of the present application, the target self-value includes the first self-value at the first moment and the second self-value at the second moment obtained through the first self-value network, and the target team value includes the first team value at the first moment and the second team value at the second moment obtained through the first team value network; Figure 19

[0420] The training module 330 is specifically configured to calculate the self-advantage estimation value according to the first self-value of each agent at the first moment, the second self-value at the second moment, and the reward information;

[0421] Calculate the team advantage estimation value according to the first team value of each agent at the first moment, the second team value at the second moment, and the reward information;

[0422] Determine the self-loss value according to the action distribution of each agent and the self-advantage estimation value;

[0423] Determine the team loss value according to the action distribution of each agent and the team-advantage estimation value;

[0424] Sum up the self-loss value and the team loss value to calculate the target loss value.

[0425] In the embodiments of the present application, a model training device is provided. By using the above device, the advantage estimation value is calculated based on the self-value and the team value respectively, and the corresponding loss value is calculated based on the PPO algorithm respectively and used to update the network, thereby improving the feasibility of the solution.

[0426] Figure 20 ​FIG. 0 is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 400 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 422 (for example, one or more processors) and a memory 432, and one or more storage media 430 (for example, one or more mass storage devices) storing application programs 442 or data 444. Among them, the memory 432 and the storage media 430 may be transient storage or persistent storage. The programs stored in the storage media 430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the computer device. Further, the central processing unit 422 may be configured to communicate with the storage media 430 and execute a series of instruction operations in the storage media 430 on the computer device 400.

[0427] The computer device 400 may further include one or more power supplies 426, one or more wired or wireless network interfaces 450, one or more input / output interfaces 458, and / or one or more operating systems 441, such as Windows Server TM , Mac OS X TM , Unix TM Linux TM , FreeBSD TM and so on.

[0428] The steps performed by the computer device in the above embodiments may be based on the Figure 20 shown computer device structure.

[0429] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, which when running on a computer, causes the computer to execute the methods described in the foregoing embodiments.

[0430] An embodiment of the present application further provides a computer program product including a program, which when running on a computer, causes the computer to execute the methods described in the foregoing embodiments.

[0431] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above may refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0432] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in the form of electricity, machinery, or other forms.

[0433] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0434] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0435] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0436] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A training method for a neural network model, characterized in that, it includes: Obtain the first global information and the first unit information corresponding to the first moment, where the first global information is used to describe the game situation of the camp at the first moment, the first unit information is used to describe the state of the interaction unit at the first moment, and the interaction unit represents an object having an interaction relationship with the game character in the game; the first global information includes at least one of the global map information and the camp confrontation information corresponding to the first moment; the first unit information includes at least one of the game character information, non-player character (NPC) information, and defense tower information corresponding to the first moment; Based on the first global information and the first unit information, obtain the action distribution of each of the K agents at the first moment through the behavior prediction network, where the K agents belong to the same camp, the agent is a game character controlled by artificial intelligence, and K is an integer greater than 1; According to the action distribution of each agent, obtain the second global information and the second unit information at the second moment, where the second moment is the next moment of the first moment, the second global information is used to describe the game situation of the camp at the second moment, and the second unit information is used to describe the state of the interaction unit at the second moment; According to the first global information, the first unit information, the second global information, and the second unit information, determine the reward information of each agent, where the reward information is the incentive obtained after triggering the action; the reward information includes a dense incentive factor, a sparse incentive factor, a survival incentive factor, a defense tower incentive factor, and a game incentive factor; Based on the first global information, the first unit information, the second global information, and the second unit information, obtain the target self-value and the target team value of each agent through the value network, where the target team value represents the value of the agent in the camp; According to the action distribution, the reward information, the target self-value, and the target team value of each agent, train the behavior prediction network and the value network.

2. The training method according to claim 1, characterized in that, the obtaining of the first global information and the first unit information corresponding to the first moment includes: Obtain the first global information corresponding to the first moment from the screen at the first moment; Obtain the first unit information corresponding to the first moment from the screen at the first moment.

3. The training method according to claim 1, characterized in that, the obtaining of the action distribution of each of the K agents at the first moment through the behavior prediction network based on the first global information and the first unit information includes: Based on the first global information and the first unit information, obtain the first action distribution of each of the K agents at the first moment through the first behavior prediction network; Based on the first global information and the first unit information, obtain the second action distribution of each of the K agents at the first moment through a second action prediction network, where the second action prediction network is generated after the first action prediction network.

4. The training method according to claim 3, wherein, the obtaining the first action distribution of each of the K agents at the first moment through the first action prediction network based on the first global information and the first unit information includes: extracting first global features based on the first global information through a global feature network included in the first action prediction network; extracting first unit features based on the first unit information through a local feature network included in the first action prediction network; obtaining the first action distribution of each of the K agents at the first moment through a fully connected layer included in the first action prediction network based on the first global features and the first unit features; the obtaining the second action distribution of each of the K agents at the first moment through the second action prediction network based on the first global information and the first unit information includes: extracting second global features based on the first global information through a global feature network included in the second action prediction network; extracting second unit features based on the first unit information through a local feature network included in the second action prediction network; obtaining the second action distribution of each of the K agents at the first moment through a fully connected layer included in the second action prediction network based on the second global features and the second unit features.

5. The training method according to claim 3, wherein, the obtaining the second global information and the second unit information at the second moment according to the action distribution of each agent includes: determining the predicted action of each agent according to the first action distribution of each agent at the first moment, where the predicted action is the action corresponding to the maximum probability in the first action distribution; interacting the predicted action of each agent with the game environment to obtain the second global information and the second unit information at the second moment.

6. The training method according to claim 1, wherein, the determining the reward information of each agent according to the first global information, the first unit information, the second global information, and the second unit information includes: obtaining the dense incentive factor of each agent according to the first unit information and the second unit information, where the dense incentive factor includes at least one of an asset incentive factor, an experience incentive factor, and a mana incentive factor; obtaining the sparse incentive factor of each agent according to the first unit information and the second unit information, where the sparse incentive factor includes at least one of a kill incentive factor, a death incentive factor, and an assist incentive factor; Based on the first unit information and the second unit information, obtain the survival incentive factor of each agent, where the survival incentive factor includes at least one of the self - survival incentive factor and the enemy - survival incentive factor; Based on the first unit information and the second unit information, obtain the defense tower incentive factor of each agent, where the defense tower incentive factor includes at least one of the quantity incentive factor and the firmness incentive factor; Based on the first global information and the second global information, obtain the game incentive factor of each agent.

7. According to the training method described in claim 1, characterized in that, Based on the first global information, the first unit information, the second global information, and the second unit information, obtaining the target self - value and the target team - value of each agent through a value network includes: Based on the first global information and the first unit information, obtain the first self - value of each agent at the first moment through the first self - value network; Based on the second global information and the second unit information, obtain the second self - value of each agent at the second moment through the first self - value network; Based on the first global information and the first unit information, obtain the third self - value of each agent at the first moment through the second self - value network, where the second self - value network is generated after the first self - value network; Based on the first global information and the first unit information, obtain the first team - value of each agent at the first moment through the first team - value network; Based on the second global information and the second unit information, obtain the second team - value of each agent at the second moment through the first team - value network; Based on the first global information and the first unit information, obtain the third team - value of each agent at the first moment through the second team - value network, where the second team - value network is generated after the first team - value network.

8. According to the training method described in claim 7, characterized in that, The obtaining of the first self - value of each agent at the first moment based on the first global information and the first unit information through the first self - value network includes: Based on the first global information and the first unit information, obtain the dense value factor of each agent at the first moment through the dense evaluation network included in the first self - value network; Based on the first global information and the first unit information, obtain the sparse value factor of each agent at the first moment through the sparse evaluation network included in the first self - value network; Based on the first global information and the first unit information, obtain the survival value factor of each agent at the first moment through the survival evaluation network included in the first self - value network; Based on the first global information and the first unit information, obtain the defense tower value factor of each agent at the first moment through the defense tower evaluation network included in the first self-value network; Based on the first global information and the first unit information, obtain the game value factor of each agent at the first moment through the game evaluation network included in the first self-value network.

9. The training method according to claim 7, wherein, the obtaining the first team value of each agent at the first moment based on the first global information and the first unit information through the first team value network includes: obtain the role value vector of each agent among the K agents, where the role value vector includes L feature elements, and L is an integer greater than or equal to 1; recombine the role value vectors of each agent among the K agents to obtain L hierarchical value vectors, where each hierarchical value vector includes the feature elements of each agent; Based on the L hierarchical value vectors, the first global information, and the first unit information, obtain L tapped value vectors through the value attention network included in the first team value network, where each tapped value vector includes T tapped values, and T is an integer greater than or equal to 1; Based on the first global information, the first unit information, and the L tapped value vectors, obtain the first team value of each agent at the first moment through the value fusion network included in the first team value network.

10. The training method according to claim 9, wherein, the obtaining L tapped value vectors based on the L hierarchical value vectors, the first global information, and the first unit information through the value attention network included in the first team value network includes: Based on the first global information, the first unit information, and the dense layer value vector among the L hierarchical value vectors, obtain the first tapped value vector through the dense value attention network included in the value attention network, where the value attention network is included in the first team value network; Based on the first global information, the first unit information, and the sparse layer value vector among the L hierarchical value vectors, obtain the second tapped value vector through the sparse value attention network included in the value attention network; Based on the first global information, the first unit information, and the survival layer value vector among the L hierarchical value vectors, obtain the third tapped value vector through the survival value attention network included in the value attention network; Based on the first global information, the first unit information, and the defense tower layer value vector among the L hierarchical value vectors, obtain the fourth tapped value vector through the defense tower value attention network included in the value attention network; Based on the first global information, the first unit information, and the in-game layer value vector among the L layer value vectors, obtain a fifth tap value vector through the in-game value attention network included in the value attention network, where the first tap value vector, the second tap value vector, the third tap value vector, the fourth tap value vector, and the fifth tap value vector are included in the L tap value vectors.

11. The training method according to claim 9, wherein, the obtaining of the first team value of each agent at the first moment based on the first global information, the first unit information, and the L tap value vectors through the value fusion network included in the first team value network includes: obtaining a first fusion value based on the first global information, the first unit information, and the first tap value vector among the L tap value vectors through the dense value fusion network included in the value fusion network, where the value fusion network is included in the first team value network; obtaining a second fusion value based on the first global information, the first unit information, and the second tap value vector among the L tap value vectors through the sparse value fusion network included in the value fusion network; obtaining a third fusion value based on the first global information, the first unit information, and the third tap value vector among the L tap value vectors through the survival value fusion network included in the value fusion network; obtaining a fourth fusion value based on the first global information, the first unit information, and the fourth tap value vector among the L tap value vectors through the turret value fusion network included in the value fusion network; obtaining a fifth fusion value based on the first global information, the first unit information, and the fifth tap value vector among the L tap value vectors through the in-game value fusion network included in the value fusion network, where the first fusion value, the second fusion value, the third fusion value, the fourth fusion value, and the fifth fusion value are included in the first team value of each agent in the same camp at the first moment.

12. The training method according to any one of claims 1 to 11, wherein, the training of the behavior prediction network and the value network according to the action distribution, the reward information, the target self-value, and the target team value of each agent includes: determining a target loss value according to the target self-value, the target team value, the action distribution, and the reward information of each agent; updating the network parameters of the second behavior prediction network using the target loss value, where the second behavior prediction network is included in the behavior prediction network; Determine a first loss value according to the third self-value of each agent at the first moment, the reward information, and the second self-value at the second moment, where the third self-value is obtained through a second self-value network, the second self-value is obtained through a first self-value network, and the second self-value network is generated after the first self-value network; Update the network parameters of the second self-value network using the first loss value, where the second self-value network is the second value network included in the value network; Determine a second loss value according to the third team value of each agent at the first moment, the reward information, and the second team value at the second moment, where the third team value is obtained through a second team value network, the second team value is obtained through a first team value network, and the second team value network is generated after the first team value network; Update the network parameters of the second team value network using the second loss value, where the second team value network is the second value network included in the value network.

13. The training method according to claim 12, wherein, The target self-value includes the first self-value at the first moment and the second self-value at the second moment obtained through the first self-value network, and the target team value includes the first team value at the first moment and the second team value at the second moment obtained through the first team value network; The determining the target loss value according to the target self-value, the target team value, the action distribution, and the reward information of each agent includes: Calculating a self-advantage estimate value according to the first self-value of each agent at the first moment, the second self-value at the second moment, and the reward information; Calculating a team-advantage estimate value according to the first team value of each agent at the first moment, the second team value at the second moment, and the reward information; Determining a self-loss value according to the action distribution of each agent and the self-advantage estimate value; Determining a team-loss value according to the action distribution of each agent and the team-advantage estimate value; Performing a summation process on the self-loss value and the team-loss value to calculate the target loss value.

14. A training method for a neural network model, wherein, comprising: Invoking a behavior prediction network to perform self-play battles to obtain training data; Obtain the global information and unit information at different times from the training data, where the global information is used to describe the game situation of the camp, the unit information is used to describe the real-time state of the interactive unit, and the interactive unit represents an object having an interactive relationship with the game character in the game; the global information includes at least one of the global map information and the camp confrontation information; the unit information includes at least one of the game character information, the non-player character (NPC) information, and the defense tower information; Determine the reward information of each agent in the same camp according to the global information and the unit information at different times, where the reward information is the incentive obtained after triggering an action, and the agent is a game character controlled by artificial intelligence; the reward information includes a dense incentive factor, a sparse incentive factor, a survival incentive factor, a defense tower incentive factor, and a game incentive factor; Based on the global information and the unit information, obtain the target self-value and the target team value of each agent through the value network, where the target team value represents the value of the agent in the camp; Train the behavior prediction network and the value network according to the reward information, the target self-value, and the target team value of each agent.

15. A model training device, characterized in that, it includes: An acquisition module, configured to acquire the first global information and the first unit information corresponding to the first moment, where the first global information is used to describe the game situation of the camp at the first moment, the first unit information is used to describe the state of the interactive unit at the first moment, and the interactive unit represents an object having an interactive relationship with the game character in the game; the first global information includes at least one of the global map information corresponding to the first moment and the camp confrontation information; the first unit information includes at least one of the game character information, the non-player character (NPC) information, and the defense tower information corresponding to the first moment; The acquisition module is further configured to, based on the first global information and the first unit information, obtain the action distribution of each of the K agents at the first moment through the behavior prediction network, where the K agents belong to the same camp, the agent is a game character controlled by artificial intelligence, and K is an integer greater than 1; The acquisition module is further configured to, according to the action distribution of each agent, obtain the second global information and the second unit information at the second moment, where the second moment is the next moment of the first moment, the second global information is used to describe the game situation of the camp at the second moment, and the second unit information is used to describe the state of the interactive unit at the second moment; A determination module, configured to determine the reward information of each agent according to the first global information, the first unit information, the second global information, and the second unit information, where the reward information is the incentive obtained after triggering an action; the reward information includes a dense incentive factor, a sparse incentive factor, a survival incentive factor, a defense tower incentive factor, and a game incentive factor; The obtaining module is further configured to obtain the target self-value and the target team value of each agent through a value network based on the first global information, the first unit information, the second global information, and the second unit information, where the target team value represents the value of the agent in the camp; A training module, configured to train the behavior prediction network and the value network according to the action distribution, the reward information, the target self-value, and the target team value of each agent.

16. The training device according to claim 15, wherein, The obtaining module is specifically configured to: Obtain the first global information corresponding to the first moment from the picture of the first moment; Obtain the first unit information corresponding to the first moment from the picture of the first moment.

17. The training device according to claim 15, wherein, The obtaining module is specifically configured to: Based on the first global information and the first unit information, obtain the first action distribution of each of the K agents at the first moment through a first behavior prediction network; Based on the first global information and the first unit information, obtain the second action distribution of each of the K agents at the first moment through a second behavior prediction network, where the second behavior prediction network is generated after the first behavior prediction network.

18. The training device according to claim 17, wherein, The obtaining module is specifically configured to: Extract a first global feature based on the first global information through a global feature network included in the first behavior prediction network; Extract a first unit feature based on the first unit information through a local feature network included in the first behavior prediction network; Based on the first global feature and the first unit feature, obtain the first action distribution of each of the K agents at the first moment through a fully connected layer included in the first behavior prediction network; The obtaining module is specifically configured to: Extract a second global feature based on the first global information through a global feature network included in the second behavior prediction network; Extract a second unit feature based on the first unit information through a local feature network included in the second behavior prediction network; Based on the second global feature and the second unit feature, obtain the second action distribution of each of the K agents at the first moment through a fully connected layer included in the second behavior prediction network.

19. The training device according to claim 17, wherein, The obtaining module is specifically configured to: Determine the predicted action of each agent according to the first action distribution of each agent at the first moment, where the predicted action is the action corresponding to the maximum probability in the first action distribution; Interact the predicted action of each agent with the game environment to obtain the second global information and the second unit information at the second moment.

20. The training device according to claim 15, wherein, The determining module is specifically configured to: Obtain the dense incentive factor of each agent according to the first unit information and the second unit information, where the dense incentive factor includes at least one of an asset incentive factor, an experience incentive factor, and a blue amount incentive factor; Obtain the sparse incentive factor of each agent according to the first unit information and the second unit information, where the sparse incentive factor includes at least one of a kill incentive factor, a death incentive factor, and an assist incentive factor; Obtain the survival incentive factor of each agent according to the first unit information and the second unit information, where the survival incentive factor includes at least one of a self-survival incentive factor and an enemy survival incentive factor; Obtain the defense tower incentive factor of each agent according to the first unit information and the second unit information, where the defense tower incentive factor includes at least one of a quantity incentive factor and a firmness incentive factor; Obtain the game incentive factor of each agent according to the first global information and the second global information.

21. The training device according to claim 15, wherein, The obtaining module is specifically configured to: Based on the first global information and the first unit information, obtain the first self-value of each agent at the first moment through a first self-value network; Based on the second global information and the second unit information, obtain the second self-value of each agent at the second moment through the first self-value network; Based on the first global information and the first unit information, obtain the third self-value of each agent at the first moment through a second self-value network, where the second self-value network is generated after the first self-value network; Based on the first global information and the first unit information, obtain the first team value of each agent at the first moment through a first team value network; Based on the second global information and the second unit information, obtain the second team value of each agent at the second moment through the first team value network; Based on the first global information and the first unit information, obtain the third team value of each agent at the first moment through a second team value network, where the second team value network is generated after the first team value network.

22. The training device according to claim 21, wherein, The obtaining module is specifically configured to: Based on the first global information and the first unit information, obtain the dense value factors of each agent at the first moment through the dense evaluation network included in the first self-value network; Based on the first global information and the first unit information, obtain the sparse value factors of each agent at the first moment through the sparse evaluation network included in the first self-value network; Based on the first global information and the first unit information, obtain the survival value factors of each agent at the first moment through the survival evaluation network included in the first self-value network; Based on the first global information and the first unit information, obtain the defense tower value factors of each agent at the first moment through the defense tower evaluation network included in the first self-value network; Based on the first global information and the first unit information, obtain the game value factors of each agent at the first moment through the game evaluation network included in the first self-value network.

23. The training device according to claim 21, wherein, the obtaining module is specifically configured to: Obtain the role value vector of each agent among the K agents, where the role value vector includes L feature elements, and L is an integer greater than or equal to 1; Recombine the role value vectors of each agent among the K agents to obtain L hierarchical value vectors, where each hierarchical value vector includes the feature elements of each agent; Based on the L hierarchical value vectors, the first global information, and the first unit information, obtain L tapped value vectors through the value attention network included in the first team value network, where each tapped value vector includes T tapped values, and T is an integer greater than or equal to 1; Based on the first global information, the first unit information, and the L tapped value vectors, obtain the first team value of each agent at the first moment through the value fusion network included in the first team value network.

24. The training device according to claim 23, wherein, the obtaining module is specifically configured to: Based on the first global information, the first unit information, and the dense layer value vector among the L hierarchical value vectors, obtain the first tapped value vector through the dense value attention network included in the value attention network, where the value attention network is included in the first team value network; Based on the first global information, the first unit information, and the sparse layer value vector among the L hierarchical value vectors, obtain the second tapped value vector through the sparse value attention network included in the value attention network; Based on the first global information, the first unit information, and the survival layer value vector among the L hierarchical value vectors, obtain the third tapped value vector through the survival value attention network included in the value attention network; Based on the first global information, the first unit information, and the value vector of the defense tower layer among the L hierarchical value vectors, obtain a fourth tapped value vector through the defense tower value attention network included in the value attention network; Based on the first global information, the first unit information, and the value vector of the game layer among the L hierarchical value vectors, obtain a fifth tapped value vector through the game value attention network included in the value attention network, where the first tapped value vector, the second tapped value vector, the third tapped value vector, the fourth tapped value vector, and the fifth tapped value vector are included in the L tapped value vectors.

25. The training device according to claim 23, wherein, the obtaining module is specifically configured to: Based on the first global information, the first unit information, and the first tapped value vector among the L tapped value vectors, obtain a first fusion value through the dense value fusion network included in the value fusion network, where the value fusion network is included in the first team value network; Based on the first global information, the first unit information, and the second tapped value vector among the L tapped value vectors, obtain a second fusion value through the sparse value fusion network included in the value fusion network; Based on the first global information, the first unit information, and the third tapped value vector among the L tapped value vectors, obtain a third fusion value through the survival value fusion network included in the value fusion network; Based on the first global information, the first unit information, and the fourth tapped value vector among the L tapped value vectors, obtain a fourth fusion value through the defense tower value fusion network included in the value fusion network; Based on the first global information, the first unit information, and the fifth tapped value vector among the L tapped value vectors, obtain a fifth fusion value through the game value fusion network included in the value fusion network, where the first fusion value, the second fusion value, the third fusion value, the fourth fusion value, and the fifth fusion value are included in the first team value of each agent in the same camp at the first moment.

26. The training device according to any one of claims 15 to 25, wherein, the training module is specifically configured to: Determine a target loss value according to the target self-value, the target team value, the action distribution, and the reward information of each agent; Update the network parameters of the second behavior prediction network by using the target loss value, where the second behavior prediction network is included in the behavior prediction network; Determine a first loss value according to the third self-value of each agent at the first moment, the reward information, and the second self-value at the second moment, where the third self-value is obtained through the second self-value network, the second self-value is obtained through the first self-value network, and the second self-value network is generated after the first self-value network; Update the network parameters of the second self-value network by using the first loss value, where the second self-value network is included in the second value network of the value network; Determine a second loss value according to the third team value of each agent at the first moment, the reward information, and the second team value at the second moment, where the third team value is obtained through a second team value network, the second team value is obtained through a first team value network, and the second team value network is generated after the first team value network; Update the network parameters of the second team value network by using the second loss value, where the second team value network is included in the second value network of the value network.

27. The training device according to claim 26, wherein, the target self-value includes the first self-value at the first moment and the second self-value at the second moment obtained through the first self-value network, and the target team value includes the first team value at the first moment and the second team value at the second moment obtained through the first team value network; The training module is specifically configured to: calculate a self-advantage estimation value according to the first self-value of each agent at the first moment, the second self-value at the second moment, and the reward information; calculate a team-advantage estimation value according to the first team value of each agent at the first moment, the second team value at the second moment, and the reward information; determine a self-loss value according to the action distribution of each agent and the self-advantage estimation value; determine a team-loss value according to the action distribution of each agent and the team-advantage estimation value; sum up the self-loss value and the team-loss value to calculate the target loss value.

28. A training device for a neural network model, wherein, it includes: a calling module, configured to call a behavior prediction network to perform self-play battles to obtain training data; an obtaining module, configured to obtain global information and unit information at different moments from the training data, where the global information is used to describe the game situation of the camp, the unit information is used to describe the real-time state of the interaction unit, and the interaction unit represents an object having an interaction relationship with the game character in the game; the global information includes at least one of global map information and camp confrontation information; the unit information includes at least one of game character information, non-player character (NPC) information, and defense tower information; a determining module, configured to determine the reward information of each agent in the same camp according to the global information and the unit information at different moments, where the reward information is the incentive obtained after triggering an action, and the agent is a game character controlled by artificial intelligence; the reward information includes a dense incentive factor, a sparse incentive factor, a survival incentive factor, a defense tower incentive factor, and a game incentive factor; The acquisition module is further configured to obtain the target self-value and the target team value of each agent through a value network based on the global information and the unit information, where the target team value represents the value of the agent in the camp. The training module is configured to train the behavior prediction network and the value network according to the reward information, the target self-value, and the target team value of each agent.

29. A computer device characterized in that it includes: a memory, a processor, and a bus system; wherein, the memory is used to store programs; the processor is configured to execute the programs in the memory, and the processor is configured to execute the training method according to any one of claims 1 to 13 or execute the training method according to claim 14 according to the instructions in the program code; the bus system is configured to connect the memory and the processor to enable communication between the memory and the processor.

30. A computer-readable storage medium includes instructions that, when running on a computer, cause the computer to execute the training method according to any one of claims 1 to 13 or execute the training method according to claim 14.

31. A computer program product includes a computer program and instructions characterized in that when the computer program / instruction is executed by a processor, it implements the training method according to any one of claims 1 to 13 or executes the training method according to claim 14.

Citation Information

Patent Citations

  • Multi-agent training method and device and electronic equipment

    CN110882544A

  • Model training method and object control method and device based on interactive application

    CN111282279A