Training Method, Device, Equipment and Medium of Artificial Intelligence AI Model

By subdividing the value estimates of game states into different value categories, and using different value calculation formulas and reinforcement learning algorithms to train AI models, the problem that AI models cannot accurately evaluate the value of game states in MOBA games is solved, and the decision-making ability and training credibility of AI models are improved.

CN112221152BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011164804.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-27
Publication Date
2025-07-11
Estimated Expiration
2040-10-27

AI Technical Summary

Technical Problem

In the prior art, the AI model of MOBA games cannot accurately estimate the value of game state because the existing methods only consider time attenuation, and the impact of different game parameters on strategy decisions varies, resulting in the value network being unable to accurately evaluate the value of the state.

Method used

The value estimate of game state is subdivided into different value categories, different value calculation formulas are used to calculate action value, and the AI model is trained through reinforcement learning algorithms to estimate state value from multiple angles.

Benefits of technology

It improves the accuracy of the AI model's state value estimate of game states, enhances its decision-making ability, and improves the credibility of training and the accuracy of strategic decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112221152B_ABST
    Figure CN112221152B_ABST
Patent Text Reader

Abstract

The present application discloses a training method, device, equipment and medium for an artificial intelligence (AI) model, relating to the field of machine learning in artificial intelligence. The method includes: invoking the AI model to play a game in a game program to obtain training data, where the training data includes a reference game state in the game, a target game action output by a decision network according to the reference game state, and a state value output by a value network according to the reference game state, the state value includes k state sub-values on k value classifications, and k is an integer greater than 1; calculating an action value of the AI model adopting the target game action in the reference game state according to the training data and k value calculation formulas corresponding to the k value classifications, the action value includes k action sub-values on the k value classifications; and training the AI model according to the difference between the state value and the action value. This method can improve the accuracy of the value network in predicting the state value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of machine learning in artificial intelligence, and particularly to a method, device, equipment and medium for training an artificial intelligence (AI) model. Background Art

[0002] Reinforcement learning is one of the paradigms and methodologies of machine learning, which is used to describe and solve the problem that an agent (also known as an "agent") maximizes the reward or achieves a specific goal by learning strategies during the interaction with the environment. An AI (Artificial Intelligence) model designed based on reinforcement learning can make game decisions to win the game.

[0003] The AI model includes a decision network and a value network. The decision network is used to determine action instructions according to the game state, and the value network is used to evaluate the value of the game state. During the training stage of the AI model, the decision network needs to be guided by the value output by the value network to make policy decisions in the direction of maximizing the value. Therefore, it is particularly important whether the value network can accurately estimate the value of the game state.

[0004] In a MOBA game, the AI model of the related technology calculates the moment value at each moment according to the change of the game state at each moment after executing an action. Among them, the moment value at a moment is calculated according to the value factors of multiple game parameters in the game state, and then the multiple moment values are weighted and summed according to the decay factor at each moment to obtain the action value. The value network is trained according to the action value to accurately estimate the state value of the game state after executing the action. The decay factor represents the influence length of the action. The influence of the action is greater in the recent period of time, and the decay factor is larger; as time goes by, the influence degree gradually decreases, and the decay factor gradually decreases.

[0005] However, due to the complex game environment of the MOBA game, the game parameters change due to the execution of actions, and the influence degrees of the changes of different game parameters on the policy decision are also different. For example, the influence length of the death of enemy minions on the game is much smaller than the influence length of the destruction of the enemy defense tower. The method in the related technology only considers the decay in time, and the influence lengths of the value factors of all game parameters are the same, so the value network cannot accurately estimate the state value. Summary of the Invention

[0006] The embodiments of the present application provide a method, device, equipment and medium for training an artificial intelligence (AI) model, which can improve the accuracy of the value network in estimating the state value. The technical solutions are as follows:

[0007] On the one hand, a training method for an artificial intelligence (AI) model is provided. The AI model includes a value network and a decision network. The method includes:

[0008] Invoking the AI model to play a game in a game program to obtain training data, where the training data includes the reference game state in the game, the target game actions output by the decision network according to the reference game state, and the state value output by the value network according to the reference game state. The state value includes k state sub-values on k value classifications, and k is an integer greater than 1;

[0009] Calculating the action value of the AI model taking the target game action in the reference game state according to the training data and the k value calculation formulas corresponding to the k value classifications. The action value includes k action sub-values on the k value classifications;

[0010] Training the AI model according to the difference between the state value and the action value.

[0011] On the other hand, a training device for an AI model is provided. The AI model includes a value network and a decision network. The device includes:

[0012] A model module for invoking the AI model to play a game in a game program to obtain training data, where the training data includes the reference game state in the game, the target game actions output by the decision network according to the reference game state, and the state value output by the value network according to the reference game state. The state value includes k state sub-values on k value classifications, and k is an integer greater than 1;

[0013] A calculation module for calculating the action value of the AI model taking the target game action in the reference game state according to the training data and the k value calculation formulas corresponding to the k value classifications. The action value includes k action sub-values on the k value classifications;

[0014] A training module for training the AI model according to the difference between the state value and the action value.

[0015] On the other hand, a computer device is provided. The computer device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the training method for the AI model as described in the above aspect.

[0016] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the training method of the artificial intelligence AI model as described in the above aspect.

[0017] On the other hand, an embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the training method of the artificial intelligence AI model provided in the above optional implementation manner.

[0018] The beneficial effects brought by the technical solution provided by the embodiment of the present application at least include:

[0019] By classifying and calculating the state value and action value of the game state, the value estimation of the state by the AI model is subdivided into different value classifications, so that the AI model estimates the game state from multiple value classification perspectives. Different value calculation formulas are used to calculate the action value, and the action value is used to train the AI model to calculate the state value of different value classifications, so that the AI model can estimate the state value from multiple perspectives, improve the accuracy of the state value of the game state by the AI model, thereby improving the credibility of the value for the training of the AI model, and further improving the decision-making ability of the AI model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0021] Figure 1 is a structural block diagram of a computer system provided by an exemplary embodiment of the present application;

[0022] Figure 2 is a method flow chart of a training method of an AI model provided by another exemplary embodiment of the present application;

[0023] Figure 3 is a schematic diagram of a game interface of a training method of an AI model provided by another exemplary embodiment of the present application;

[0024] Figure 4It is a flowchart of a method for training an AI model provided by another exemplary embodiment of the present application;

[0025] Figure 5 It is a flowchart of a method for training an AI model provided by another exemplary embodiment of the present application;

[0026] Figure 6 It is a schematic diagram of game state extraction in a method for training an AI model provided by another exemplary embodiment of the present application;

[0027] Figure 7 It is a schematic diagram of game state extraction in a method for training an AI model provided by another exemplary embodiment of the present application;

[0028] Figure 8 It is a schematic diagram of a value network in a method for training an AI model provided by another exemplary embodiment of the present application;

[0029] Figure 9 It is a schematic diagram of the relationship between the decay factor and time in a method for training an AI model provided by another exemplary embodiment of the present application;

[0030] Figure 10 It is a flowchart of a method for training an AI model provided by another exemplary embodiment of the present application;

[0031] Figure 11 It is a block diagram of a training device for an AI model provided by another exemplary embodiment of the present application;

[0032] Figure 12 It is a block diagram of a terminal provided by another exemplary embodiment of the present application;

[0033] Figure 13 It is a block diagram of a server provided by another exemplary embodiment of the present application. Detailed implementation manners

[0034] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0035] First, a brief introduction is given to several nouns related to the embodiments of the present application.

[0036] Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.

[0037] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0038] Machine Learning (ML) is an interdisciplinary subject that involves multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.

[0039] Reinforcement Learning (RL), also known as reward learning, evaluation learning, or enhancement learning, is one of the paradigms and methodologies of machine learning, used to describe and solve the problem of an agent achieving maximum reward or specific goals through learning strategies during the interaction with the environment. Reinforcement learning is a way for an agent to learn by "trial and error", guiding its behavior through the rewards obtained from interacting with the environment. The goal is to enable the agent to obtain the maximum reward. Reinforcement learning is different from supervised learning in connectionist learning, mainly in terms of the reinforcement signal. In reinforcement learning, the reinforcement signal provided by the environment is an evaluation of the quality of the generated action (usually a scalar signal), rather than telling the Reinforcement Learning System (RLS) how to generate the correct action. Since the information provided by the external environment is scarce, the RLS must learn from its own experiences. In this way, the RLS acquires knowledge in an action-evaluation environment and improves its action plan to adapt to the environment.

[0040] Multiplayer Online Battle Arena (MOBA) refers to: in a virtual environment, different virtual teams belonging to at least two opposing camps each occupy their own map areas and compete with a certain victory condition as the goal. The victory condition includes but is not limited to: occupying a stronghold or destroying the stronghold of the opposing camp, killing virtual characters of the opposing camp, ensuring one's own survival within a specified scenario and time, seizing a certain resource, and having a score higher than the other party within a specified time, among others. The battle arena can be carried out in units of rounds, and the maps for each round of the battle arena can be the same or different. Each virtual team includes one or more virtual characters, such as 1, 2, 3, or 5.

[0041] MOBA game: It is a game that provides several strongholds in a virtual environment. Users in different camps control virtual characters to fight in the virtual environment and occupy or destroy the strongholds of the opposing camp. For example, a MOBA game can divide users into two opposing camps, disperse the virtual characters controlled by users in the virtual environment to compete with each other, and take destroying or occupying all the strongholds of the enemy as the victory condition. A MOBA game is carried out in units of rounds, and the duration of one round of a MOBA game is from the start time of the game to the time when the victory condition is achieved.

[0042] Figure 1 The block diagram of the computer system provided by an exemplary embodiment of the present application is given. The computer system 100 includes: a computer device 110 and a server 120.

[0043] Exemplarily, the computer device 110 may be a terminal or a server. The terminal includes at least one of a smart phone, a tablet computer, an e-book reader, an MP3 player, an MP4 player, a laptop computer, and a desktop computer; the server includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center.

[0044] The computer device 110 installs and runs a client 111 of a game program, and the client 111 may be a multiplayer online battle program. Exemplarily, if the computer device has a display screen, when the computer device runs the client 111, a user interface of the client 111 is displayed on the screen of the computer device 110. The game program may be any one of a Multiplayer Online Battle Arena Games (MOBA) program, a battle royale shooting game program, a Virtual Reality (VR) application program, an Augmented Reality (AR) program, a three-dimensional map program, a virtual reality game, an augmented reality game program, a First-Person Shooting Game (FPS) program, a Third-Person Shooting Game (TPS) program, and a Simulation Game (SLG) program. In this embodiment, the client is taken as an example of a MOBA game for illustration. Exemplarily, a program of an AI model also runs on the computer device 110, and the AI model is used to control the main controlled virtual character in the game program to play a game round, that is, the AI model obtains game information from the game program and sends a control instruction to the game program to implement the operation control of the game program. Schematically, the main controlled virtual character is a first virtual character, such as a simulated character or an anime character.

[0045] The computer device 110 is connected to the server 120 through a wireless network or a wired network.

[0046] The server 120 is used to provide background services for the client of the game program. Optionally, the server 120 undertakes the main computing work, and the terminal undertakes the secondary computing work; or, the server 120 undertakes the secondary computing work, and the terminal undertakes the main computing work; or, a distributed computing architecture is adopted between the server 120 and the terminal for collaborative computing.

[0047] In a schematic example, the server 120 includes a processor 122, a user account database 123, a battle service module 124, and a user-oriented input / output interface (I / O interface) 125. Among them, the processor 122 is used to load the instructions stored in the server 120 and process the data in the user account database 123 and the battle service module 124; the user account database 123 is used to store the data of the user accounts used by the computer device 110, such as the avatars of the user accounts, the nicknames of the user accounts, the combat power indexes of the user accounts, and the service areas where the user accounts are located; the battle service module 124 is used to provide multiple battle rooms for users to conduct battles, such as 1V1 battles, 3V3 battles, 5V5 battles, etc.; the user-oriented I / O interface 125 is used to establish communication with the computer device 110 through a wireless network or a wired network to exchange data.

[0048] Combined with the above introduction to the virtual environment and the description of the implementation environment, the training method of the AI model provided by the embodiments of the present application will be described, taking the AI model running on the Figure 1 computer device shown as an example. The computer device also runs a game program.

[0049] Figure 2 The flowchart of the training method of the artificial intelligence AI model provided by an exemplary embodiment of the present application is shown. The method can be executed by the above Figure 1 computer device, and the AI model and the game program are running on the computer device. The method includes:

[0050] Step 201, call the AI model to conduct a game session in the game program to obtain training data, where the training data includes the reference game state in the game session, the target game action output by the decision network according to the reference game state, and the state value output by the value network according to the reference game state. The state value includes k state sub-values on k value classifications, and k is an integer greater than 1.

[0051] Exemplarily, the AI model includes a value network and a decision network. Exemplarily, the AI model is used to control the main virtual character in the game program to move in the virtual environment. Among them, the decision network is used to determine the action of the main virtual character at the next moment according to the game state at the current moment; the value network is used to estimate the value of the game state at the current moment. Exemplarily, the decision network and the value network are two neural networks, and their network structures can be arbitrary. This embodiment does not limit this, for example, the value network and the decision network include an input layer, a hidden layer, and an output layer, and the hidden layer can be composed of at least one of: a convolutional layer, a BN layer, an activation layer, a pooling layer, and a fully connected layer.

[0052] Exemplarily, the game program is a program provided by a game developer or a game operator, and the game program is used to implement a game session and perform logical operations in the game session. Exemplarily, some of the logical operations in the game session are performed by the client of the game program installed in the computer device, and some of the logical operations are performed by the server of the game program. Exemplarily, the game program performs logical operations of the game session according to the actions (action instructions / control instructions) output by the AI model, and generates a new game state. For example, when the action output by the AI model is to control the main controlled virtual character to move forward 1 meter, the game program controls the main controlled virtual character to move forward 1 meter in the virtual environment according to the action output by the AI model. If the computer device has a display, the movement will be displayed on the display.

[0053] Exemplarily, in the training phase, the AI model is called to perform a game session in the game program to obtain training data, and the training data includes data generated by at least one game session. Exemplarily, the AI model completes the game session in a self-play manner, that is, when a game session requires multiple players to complete together, multiple AI models are used for self-play to jointly complete the session. Exemplarily, during self-play, an old version of the AI model is used to play against the latest version of the AI model to collect the training data generated by the latest version of the AI model during the session, where the old version of the AI model is an AI model with an iteration update count less than that of the latest version of the AI model. Exemplarily, when training the AI model, each time the AI model is updated, the updated AI model is stored in the version pool of the AI model. During self-play, an old version of the AI model is randomly selected from the version pool to play against the latest version of the AI model. Exemplarily, when the game program is a MOBA game, usually a game session is jointly completed by multiple virtual characters. For example, when jointly completed by ten virtual characters, the AI model is called to control one of the virtual characters, and the other virtual characters are controlled by randomly selecting any number of old versions of the AI model from the version pool. The training data is the data generated by the AI model (the latest version of the AI model, that is, the AI model to be trained). Exemplarily, the training data can also be data generated by a human-machine battle, that is, data generated by the AI model playing against a real person.

[0054] Exemplarily, the training data includes data generated from multiple game matches. For example, the AI model generates data from 10,000 self-play game matches. The training data includes the game state at each moment in the game match, as well as each action output by the decision network based on these game states, and the state value estimated by the value network for each game moment. Exemplarily, in this embodiment, a game state in the training data: the reference game state is taken as an example to illustrate the training process. For other game states in the training data, the AI model can be iteratively trained by referring to the training method based on the reference game state.

[0055] The game state is the state of the game match at a moment, and each moment corresponds to a game state. Exemplarily, the game state refers to the state information of the game match that the main controlled virtual character controlled by the AI model can obtain at a moment. Exemplarily, the game state includes all game information displayed in the game interface of the game program that controls the main controlled virtual character. The game information includes at least one of data information, picture information, icon information, and UI control information displayed on the game interface (including the game interface and small windows that can be displayed on the game interface).

[0056] For example, such as Figure 3As shown, a game interface of the main controlled virtual character in a MOBA game is given. All the game information that can be obtained in the game interface belongs to the game state at the current moment. For example, the game information of the game state at a certain moment includes: the information in the mini-map 301 displayed in the upper left corner of the game interface (the position and orientation of the main controlled virtual character, the position and orientation of our virtual characters, the position and orientation of the enemy virtual characters, the position of our army line, the position of the enemy army line, the existing quantity and status of our defense towers, the existing quantity and status of the enemy defense towers, the status of the Lord, the status of the Tyrant, etc.); the information in the status bar 302 of our virtual characters on the right side of the mini-map 301 (the heroes, health, ultimate skill status, etc. of the teammates); the gold information in the economic status bar 303; the information in the KDA (Kill Death Assist) bar 304 in the upper right corner of the game interface (the KDA of the main controlled virtual character, the number of kills of our side, the number of kills of the enemy side, the duration of the game round, etc.); the status information of the UI (User Interface) control 305 on the game interface (skill level, available skills, skill cooldown status, skill type, skill release method, etc.); at least one of the information displayed in the virtual environment screen 306 (the level, health, and mana of the main controlled virtual character; the position, health, attack range, and attack status of the defense towers in the virtual environment; the position, health, mana, level, hero model, and equipment special effects of our virtual characters or enemy virtual characters; the relative position relationship of each unit; the health and type of our or enemy minions; the health and type of the wild monsters in the virtual environment; the damage value caused by mutual attacks between each unit, the health value restored by each unit, etc.).

[0057] Exemplarily, the game information of the game state at a certain moment not only includes the game information displayed in the game interface, but also includes the game information in the window that can be opened on the game interface. For example, clicking on the economic status bar 303 can open the equipment purchase interface, and in the equipment purchase interface, the equipment information that the main controlled virtual character has currently purchased and the equipment information that can be purchased can be obtained, etc.; another example is that clicking on the game information control 307 can open the game information window, and in the window, the equipment information, KDA, gold information, hero type, summoner skills, hero attribute information, etc. of all virtual characters participating in the game can be obtained; clicking on the main controlled virtual character status bar 308 can open the main controlled virtual character status window, and in the window, the hero attributes of the main controlled virtual character can be obtained: health, spell attack, spell defense, physical attack, physical defense, movement speed, cooldown reduction, tenacity, spell penetration, physical penetration, etc. information.

[0058] In summary, the game state includes all game information that can be obtained in the game round with the permissions of the main controlled virtual character. Exemplarily, the AI model can obtain some game information through the game interface generated by the game program, or directly obtain some game information at the specified interface of the game program. For example, the AI model can obtain the health value by identifying the health value of the main controlled virtual character displayed in the game round interface, or obtain the health value of the main controlled virtual character through the interface of the game program.

[0059] The value network needs to estimate the state value Q(s) of the current game state based on multiple game information of the game state, where Q is the state value and s is the game state. The state value is equivalent to the average value of the current game state, and is used to compare with the action value Rt(s, a) of the game action taken in the current game state to determine the quality of the game action taken. Since the values between different game information will affect and offset each other. For example, after the main controlled virtual character executes the tower-pushing action and destroys the enemy's defense tower, but at the same time the main controlled virtual character dies. Then, the destruction of the enemy's defense tower will generate a positive value, while the death of the main controlled virtual character will generate a negative value. If the method in the related technology is used to accumulate these two values, the positive value and the negative value will offset each other, and the action value of the tower-pushing action of the main controlled virtual character will be very small. Similarly, the state value obtained by the value network trained in this calculation method will also be very small, making the AI model unable to distinguish the value of this action at each level, and can only learn the cumulative value at each level, reducing the short-term planning and long-term planning capabilities of the AI model.

[0060] Therefore, the method provided in this embodiment divides the estimation and calculation of the value into different value classifications, estimates the state value of the game state and calculates the action value of the game action from different levels respectively, so that the values generated by different game information can be calculated separately, reducing the influence of the values between each game information, and also enabling the AI model to perform more refined and accurate strategy estimation based on the different state sub-values and action sub-values obtained from different value classifications, improving the planning ability of the AI model.

[0061] Exemplarily, the value classification is to classify the game information in the game state, and divide the game information with the same decay trend into the same value classification. When calculating the action value, according to the decay factor corresponding to the decay trend of each value classification, calculate the action sub-value on each value classification respectively, and then train the value network according to the action sub-values of each value classification, so that the value network can also estimate the state value of the game state according to the value classification.

[0062] For example, the value classifications include at least two of the following: dense value classification, sparse value classification, hero value classification, turret value classification, and game win / loss value classification. Among them, the dense value classification and the sparse value classification are divided according to the length of the influence time of game information on the game match; the hero value classification and the turret value classification are divided according to the importance of the influence of game information on the game match; the game win / loss value classification includes game information that determines whether the game match ends.

[0063] For example, the game information of the game state in MOBA includes: the number of defense turrets on both sides, the health of the defense turrets, the health of the main controlled virtual character, whether the main controlled virtual character is dead, the amount of gold of the main controlled virtual character, the mana of the main controlled virtual character, the KDA of the main controlled virtual character, the attack power of the main controlled virtual character, game win / loss, etc. According to the value classification, the above game information can be divided into five value classifications.

[0064] The dense value classification includes: the amount of gold of the main controlled virtual character, the mana of the main controlled virtual character. The change of the game information in the dense value classification has a short-term impact on the game state and only has a greater impact on the most recent period of time when the change occurs.

[0065] The sparse value classification includes: the KDA of the main controlled virtual character and the number of defense turrets on both sides. The change of the KDA of the main controlled virtual character and the number of defense turrets on both sides has a long-term impact on the game state and will affect the game state for a long time after the change occurs.

[0066] The hero value classification includes: whether the main controlled virtual character is dead, health, attack power, etc. The turret value classification includes: the health of the defense turrets, etc. Since the hero state and the turret state are relatively important factors affecting the changes in the MOBA match, the hero value classification and the turret value classification can be set separately for the hero state and the turret state in order to calculate the value of these two important factors to the game state.

[0067] The game win / loss value classification includes game win / loss. Since the game win / loss directly affects whether the match ends, a separate value classification can be set for value calculation.

[0068] It should be noted that the specific types of game information listed above are for illustrative purposes. Since the game information in MOBA games is numerous and complex, other game information can be assigned to different value classifications according to the above ideas for action value calculation. Exemplarily, in other games, such as shooting games and racing games, the way of value classification may not be limited to the above five classifications, and the game information can be classified according to the influence degree of the game information in different games.

[0069] Step 202: Calculate the action value of the AI model adopting the target game action in the reference game state according to the training data and the k value calculation formulas corresponding to the k value classifications. The action value includes k action sub-values on the k value classifications.

[0070] Exemplarily, according to the game information corresponding to each value classification in the game state of the training data, calculate the action sub-value of the game action in this value classification. k action sub-values can be calculated for the k value classifications. Exemplarily, since the game information in each value classification is different, the value calculation formulas for calculating the action sub-values are also different.

[0071] Step 203: Train the AI model according to the difference between the state value and the action value.

[0072] Exemplarily, calculate the loss between the state value of the reference game state and the action value of the target game action, and train the AI model according to the loss value.

[0073] In summary, the method provided in this embodiment calculates the state value and action value of the game state by classification, breaks down the value estimation of the state by the AI model into different value classifications, and enables the AI model to estimate the game state from multiple value classification perspectives. Different value calculation formulas are used to calculate the action value, and the action value is used to train the AI model to calculate the state value of different value classifications, enabling the AI model to estimate the state value from multiple perspectives, improving the accuracy of the state value of the game state by the AI model, thereby enhancing the credibility of the value for the training of the AI model, and further improving the decision-making ability of the AI model.

[0074] Exemplarily, the AI model includes a feature extraction network, a value network, and a decision network. An exemplary embodiment of obtaining training data and training the AI model based on the training data is given.

[0075] Figure 4 Shows the flowchart of the training method of the artificial intelligence AI model provided by an exemplary embodiment of the present application. Based on Figure 2 the embodiment shown, step 201 includes steps 2011 to 2015, step 202 includes steps 2021 to 2022, and step 203 includes step 2031.

[0076] Step 2011: Obtain the reference game state from the game program.

[0077] Exemplarily, the AI model includes a feature extraction network, a value network, and a decision network. The feature extraction network is used to extract the features of the game state.

[0078] Exemplarily, as Figure 5As shown, a flowchart of an AI model playing a game in a game program is given. The AI model obtains the current game state 505 of the game from the game program 501, inputs the game state 505 into the feature extraction network 502 for feature extraction to obtain game state features, the value network 503 outputs a state value according to the game state features, the decision network 504 outputs a decision result according to the game state features, determines the game action 506 according to the decision result, and generates an instruction to execute the game action 506 to the game program 501, so as to control the game program 501 to execute the game action 506. After the game program 501 executes the game action 506, the game state 505 will change, and the AI model obtains a new game state 505 from the game program 501 again, and outputs the next game action according to the new game state 505. In this way, the AI model can continuously send instructions to the game program, so as to play a game in the game program.

[0079] Exemplarily, in this embodiment, an example is given in which the AI model outputs a target game action according to a reference game state, and the AI model can continuously repeat this process to play a game.

[0080] Step 2012, call the feature extraction network to perform feature extraction on the reference game state to obtain reference game state features.

[0081] Exemplarily, the feature extraction network is a neural network model, and the way it performs feature extraction can be arbitrary. For example, the feature extraction layer can be composed of at least one of a fully connected layer, a convolutional layer, a BN layer, an activation layer, and a pooling layer. In this embodiment, the network composition of the feature extraction network is not limited.

[0082] Exemplarily, since the game state is composed of various game information, the AI model can adopt various ways to represent this game information, so as to input the game information (game state) into the feature extraction network for feature extraction. That is, the AI model can directly or indirectly obtain the game state from the game program in various ways. For example, the game state can be obtained in at least one of an image-based, quasi-image-based, or vector-based manner.

[0083] For example, an image area is intercepted from the game interface of the game program as a reference game state. Exemplarily, image conversion refers to directly obtaining an image generated by the game program from the game program and inputting the image into a feature extraction network for feature extraction. For example, the game interface displayed by the game program is input into the feature extraction network in the form of an image for feature extraction (image-based feature extraction), or the game interface displayed by the game program is divided into multiple small images and input into the feature extraction network in the form of images for feature extraction. For example, the image-based game state is an image of 64 pixels * 64 pixels, and the value of each pixel is a pixel value, or each pixel corresponds to three values, which are the pixel values of the R channel, G channel, and B channel respectively. Exemplarily, the reference game state extracted by image extraction includes but is not limited to: the entire game interface (the complete interface of the game program displayed on the terminal), the image of the mini-map, the image of the virtual environment screen, and at least one of the images of any part intercepted from the game interface.

[0084] For another example, an image region is intercepted from the game interface of a game program; the image region is subjected to image-like processing to obtain a simplified image, and the simplified image is used as a reference game state. Exemplarily, image-like processing means simplifying the image region intercepted from the game interface of the game program, and inputting the simplified image after simplification into a feature extraction network for feature extraction (image-like feature extraction). Of course, for image-like feature extraction, the complete game interface can also be used as the image region, and the image region of the complete game interface is simplified to obtain a simplified image. Image-like processing can retain important information in the image and simplify redundant information, facilitating calculation and processing. Exemplarily, the image-like processing method can be to assign the pixel points at the positions of units in the game interface a value of 1 and other pixel points a value of 0. Through this image-like processing method, the positional relationship of each unit in the game interface can be extracted. Exemplarily, in a MOBA game, a unit can include at least one of virtual characters (including our side and the enemy side), wild monsters, minions, and defense towers. Of course, for different units, the pixel point assignment can be different. For example, the main controlled virtual character can be assigned 01, the enemy virtual character can be assigned 11, and the minion can be assigned 10 to distinguish different units, so as to determine the relative positional relationship of different units. For another example, taking the image-like processing of a small image of the mini-map in a MOBA game as an example, the image-like processing method can also be to subtract the pixel value of the original mini-map from the pixel value of the mini-map in the obtained reference game state. Here, the pixel value of the original mini-map is the mini-map in the initial state where the positions of virtual characters and minions are not displayed on the mini-map. Subtracting the original mini-map from the mini-map in the reference game state can obtain the change amount of the mini-map in the reference game state, and the image of this change amount is input into the feature extraction network for feature extraction. Of course, other methods can also be used to simplify the image to obtain the game state after image-like processing, so as to input the game state after image-like processing into the feature extraction network for feature extraction.

[0085] For another example, status data is obtained from a game program through a data interface, and the status data includes data of at least one game information; a vector obtained by splicing according to the status data is used as a reference game state. Exemplarily, the status data includes all game information data presented in data form. Exemplarily, vectorization means that for game states (game information) with numerical values, these numerical values can be directly spliced into a vector, and the vector is input into a feature extraction network for feature extraction. For example, for the status information (health value, mana, spell attack, physical attack, etc.) of the main control virtual character, the numerical values of multiple status information of the main control virtual character can be directly spliced into a vector according to a preset order, and this vector is used as the game state and input into the feature extraction network for feature extraction. Exemplarily, in this embodiment, the splicing method of each status data is not limited. For example, multiple status data (such as health value, mana, attack power, defense power, etc.) of the same unit can be sequentially spliced into a vector as the reference game state, and the status data of multiple units can be spliced to obtain multiple vector-based reference game states. For another example, the same status data of different units can also be spliced into a vector-based reference game state. For example, the health values of the main control virtual character, our virtual character, and the enemy virtual character are sequentially spliced to obtain a vector-based reference game state of the health value, and multiple status data can be spliced to obtain multiple vector-based reference game states.

[0086] For example, as Figure 6As shown, for different game information in the game state, different forms can be used to express this game information. For game information with specific numerical values, for example, the KDA value in the KDA bar 304 in the upper right corner of the game interface, it can be vectorized and extracted in a vectorized form to obtain the vector feature 601. For image-based game information. For example, the KDA bar in the upper right corner of the game interface is intercepted as a KDA image, and an optical character recognition model is called to perform optical character recognition on the KDA image to obtain the KDA value. For example, if the KDA values are 1, 2, and 3, the KDA values are concatenated into the vector feature (1, 2, 3) of KDA, and the vector feature of KDA is input into the feature extraction layer for feature extraction. Another example is that through the interface of the game program, the AI model can obtain the information parameters of the main controlled virtual character in the current game state: the health value of the main controlled virtual character is 1000, the blue amount is 2000, the attack power is 100, the defense value is 200, KDA is 1 / 2 / 3, the cooling countdown of the first skill is 3s, the cooling countdown of the second skill is 0s, the cooling countdown of the third skill is 10s, etc. According to the order of health value, blue amount, attack power, defense value, KDA, cooling countdown of the first skill, cooling countdown of the second skill, cooling countdown of the third skill, the vector feature (1000, 2000, 100, 200, 1, 2, 3, 3, 0, 10) of the information parameters of the main controlled virtual character can be concatenated. The vector feature is input into the feature extraction layer for feature extraction.

[0087] For example, as Figure 6 shown, the image of the mini-map 301 in the upper left corner of the game interface, or, the image 602 of the virtual environment screen occupying most of the area in the game interface, can be class-imageified and extracted in a class-imageified form to obtain class-image features. For example, the mini-map image is class-imageified and extracted to obtain the mini-map image feature 603, and the virtual environment screen is class-imageified and extracted to obtain the local image feature 604. Taking the mini-map as an example, the image of the upper left corner mini-map is intercepted from the battle interface, and the mini-map image (1) as shown in Figure 7 can be obtained. The mini-map image (1) can be used as the mini-map image feature and input into the feature extraction layer for feature extraction. Exemplarily, the mini-map image (1) can also be class-imageified to obtain the class-image feature of the mini-map. A method for class-imageification is given. As shown in (2) in Figure 7 , first, the mini-map image (1) is equally divided into multiple small squares, and each small square is equivalent to a pixel point. The value of each small square is determined according to the positions of the virtual characters in the mini-map. Among them, the main controlled virtual character is 10, the friendly virtual character is 01, the enemy virtual character is 11, and the blank position is 00. Then, the class-image feature (3) of the mini-map can be obtained, and the class-image feature (3) is input into the feature extraction layer for feature extraction. Exemplarily, Figure 7The mini - map image is equally divided into 9 * 9 small grids. To improve the position accuracy, of course, the mini - map image can also be divided more finely, for example, into 1000 * 1000 small grids. The value and meaning of each small grid can also be set arbitrarily. For example, when the main controlled virtual character is the first hero, the value of the small grid representing its position is 0001, and when the main controlled virtual character is the second hero, the value of the small grid representing its position is 0002, and so on. Exemplarily, the method of determining the value of the small grid can also be arbitrary: for example, the value of the small grid is determined according to the content occupying more than 50% of the small grid. When the position icon of the virtual character occupies more than 50% of the area of the small grid, the value of the small grid is determined as the value corresponding to the virtual character. When the map occupies more than 50% of the area of the small grid, the value of the small grid is determined as the value of the blank position.

[0088] Exemplarily, the reference game state is input into the feature extraction network in a vectorized and image - like form for feature extraction to obtain the reference game state feature.

[0089] Step 2013, call the decision network to make a decision prediction on the reference game state feature to obtain a decision result. The decision result includes at least one probability value corresponding to at least one candidate action; execute the target game action with the largest probability value in the decision result.

[0090] Exemplarily, the decision network is a classification network. The decision network outputs the probability value of each candidate action according to the input game state feature (reference game state feature). The candidate action with the largest probability value is the game action (target game action) that needs to be executed.

[0091] Exemplarily, the candidate actions include all actions that can be executed in the game program. For example, when the game program is a MOBA program, the candidate actions can include actions such as controlling the main controlled virtual character to move in the virtual environment, using skills, attacking targets, buying equipment, selling equipment, etc. Among them, if an action has multiple possibilities, the candidate actions include all possibilities of this action. If the main controlled virtual character moves in one of 360 directions, then the candidate actions include 360 movement operations, and each movement operation corresponds to a movement direction.

[0092] For each candidate action, the decision network will output a probability value. The candidate action with the largest probability value is the target game action. The AI model sends an instruction of the target game action to the game program so that the game program can execute the target game action.

[0093] Step 2014, call the value network to make a value prediction on the reference game state feature to obtain the state value.

[0094] Exemplarily, the value network includes k value branches corresponding to k value classifications. The j-th value branch is used to estimate the state sub-value of the game state in the j-th value classification. The value network is called, and the k state sub-values of the reference game state in the k value classifications are output through the k value branches, where k is an integer greater than 1.

[0095] Exemplarily, as Figure 8 shown in (2) of [], the structure of the value network is a branched structure. Each value classification corresponds to a value branch. For example, the first value branch corresponding to the first value classification is used to output the first state sub-value Q1 of the game state, the second value branch corresponding to the second value classification is used to output the second state sub-value Q2 of the game state, and the k-th value branch corresponding to the k-th value classification is used to output the k-th state sub-value Qk of the game state. As Figure 8 shown in (1) of [], a value branch includes a multi-layer network structure 701. After the reference game state features are input into the value network, through the multi-layer network structure, the state sub-values of each value classification are output after each value branch. Exemplarily, as Figure 8 shown in (2) of [], in a network structure of a value network, the network structure of the value network includes a common network structure and a branch network structure. The common network structure is the structure shared by different value branches, and the branch network structure is the structure unique to different value branches. The game state features obtain common features after passing through the common network structure, and the branch network structures of different value branches perform value estimation based on the common features to obtain the state sub-values in different value classifications. Exemplarily, different branch network structures will focus on different feature parts in the common features, so as to calculate the state sub-values in each value classification. The focus of the branch network structure on the feature parts is guided by the action values of different value classifications during the training phase, so that the branch network structure can pay more attention to the feature parts belonging to this value classification.

[0096] Step 2015, repeat the above steps to call the AI model to conduct game matches in the game program and obtain training data.

[0097] Repeat steps 2011 to 2014. Take each game state in the game match as the reference game state, output the state value and the target game action of each game state, and let the AI model conduct game matches. By calling the AI model to conduct a large number of game matches, a large amount of training data can be obtained. Exemplarily, the training data includes all the data generated by the AI model during the game match process.

[0098] Step 2021, obtain the game states from time t0 to time t in the training data. The reference game state is the game state at time t0, t n moment, t nThe moment is the end moment of the game round, and n is a positive integer.

[0099] Exemplarily, the reference game state includes at least k pieces of game information. The k value classifications are divided according to the influence of the game information on the game round. The game information belonging to the same value classification has the same influence attenuation trend.

[0100] The calculation formula of the action value is as follows.

[0101]

[0102] Among them, the above formula is the formula for calculating the action sub-value in one value classification. r t is the moment value of the game action at moment t, r l is the value factor of the l-th game information belonging to this value classification at moment t, w l is the weight of the l-th game information, and L is the total number of game information belonging to this value classification. R t is the action sub-value of the game action in this value classification, r (t,i) is the moment value of the i-th moment, γ i is the attenuation factor of the i-th moment, and i = 0 is the moment before executing this game action (the game action is determined by the decision network according to the game state at the 0-th moment).

[0103] It can be seen from the formula that the calculation of the action sub-value is to first calculate the moment value of the target game action at each moment, and then perform the summation operation on the moment value of each moment according to the attenuation factor. That is, the action sub-value of a game action in one value classification is equal to the weighted sum of the moment values of each moment from the moment of the reference game state to the end moment of the game round. Then, to calculate the action sub-value of an action in one value classification, it is necessary to obtain multiple game states between the game state from the reference game state (the game state at t0 moment) to the end moment of the game round.

[0104] Step 2022, for the j-th value classification among the k value classifications, according to the game information belonging to the j-th value classification in the game state from t0 moment to t n moment, calculate the action sub-value of the target game action in the j-th value classification. j is a positive integer less than or equal to k, and k is an integer greater than 1; repeat this step to calculate the k action sub-values of the target game action on the k value classifications.

[0105] Exemplarily, there are multiple pieces of game information in the game state at one moment. When calculating the action sub-value of one value classification, only the game information belonging to this value classification in the game state needs to be used.

[0106] Exemplarily, for the j-th value category among the k value categories, according to the game states at time t i and time t i+1 , obtain the value factors of the game information belonging to the j-th value category, calculate the weighted sum of the value factors to obtain the moment value at time t i ; repeat this step to calculate the n moment values at n moments from time t0 to time t n-1 ; according to the n decay factors corresponding to the n moments in the j-th value category, calculate the weighted sum of the n moment values at the n moments to obtain the action sub-value of the target game action in the j-th value category; wherein, the decay factor at time t i is used to describe the decay degree of the moment value at time t i , j is a positive integer less than or equal to k, k is an integer greater than 1, i is a non-negative integer less than n, and n is a positive integer.

[0107] Exemplarily, for the calculation of the moment value at time t i of the target game action in the j-th value category, according to the changes in the two game states before and after time t i and time t i+1 , calculate the value factor of each game information. For example, if the game information is the health value of the main controlled virtual character, then according to the difference in the health value between the two game states: -200, calculate the value factor of the health value. Exemplarily, for different game information, the calculation method can be set separately to calculate its value factor. For example, for the value factor of the health value, every 100 health points correspond to 1 value factor, so the value factor corresponding to -200 health points is -2. Another example is that for the value factor of the defense tower, one defense tower corresponds to 30 value factors. If the number of our defense towers in the two game states before and after is -1, then the value factor corresponding to the defense tower is -30. In summary, the value factor of a game information at a moment is calculated based on the difference in this game information between the two game states before and after, or is obtained based on the difference in this game information between the two game states before and after.

[0108] For example, assume that the dense value categories include: the gold amount of the main controlled virtual character, the blue amount of the main controlled virtual character. At the game state at time t i , the gold amount of the main controlled virtual character is 0 and the blue amount of the main controlled virtual character is 200. At the game state at time t i+1 , the gold amount of the main controlled virtual character is 100 and the blue amount of the main controlled virtual character is 100. Among them, every 100 gold coins correspond to 1 value factor for the gold amount, and every 100 corresponds to 2 value factors for the blue amount. Then according to time t i and time t i+1The difference in the number of gold coins at a moment + 100, and the difference in the amount of blue - mana - 100. The value factor of the number of gold coins can be obtained as 100 / 100 * 1 = 1, and the value factor of the amount of blue - mana is - 100 / 100 * 2 = - 2. Then in the dense value classification at time t i The moment value at time t is equal to the value factor of the number of gold coins plus the value factor of the amount of blue - mana, which is equal to - 1.

[0109] For another example, assume that the sparse value classification includes: the KDA of the main - controlled virtual character and the number of defense towers on both sides. Among them, for every 1 increase in the number of kills K of the main - controlled virtual character, the corresponding value factor is 2; for every 1 increase in the number of deaths D, the corresponding value factor is 2; for every 1 increase in the number of assists A, the corresponding value factor is 1; for every 1 decrease in the number of our defense towers, the corresponding value factor is - 3; for every 1 decrease in the number of enemy defense towers, the corresponding value factor is 3. Then assume that at time t i The game state is that the KDA of the main - controlled virtual character is 0 / 0 / 0, the number of our defense towers is 9, and the number of enemy defense towers is 9. At time t i+1 The game state is that the KDA of the main - controlled virtual character is 1 / 0 / 0, the number of our defense towers is 8, and the number of enemy defense towers is 9. Then the value factor of the KDA of the main - controlled virtual character is 2, the value factor of the number of our defense towers is - 3, and the value factor of the number of enemy defense towers is 0. Then in the sparse value classification at time t i The moment value at time t is equal to the value factor of the KDA of the main - controlled virtual character plus the value factor of the number of our defense towers plus the value factor of the number of enemy defense towers, which is equal to - 1.

[0110] For another example, assume that the hero value classification includes: whether the main - controlled virtual character is dead, health points, and attack power. Among them, when the main - controlled virtual character changes from alive to dead, the corresponding value factor is - 2; when it changes from dead to alive, the corresponding value factor is 1; for every 100 health points, the corresponding value factor is 1; for every 100 attack power, the corresponding value factor is 1. Then if from time t i to time t i+1 the main - controlled virtual character changes from alive to dead, the health points change from 100 to 0, and the attack power changes from 3000 to 3200. Then in the hero value classification at time t i The moment value at time t is equal to the value factor of whether it is dead plus the value factor of health points plus the value factor of attack power, which is equal to - 2 - 1+2=-1.

[0111] For another example, assume that the defense - tower value classification includes: the health of the defense tower. Among them, for every 100 health points, the corresponding value factor is 1. Then if from time t i to time t i+1 the health of our defense tower decreases by 200, and the health of the enemy defense tower decreases by 300. Then in the defense - tower value classification at time t iThe moment value at a moment is equal to the value factor of our defense tower's health plus the value factor of the enemy's defense tower's health, which is -2 + 3 = 1.

[0112] For another example, assume that the game win-loss value classification includes game win and loss. Among them, the value factor for a game win is 100, and the value factor for a game loss is -100. Then, if from time t i to time t i+1 our side wins, then in the game win-loss value classification, the moment value at time t i is equal to 100; if there is no game win or loss (the game continues), then in the game win-loss value classification, the moment value at time t i is 0. Exemplarily, in a MOBA game, game win and loss are determined by the destruction of the base. That is, if our base is destroyed, the value factor is -100, and if the enemy base is destroyed, the value factor is 100.

[0113] Exemplarily, the above embodiments are mainly for listing the method of obtaining the value factor of the game information based on the changes in the game information between two consecutive game states. In the above embodiments, the moment value is calculated according to the weight of each game information being 1. Of course, in practical applications, different weights can be set for different game information. For example, in the first value classification: the weight of the first game information is 1, and the weight of the second game information is 2. Then, the moment value of the first value classification is equal to the value factor of the first game information * 1 + the value factor of the second game information * 2.

[0114] After calculating the value factor of each game information, the value factors are weighted and summed according to the weight corresponding to each game information to obtain the moment value of the target game action at a moment in a value classification. After calculating the n moment values from time t0 to time t n-1 weighted and summing the n moment values can obtain the action sub-value of the target game action in a value classification.

[0115] Exemplarily, the weighted summation of the moment values is weighted according to the time distribution of the decay factor of each value classification. Exemplarily, different value classifications correspond to different time distributions of the decay factor. That is, for game information belonging to different value classifications, the decay trend of its influence on subsequent game states is different.

[0116] For example, as Figure 9 shown, the decay factor and the moment when the target game action occurs (from the unoccurred t0 moment to the end of the game at t nExample of the correspondence between the (time) moments. In one value classification, the attenuation trend of the impact of the change in game information on the subsequent game state presents a first curve type 801 decrease. The impact of this part of the game information on the game state is relatively large at the nearest moment when the target game action occurs. In another value classification, the attenuation trend of the impact of the change in game information on the subsequent game state presents a linear type 802 decrease. The degree of impact of this part of the game information on the game state gradually decreases after the target game action occurs. In another value classification, the attenuation trend of the impact of the change in game information on the subsequent game state presents a broken line type 803 decrease. The position of this part of the game information is at the same impact level within a period of time after the target game action occurs, and the impact degree becomes 0 after time t m moment. In another value classification, the attenuation trend of the impact of the change in game information on the subsequent game state presents a second curve type 804 decrease. Compared with the first curve type 801, the attenuation of its impact degree is slower.

[0117] After obtaining the moment value at each moment, according to the correspondence between the attenuation factor corresponding to each value classification and the moment, the attenuation factor can be used as a weight to perform weighted (attenuation factor multiplied by the moment value) summation on the moment value to obtain the action sub-value of the target game action in this value classification.

[0118] According to the above method, multiple action sub-values of the target game action in multiple value classifications can be obtained.

[0119] Step 2031, call the reinforcement learning algorithm to train the AI model according to the difference between the state value and the action value.

[0120] Exemplarily, train the AI model according to the difference between the state sub-value and the action sub-value of each value classification.

[0121] Exemplarily, any reinforcement learning algorithm can be used to train the AI model. For example, at least one of the Proximal Policy Optimization (PPO) algorithm, the A3C algorithm, and the Deep Deterministic Policy Gradient (DDPG) algorithm can be used.

[0122] Exemplarily, such as Figure 10As shown in the figure, a flowchart for training an AI model using PPO is given. Taking the reference game state s, target game action, state value Q of the reference game state s, and the old policy (decision result) output according to the reference game state s included in the training data as examples, the action value Rt of the target game action is calculated based on the game states from the reference game state s to the end of the game in the training data. The first loss 901 is obtained from the difference between the state value Q and the action value Rt. The feature extraction network and the value network are trained and updated according to the first loss 901 to obtain a new feature extraction network and a new value network. Then, the new feature extraction network and the original decision network are used to output a new policy according to the reference game state s again (since the feature extraction network is updated, the output policy will be different from the old policy). The ratio of the new policy to the old policy is calculated, and the ratio is limited within the range of (0.8 - 1.2) to obtain a limited ratio. The state value Q and the action value Rt are substituted into the advantage function (advatage function) to obtain the advantage. Then, the first product of the ratio and the advantage and the second product of the limited ratio and the advantage are calculated. The smaller one of the first product and the second product is used as the second loss 902, and the second loss 902 is used to train the feature extraction network and the decision network. Among them, the advantage function is used to calculate the advantage of the action value Rt relative to the state value Q. If the target game action is a better game action in the reference game state, the advantage is a positive number, so as to train the AI model to make decisions in the direction of the target game action in the case of the reference game state; if the target game action is a bad game action in the reference game state, the advantage is a negative number, so as to train the AI model not to make decisions in the direction of the target game action in the case of the reference game state.

[0123] Exemplarily, after training an AI model using the training method of the AI model provided in the above embodiment, the AI model can be called to control the main control virtual character to play a game in the game program.

[0124] In summary, the method provided in this embodiment distributes multiple game information in the reference game state to different value classifications, calculates the action sub-values respectively, and uses the action sub-value of a certain value classification to independently train the state sub-value of this value classification in the value network, reducing the interference caused by the mutual influence between different game information on value estimation, and improving the learning speed and ability ceiling of the AI model.

[0125] The method provided in this embodiment configures different decay factors for different value classifications, thereby distinguishing the influence lengths of different game information on the game round, improving the accuracy of action values, and thus better training the value network to estimate the values of different value classifications, enhancing the stability of the value network. Moreover, configuring decay factors separately for each value classification enables the AI model to take into account learning game information with different influence lengths, improving the AI model's instant micro-operation and long-term planning capabilities.

[0126] Exemplarily, an exemplary embodiment of training an AI model by applying the training method of the AI model provided in this application in a MOBA is given.

[0127] First, call the training data obtained by the AI model in the MOBA game program. For example, the training data includes the game data of the AI model in a game round. This game round lasts for 6 minutes, and the minimum time unit of the game round is 0.01 s. Then, the 6-minute game round contains 36,000 game states. During the game round, the AI model outputs the game actions to be executed in the next 0.01 s according to the game state every 0.01 s. At the same time, the AI model also outputs the state value of the current game state. For example, for the game state at 5 minutes and 59.98 seconds, the five value branches of the value network respectively output five state sub-values on the five value classifications: the state sub-value 1 of the dense value classification, the state sub-value 2 of the sparse value classification, the state sub-value 3 of the hero value classification, the state sub-value 4 of the defense tower value classification, and the state sub-value 5 of the game win-loss value classification.

[0128] Then, calculate the action value of the game actions executed in each game state. For example, if the AI model determines to execute the game action of moving right according to the game state at 5 minutes and 59.98 seconds, then the action value of the game action of moving right executed at 5 minutes and 59.98 seconds includes five action sub-values in five value classifications: dense value classification, sparse value classification, hero value classification, turret value classification, and game win-loss value classification. The action sub-value of each value classification is obtained by calculating the moment value based on the game information belonging to the current value classification and then weighted summing the moment values with the decay factor. Since there are only three game states at 5 minutes and 59.98 seconds, 5 minutes and 59.99 seconds, and 6 minutes from 5 minutes and 59.98 seconds to the end of the game at 6 minutes, the action sub-value of each value classification is the sum of the moment values of their respective game information from 5 minutes and 59.98 seconds to 5 minutes and 59.99 seconds and from 5 minutes and 59.99 seconds to 6 minutes. Therefore, for the action sub-value of each value classification, it is necessary to calculate the moment value from 5 minutes and 59.98 seconds to 5 minutes and 59.99 seconds according to the game state at 5 minutes and 59.98 seconds and the game state at 5 minutes and 59.99 seconds, and calculate the moment value from 5 minutes and 59.99 seconds to 6 minutes according to the game state at 5 minutes and 59.99 seconds and the game state at 6 minutes.

[0129] Since examples of calculating the moment values of the five value classifications have been given respectively in the above embodiments, only an example of calculating the moment value of the win-loss value classification is given here, and the calculation methods in this embodiment and the above embodiments can be referred to for other value classifications. Assume that in the game state at 6 minutes, our virtual character destroyed the enemy base and won the game. Then, in the game state from 5 minutes and 59.98 seconds to 5 minutes and 59.99 seconds, no base was destroyed in the virtual environment, so the moment value at 5 minutes and 59.98 seconds is 0; in the game state from 5 minutes and 59.99 seconds to 6 minutes, our side destroyed the enemy base, so the moment value at 5 minutes and 59.99 seconds is 100. Then, according to the moment distribution of the decay factor of the win-loss value classification: the decay factor at the first moment is 3, the decay factor at the second moment is 2, the decay factor at the third moment is 1... The moment values at 5 minutes and 59.98 seconds and 5 minutes and 59.99 seconds are weighted and summed to get 0*3 + 2*100 = 200. Then, the action sub-value of the game action of moving right executed at the game state of 5 minutes and 59.98 seconds in the win-loss value classification is 200. Similarly, the action sub-values of the other four value classifications are calculated as follows: the action sub-value of the dense value classification is 1.2, the action sub-value of the sparse value classification is 3, the action sub-value of the hero value classification is 2, and the action sub-value of the turret value classification is 5.

[0130] Then, based on the game state at 5 minutes, 59.98 seconds, the difference between the action sub-value and the state sub-value in the dense value classification is 1.2 - 1 = 0.2, the difference between the action sub-value and the state sub-value in the sparse value classification is 3 - 2 = 1, the difference between the action sub-value and the state sub-value in the hero value classification is 2 - 3 = -1, the difference between the action sub-value and the state sub-value in the turret value classification is 5 - 4 = 1, and the difference between the action sub-value and the state sub-value in the win-loss value classification is 200 - 5 = 195. The PPO algorithm is called to train the AI model, and a new AI model is obtained. In this way, 35,999 iterative trainings can be performed based on 36,000 game states in a game match, continuously updating the AI model and improving the decision-making ability of the AI model.

[0131] After the training is completed, the trained AI model can be used for game matches of MOBA games.

[0132] The following is an apparatus embodiment of the present application. For details not described in detail in the apparatus embodiment, reference may be made to the above method embodiment.

[0133] Figure 11 It is a block diagram of a training apparatus for an artificial intelligence AI model provided by an exemplary embodiment of the present application. The artificial intelligence AI model includes a value network and a decision network, and the apparatus includes:

[0134] A model module 1001, configured to call the artificial intelligence AI model to perform a game match in a game program to obtain training data, where the training data includes a reference game state in the game match, a target game action output by the decision network according to the reference game state, and a state value output by the value network according to the reference game state. The state value includes k state sub-values on k value classifications, and k is an integer greater than 1;

[0135] A calculation module 1002, configured to calculate an action value of the artificial intelligence AI model adopting the target game action in the reference game state according to the training data and k value calculation formulas corresponding to the k value classifications. The action value includes k action sub-values on the k value classifications;

[0136] A training module 1003, configured to train the artificial intelligence AI model according to the difference between the state value and the action value.

[0137] In an optional embodiment, the reference game state includes at least k game information, and the k value classifications are divided according to the influence of the game information on the game match. The game information belonging to the same value classification has the same influence attenuation trend; the apparatus further includes:

[0138] An acquisition module 1004, configured to acquire the game states from time t0 to time t in the training data, where the reference game state is the game state at time t0, and time t is the end time of the game session, and n is a positive integer; n The reference game state is the game state at time t0, and time t n is the end time of the game session, and n is a positive integer;

[0139] The calculation module 1002 is further configured to, for the j-th value classification among the k value classifications, calculate the action sub-value of the target game action in the j-th value classification according to the game information belonging to the j-th value classification in the game state from time t0 to time t n where j is a positive integer less than or equal to k, and k is an integer greater than 1; repeating this step to calculate the k action sub-values of the target game action on the k value classifications.

[0140] In an alternative embodiment, the calculation module 1002 is further configured to, for the j-th value classification among the k value classifications, obtain the value factor of the game information belonging to the j-th value classification according to the game states at time t i and time t i+1 calculate the weighted sum of the value factors to obtain the time value at time t i where j is a positive integer less than or equal to k, k is an integer greater than 1, i is a non-negative integer less than n, and n is a positive integer; repeating this step to calculate n time values for a total of n time points from time t0 to time t n-1 ;

[0141] The calculation module 1002 is further configured to calculate the weighted sum of the n time values at the n time points according to the n decay factors corresponding to the n time points in the j-th value classification, to obtain the action sub-value of the target game action in the j-th value classification;

[0142] where the decay factor at time t i is used to describe the decay degree of the time value at time t i where j is a positive integer less than or equal to k, k is an integer greater than 1, i is a non-negative integer less than n, and n is a positive integer.

[0143] In an alternative embodiment, the artificial intelligence AI model further includes: a feature extraction network; the model module 1001 includes: an acquisition sub-module 1005, a feature extraction sub-module 1006, a decision-making sub-module 1007, and a value sub-module 1008;

[0144] The acquisition sub-module 1005 is configured to acquire the reference game state from the game program;

[0145] The feature extraction sub-module 1006 is used to call the feature extraction network to extract features from the reference game state to obtain reference game state features;

[0146] The decision-making sub-module 1007 is used to call the decision-making network to perform decision-making estimation on the reference game state features to obtain a decision result, where the decision result includes at least one probability value corresponding to at least one candidate action; execute the target game action with the largest probability value in the decision result;

[0147] The value sub-module 1008 is used to call the value network to perform value estimation on the reference game state features to obtain the state value;

[0148] The model module 1001 is further used to repeat the above steps to call the artificial intelligence AI model to conduct a game session in the game program to obtain the training data.

[0149] In an alternative embodiment, the value network includes k value branches corresponding to the k value classifications, and the j-th value branch is used to estimate the state sub-value of the game state in the j-th value classification;

[0150] The value sub-module 1008 is further used to call the value network to output the k state sub-values of the reference game state in the k value classifications through the k value branches, where k is an integer greater than 1.

[0151] In an alternative embodiment, the training module 1003 is further used to call a reinforcement learning algorithm to train the artificial intelligence AI model according to the difference between the state value and the action value.

[0152] In an alternative embodiment, the reference game state includes game information; the value classifications include at least two of: dense value classification, sparse value classification, hero value classification, defense tower value classification, game win-loss value classification;

[0153] Among them, the dense value classification and the sparse value classification are divided according to the influence time length of the game information on the game session; the hero value classification and the defense tower value classification are divided according to the influence importance degree of the game information on the game session; the game win-loss value classification includes the game information that determines whether the game session ends.

[0154] In an alternative embodiment, the model module 1001 is further used to call the artificial intelligence AI model to control the main control virtual character to conduct the game session in the game program.

[0155] It should be noted that: for the training device of the artificial intelligence AI model provided in the above embodiments, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the training device of the artificial intelligence AI model provided in the above embodiments and the embodiments of the training method of the artificial intelligence AI model belong to the same concept. For the specific implementation process, please refer to the method embodiments and will not be elaborated here.

[0156] This application also provides a terminal, which includes a processor and a memory. At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the training method of the artificial intelligence AI model provided in each of the above method embodiments. It should be noted that the terminal may be as follows Figure 12 the provided terminal.

[0157] Figure 12 shows the structural block diagram of the terminal 1100 provided in an exemplary embodiment of this application. The terminal 1100 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer, or a desktop computer. The terminal 1100 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.

[0158] Generally, the terminal 1100 includes: a processor 1101 and a memory 1102.

[0159] The processor 1101 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1101 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1101 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU. The coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1101 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1101 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0160] The memory 1102 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1102 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1102 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1101 to implement the training method of the artificial intelligence AI model provided in the method embodiments of the present application.

[0161] In some embodiments, the terminal 1100 may further optionally include: a peripheral device interface 1103 and at least one peripheral device. The processor 1101, the memory 1102, and the peripheral device interface 1103 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1103 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1104, a display screen 1105, a camera 1106, an audio circuit 1107, and a power supply 1108.

[0162] The peripheral device interface 1103 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1101 and the memory 1102. In some embodiments, the processor 1101, the memory 1102, and the peripheral device interface 1103 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1101, the memory 1102, and the peripheral device interface 1103 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0163] The radio frequency circuit 1104 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1104 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1104 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1104 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 1104 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: the World Wide Web, a metropolitan area network, an intranet, each generation of mobile communication network (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1104 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0164] The display screen 1105 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1105 is a touch display screen, the display screen 1105 also has the ability to collect touch signals on or above the surface of the display screen 1105. The touch signals can be input to the processor 1101 as control signals for processing. At this time, the display screen 1105 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there may be one display screen 1105, which is provided on the front panel of the terminal 1100; in other embodiments, there may be at least two display screens 1105, which are respectively provided on different surfaces of the terminal 1100 or are in a foldable design; in still other embodiments, the display screen 1105 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 1100. Even, the display screen 1105 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1105 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0165] The camera module 1106 is used to capture images or videos. Optionally, the camera module 1106 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera, so as to realize the function of background blurring by fusing the main camera and the depth-of-field camera, the function of panoramic shooting by fusing the main camera and the wide-angle camera, and the VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera module 1106 may further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, and can be used for light compensation under different color temperatures.

[0166] The audio circuit 1107 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1101 for processing, or input to the radio frequency circuit 1104 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1100. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1101 or the radio frequency circuit 1104 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1107 may further include a headphone jack.

[0167] The power supply 1108 is used to supply power to each component in the terminal 1100. The power supply 1108 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1108 includes a rechargeable battery, the rechargeable battery may be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery charged through a wired line, and a wireless rechargeable battery is a battery charged through a wireless coil. The rechargeable battery may also be used to support fast charging technology.

[0168] In some embodiments, the terminal 1100 further includes one or more sensors 1110. The one or more sensors 1110 include but are not limited to: an acceleration sensor 1111, a gyroscope sensor 1112, a pressure sensor 1113, an optical sensor 1114, and a proximity sensor 1115.

[0169] The acceleration sensor 1111 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established with the terminal 1100. For example, the acceleration sensor 1111 can be used to detect the components of the gravitational acceleration on the three coordinate axes. The processor 1101 can control the display screen 1105 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1111. The acceleration sensor 1111 can also be used for collecting game or user's motion data.

[0170] The gyroscope sensor 1112 can detect the body direction and rotation angle of the terminal 1100. The gyroscope sensor 1112 can cooperate with the acceleration sensor 1111 to collect the 3D actions of the user on the terminal 1100. According to the data collected by the gyroscope sensor 1112, the processor 1101 can achieve the following functions: motion sensing (such as changing the UI according to the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.

[0171] The pressure sensor 1113 can be disposed on the side frame of the terminal 1100 and / or the lower layer of the display screen 1105. When the pressure sensor 1113 is disposed on the side frame of the terminal 1100, it can detect the holding signal of the user on the terminal 1100, and the processor 1101 can perform left / right hand recognition or shortcut operations according to the holding signal collected by the pressure sensor 1113. When the pressure sensor 1113 is disposed on the lower layer of the display screen 1105, the processor 1101 can control the operable controls on the UI interface according to the pressure operation of the user on the display screen 1105. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.

[0172] The optical sensor 1114 is used to collect the ambient light intensity. In one embodiment, the processor 1101 can control the display brightness of the display screen 1105 according to the ambient light intensity collected by the optical sensor 1114. Specifically, when the ambient light intensity is high, the display brightness of the display screen 1105 is increased; when the ambient light intensity is low, the display brightness of the display screen 1105 is decreased. In another embodiment, the processor 1101 can also dynamically adjust the shooting parameters of the camera module 1106 according to the ambient light intensity collected by the optical sensor 1114.

[0173] The proximity sensor 1115, also known as the distance sensor, is usually disposed on the front panel of the terminal 1100. The proximity sensor 1115 is used to collect the distance between the user and the front of the terminal 1100. In one embodiment, when the proximity sensor 1115 detects that the distance between the user and the front of the terminal 1100 is gradually decreasing, the processor 1101 controls the display screen 1105 to switch from the lit state to the screen-off state; when the proximity sensor 1115 detects that the distance between the user and the front of the terminal 1100 is gradually increasing, the processor 1101 controls the display screen 1105 to switch from the screen-off state to the lit state.

[0174] Those skilled in the art can understand that Figure 12 the structure shown in does not constitute a limitation on the terminal 1100, and it may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements.

[0175] The memory further includes one or more programs, and the one or more programs are stored in the memory. The one or more programs include methods for training the artificial intelligence AI model provided in the embodiments of the present application.

[0176] Figure 13It is a schematic structural diagram of a server provided by an embodiment of the present application. Specifically: The server 1200 includes a Central Processing Unit (CPU) 1201, a system memory 1204 including a Random Access Memory (RAM) 1202 and a Read-Only Memory (ROM) 1203, and a system bus 1205 connecting the system memory 1204 and the central processing unit 1201. The server 1200 also includes a Basic Input / Output System (I / O system) 1206 for facilitating information transfer between various components within the computer, and a mass storage device 1207 for storing an operating system 1213, application programs 1214, and other program modules 1215.

[0177] The Basic Input / Output System 1206 includes a display 1208 for displaying information and input devices 1209 such as a mouse, keyboard, etc. for user input of information. Among them, both the display 1208 and the input devices 1209 are connected to the central processing unit 1201 through an Input / Output Controller 1212 connected to the system bus 1205. The Basic Input / Output System 1206 may also include an Input / Output Controller 1212 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the Input / Output Controller 1212 also provides output to a display screen, printer, or other types of output devices.

[0178] The mass storage device 1207 is connected to the central processing unit 1201 through a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1207 and its associated computer-readable medium provide non-volatile storage for the server 1200. That is to say, the mass storage device 1207 may include computer-readable media (not shown) such as a hard disk or a Compact Disc Read-Only Memory (CD-ROM) drive.

[0179] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, magnetic tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that computer storage media is not limited to the above several types. The above system memory 1204 and mass storage device 1207 can be collectively referred to as memory.

[0180] According to various embodiments of the present application, the server 1200 can also be run by a remote computer on the network connected through a network such as the Internet. That is, the server 1200 can be connected to the network 1212 through the network interface unit 1211 connected to the system bus 1205, or rather, the network interface unit 1211 can also be used to connect to other types of networks or remote computer systems (not shown).

[0181] The present application provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by the processor to implement the training method of the artificial intelligence AI model provided by each of the above method embodiments.

[0182] The present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the training method of the artificial intelligence AI model provided in the above optional implementation manners.

[0183] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.

[0184] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk, an optical disc, etc.

[0185] The above are only alternative embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A training method for an artificial intelligence AI model, characterized in that, The artificial intelligence AI model includes a value network and a decision network, and the method includes: Invoking the artificial intelligence AI model to conduct a game session in a game program to obtain training data, where the training data includes the reference game state in the game session, the target game action output by the decision network according to the reference game state, and the state value output by the value network according to the reference game state. The state value includes k state sub-values on k value classifications, and k is an integer greater than 1; the reference game state includes at least k game information, and the k value classifications are divided according to the influence of the game information on the game session. The game information belonging to the same value classification has the same influence attenuation trend; Obtain the game state from time t0 to time t in the training data, where the reference game state is the game state at time t0, and time t n is the end time of the game session, and n is a positive integer; n is the end time of the game session, and n is a positive integer; For the j-th value category among the k value categories, based on the game information belonging to the j-th value category in the game state from the time t0 to the time t n calculate the action sub-value of the target game action in the j-th value category. j is a positive integer less than or equal to k, and k is an integer greater than 1. Repeat this step to calculate the k action sub-values of the target game action for the k value categories. The action value includes the k action sub-values for the k value categories; Training the artificial intelligence AI model according to the difference between the state value and the action value.

2. The method according to claim 1, wherein For the j-th value classification among the k value classifications, according to the game information belonging to the j-th value classification in the game state from the time t0 to the time t n calculate the action sub-value of the target game action in the j-th value classification, including: For the j-th value classification among the k value classifications, according to the game states at time t i and time t i+1 , obtain the value factor of the game information belonging to the j-th value classification, calculate the weighted sum of the value factors to obtain the time value at time t i ; repeat this step to calculate n time values at n times from time t0 to time t n-1 ; Calculating the weighted sum of the n moment values at the n moments corresponding to the j-th value classification according to the n attenuation factors at the n moments, to obtain the action sub-value of the target game action in the j-th value classification; Among them, the decay factor at time t i is used to describe the decay degree of the time value at time t i . j is a positive integer less than or equal to k, k is an integer greater than 1, i is a non-negative integer less than n, and n is a positive integer.

3. The method according to claim 1 or 2, characterized in that, The reference game state includes game information; the value classifications include at least two of the following: dense value classification, sparse value classification, hero value classification, defense tower value classification, game win-loss value classification; Among them, the dense value classification and the sparse value classification are divided according to the influence time length of the game information on the game session; the hero value classification and the defense tower value classification are divided according to the influence importance of the game information on the game session; the game win-loss value classification includes the game information that determines whether the game session ends.

4. The method according to claim 1 or 2, characterized in that, The artificial intelligence AI model further includes: a feature extraction network; The step of invoking the artificial intelligence AI model to conduct a game session in a game program to obtain training data includes the following steps: Obtaining the reference game state from the game program; Invoking the feature extraction network to perform feature extraction on the reference game state to obtain reference game state features; Invoking the decision network to perform decision estimation on the reference game state features to obtain a decision result, where the decision result includes at least one probability value corresponding to at least one candidate action; executing the target game action with the largest probability value in the decision result; Invoking the value network to perform value estimation on the reference game state features to obtain the state value; Repeating the above steps to invoke the artificial intelligence AI model to conduct a game session in the game program to obtain the training data.

5. The method according to claim 4, wherein The value network includes k value branches corresponding to the k value classifications, and the j-th value branch is used to estimate the state sub-value of the game state on the j-th value classification; The step of invoking the value network to perform value estimation on the reference game state features to obtain the state value includes: Invoking the value network to output the k state sub-values of the reference game state features on the k value classifications through the k value branches, where k is an integer greater than 1.

6. The method according to claim 4, wherein The obtaining the reference game state from the game program includes: An image area is captured from the game interface of the game program as the reference game state.

7. The method according to claim 4, wherein The obtaining the reference game state from the game program includes: Capturing an image area from the game interface of the game program; The image region is subjected to image-like processing to obtain a simplified image, and the simplified image is used as the reference game state.

8. The method according to claim 4, wherein The obtaining the reference game state from the game program includes: Acquiring state data from the game program through a data interface, wherein the state data includes data of at least one game information; The vector obtained by splicing the state data is used as the reference game state.

9. The method according to claim 1 or 2, characterized in that, The step of training the artificial intelligence AI model according to the difference between the state value and the action value includes: The reinforcement learning algorithm is called to train the artificial intelligence AI model according to the difference between the state value and the action value.

10. The method according to claim 1 or 2, characterized in that The method further comprises: The artificial intelligence AI model is called to control the main virtual character in the game program to play the game.

11. A training device for an artificial intelligence (AI) model, characterized in that, The artificial intelligence AI model includes a value network and a decision network, and the device includes: A model module, used to call the artificial intelligence AI model to play a game in a game program to obtain training data, wherein the training data includes a reference game state in the game, a target game action output by the decision network according to the reference game state, and a state value output by the value network according to the reference game state, wherein the state value includes k state sub-values ​​on k value categories, where k is an integer greater than 1; the reference game state includes at least k game information, and the k value categories are divided according to the influence of the game information on the game, and the game information belonging to the same value category has the same influence attenuation trend; An acquisition module for acquiring the game states from time t0 to time t in the training data, where the reference game state is the game state at time t0, and time t is the end time of the game session, and n is a positive integer; n The game state at time t, where the reference game state is the game state at time t0, and time t n is the end time of the game session, and n is a positive integer; A calculation module, for the j-th value classification among the k value classifications, according to the game information belonging to the j-th value classification in the game state from the t0 moment to the t n moment, calculate the action sub-value of the target game action in the j-th value classification, where j is a positive integer less than or equal to k, and k is an integer greater than 1; repeat this step to calculate the k action sub-values of the target game action in the k value classifications; the action value includes the k action sub-values in the k value classifications; A training module is used to train the artificial intelligence AI model according to the difference between the state value and the action value.

12. A computer device, characterized in that, The computer device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the training method of the artificial intelligence AI model as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The readable storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the training method of the artificial intelligence AI model as described in any one of claims 1 to 10.

14. A computer program product, characterized in that, The computer program product includes computer instructions, which are stored in a computer-readable storage medium. The processor reads and executes the computer instructions from the computer-readable storage medium to implement the training method of the artificial intelligence AI model as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Self-playing model training method and device for multi-player battle game and computer equipment

    CN111111220A

  • Virtual robot training method and device, electronic equipment and medium

    CN111389010A