Action instruction generation method, training method and device for action decision-making model
By obtaining and splicing the private situation information of the participating objects and the target objects, more accurate action instructions are generated, which solves the problem of low accuracy of action instructions in the existing technology and improves the level of artificial intelligence.
Patent Information
- Application Number
- CN202111401276.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-24
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-11-24
AI Technical Summary
The existing teaching objects based on behavior tree AI have low accuracy in multi-participation objects and complex operation scenarios, resulting in a decrease in the level of artificial intelligence.
By obtaining the public situation information of multiple participants participating in the game game and the private situation information of the target object, input it into the action decision model for splicing and processing, and generating more accurate action instructions.
It improves the accuracy and rationality of action commands, improves the artificial intelligence level of the target object, and adapts to complex scenarios.
Smart Images

Figure CN114344912B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method for generating action instructions, a method for training an action decision-making model, and an apparatus therefor. Background Art
[0002] With the development of computer technology, electronic devices can implement more rich and vivid virtual scenarios. A virtual scenario refers to a digital scenario outlined by a computer through digital communication technology. Multiple participating objects can participate in a game session in the same virtual scenario. In the game session, the virtual character can be operated through the action instructions of the participating objects. For example, according to the character configuration instructions of the participating object, the virtual character owned by the participating object can be configured in the virtual scenario. Then, the game application providing the virtual scenario can automatically control the virtual character to interact and output the game result.
[0003] Currently, game applications providing virtual scenarios usually set up teaching objects based on Artificial Intelligence (AI) to automatically play game sessions with participating objects to achieve the effect of teaching or training the participating objects. In related technologies, generally, the corresponding action instructions of the teaching object are predicted based on a behavior tree AI. However, the behavior tree AI simply predicts action instructions according to preset rules. When the number of participating objects in the game session is large and the operations that the participating objects can perform on the virtual character are relatively complex, the accuracy of the predicted action instructions is low, resulting in a reduction in the artificial intelligence level of the teaching object. Summary of the Invention
[0004] The following is an overview of the subject matter described in detail in this document. This overview is not intended to limit the scope of protection of the claims.
[0005] Embodiments of the present invention provide a method for generating action instructions, a method for training an action decision-making model, and an apparatus therefor, which can improve the accuracy and rationality of the generated action instructions and improve the artificial intelligence level of the target object.
[0006] On the one hand, an embodiment of the present invention provides a method for generating action instructions, including:
[0007] Obtain the first object features of multiple participating objects participating in a game session, and obtain the second object features of a target object among the multiple participating objects. The first object features are used to represent the public situation information of the participating objects in the game session, and the second object features are used to represent the private situation information of the target object in the game session;
[0008] Input the first object feature and the second object feature into an action decision model, perform splicing processing on the first object feature and the second object feature to obtain a spliced object feature, and obtain a first label and a second label according to the spliced object feature. The first label is used to characterize the action type of the target object, and the second label is used to characterize the action object corresponding to the action type;
[0009] Generate a target action instruction for the target object according to the first label and the second label.
[0010] On the other hand, an embodiment of the present invention further provides a training method for an action decision model, including:
[0011] Obtain the first sample feature of multiple participating objects participating in a game session, and obtain the second sample feature of the target object among the multiple participating objects. The first sample feature is used to characterize the public situation information of the participating objects in the game session, and the second sample feature is used to characterize the private situation information of the target object in the game session;
[0012] Input the first sample feature and the second sample feature into an action decision model, perform splicing processing on the first sample feature and the second sample feature to obtain a spliced sample feature, and obtain a third label and a fourth label according to the spliced sample feature. The third label is used to characterize the action type of the target object, and the fourth label is used to characterize the action object corresponding to the action type;
[0013] Generate a training action instruction for the target object according to the third label and the fourth label;
[0014] Obtain the target result information after the target object executes the training action instruction, determine the target reward value of the target function according to the target result information, and correct the parameters of the action decision model according to the target reward value.
[0015] On the other hand, an embodiment of the present invention further provides an action instruction generation device, including:
[0016] An object feature acquisition module, configured to acquire the first object feature of multiple participating objects participating in a game session, and acquire the second object feature of the target object among the multiple participating objects. The first object feature is used to characterize the public situation information of the participating objects in the game session, and the second object feature is used to characterize the private situation information of the target object in the game session;
[0017] The first model processing module is configured to input the first object feature and the second object feature into an action decision model, perform splicing processing on the first object feature and the second object feature to obtain a spliced object feature, and obtain a first label and a second label according to the spliced object feature. The first label is used to represent the action type of the target object, and the second label is used to represent the action object corresponding to the action type.
[0018] The first instruction generation module is configured to generate a target action instruction for the target object according to the first label and the second label.
[0019] Further, a first area, a second area, and a third area are set in the virtual scene corresponding to the game session. The first object feature includes the first character feature of the virtual character of the participating object in the first area and the second area, and the second object feature includes the second character feature of the virtual character of the target object in the third area. The above object feature acquisition module is specifically configured to:
[0020] Determine the number of types of virtual characters in the game session, determine the number of slots for configuring the virtual characters in the first area, the second area, and the third area, determine the feature dimensions corresponding to the virtual characters in the first area, the second area, and the third area according to the number of types and the number of slots, and determine the first character feature and the second character feature according to the feature dimensions.
[0021] Alternatively, determine the number of types of virtual characters in the game session, determine the maximum number of characters that the participating object controls each virtual character, determine the feature dimensions of the virtual characters in the first area, the second area, and the third area according to the number of types and the maximum number of characters, and determine the first character feature and the second character feature according to the feature dimensions.
[0022] Further, the above first model processing module is specifically configured to:
[0023] Perform pooling processing on the first object features of other participating objects except the target object to obtain a third object feature;
[0024] Perform splicing processing on the third object feature, the first object feature of the target object, and the second object feature to obtain a spliced object feature.
[0025] Further, the above first model processing module is specifically configured to:
[0026] The first classifier based on the action decision model classifies the splicing object features to obtain a first label, and multiple second classifiers based on the action decision model classify the splicing object features to obtain multiple second labels;
[0027] Alternatively, the first classifier based on the action decision model classifies the splicing object features to obtain a first label, determines a target classifier from multiple second classifiers of the action decision model according to the first label, and classifies the splicing object features based on the target classifier to obtain a second label.
[0028] Further, the above-mentioned first model processing module is specifically configured to:
[0029] Classify the splicing object features by the first classifier based on the action decision model to obtain multiple candidate action types;
[0030] Determine the current running stage of the game session, and determine a target action type from the candidate action types according to the running stage to obtain a first label corresponding to the target action type.
[0031] Further, the number of the second labels is multiple, and the above-mentioned first model processing module is specifically configured to:
[0032] Determine a target action type according to the first label;
[0033] Determine a target label from multiple second labels according to the target action type, and determine an action object corresponding to the target action type according to the target label;
[0034] Generate a target action instruction for the target object according to the target action type and the action object corresponding to the target action type.
[0035] Further, the target action type includes selecting a virtual item, the action object includes a target virtual item, and the above-mentioned first instruction generation module is further configured to:
[0036] Determine a first target virtual character from the virtual characters controlled by the target object;
[0037] Generate a prop equipment instruction for the target object according to the matching relationship between the character attributes of the first target virtual character and the prop attributes of the target virtual item.
[0038] Further, the target action type includes selecting a virtual character configured in a first area, the action object includes a second target virtual character, and multiple candidate slots for configuring the virtual character are provided in the first area. The above-mentioned first instruction generation module is further configured to:
[0039] Randomly determine an initial slot from multiple candidate slots;
[0040] Determine a target formation from multiple preset configuration formations, where the target formation is used to represent the configuration slots of the second target virtual character;
[0041] Generate a character configuration instruction for the target object according to the initial slot and the target formation.
[0042] On the other hand, an embodiment of the present invention further provides a training device for an action decision model, including:
[0043] A sample feature acquisition module, configured to acquire first sample features of multiple participating objects participating in a game session, and acquire second sample features of a target object among the multiple participating objects. The first sample features are used to represent the public situation information of the participating objects in the game session, and the second sample features are used to represent the private situation information of the target object in the game session;
[0044] A second model processing module, configured to input the first sample features and the second sample features into an action decision model, perform splicing processing on the first sample features and the second sample features to obtain spliced sample features, and obtain a third label and a fourth label according to the spliced sample features. The third label is used to represent the action type of the target object, and the fourth label is used to represent the action object corresponding to the action type;
[0045] A second instruction generation module, configured to generate a training action instruction for the target object according to the third label and the fourth label;
[0046] A parameter correction module, configured to acquire target result information after the target object executes the training action instruction, determine a target reward value of a target function according to the target result information, and correct parameters of the action decision model according to the target reward value.
[0047] Further, the parameter correction module is specifically configured to:
[0048] Determine a first reward value corresponding to the target result information, where the target result information includes at least one of the remaining health value of the target object, the character level of the virtual character controlled by the target object, the interest value of the virtual currency of the target object, the ranking result of the target object in the current round, and the association attribute between the virtual characters controlled by the target object;
[0049] Obtain the target reward value of the target function according to the first reward value.
[0050] Further, the above parameter correction module is specifically used for at least one of the following:
[0051] Determine a first change value of the remaining health value of the target object at the end of the current round, and obtain the first reward value according to the first change value;
[0052] Alternatively, determine a second change value of the character level of the virtual character controlled by the target object in the current data frame, and obtain the first reward value according to the second change value;
[0053] Alternatively, determine a third change value of the interest value of the virtual currency of the target object at the end of the current round, and obtain the first reward value according to the third change value;
[0054] Alternatively, determine a first preset score corresponding to the ranking result of the target object in the current round, and obtain the first reward value by obtaining the first preset score;
[0055] Alternatively, determine a second preset score corresponding to the association attribute between the virtual characters controlled by the target object, and obtain the first reward value according to the second preset score.
[0056] Further, the above parameter correction module is specifically used for:
[0057] Obtain a second reward value and a third reward value output by the value model, where the second reward value is the reward value output by the value model after the target object executes the training action corresponding to the training action instruction, and the third reward value is the reward value output by the value model before the target object executes the training action;
[0058] Obtain a fourth reward value according to the sum of the first reward value and the second reward value;
[0059] Obtain the target reward value of the target function according to the difference between the fourth reward value and the third reward value.
[0060] Further, the number of the value models is multiple, and different value models correspond to different target result information. The above parameter correction module is specifically used for:
[0061] Obtain a fifth reward value corresponding to each target result information according to the difference between the fourth reward value and the third reward value corresponding to each target result information;
[0062] Perform weighted processing on the fifth reward value to obtain the target reward value of the target function.
[0063] On the other hand, an embodiment of the present invention further provides an electronic device, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned action instruction generation method or the training method of the action decision model is implemented.
[0064] On the other hand, an embodiment of the present invention further provides a computer-readable storage medium. The storage medium stores a program, and when the program is executed by a processor, the above-mentioned action instruction generation method or the training method of the action decision model is implemented.
[0065] On the other hand, an embodiment of the present invention further provides a computer program product. The computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device implements the above-mentioned action instruction generation method or the training method of the action decision model.
[0066] The embodiments of the present invention at least include the following beneficial effects: By obtaining the first object features of multiple participating objects in a game session, obtaining the second object features of a target object among the multiple participating objects, and inputting both the first object features and the second object features into the action decision model. Since the first object features are used to represent the public situation information of the participating objects in the game session, and the second object features are used to represent the private situation information of the target object in the game session, the situation state of the target object in the game session can be better and more comprehensively described, making the subsequent generated action instructions more accurate. Moreover, the first label and the second label generated by the action decision model are respectively used to represent the action type and the action object, which can achieve multi-level action prediction, is beneficial to reducing the prediction difficulty of complex actions, and further improves the accuracy of the generated action instructions. Therefore, the action instruction generation method provided by the embodiments of the present invention can better adapt to the scenario with a large number of participating objects and complex operations executable for virtual characters through the comprehensive and reasonable description of the features of the target object and the hierarchical action labels, thereby improving the accuracy and rationality of the generated action instructions and the artificial intelligence level of the target object. In addition, when training the action decision model, by obtaining the target result information to determine the target reward value of the objective function, the effect of reinforcement learning is achieved, and there is no need to perform the annotation of supervision labels, which is beneficial to improving the training efficiency of the action decision model.
[0067] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures specifically pointed out in the specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] The accompanying drawings are used to provide a further understanding of the technical solution of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the technical solution of the present invention, and do not constitute a limitation to the technical solution of the present invention.
[0069] Figure 1 It is a schematic diagram of an implementation environment provided for an embodiment of the present invention;
[0070] Figure 2 It is a schematic diagram of a game AI service architecture provided for an embodiment of the present invention;
[0071] Figure 3 It is a flowchart of a method for generating action instructions provided for an embodiment of the present invention;
[0072] Figure 4 It is a schematic diagram of the situation information corresponding to the first object feature and the second object feature provided for an embodiment of the present invention;
[0073] Figure 5 It is a schematic diagram of the main interface of a self-playing chess game provided for an embodiment of the present invention;
[0074] Figure 6 It is a schematic diagram of the interface of a purchase area provided for an embodiment of the present invention;
[0075] Figure 7 It is a schematic diagram of the structure of an action decision-making model provided for an embodiment of the present invention;
[0076] Figure 8 It is a schematic diagram of the hierarchical structure of the first label and the second label provided for an embodiment of the present invention;
[0077] Figure 9 It is a schematic diagram of the first classifier classifying and processing the spliced object features according to the running stage provided for an embodiment of the present invention;
[0078] Figure 10 It is a schematic diagram of the complete process of a method for generating action instructions provided for an embodiment of the present invention;
[0079] Figure 11 It is a flowchart of a method for training an action decision-making model provided for an embodiment of the present invention;
[0080] Figure 12 It is a schematic diagram of the prediction method of the server for each round provided for an embodiment of the present invention;
[0081] Figure 13 It is a schematic diagram of the complete process of a training method provided for an embodiment of the present invention;
[0082] Figure 14 It is a schematic diagram of the training framework of an action decision-making model provided for an embodiment of the present invention;
[0083] Figure 15 It is a schematic structural diagram of the action instruction generation device provided by the embodiment of the present invention;
[0084] Figure 16 It is a schematic structural diagram of the training device of the action decision-making model provided by the embodiment of the present invention;
[0085] Figure 17 It is a partial structural block diagram of the server provided by the embodiment of the present invention. Specific embodiments
[0086] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0087] Before further elaborating on the embodiments of the present invention, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are applicable to the following explanations:
[0088] Participating object: It may refer to a user participating in a game session, or it may refer to a game account participating in a game session.
[0089] Teaching object: An intelligent agent provided by a game application program, which can participate in a game session as a participating object and can independently execute action decisions.
[0090] Chessboard: The area in the self-play chess game session interface for preparing for and conducting battles, which can be any one of a two-dimensional virtual chessboard, a 2.5-dimensional virtual chessboard, and a three-dimensional virtual chessboard. The embodiments of the present invention do not make limitations. The chessboard is divided into a battle area and a preparation area. Among them, the battle area includes several battle slots of the same size, and the battle slots are used to configure the battle chess pieces during the battle; the preparation area includes several preparation slots, and the preparation slots are used to configure the preparation slots. The preparation chess pieces do not participate in the battle during the battle, but can be dragged and configured in the battle area during the preparation stage. Regarding the setting method of the slots in the battle area, in a possible implementation, the battle area includes n (rows) × m (columns) battle slots, where n is an integer multiple of 2, and the adjacent two rows of slots are aligned, or the adjacent two rows of slots are staggered. In addition, the battle area is evenly divided into two parts according to rows, namely the friendly battle area and the enemy battle area, and during the preparation stage, only chess pieces can be configured in the friendly battle area.
[0091] Virtual Character: A character in the game session controlled by the participating or teaching object. For example, in an auto chess game, the pieces placed on the chessboard are virtual characters. Correspondingly, virtual characters can include battle virtual characters and standby virtual characters. Among them, a battle virtual character is a virtual character located in the battle area, and a standby virtual character is a virtual character located in the standby area. Among them, virtual characters can be virtual pieces, virtual characters, virtual animals, anime characters, etc., and virtual characters can be displayed using 3D models. Among them, the position of the virtual character on the chessboard can be changed. In the preparation stage, the position of the battle virtual character in the battle area can be adjusted, the position of the standby virtual character in the standby area can be adjusted, the battle virtual character can be moved to the standby area (when there is an idle standby chess square in the standby area), or the standby virtual character can be moved to the battle area. It should be noted that in the battle stage, the position of the standby virtual character in the standby area can still be adjusted. In the battle stage, the position of the battle virtual character in the battle area is different from that in the preparation stage. For example, in the battle stage, the battle virtual character can automatically move from its own battle area to the enemy's battle area and interact with the enemy's battle virtual character; or the battle virtual character can automatically move within its own battle area. In addition, in the preparation stage, the battle virtual character can only be set in its own battle area, and the battle virtual characters set by the enemy are invisible on the chessboard. Regarding the acquisition method of virtual characters, in one possible implementation, during the game session, virtual characters can be purchased using virtual currency in the purchase area in the preparation stage. In addition, virtual characters can be upgraded, or synthesized. Multiple identical virtual characters can be upgraded to one virtual character, and the combat power value of the upgraded virtual character can be increased. And the participating or teaching object can discard the virtual characters it owns, and the discarded virtual characters will be placed in the abandonment area.
[0092] Attributes: Virtual characters all have their own attributes. For example, in an auto chess game, each piece has its own attributes, which can be: the camp to which the piece belongs (such as Alliance A, Alliance B, neutral faction, etc.), the occupation of the piece (such as warrior, shooter, mage, assassin, guard, swordsman, gunner, fighter, etc.), the attack type of the piece (such as magic, physical, etc.), the identity of the piece (such as noble, demon, elf, etc.), etc. The embodiments of the present invention do not make limitations.
[0093] Associated Attribute: An attribute relationship that can bring a buff effect to a virtual character, such as "synergy" in an auto chess game. In the battle area, when different chess pieces have an associated attribute (including different chess pieces having the same attribute or different chess pieces having complementary types of attributes), and the quantity reaches the quantity threshold, the chess pieces with this attribute or all the chess pieces in the battle area can obtain the buff effect corresponding to this attribute. For example, when there are 2 chess pieces with the attribute of "Warrior" in the battle area at the same time, all the chess pieces in the battle area obtain a 10% defense bonus; when there are 4 chess pieces with the attribute of "Warrior" in the battle area at the same time, all the chess pieces in the battle area obtain a 20% defense bonus; when there are 3 chess pieces with the attribute of "Elf" in the battle area at the same time, all the chess pieces in the battle area obtain a 20% dodge probability bonus.
[0094] Virtual Item: A virtual item used to enhance the combat power value of a virtual character, such as equipment in an auto chess game. When a chess piece carries equipment, its combat power value will be correspondingly enhanced according to the attributes of the equipment. The combat power value can include attack power, defense power, dodge rate, etc.
[0095] Population: A game term used to represent the number of virtual characters controlled by a participating entity. Each virtual character can occupy one or more populations.
[0096] Talent: A buff effect carried by a participating entity in a game round, which can be obtained by upgrading the level of the participating entity using virtual currency.
[0097] Formation: The formation formed by each virtual character, used to indicate the positions of each virtual character configured in the virtual scene, such as indicating the positions of chess pieces configured in the battle area in an auto chess game.
[0098] Interest of Virtual Currency: In a game round, the system distributes a certain amount of virtual currency to a participating entity according to the quantity of virtual currency currently owned by the participating entity.
[0099] Artificial Intelligence (AI) is the theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence is also the study of the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning and decision-making.
[0100] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, autonomous driving, and intelligent transportation.
[0101] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0102] Reinforcement Learning (RL), also known as reward learning, evaluation learning, or enhancement learning, is one of the paradigms and methodologies of machine learning, used to describe and solve the problem of an intelligent agent maximizing rewards (or incentives) or achieving specific goals through learning strategies during the interaction with the environment.
[0103] Supervised Learning (SL) is the process of using a set of samples with known categories to adjust the parameters of a classifier to achieve the required performance, also known as supervised training or learning with a teacher.
[0104] With the development of computer technology, electronic devices can implement more rich and vivid virtual scenarios. A virtual scenario refers to a digital scenario outlined by a computer through digital communication technology. Multiple participating objects can participate in a game session in the same virtual scenario. In the game session, the virtual character can be operated through the action instructions of the participating objects. For example, the virtual character owned by the participating object can be configured in the virtual scenario according to the role configuration instructions of the participating object. Then, the game application that provides the virtual scenario can automatically control the virtual character to interact and output the game result.
[0105] Taking the currently common auto chess game as an example, the chessboard in the auto chess game belongs to a virtual scenario. The participating object can configure the virtual character it owns in the virtual scenario, and the application of the auto chess game can automatically control the virtual character to interact and output the game result.
[0106] At present, game applications that provide virtual scenarios usually set up teaching objects based on Artificial Intelligence (AI) to automatically play game rounds with participating objects, so as to achieve the effect of teaching or training the participating objects. In related technologies, generally, a behavior tree AI is used to predict the corresponding action instructions of the teaching object. However, the behavior tree AI simply predicts action instructions according to preset rules. When the number of participating objects in the game round is large and the operations that the participating objects can perform on the virtual character are complex, the accuracy of the predicted action instructions is low, resulting in a decrease in the artificial intelligence level of the teaching object.
[0107] For example, there are usually many participating objects in a self-chess game, generally up to 8. In the game round, different participating objects will perform different operations on the virtual characters they already have according to their current situation. For the teaching object, it will predict the next operation to be performed on the virtual character. Moreover, the types of operations that can be performed on the virtual character are also diverse, such as buying chess pieces, discarding chess pieces, upgrading chess pieces, buying equipment, and so on.
[0108] Table 1 Comparison of the complexity of different games
[0109]
[0110] Referring to Table 1, Table 1 shows the comparison of the complexity of different games. For currently common Go or card games, the number of participating objects in the game round, the number of cards involved, the types of decision-making actions, and the number of participating objects to make decisions per round are all less than those in the self-chess game. Their state complexity is much lower than that of the self-chess game. It can be seen that for games of the self-chess type, a higher artificial intelligence level is required for the teaching object, and the artificial intelligence level of the teaching object directly affects the teaching or training effect of the teaching object on the participating objects. However, the behavior tree AI in related technologies cannot meet the requirements of a high artificial intelligence level.
[0111] Based on this, the embodiments of the present invention provide a method for generating action instructions, a method for training an action decision-making model, and a device, which can improve the accuracy and rationality of the generated action instructions and improve the artificial intelligence level of the target object.
[0112] Referring to Figure 1 , Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present invention. The implementation environment includes a terminal 101 and a server 102. Among them, the terminal 101 and the server 102 are connected through a communication network 103.
[0113] The server 102 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0114] In addition, the server 102 can also be a node server in a blockchain network.
[0115] The terminal 101 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication means, and the embodiments of the present invention do not limit this here.
[0116] The solution provided by the embodiments of the present invention can be applied to various technical fields, including but not limited to technical fields such as cloud technology and artificial intelligence. In the embodiments of the present invention, the auto chess game is used as an example for detailed description.
[0117] Refer to Figure 2 , Figure 2 which is a schematic diagram of the game AI service architecture provided by the embodiments of the present invention. Figure 1 In the shown terminal 101, a game application program (or called a game client) can be installed. Users can use the game application program for game teaching or training. When teaching or training, the terminal will send the frame data in the game round to the game AI service. The frame data is the situation data in the game round and is used to represent the situation information of each participating object in the game round. For example, the remaining health value of the participating object, the number of chess pieces, the level of the chess pieces, the total number of virtual currencies, the equipment of the chess pieces, etc., are not listed one by one here. After receiving the frame data sent by the terminal, the game AI service generates corresponding action instructions according to the frame data and returns them to the terminal. The teaching object in the game application program can then automatically execute corresponding interaction actions according to the action instructions. It can be understood that the above game AI service can be deployed in Figure 1 the shown server 102.
[0118] Based on Figure 1 the shown implementation environment and Figure 2 the shown game AI service architecture, the embodiments of the present invention provide a method for generating action instructions. This method for generating action instructions can be executed by Figure 1 the shown server 102, or can be executed by Figure 1 the shown terminal 101, or by Figure 1The terminal 101 and the server 102 shown cooperate to execute. In the embodiments of the present invention, this action instruction generation method is described by taking the execution by the Figure 1 server 102 shown as an example.
[0119] Refer to Figure 3 , Figure 3 which is a flowchart of the action instruction generation method provided by the embodiments of the present invention. The action instruction generation method includes but is not limited to the following steps 301 to 303.
[0120] Step 301: Obtain the first object features of multiple participating objects participating in the game session, and obtain the second object features of the target object among the multiple participating objects;
[0121] Among them, the first object features are used to represent the public situation information of the participating objects in the game session. The public situation information includes the situation information of each participating object that is pushed or displayed to all participating objects, and each participating object can obtain it. Taking the auto chess game as an example, the public situation information may include information such as the pieces, synergies, equipment, talents, health points, virtual currency quantity, and whether being eliminated of the participating objects.
[0122] The target object is the object for which an action instruction needs to be generated currently. The target object can be a teaching object, and the target object is one of the participating objects.
[0123] The second object features are used to represent the private situation information of the target object in the game session. The private situation information is the situation information extracted additionally for the target object. Moreover, the private situation information includes the situation information of the target object that is not pushed or displayed to other participating objects, and only the target object can obtain it. Taking the auto chess game as an example, the private situation information may include information such as the pieces information in the purchase area, the pieces information in the discard area, and the upgrade cost of the target object. These private situation information belong to the objective type of situation information. In addition, estimated situation information can be further added, such as the number of missing pieces for synergy upgrade, the number of missing pieces for piece upgrade, etc., so that the private situation information included in the second object features can be more abundant.
[0124] In addition, the above private situation information can further include the game session status information, so as to further enrich the private situation information included in the second object features. For example, the game session status information may include information such as the current round number, the running stage, and the remaining operation time. It can be understood that these game session status information are the situation information depicted based on the game session.
[0125] It can be seen that from the perspective of the target object, by obtaining the first object feature and the second object feature, the situation state of the target object in the game can be better depicted, the comprehensiveness of the features is higher, and inputting both the first object feature and the second object feature into the action decision-making model can make the subsequent generated action instructions more accurate.
[0126] The following takes the auto chess game as an example to elaborate on the above-mentioned first object feature and second object feature in detail.
[0127] Refer to Figure 4 , Figure 4 is a schematic diagram of the situation information corresponding to the first object feature and the second object feature provided by the embodiment of the present invention. Among them, the first object feature may include chess piece information, synergy information, and other information. The chess piece information corresponding to the first object feature may include the attributes of the chess pieces in the battle area and the slots in the preparation area, etc.; the synergy information may include each synergy level, etc.; the other information corresponding to the first object feature may include equipment, talents, health points, population, etc. The second object feature may include chess piece information, discard area information, and other information. The chess piece information corresponding to the second object feature includes the attributes of the chess pieces in the purchase area slots, etc., the discard area information includes the discarded chess piece information of different professional camps, etc., and the other information corresponding to the second object feature includes the upgrade cost of the target object, the inventory quantity in the chess piece library, the round, etc.
[0128] Refer to Figure 5 , Figure 5 is a schematic diagram of the main interface of the auto chess game provided by the embodiment of the present invention. This interface displays a battle area 501, a preparation area 502, a population information area 503, a virtual currency information area 504, a synergy information area 505, an object information area 506, and a game state information area 507. Among them, up to 9 chess pieces can be configured in the battle area 501, and up to 8 chess pieces can be configured in the preparation area 502. The camp, profession, etc. information of the corresponding chess pieces can be displayed in both the battle area 501 and the preparation area 502. For example, "Guardian", "Wu", etc. represent the camp of the chess pieces, and "Archer", "Mage", etc. represent the profession of the chess pieces. The population information area 503 is used to display the current population information and the upgrade cost of the target object. The virtual currency information area 504 is used to display the quantity of virtual currency currently held by the target object. The synergy information area 505 is used to display the synergies triggered by the chess pieces in the battle area of the target object. The object information area 506 is used to display the information of other participating objects, such as the quantity of virtual currency currently held by other participating objects and the remaining health points, etc. The game state information area 507 is used to display information such as the current round, operation stage, remaining operation time, etc. It can be understood that Figure 5 All the various situation information shown can be used to obtain the above-mentioned first object feature.
[0129] Refer toFigure 6 , Figure 6 This is a schematic diagram of the interface of the purchase area provided by the embodiment of the present invention. Five candidate chess pieces are provided in the purchase area for the target object to purchase. Moreover, the attributes of each candidate chess piece, such as camp, occupation, and price, etc., will be displayed in the purchase area 601. In addition, a refresh button 602 is also set in the purchase area 601. The refresh button 602 is used to refresh the types of candidate chess pieces provided in the purchase area 601. Correspondingly, the quantity of virtual currency spent on the refresh will also be displayed in the purchase area 601. It can be understood that Figure 6 All kinds of situation information shown can be used to obtain the above-mentioned second object feature.
[0130] It can be understood that the specific situation information represented by the first object feature and the second object feature is only for exemplary illustration, and the embodiments of the present invention do not make limitations.
[0131] Step 302: Input the first object feature and the second object feature into the action decision model, perform splicing processing on the first object feature and the second object feature to obtain a spliced object feature, and obtain a first label and a second label according to the spliced object feature;
[0132] In a possible implementation manner, the above-mentioned first object feature is represented by a first feature vector, the second object feature is represented by a second feature vector, and the spliced object feature is represented by a spliced feature vector. The action decision model performs splicing processing on the first object feature and the second object feature to obtain a spliced object feature, which may be to perform head-to-tail splicing on the first feature vector and the second feature vector to obtain a spliced feature vector.
[0133] Specifically, performing splicing processing on the first object feature and the second object feature to obtain a spliced object feature may be to perform pooling processing on the first object feature of other participating objects except the target object to obtain a third object feature, and perform splicing processing on the third object feature, the first object feature of the target object, and the second object feature to obtain a spliced object feature. Correspondingly, the third object feature is represented by a third feature vector.
[0134] For example, after the first feature vector and the second feature vector are input into the action decision model, they are first processed through multiple layers of fully connected layers, referring to Figure 7 , Figure 7 This is a schematic diagram of the structure of the action decision model provided by the embodiment of the present invention. Exemplarily, the action decision model first processes the first feature vector and the second feature vector through two layers of fully connected layers. Among them, the weights of the first two layers of fully connected layers of the action decision model are shared for all participating objects, mainly used for performing feature space transformation on the first feature vector and the second feature vector to achieve the effect of information extraction and integration.
[0135] Moreover, the action decision-making model further processes the first feature vector through the third fully connected layer to perform a feature space transformation on the first feature vector. Among them, the third fully connected layer differentiates weights for the target object and other participating objects. For example, the weight of the target object can be greater than that of other participating objects, making the overall feature characterization effect of the feature vector more reasonable.
[0136] Furthermore, the action decision-making model performs pooling processing on the first feature vectors of other participating objects except the target object through the max pooling layer to obtain the third feature vector, thereby achieving the effect of feature screening, eliminating some unimportant feature information, and making the feature vector more accurate.
[0137] Next, the action decision-making model concatenates the third feature vector, the first feature vector processed by the fully connected layer, and the second feature vector through the concatenation layer to obtain the concatenated feature vector.
[0138] Among them, the action decision-making model can process the concatenated feature vector through the fully connected layer again to perform a feature space transformation on the concatenated feature vector, achieving the effect of information extraction and integration.
[0139] In addition, after processing the concatenated feature vector through the fully connected layer again, it can further process the concatenated feature vector through at least one of the LSTM (Long Short-Term Memory) layer or the attention layer. The purpose of processing through the LSTM layer is to introduce historical features, making the number of features more and the feature expression more abundant; while the purpose of processing through the attention layer is to promote feature interaction, which can also make the feature expression more abundant.
[0140] Finally, the concatenated feature vector processed by the fully connected layer can be classified based on the classifier of the action decision-making model (which can be implemented by the fully connected layer) to obtain the first label and the second label. Among them, the first label is used to represent the action type of the target object, and the second label is used to represent the action object corresponding to the action type.
[0141] The first label and the second label can be represented in the form of vectors. The specific number of dimensions of the vector can be set according to the actual situation. Taking the first label as an example, assuming that there are three types of action types in the game session, including "purchase virtual character", "configure virtual character", and "select virtual item", then the vector dimension corresponding to the first label is three-dimensional, and each dimension of the vector element represents an action type. When the action type corresponding to the first label is "purchase virtual character", the vector element can be assigned in the corresponding dimension. Similarly, the second label can also be vectorized in the above manner. For example, if there are five purchasable virtual characters in the game session, then the vector dimension of the second label is five-dimensional.
[0142] It can be understood that the above examples are only used to exemplarily illustrate the vector dimensions of the first label and the second label and the specific vector expression principle. The vector expressions of the first label and the second label can be determined according to the actual situation of the game session, and the embodiments of the present invention do not make any limitations.
[0143] The classifier classifies the spliced feature vector processed by the fully connected layer, that is, calculates the probabilities of different action types or action objects, and finally selects the action type or action object with the highest probability for output. For example, taking the first label as an example, assuming there are three action types, including "purchase virtual character", "configure virtual character", and "select virtual item", and the probability of "purchase virtual character" is the highest, then the first label finally output can be [1, 0, 0].
[0144] Specifically, the classifier of the action decision model may include a first classifier and multiple second classifiers. The first classifier is used to output the above first label, for example Figure 7 the classifier corresponding to the "action type" as shown; the second classifier is used to output the above second label, and different second classifiers correspond to different action types, for example Figure 7 the classifier corresponding to "purchase virtual character" and the classifier corresponding to "configure virtual character to the battle area" as shown. It should be added that Figure 7 the classifier corresponding to the "reward average" in is used to train the action decision model, and its specific function will be described in detail in the embodiments of the subsequent training method of the action decision model.
[0145] Among them, there are at least the following two situations where the action decision model obtains the first label and the second label according to the spliced object features:
[0146] One situation is that the first classifier of the action decision model classifies the spliced object features to obtain the first label, and the multiple second classifiers of the action decision model classify the spliced object features to obtain multiple second labels. At this time, the first classifier and the multiple second classifiers form parallel tasks, separately process the spliced feature vector and output the corresponding first label or second label.
[0147] It can be seen that in this situation, the action decision model outputs the first label and multiple second labels. Subsequently, when generating the target action instruction of the target object according to the first label and the second label, the corresponding second label is filtered out according to the first label.
[0148] Another situation is that the first classifier of the action decision model classifies the spliced object features to obtain the first label, determines the target classifier from the multiple second classifiers of the action decision model according to the first label, and classifies the spliced object features based on the target classifier to obtain the second label.
[0149] It can be seen that in this case, the action decision model only outputs a first label and a second label, that is, the selection of the second label corresponding to the action type from multiple second labels is performed by the action decision model.
[0150] Step 303: Generate a target action instruction for the target object according to the first label and the second label.
[0151] In a possible implementation manner, the terminal continuously sends frame data to the server. For example, the terminal can send frame data to the server every 0.2 seconds. The server can generate the first label and the second label according to a preset frequency. For example, the server can generate the first label and the second label every 5 frames.
[0152] Among them, if the action decision model outputs a first label and multiple second labels, then when the server generates a target action instruction for the target object according to the first label and the second label, it specifically determines the target action type according to the first label, determines the target label from multiple second labels according to the target action type, determines the action object corresponding to the target action type according to the target label, and generates a target action instruction for the target object according to the target action type and the action object corresponding to the target action type.
[0153] For example, if the target action type corresponding to the first label is "purchase virtual character", the corresponding classifier can be determined according to this target action type, and the second label output by the classifier is the target label. Suppose the action object corresponding to the target label is "virtual character A", then the action object corresponding to the target action type of "purchase virtual character" is "virtual character A".
[0154] Then, obtaining the first label and the second label is equivalent to determining the specific action to be performed by the target object. Based on the above example, it can be determined that the action to be performed by the target object according to the first label and the second label is "purchase virtual character A". At this time, the server can generate a target action instruction corresponding to "purchase virtual character A" and send the target action instruction to the terminal, and the teaching object in the game application can perform the corresponding action according to the target action. In the above manner, the server continuously generates corresponding target action instructions according to the first object feature and the second object feature and sends the target action instructions to the terminal, so that the teaching object in the terminal can automatically conduct a game session to achieve the effect of teaching or training other participating objects.
[0155] In the related art, if a label is directly generated through a model, for example, the model directly outputs "purchase virtual character A", and a corresponding label is set for all possible actions, that is, the one-hot encoding method is adopted, then the total number of labels that can be generated is the product of the total number of action types and the total number of action objects. It can be seen that the total number of labels in this way is relatively large, which increases the difficulty of complex action prediction. In the embodiments of the present invention, the first label and the second label generated by the action decision model are respectively used to represent the action type and the action object, which can realize multi-level action prediction, is beneficial to reducing the difficulty of complex action prediction, and further improves the accuracy of the generated action instruction. Moreover, through the parallel classification processing of the first classifier and multiple second classifiers, it is beneficial to improve the output efficiency of the labels.
[0156] Therefore, based on the above steps 301 to 303, the action instruction generation method provided by the embodiments of the present invention can better adapt to the scenario where the number of participating objects is large and the operations executable on the virtual character are relatively complex through the comprehensive and reasonable description of the features of the target object and the hierarchicalization of the action labels, thereby improving the accuracy and rationality of the generated action instruction and the artificial intelligence level of the target object.
[0157] Based on Figure 2 In the service architecture shown, in a possible implementation manner, after receiving the frame data sent by the terminal, the game AI service first performs vectorization processing on the received frame data, and then obtains the first object feature and the second object feature. The above frame data is the original data of the above public situation information and private situation information. It can be understood that the vectorization processing of the above frame data can actually be completed on the terminal side, that is, the first object feature and the second object feature obtained after the vectorization processing of the frame data are sent to the game AI service side by the terminal.
[0158] Among them, the vectorization processing of the received frame data can be based on a preset vector format. The vector format can segment the vector according to different frame data types, and finally the received frame data can be mapped to the corresponding segment to obtain the vector corresponding to the frame data.
[0159] Further, a first area, a second area, and a third area may be set in the virtual scene corresponding to the game session. The first object feature includes the first character feature of the virtual characters of the participating objects in the first area and the second area. The second object feature includes the second character feature of the virtual character of the target object in the third area. Correspondingly, taking the auto chess game as an example, the above-mentioned first area is the battle area, the second area is the preparation area, and the third area is the purchase area. The chess pieces in different areas (the battle area, the preparation area, and the purchase area) have corresponding features. The first character feature is one of the first object features, correspondingly the features of the chess pieces in the battle area and the preparation area. The second character feature is one of the second object features, correspondingly the features of the chess pieces in the purchase area.
[0160] In the embodiments of the present invention, there are at least two different ways to characterize the above-mentioned first character feature and second character feature:
[0161] One way is to characterize based on slots, that is, to determine the above-mentioned first character feature and second character feature according to the types of virtual characters corresponding to different slots. Specifically, the number of types of virtual characters in the game session can be determined, the number of slots for configuring virtual characters in the first area, the second area, and the third area can be determined, the feature dimensions corresponding to the virtual characters in the first area, the second area, and the third area can be determined according to the number of types and the number of slots, and the first character feature and the second character feature can be determined according to the feature dimensions.
[0162] Taking the auto chess game as an example, assuming that the number of types of chess pieces in the game session is 60, the number of slots in the battle area is 9, the number of slots in the preparation area is 8, and the number of slots in the purchase area is 5, then the vector dimension of the first character feature of the chess pieces in the battle area is 9 * 60, the vector dimension of the first character feature of the chess pieces in the preparation area is 8 * 60, and the vector dimension of the second character feature of the chess pieces in the purchase area is 5 * 60. After determining the feature dimensions of the first character feature and the second character feature, the vector elements corresponding to the first character feature and the second character feature can be determined according to the actual frame data, and then the vector expressions of the first character feature and the second character feature can be determined.
[0163] The above-mentioned way of characterizing based on slots can more finely distinguish the virtual characters in different areas, so that the first character feature and the second character feature of the virtual characters can be more finely characterized, making the state space of the first character feature and the second character feature larger and the feature expression more abundant.
[0164] Another way is to characterize based on the virtual character itself, that is, to determine the above first character feature and second character feature according to the type of virtual character controlled by the participating object. Specifically, the number of types of virtual characters in the game session can be determined, the maximum number of characters of each virtual character controlled by the participating object can be determined, the feature dimensions of the virtual characters in the first area, the second area and the third area can be determined according to the number of types and the maximum number of characters, and the first character feature and the second character feature can be determined according to the feature dimensions.
[0165] Taking the auto chess game as an example, in this way, the area where the chess piece is located is no longer considered. For a certain type of chess piece, assuming that the number of types of chess pieces in the game session is 60 and the maximum number of chess pieces that the participating object can have is 9, then the first character feature of the chess pieces in the battle area and the preparation area and the second character feature of the chess pieces in the purchase area can all have a feature dimension of 9 * 60. The chess pieces in the battle area, the preparation area and the purchase area all belong to the chess pieces that the participating object can control. Therefore, the fact that the chess pieces are in different areas will not affect the maximum number of characters of each chess piece controlled by the participating object.
[0166] The above way of characterizing based on the virtual character itself can reduce the feature changes caused by the change of the area where the virtual character is located, make the state space of the first character feature and the second character feature smaller, can improve the processing efficiency of the subsequent action decision model for the first character feature and the second character feature, and improve the training efficiency of the subsequent action decision model.
[0167] It can be understood that the above two ways of characterizing the first character feature and the second character feature each have their own advantages, and the corresponding way of characterizing features can be selected according to the actual situation in application.
[0168] In a possible implementation, there are different running stages in the game session, and the action types corresponding to different running stages will also be different. Taking the auto chess game as an example, referring to Figure 8 , Figure 8Schematic diagram of the hierarchical structure of the first label and the second label provided by the embodiments of the present invention. The running stages of the game round may at least include a cultivation stage and a battle stage. In the cultivation stage, the possible action types that may be executed may include: purchasing chess pieces, selling chess pieces, selecting talents, selecting equipment, resetting chess pieces (reconfiguring the chess pieces in the discard area back into the chess piece library, and the chess pieces in the chess piece library will reappear in the purchase area for the target object to purchase), recasting talents, refreshing chess pieces (refreshing the chess pieces in the purchase area), upgrading the level of the target object (affecting the types of chess pieces that appear in the purchase area or affecting the number of chess pieces that can be configured in the battle area), stopping actions, and so on. In the battle stage, the possible action types that may be executed may include: configuring chess pieces to the battle area, configuring chess pieces to the preparation area, selecting a formation, selecting key chess pieces, stopping actions, and so on.
[0169] Among them, the action types that may be executed in the cultivation stage and the battle stage include "stopping actions", enabling the target object to automatically stop executing actions, reducing the occurrence of repeated execution of actions, and improving the rationality of actions. In addition, an "empty" action type ( Figure 8 not shown) may also be added in the cultivation stage and the battle stage. An action type of "empty" means that the target object does not execute any actions. By setting the "empty" action type, the target object can be more reasonable when actually executing actions, and the comprehensiveness of the action types is higher.
[0170] It can be seen that Figure 8 the "cultivation actions" and "battle actions" in correspond to the first label and are used to determine the corresponding action types. The "cultivation actions" may be 10-dimensional vectors, and the "battle actions" may be 6-dimensional vectors; the rest, such as "purchasing chess pieces" and "configuring chess pieces to the battle area", etc., correspond to the second label and are used to determine the corresponding action objects. "Purchasing chess pieces" may be a 61-dimensional vector, and "configuring chess pieces to the battle area" may be a 61-dimensional vector, and so on. It can be understood that the dimensions of each first label and second label can be set according to the actual situation of the game round, and the embodiments of the present invention do not make limitations.
[0171] On this basis, when classifying the spliced object features by the first classifier of the action decision model to obtain the first label, the spliced object features may first be classified by the first classifier of the action decision model to obtain multiple candidate action types, the running stage in which the game round currently is located is determined, and the target action type is determined from the candidate action types according to the running stage, and the first label corresponding to the target action type is obtained.
[0172] It should be added that the running stage in which the game round currently is located may be included in the frame data sent by the terminal.
[0173] Specifically, referring to Figure 9, Figure 9 This is a schematic diagram of the first classifier provided by the embodiment of the present invention for classifying the stitching object features according to the operation stage. After the first classifier classifies the stitching object features, probabilities corresponding to different candidate action types will be obtained. Different candidate action types can be pre-grouped according to the operation stage corresponding to the game round. When the first classifier outputs the first label, it selects the candidate action type with the highest probability in the corresponding group according to the current operation stage as the target action type, and then outputs the corresponding first label.
[0174] For example, if the current operation stage of the current game round is operation stage 1, and the action type with the highest probability in operation stage 1 is action type 2, then the target action type finally output by the first classifier is action type 2.
[0175] In a possible implementation manner, the operation stage corresponding to the game round can be judged by a preset duration. For example, the first M seconds of a certain round is an operation stage, and the next N seconds is an operation stage. M and N are constants. M can be equal to N. The values of M and N can be set according to the actual situation of the game round, and the embodiments of the present invention do not make limitations.
[0176] In a possible implementation manner, the target action type may include selecting a virtual item, and the action object may include the target virtual item. After the server generates an action instruction to select the target virtual item, the target object will select the target virtual item according to the corresponding action instruction. The selected target virtual item needs to be equipped on the corresponding virtual character. When the number of target virtual items is multiple and the number of virtual characters controlled by the target object is also multiple, then equipping the target virtual items on the corresponding virtual characters belongs to a many-to-many problem.
[0177] In the embodiment of the present invention, the first target virtual character can be determined from the virtual characters controlled by the target object, and the item equipment instruction of the target object can be generated according to the matching relationship between the character attribute of the first target virtual character and the item attribute of the target virtual item.
[0178] Specifically, there are at least the following two ways to determine the first target virtual character from the virtual characters controlled by the target object:
[0179] One way is that the first target virtual character is output by the action decision model. Taking the auto chess game as an example, referring to Figure 8 , "selecting the key chess piece" in the battle stage corresponds to the first target virtual character. The server can combine the first target virtual character output by the action decision model and generate the item equipment instruction of the target object according to the matching relationship between the character attribute of the first target virtual character and the item attribute of the target virtual item.
[0180] Another way is that the server directly determines the first target virtual character based on the frame data sent by the terminal. For example, the first target virtual character can be determined according to the virtual character level, virtual character combat power, and so on.
[0181] In a possible implementation manner, the target action type may also include selecting a virtual character configured in the first area, and the action object may also include the second target virtual character. There are multiple candidate slots for configuring virtual characters in the first area. After the server generates an action instruction to select a virtual character configured in the first area, the target object will select the second target virtual character according to the corresponding action instruction, and the selected second target virtual character needs to be configured in the first area. When the number of second target virtual characters is multiple and the number of candidate slots in the first area is also multiple, then configuring the second target virtual characters in the corresponding candidate slots in the first area also belongs to a many-to-many problem.
[0182] In the embodiments of the present invention, an initial slot can be randomly determined from multiple candidate slots, a target formation can be determined from a preset multiple configuration formations, and a role configuration instruction for the target object can be generated according to the initial slot and the target formation.
[0183] Among them, the target formation is used to represent the configuration slots of the second target virtual character. When configuring the second target virtual character, the second target virtual character can be first configured in the randomly determined initial slot, and then the slot of the second target virtual character can be adjusted according to the target formation.
[0184] Specifically, there are at least the following two ways to determine the target formation from a preset multiple configuration formations:
[0185] One way is that the target formation is output by an action decision model. Taking the auto chess game as an example, referring to Figure 8 , the "select formation" in the battle stage corresponds to the target formation. The server can combine the target formation output by the action decision model and generate a role configuration instruction for the target object according to the initial slot and the target formation.
[0186] Another way is that the server determines the target formation from a preset multiple configuration formations according to the role attributes of the second target virtual character. For example, the role attribute of the second target virtual character can be the class type. If most of the second target virtual characters are virtual characters of the mage class, then a target formation matching the mage is selected.
[0187] Referring to Figure 10 , Figure 10 FIG.
[0188] The terminal collects frame data in the game session and sends the frame data to the server;
[0189] The server performs vectorization processing on the frame data to obtain the first object features of multiple participating objects participating in the game session and the second object features of the target object;
[0190] The server inputs the first object features and the second object features into an action decision model, performs splicing processing on the first object features and the second object features to obtain spliced object features, performs classification processing on the spliced object features based on the first classifier of the action decision model to obtain multiple candidate action types, determines the current running stage of the game session, determines the target action type from the candidate action types according to the running stage, obtains the first label corresponding to the target action type, and performs classification processing on the spliced object features based on multiple second classifiers of the action decision model to obtain multiple second labels;
[0191] The server determines the target action type according to the first label, determines the target label from the multiple second labels according to the target action type, determines the action object corresponding to the target action type according to the target label, and generates a target action instruction for the target object according to the target action type and the action object corresponding to the target action type;
[0192] The server sends the target action instruction to the terminal;
[0193] The target object in the game application program running on the terminal executes the corresponding action according to the target action instruction.
[0194] The action instruction generation method provided by the embodiments of the present invention can better adapt to the scenario where the number of participating objects is large and the operations executable for virtual characters are complex through the comprehensive and reasonable characterization of the features of the target object and the hierarchicalization of action labels, thereby improving the accuracy and rationality of the generated action instructions and the artificial intelligence level of the target object.
[0195] The action instruction generation method provided by the embodiments of the present invention can be used not only for the teaching object in the game application program to predict action instructions, but also for the balance test or fault (bug) test of the game application program according to the generated target action instruction, thereby improving the test effect of the game application program.
[0196] The following details the training process of the action decision model provided by the embodiments of the present invention. Refer to Figure 11 , Figure 11 is the flowchart of the training method of the action decision model provided by the embodiments of the present invention. This training method can be executed by the server 102 shown in Figure 1 , or can be executed by the terminal 101 shown in Figure 1 , or by Figure 1The terminal 101 and the server 102 shown cooperate to execute. In this embodiment of the present invention, this training method is Figure 1 illustrated by taking the execution by the server 102 shown as an example. This training method includes but is not limited to the following steps 1101 to step 1104.
[0197] Step 1101: Obtain the first sample features of multiple participating objects participating in the game session, and obtain the second sample features of the target object among the multiple participating objects;
[0198] Step 1102: Input the first sample features and the second sample features into the action decision model, perform splicing processing on the first sample features and the second sample features to obtain spliced sample features, and obtain the third label and the fourth label according to the spliced sample features;
[0199] Step 1103: Generate a training action instruction for the target object according to the third label and the fourth label;
[0200] Step 1104: Obtain the target result information after the target object executes the training action instruction, determine the target reward value of the target function according to the target result information, and correct the parameters of the action decision model according to the target reward value.
[0201] Among them, the first sample features are used to represent the public situation information of the participating objects in the game session, the second sample features are used to represent the private situation information of the target object in the game session, the third label is used to represent the action type of the target object, and the fourth label is used to represent the action object corresponding to the action type. The principles of steps 1101 to 1103 are similar to steps 301 to 303 of the above action instruction generation method, and will not be elaborated here.
[0202] In step 1104, the target result information is the situation information generated after the target object executes the training action instruction. In this embodiment of the present invention, the action decision model is trained in a reinforcement learning manner. In this embodiment of the present invention, the target function is obtained based on the reward function, the target reward value of the target function is determined according to the target result information, and the parameters of the action decision model are corrected according to the target reward value. Among them, the parameters of the action decision model can be the weights of the fully connected layer, the number of layers of the fully connected layer, and so on.
[0203] The training method provided by this embodiment of the present invention, when training the action decision model, determines the reward value of the target function by obtaining the target result information, achieves the effect of reinforcement learning, and does not require the annotation of supervised labels, which is beneficial to improving the training efficiency of the action decision model.
[0204] In a possible implementation, the target result information may include at least one of the remaining health value of the target object, the character level of the virtual character controlled by the target object, the interest value of the virtual currency of the target object, the win / loss result of the target object in the current round, and the association attribute between the virtual characters controlled by the target object. Determining the target reward value of the target function according to the target result information may specifically be determining the first reward value corresponding to the target result information, and obtaining the target reward value of the target function based on the first reward value.
[0205] Wherein, when the number of pieces of target result information is one, the number of first reward values is also one correspondingly. In this case, the first reward value can be directly used as the target reward value of the target function.
[0206] When the number of pieces of target result information is at least two, the number of first reward values is also at least two correspondingly. In this case, the at least two first reward values can be weighted to obtain the target reward value of the target function. The weights of different first reward values can be set according to the actual situation, or the weights of different first reward values can all be 1, that is, directly summing the at least two first reward values to obtain the target reward value of the target function.
[0207] First, taking the auto chess game as an example, the prediction method of the server in each round of the game in the embodiments of the present invention will be introduced below. Refer to Figure 12 , Figure 12 which is a schematic diagram of the prediction method of the server in each round provided by the embodiments of the present invention. Exemplarily, the duration of a game session can be about 20 minutes to 30 minutes, a game session can include about 20 to 30 rounds, and each round has multiple operation stages. For example, it can include a preparation stage, a battle stage, and a settlement stage.
[0208] In the preparation stage, the participating objects perform corresponding actions for operation. For example, it can be buying chess pieces, selling chess pieces, selecting equipment, arranging the chess pieces into the battle area, selecting the corresponding target formation, and so on.
[0209] In the battle stage, the participating objects no longer perform actions, and the chess pieces in the battle area automatically interact to determine the game result of this round.
[0210] In the settlement stage, the health value of the participating objects will change according to the game result. The participating objects with a failed game result in this round will have a certain amount of health value deducted. When the remaining health value of the participating object becomes 0, the participating object will be eliminated. In this stage, the participating objects can also perform the actions that can be performed in the preparation stage.
[0211] Finally, ranking is performed according to the elimination order of the participating objects, and the final game result is output.
[0212] In the embodiments of the present invention, the server may generate training action instructions in the preparation stage and the settlement stage and send them to the terminal, so that the terminal executes corresponding actions according to the training action instructions. The frequency of generating training action instructions may be 5 frames. Figure 12 The prediction shown represents generating training action instructions. In the battle stage, training action instructions may not be generated, thereby reducing the amount of data processing. In addition, in the preparation stage, the server may stop generating training action instructions at the end of this stage and only execute the established rules related to the game session. The established rules are actions that do not require the participation of the participating objects. It can be understood that the server may stop generating training action instructions 5 seconds before the end of this stage. The embodiments of the present invention do not limit the specific time.
[0213] The following uses several actual examples to illustrate the calculation method of the target reward value in the embodiments of the present invention.
[0214] Example 1
[0215] The target result information may be the remaining health value of the target object. In the game session, the remaining health value of the participating object is an indicator that directly affects the outcome of the game session. Therefore, the remaining health value of the target object may be used to determine the reward value of the target function to evaluate the action decision-making model, and then modify the parameters of the action decision-making model to achieve the training effect.
[0216] Referring to Table 2, Table 2 shows the relevant parameters when calculating the target reward value according to the remaining health value of the target object. Specifically, when determining the first reward value corresponding to the remaining health value of the target object, it may be to determine the first change value of the remaining health value of the target object at the end of the current round, and obtain the first reward value according to the first change value. The first change value is the difference between the remaining health value at the end of the current round and the remaining health value at the end of the previous round. In addition, the first change value may be multiplied by the weight value corresponding to the remaining health value of the target object to obtain the first reward value. The weight value corresponding to the remaining health value of the target object may be set according to the actual situation. For example, it may be set according to the value range of the first reward value or human experience, etc. The embodiments of the present invention use 0.1 as the weight value corresponding to the remaining health value of the target object as an example for illustration. In addition, since the remaining health value of the target object changes according to the rounds, the corresponding first change value may be determined at the end of the current round.
[0217] Table 2 A parameter setting method for calculating the target reward value
[0218] Target result information Implementation logic Change cycle Weight Value range Remaining health points Winning or losing health points Changing according to rounds 0.1 -5.0 to 5.0
[0219] Example 2
[0220] The target result information may also include the remaining health value of the target object, the character level of the virtual character controlled by the target object, and the interest value of the virtual currency of the target object. By introducing the character level of the virtual character controlled by the target object and the interest value of the virtual currency of the target object, it is beneficial to improve the training effect.
[0221] When determining the first reward value corresponding to the character level of the virtual character controlled by the target object, it may be to determine the second change value of the character level of the virtual character controlled by the target object in the current data frame, and obtain the first reward value according to the second change value; the second change value is the difference in character level between the current data frame and the previous data frame. When determining the first reward value corresponding to the interest value of the virtual currency of the target object, it may be to determine the third change value of the interest value of the virtual currency of the target object at the end of the current round, and obtain the first reward value according to the third change value. The third change value is the difference in interest value between the current data frame and the previous data frame.
[0222] Then, the target reward value of the objective function can be obtained by performing weighted processing on the first reward values corresponding to the remaining health value of the target object, the character level of the virtual character controlled by the target object, and the interest value of the virtual currency of the target object.
[0223] Obtaining the first reward value according to the second change value and the third change value is similar to obtaining the first reward value according to the first change value, and can be determined according to the product with the corresponding weight value. Similarly, the weight values of the character level of the virtual character controlled by the target object and the interest value of the virtual currency of the target object can be set according to the actual situation. In addition, since the character level of the virtual character controlled by the target object can change multiple times in a round, accordingly its change period can be determined according to the data frame, and since the interest value of the virtual currency of the target object changes according to the round, accordingly the third change value can be determined at the end of the current round.
[0224] Among them, the character level of the virtual character controlled by the target object can be further refined, and the higher the character level, the higher the weight when calculating the first reward value. Referring to Table 3, Table 3 is the relevant parameters when calculating the target reward value according to the remaining health value of the target object, the character level of the virtual character controlled by the target object, and the interest value of the virtual currency of the target object. Among them, the character level of the virtual character is subdivided into two-star pieces and three-star pieces, the weight corresponding to the two-star pieces is 0.5, and the weight corresponding to the three-star pieces is 1.0. By refining the character level of the virtual character controlled by the target object and configuring different weights according to different character levels, the rationality of the first reward value can be improved, and the training effect of the action decision-making model can be improved.
[0225] Table 3 Another parameter setting method for calculating the target reward value
[0226]
[0227]
[0228] In a possible implementation, since the server generates training action instructions based on data frames, the parameters of the action decision model can be adjusted accordingly based on the data frames. Based on this, when the target result information changes according to the round, the first reward value corresponding to the target result information can be allocated to the corresponding data frames according to a preset ratio, thereby increasing the number of samples. The preset ratio can be determined according to the number of data frames included in each round.
[0229] Referring to Table 4, Table 4 shows the comparison of the training effects of Example 1 and Example 2 above. The specific parameters compared are the number of two-star chess pieces, the number of three-star chess pieces, the total value of the chess pieces, and the average interest value per round controlled by the target object in the last round of the game. It can be seen that the target result information adopted in Example 2 can make the target object more inclined to synthesize two-star and three-star chess pieces. The total value of the chess pieces controlled by the target object is higher, the average interest value per round is also higher, the training effect is better, and the artificial intelligence level of the target object is higher.
[0230] Table 4 Comparison of the training effects of Example 1 and Example 2
[0231] Number of two-star chess pieces Number of three-star chess pieces Total value of chess pieces Average interest value per round Example 1 6 0.1 60 1.9 Example 2 7.5 1.1 87 3.5
[0232] It should be noted that in Example 2, since upgrading the virtual character requires using virtual currency to purchase the virtual character, and the interest value is determined based on the virtual currency owned at the end of the round, and the interest value affects the subsequent purchase of the virtual character, there is an association relationship between the character level of the virtual character controlled by the target object and the interest value of the virtual currency of the target object. By simultaneously using the character level of the virtual character controlled by the target object and the interest value of the virtual currency of the target object as the target result information in the embodiments of the present invention, the connection between different target result information can be strengthened, which is beneficial to improving the training effect.
[0233] Example 3
[0234] The target result information may also include the remaining health value of the target object, the character level of the virtual character controlled by the target object, the interest value of the virtual currency of the target object, the ranking result of the target object in the current round, and the association attribute between the virtual characters controlled by the target object. That is, on the basis of Example 2, the ranking result of the target object in the current round and the association attribute between the virtual characters controlled by the target object are added as the target result information. Then, the target reward value of the target function can be obtained by performing weighted processing on the first reward value corresponding to the remaining health value of the target object, the character level of the virtual character controlled by the target object, the interest value of the virtual currency of the target object, the ranking result of the target object in the current round, and the association attribute between the virtual characters controlled by the target object.
[0235] By adding the ranking result of the target object in the current round and the association attribute between the virtual characters controlled by the target object as the target result information on the basis of Example 2, the target result information can be made more abundant, thereby improving the training effect of the action decision-making model.
[0236] When determining the first reward value corresponding to the ranking result of the target object in the current round, it may be to determine the first preset score corresponding to the ranking result of the target object in the current round, and obtain the first reward value by obtaining the first preset score. For example, for the participating objects ranked 1 to 8, their first preset scores are 4, 3, 2, 1, -1, -2, -3, -4 in sequence. It can be understood that the first preset score corresponding to the ranking result can be set according to the actual situation, and the embodiments of the present invention do not make any limitations.
[0237] When determining the first reward value corresponding to the association attribute between the virtual characters controlled by the target object, it may be to determine the second preset score corresponding to the association attribute between the virtual characters controlled by the target object, and obtain the first reward value according to the second preset score. For example, for the virtual character controlled by the target object with the association attribute X, the second preset score that can be obtained is 1, and for the virtual character controlled by the target object with the association attribute Y, the second preset score that can be obtained is 2. It can be understood that the second preset score corresponding to the association attribute can be set according to the actual situation, and the embodiments of the present invention do not make any limitations.
[0238] It can be understood that the above Examples 1, 2, and 3 are only used to exemplarily illustrate the calculation method of the target reward value of the target function in the embodiments of the present invention. In fact, the target reward value can be calculated according to any combination of one or more of the remaining health value of the target object, the character level of the virtual character controlled by the target object, the interest value of the virtual currency of the target object, the ranking result of the target object in the current round, and the association attribute between the virtual characters controlled by the target object.
[0239] In the embodiments of the present invention, an action decision-making model is trained based on reinforcement learning, and the objective function adopted aims to maximize the expected value of the cumulative reward. The objective function can be expressed by the following formula:
[0240]
[0241] Where J represents the objective function, π represents the action decision-making model, τ represents the action trajectory of the game session, P represents probability, R represents the reward function, and E represents the expected value. The meaning represented by this formula is the expected value of the cumulative reward obtained by the action trajectory output by the action decision-making model in the game session, that is, the target reward value. The larger the target reward value, the better the performance of the action decision-making model.
[0242] In a possible implementation manner, the above reward function R can adopt the advantage function A, so as to reduce the variance of the objective function estimation, improve the stability of training, and improve the training effect.
[0243] Based on this, when obtaining the target reward value of the objective function according to the first reward value, specifically, the second reward value and the third reward value output by the value model can be obtained, the fourth reward value can be obtained according to the sum of the first reward value and the second reward value, and the target reward value of the objective function can be obtained according to the difference between the fourth reward value and the third reward value.
[0244] Where the second reward value is the reward value output by the value model after the target object executes the training action corresponding to the training action instruction, and the third reward value is the reward value output by the value model before the target object executes the training action.
[0245] Specifically, the above manner can be expressed by the following formula:
[0246] A π (S t , a t ) = r t + V π (S t+1 ) - V π (S t )
[0247] Where A represents the advantage function, π represents the action decision-making model, S t represents the current situation state, a t represents the action, r t represents the actual reward obtained by executing the a t action in the current situation state S t , that is, the first reward value at the current moment, t represents the time factor, t = 0, 1, 2..., V represents the value model, the output of the value model represents the possible average reward value of the current situation state, S t+1 represents the situation state att Execute a next t The next situation state after the action, V π (S t+1 ) is to execute a t The possible average reward value of the next situation state after the action, that is, the second reward value, V π (S t ) is the possible average reward value of the current situation state, that is, the third reward value. Since the value model is used to output the average reward value and its change is less drastic, the second reward value can be used to fit the first reward value, which can reduce the drastic change of the actual reward value. Let r t +V π (S t+1 ) be used as the true reward value of the current situation state, that is, the fourth reward value. Therefore, the variance of the target function estimation can be reduced, the training stability can be improved, and the training effect can be enhanced. Then, subtract V t +V π (S t+1 ) from r π (S t ) to obtain the advantage of the true reward value obtained by executing the current a t action compared to the average reward value.
[0248] In a possible implementation, the value model can be a separate model. For example, a neural network structure can be used to construct the value model, and the object features of the participating objects and the corresponding sample reward values can be used to train the value model separately. In addition, the value model can also be integrated into the action decision model. For example, referring to Figure 7 the output value of the classifier corresponding to "average reward" in
[0249] is the output value of the value model. By integrating the value model into the action decision model, the value model and the action decision model can reuse the sample data during training, reduce the duplicate processing of features, and thus reduce the data processing cost.
[0250] It can be understood that the weight of the fifth reward value corresponding to each target result information can be set according to the actual situation, and the weight of the fifth reward value corresponding to each target result information can also be 1. In this case, the target reward value is the sum of the fifth reward values corresponding to each target result information.
[0251] Among them, the number of value models is determined according to the number of types of target result information. For example, if there are three types of target result information, then there are also three value models, that is, the value models correspond one by one to each type of target result information. When using a single value model to determine the target reward value, when estimating the average reward value of a target result information with a relatively low degree of change, the estimation accuracy may decrease. By setting multiple value models to determine the target reward value, the influence of a target result information with a relatively low degree of change on the estimation of the average reward value can be reduced, the variance of the target function estimation can be better reduced, so that the target reward value is more accurate and the training effect is improved.
[0252] Correspondingly, the calculation method of the advantage function is adjusted to calculate the advantage function corresponding to each value model and then perform weighted processing, which can be specifically expressed by the following formula:
[0253]
[0254]
[0255] Among them, A i represents the advantage function corresponding to the value model, that is, the fifth reward value, and i and k are constants greater than 0, where i = 1, 2, 3... k.
[0256] Taking the auto chess game as an example for illustration, referring to Table 5, Table 5 shows the setting methods of multiple value models. The value models V1, V2, and V3 correspond to the remaining health value, the piece level, and the interest value respectively.
[0257] Table 5 Setting methods of multiple value models
[0258] Target result information Value model Function of the value model Remaining health points V1 Predicting possible health point gains Chess piece level V2 Predicting the possible number of two- and three-star chess pieces to be synthesized Interest value V3 Predicting possible interest gains
[0259] Referring to Table 6, Table 6 shows the effect comparison between a single value model and multiple value models. It can be seen that after the target object performs actions according to the action decision model trained with multiple value models, the number of two-star pieces, the number of three-star pieces, and the total value of the pieces controlled by the target object are higher. Therefore, using multiple value models to determine the target reward value, the artificial intelligence level of the trained action decision model is higher and the training effect is better.
[0260] Table 6 Effect comparison between a single value model and multiple value models
[0261] Number of two-star chess pieces Number of three-star chess pieces Total value of chess pieces Average interest value per round Single value model 7.5 1.1 87 3.5 Multiple value models 7.6 1.7 100 3.5
[0262] Refer to Figure 13 , Figure 13 which is a complete flowchart of the training method provided by the embodiments of the present invention. The following introduces the overall process of the training method in the embodiments of the present invention:
[0263] The terminal collects the frame data in the game session and sends the frame data to the server;
[0264] The server performs vectorization processing on the frame data to obtain the first sample features of multiple participating objects participating in the game session and the second sample features of the target object;
[0265] The server inputs the first sample features and the second sample features into the action decision model, splices the first sample features and the second sample features to obtain the spliced sample features, classifies the spliced sample features based on the first classifier of the action decision model to obtain multiple candidate action types, determines the current running stage of the game session, determines the target action type from the candidate action types according to the running stage to obtain the third label corresponding to the target action type, and classifies the spliced sample features based on multiple second classifiers of the action decision model to obtain multiple fourth labels;
[0266] The server determines the target action type according to the third label, determines the target label from multiple fourth labels according to the target action type, determines the action object corresponding to the target action type according to the target label, and generates a training action instruction for the target object according to the target action type and the action object corresponding to the target action type;
[0267] The server sends the training action instruction to the terminal;
[0268] The target object in the game application program running on the terminal executes the corresponding action according to the training action instruction.
[0269] The terminal collects the target result information after the target object executes the action and sends the target result information to the server;
[0270] The server determines the target reward value of the target function according to the target result information and corrects the parameters of the action decision model according to the target reward value.
[0271] The training method provided by the embodiments of the present invention determines the target reward value of the target function by obtaining the target result information, achieves the effect of reinforcement learning, and does not require the annotation of supervised labels, which is beneficial to improving the training efficiency of the action decision model.
[0272] In addition, when training the action decision-making model, a large amount of data resources are required. To improve the training efficiency, the embodiments of the present invention adopt a distributed training method. Referring to Figure 14 , Figure 14 FIG. Figure 14 is a schematic diagram of the training framework of the action decision-making model provided by the embodiments of the present invention. The training framework includes a Docker cluster and a GPU server cluster. The Docker cluster includes multiple Docker machines, and each Docker machine runs a Docker image. Each Docker image runs a game application program and can start a game session independently. The GPU server cluster includes multiple GPU machines, and each GPU machine runs a training service. Among them, each GPU machine can perform parallel training for different parts of the action decision-making model respectively, or can perform parallel training for the sample data of different target objects respectively.
[0273] The specific training steps are as follows:
[0274] Each Docker machine starts the game application program, connects to the game AI service to start a game session, and adopts the method of automatically playing the game session by multiple participating objects. The game AI service loads the latest action decision-making model from the model pool according to a preset frequency to play the game session, generates an action instruction according to the situation state of the game session and the action decision-making model, so that the participating objects in the game session execute the corresponding actions. After the game session ends, the sample data is sent to the GPU server cluster. Among them, Figure 14 Exemplarily shows the situation of 8 participating objects. Participating object 1 is the target object ( Figure 14 The action of loading the action decision-making model is omitted in
[0275] ), and participating objects 2 to 8 are the opponents of participating object 1.
[0276] The GPU machine performs model training according to the sample data. For a single GPU machine, the received sample data can be stored in different sample pools respectively, and each single GPU machine runs multiple training services. Each training service extracts sample data from the corresponding sample pool for training, and then updates the parameters of the action decision-making model. After the GPU machine iteratively trains for a certain number of rounds, the latest action decision-making model is synchronized to the opponent model pool of each Docker machine in a point-to-point manner.
[0277] Figure 14 Under the training framework shown in , combined with the comprehensive and reasonable description of the characteristics of the target object and the hierarchical action labels, it is possible to realize the automatic game session of multiple participating objects without manual intervention, and multiple training tasks can be executed in parallel, which is beneficial to improving the training effect and training efficiency.
[0278] It can be understood that although the steps in the above various flowcharts are sequentially shown according to the indications of the arrows, these steps do not necessarily need to be sequentially executed according to the order indicated by the arrows. Unless there is a clear indication in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowcharts may include multiple steps or multiple stages. These steps or stages do not necessarily need to be completed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0279] It can be understood that when the above embodiments of the present invention are applied to specific products or technologies, data related to user identity or characteristics such as frame data is involved. When the embodiments of the present invention are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards in relevant countries and regions.
[0280] Referring to Figure 15 , Figure 15 is a schematic structural diagram of an action instruction generation device provided by an embodiment of the present invention. The action instruction generation device 1500 includes:
[0281] An object feature acquisition module 1501, configured to acquire first object features of multiple participating objects participating in a game session, and acquire second object features of a target object among the multiple participating objects. The first object features are used to characterize the public situation information of the participating objects in the game session, and the second object features are used to characterize the private situation information of the target object in the game session;
[0282] A first model processing module 1502, configured to input the first object features and the second object features into an action decision model, perform splicing processing on the first object features and the second object features to obtain spliced object features, and obtain a first label and a second label according to the spliced object features. The first label is used to characterize the action type of the target object, and the second label is used to characterize the action object corresponding to the action type;
[0283] A first instruction generation module 1503, configured to generate a target action instruction for the target object according to the first label and the second label.
[0284] Further, a first area, a second area, and a third area are set in the virtual scene corresponding to the game session. The first object feature includes the first character features of the virtual characters of the participating objects in the first area and the second area, and the second object feature includes the second character features of the virtual characters of the target object in the third area. The above object feature acquisition module 1501 is specifically configured to:
[0285] Determine the number of types of virtual characters in the game session, determine the number of slots for configuring virtual characters in the first area, the second area, and the third area, determine the feature dimensions corresponding to the virtual characters in the first area, the second area, and the third area according to the number of types and the number of slots, and determine the first character feature and the second character feature according to the feature dimensions;
[0286] Or, determine the number of types of virtual characters in the game session, determine the maximum number of characters that the participating objects control each virtual character, determine the feature dimensions of the virtual characters in the first area, the second area, and the third area according to the number of types and the maximum number of characters, and determine the first character feature and the second character feature according to the feature dimensions.
[0287] Further, the above first model processing module 1502 is specifically configured to:
[0288] Perform pooling processing on the first object features of other participating objects except the target object to obtain third object features;
[0289] Perform splicing processing on the third object features, the first object features of the target object, and the second object features to obtain spliced object features.
[0290] Further, the above first model processing module 1502 is specifically configured to:
[0291] Perform classification processing on the spliced object features based on the first classifier of the action decision model to obtain a first label, and perform classification processing on the spliced object features based on multiple second classifiers of the action decision model to obtain multiple second labels;
[0292] Or, perform classification processing on the spliced object features based on the first classifier of the action decision model to obtain a first label, determine a target classifier from multiple second classifiers of the action decision model according to the first label, and perform classification processing on the spliced object features based on the target classifier to obtain a second label.
[0293] Further, the above first model processing module 1502 is specifically configured to:
[0294] Perform classification processing on the spliced object features based on the first classifier of the action decision model to obtain multiple candidate action types;
[0295] Determine the current running stage of the game session, and determine the target action type from the candidate action types according to the running stage, so as to obtain the first label corresponding to the target action type.
[0296] Furthermore, the number of the second labels is multiple, and the first model processing module 1502 is specifically configured to:
[0297] Determine the target action type according to the first label;
[0298] Determine the target label from the multiple second labels according to the target action type, and determine the action object corresponding to the target action type according to the target label;
[0299] Generate a target action instruction for the target object according to the target action type and the action object corresponding to the target action type.
[0300] Furthermore, the target action type includes selecting a virtual item, the action object includes the target virtual item, and the first instruction generation module 1503 is further configured to:
[0301] Determine the first target virtual character from the virtual characters controlled by the target object;
[0302] Generate a prop equipment instruction for the target object according to the matching relationship between the character attributes of the first target virtual character and the prop attributes of the target virtual item.
[0303] Furthermore, the target action type includes selecting a virtual character configured in the first area, the action object includes the second target virtual character, and multiple candidate slots for configuring virtual characters are set in the first area. The first instruction generation module 1503 is further configured to:
[0304] Randomly determine an initial slot from the multiple candidate slots;
[0305] Determine a target formation from the multiple preset configuration formations, where the target formation is used to represent the configuration slots of the second target virtual character;
[0306] Generate a character configuration instruction for the target object according to the initial slot and the target formation.
[0307] The action instruction generation device provided by the embodiments of the present invention is based on the same inventive concept as the above-mentioned action instruction generation method. Therefore, through the comprehensive and reasonable description of the characteristics of the target object and the hierarchical structure of the action labels, it can better adapt to the scenario where the number of participating objects is large and the operations executable on virtual characters are complex, thereby improving the accuracy and rationality of the generated action instructions and the artificial intelligence level of the target object.
[0308] Refer to Figure 16 , Figure 16Schematic structural diagram of a training device for an action decision-making model provided by an embodiment of the present invention. The training device 1600 includes:
[0309] A sample feature acquisition module 1601, configured to acquire first sample features of a plurality of participating objects participating in a game session, and acquire second sample features of a target object among the plurality of participating objects. The first sample features are used to represent the public situation information of the participating objects in the game session, and the second sample features are used to represent the private situation information of the target object in the game session;
[0310] A second model processing module 1602, configured to input the first sample features and the second sample features into the action decision-making model, perform splicing processing on the first sample features and the second sample features to obtain spliced sample features, and obtain a third label and a fourth label according to the spliced sample features. The third label is used to represent the action type of the target object, and the fourth label is used to represent the action object corresponding to the action type;
[0311] A second instruction generation module 1603, configured to generate a training action instruction for the target object according to the third label and the fourth label;
[0312] A parameter correction module 1604, configured to acquire target result information after the target object executes the training action instruction, determine a target reward value of a target function according to the target result information, and correct the parameters of the action decision-making model according to the target reward value.
[0313] Further, the parameter correction module 1604 is specifically configured to:
[0314] Determine a first reward value corresponding to the target result information, where the target result information includes at least one of the remaining health value of the target object, the character level of the virtual character controlled by the target object, the interest value of the virtual currency of the target object, the ranking result of the target object in the current round, and the association attribute between the virtual characters controlled by the target object;
[0315] Obtain the target reward value of the target function according to the first reward value.
[0316] Further, the parameter correction module 1604 is specifically used for at least one of the following:
[0317] Determine a first change value of the remaining health value of the target object at the end of the current round, and obtain the first reward value according to the first change value;
[0318] Or, determine a second change value of the character level of the virtual character controlled by the target object in the current data frame, and obtain the first reward value according to the second change value;
[0319] Or, determine a third change value of the interest value of the virtual currency of the target object at the end of the current round, and obtain the first reward value according to the third change value;
[0320] Alternatively, determine a first preset score corresponding to the ranking result of the target object in the current round, and obtain a first reward value from the first preset score.
[0321] Alternatively, determine a second preset score corresponding to the association attribute between the virtual characters controlled by the target object, and obtain a first reward value according to the second preset score.
[0322] Further, the parameter correction module 1604 is specifically configured to:
[0323] Obtain a second reward value and a third reward value output by the value model, where the second reward value is the reward value output by the value model after the target object executes the training action corresponding to the training action instruction, and the third reward value is the reward value output by the value model before the target object executes the training action.
[0324] Obtain a fourth reward value according to the sum of the first reward value and the second reward value.
[0325] Obtain the target reward value of the objective function according to the difference between the fourth reward value and the third reward value.
[0326] Further, the number of value models is multiple, and different value models correspond to different target result information. The parameter correction module 1604 is specifically configured to:
[0327] Obtain a fifth reward value corresponding to each target result information according to the difference between the fourth reward value and the third reward value corresponding to each target result information.
[0328] Perform a weighted process on the fifth reward value to obtain the target reward value of the objective function.
[0329] The training device of the action decision model provided by the embodiments of the present invention and the above-mentioned training method of the action decision model are based on the same inventive concept. Therefore, by obtaining the target result information to determine the target reward value of the objective function, the effect of reinforcement learning is achieved, and there is no need to perform the annotation of supervision labels, which is beneficial to improving the training efficiency of the action decision model.
[0330] The electronic device provided by the embodiments of the present invention for executing the above-mentioned action instruction generation method or the training method of the action decision model may be a server. Refer to Figure 17 , Figure 17The following is a block diagram of a part of the server provided by an embodiment of the present invention. The server 1700 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 1722 (for example, one or more processors) and a memory 1732, and one or more storage media 1730 (for example, one or more mass storage devices) for storing application programs 1742 or data 1744. Among them, the memory 1732 and the storage medium 1730 may be transient storage or persistent storage. The program stored in the storage medium 1730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 1700. Further, the central processing unit 1722 may be configured to communicate with the storage medium 1730 and execute a series of instruction operations in the storage medium 1730 on the server 1700.
[0331] The server 1700 may further include one or more power supplies 1726, one or more wired or wireless network interfaces 1750, one or more input / output interfaces 1758, and / or one or more operating systems 1741, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0332] The processor in the server 1700 may be used to execute the action instruction generation method or the training method of the action decision model.
[0333] An embodiment of the present invention further provides a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the action instruction generation method or the training method of the action decision model in the foregoing respective embodiments.
[0334] An embodiment of the present invention further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the action instruction generation method or the training method of the action decision model described above.
[0335] In the description of the present invention and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0336] It should be understood that in the present invention, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.
[0337] It should be understood that in the description of the embodiments of the present invention, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as greater than, less than, exceeding, etc. do not include the present number, and understandings such as above, below, within, etc. include the present number.
[0338] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be an indirect coupling or communication connection through some interfaces, devices or units, and can be in an electrical, mechanical or other form.
[0339] The unit described as a separation component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0340] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0341] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0342] It should also be understood that the various embodiments provided in the embodiments of the present invention can be combined arbitrarily to achieve different technical effects.
[0343] The above is a specific description of the preferred embodiments of the present invention, but the present invention is not limited to the above-mentioned embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present invention.
Claims
1. A method for generating action instructions, characterized in that, Including: Obtain the first object features of multiple participating objects in a game session, and obtain the second object features of a target object among the multiple participating objects. The first object features are used to represent the public situation information of the participating objects in the game session, and the second object features are used to represent the private situation information of the target object in the game session. In the virtual scene corresponding to the game session, a first area, a second area, and a third area are set. The first object features include the first role features of the virtual roles of the participating objects in the first area and the second area, and the second object features include the second role features of the virtual role of the target object in the third area. Input the first object features and the second object features into an action decision model, perform splicing processing on the first object features and the second object features to obtain spliced object features, and obtain a first label and a second label according to the spliced object features. The first label is used to represent the action type of the target object, and the second label is used to represent the action object corresponding to the action type. Generate a target action instruction for the target object according to the first label and the second label. The obtaining the first object features of multiple participating objects in a game session and obtaining the second object features of a target object among the multiple participating objects includes: Determine the number of types of virtual roles in the game session, determine the number of slots for configuring the virtual roles in the first area, the second area, and the third area, determine the feature dimensions corresponding to the virtual roles in the first area, the second area, and the third area according to the number of types and the number of slots, and determine the first role features and the second role features according to the feature dimensions. Alternatively, determine the number of types of virtual roles in the game session, determine the maximum number of roles that each participating object controls for each virtual role, determine the feature dimensions of the virtual roles in the first area, the second area, and the third area according to the number of types and the maximum number of roles, and determine the first role features and the second role features according to the feature dimensions.
2. The action instruction generation method according to claim 1, wherein The performing splicing processing on the first object features and the second object features to obtain spliced object features includes: Perform pooling processing on the first object features of other participating objects except the target object to obtain third object features. Perform splicing processing on the third object features, the first object features of the target object, and the second object features to obtain spliced object features.
3. The method for generating an action instruction according to any one of claims 1 to 2, characterized in that The obtaining the first label and the second label according to the spliced object features includes: Perform classification processing on the spliced object features based on the first classifier of the action decision model to obtain the first label, and perform classification processing on the spliced object features based on multiple second classifiers of the action decision model to obtain multiple second labels. Alternatively, classify the spliced object features using a first classifier of the action decision model to obtain a first label, determine a target classifier from multiple second classifiers of the action decision model according to the first label, and classify the spliced object features using the target classifier to obtain a second label.
4. The action instruction generation method according to claim 3, characterized in that, Classifying the spliced object features using a first classifier of the action decision model to obtain a first label includes: Classifying the spliced object features using a first classifier of the action decision model to obtain multiple candidate action types; Determine the current running stage of the game round, and determine a target action type from the candidate action types according to the running stage to obtain a first label corresponding to the target action type.
5. The action instruction generation method according to claim 1, wherein The number of the second labels is multiple. Generating a target action instruction for the target object according to the first label and the second labels includes: Determine a target action type according to the first label; Determine a target label from multiple second labels according to the target action type, and determine an action object corresponding to the target action type according to the target label; Generate a target action instruction for the target object according to the target action type and the action object corresponding to the target action type.
6. The action instruction generation method according to claim 5, wherein The target action type includes selecting a virtual item, and the action object includes a target virtual item. The method for generating an action instruction further includes: Determine a first target virtual character from the virtual characters controlled by the target object; Generate a prop equipment instruction for the target object according to the matching relationship between the character attributes of the first target virtual character and the prop attributes of the target virtual item.
7. The action instruction generation method according to claim 5, characterized in that The target action type includes selecting a virtual character configured in a first area, and the action object includes a second target virtual character. Multiple candidate slots for configuring the virtual character are provided in the first area. The method for generating an action instruction further includes: Randomly determine an initial slot from multiple candidate slots; Determine a target formation from multiple preset configuration formations, where the target formation is used to represent the configuration slots of the second target virtual character; Generate a character configuration instruction for the target object according to the initial slot and the target formation.
8. A training method for an action decision-making model, characterized in that including: Obtain first sample features of multiple participating objects participating in a game round, and obtain second sample features of a target object among the multiple participating objects. The first sample features are used to represent the public situation information of the participating objects in the game round, and the second sample features are used to represent the private situation information of the target object in the game round. A first area, a second area, and a third area are provided in the virtual scene corresponding to the game round. The first sample features include first character features of the virtual characters of the participating objects in the first area and the second area, and the second sample features include second character features of the virtual characters of the target object in the third area; Input the first sample feature and the second sample feature into an action decision model, perform splicing processing on the first sample feature and the second sample feature to obtain a spliced sample feature, and obtain a third label and a fourth label according to the spliced sample feature. The third label is used to characterize the action type of the target object, and the fourth label is used to characterize the action object corresponding to the action type; Generate a training action instruction for the target object according to the third label and the fourth label; Obtain the target result information after the target object executes the training action instruction, determine the target reward value of the target function according to the target result information, and correct the parameters of the action decision model according to the target reward value; The obtaining of the first sample feature of multiple participating objects participating in the game session and the second sample feature of the target object among the multiple participating objects includes: Determine the number of types of virtual characters in the game session, determine the number of slots for configuring the virtual characters in the first area, the second area, and the third area, determine the feature dimensions corresponding to the virtual characters in the first area, the second area, and the third area according to the number of types and the number of slots, and determine the first character feature and the second character feature according to the feature dimensions; Alternatively, determine the number of types of virtual characters in the game session, determine the maximum number of characters that each participating object controls for each virtual character, determine the feature dimensions of the virtual characters in the first area, the second area, and the third area according to the number of types and the maximum number of characters, and determine the first character feature and the second character feature according to the feature dimensions.
9. The training method according to claim 8, wherein The determining of the target reward value of the target function according to the target result information includes: Determine a first reward value corresponding to the target result information, where the target result information includes at least one of the remaining health value of the target object, the character level of the virtual character controlled by the target object, the interest value of the virtual currency of the target object, the ranking result of the target object in the current round, and the association attribute between the virtual characters controlled by the target object; Obtain the target reward value of the target function according to the first reward value.
10. The training method according to claim 9, characterized in that, The determining of the first reward value corresponding to the target result information includes at least one of the following: Determine a first change value of the remaining health value of the target object at the end of the current round, and obtain the first reward value according to the first change value; Alternatively, determine a second change value of the character level of the virtual character controlled by the target object in the current data frame, and obtain the first reward value according to the second change value; Alternatively, determine a third change value of the interest value of the virtual currency of the target object at the end of the current round, and obtain the first reward value according to the third change value; Alternatively, determine a first preset score corresponding to the ranking result of the target object in the current round, and obtain the first reward value by obtaining the first preset score; Alternatively, determine a second preset score corresponding to the association attribute between the virtual characters controlled by the target object, and obtain the first reward value according to the second preset score.
11. The training method according to claim 9 or 10, characterized in that, The obtaining the target reward value of the target function according to the first reward value includes: Obtain a second reward value and a third reward value output by the value model, where the second reward value is the reward value output by the value model after the target object executes the training action corresponding to the training action instruction, and the third reward value is the reward value output by the value model before the target object executes the training action; Obtain a fourth reward value according to the sum of the first reward value and the second reward value; Obtain the target reward value of the target function according to the difference between the fourth reward value and the third reward value.
12. The training method according to claim 11, wherein The number of the value models is multiple, and different value models correspond to different target result information. The obtaining the target reward value of the target function according to the difference between the fourth reward value and the third reward value includes: Obtain a fifth reward value corresponding to each target result information according to the difference between the fourth reward value and the third reward value corresponding to each target result information; Perform a weighted process on the fifth reward value to obtain the target reward value of the target function.
13. An action instruction generation device, characterized in that, Including: An object feature acquisition module, configured to acquire first object features of multiple participating objects participating in a game session, and acquire second object features of a target object among the multiple participating objects. The first object features are used to characterize the public situation information of the participating objects in the game session, and the second object features are used to characterize the private situation information of the target object in the game session. A first area, a second area, and a third area are set in the virtual scene corresponding to the game session. The first object features include first character features of the virtual characters of the participating objects in the first area and the second area, and the second object features include second character features of the virtual character of the target object in the third area; A first model processing module, configured to input the first object features and the second object features into an action decision model, perform a splicing process on the first object features and the second object features to obtain spliced object features, and obtain a first label and a second label according to the spliced object features. The first label is used to characterize the action type of the target object, and the second label is used to characterize the action object corresponding to the action type; A first instruction generation module, configured to generate a target action instruction of the target object according to the first label and the second label; The acquiring the first object features of the multiple participating objects participating in the game session and the second object features of the target object among the multiple participating objects includes: Determine the number of types of virtual characters in the game session, determine the number of slots for configuring the virtual characters in the first area, the second area, and the third area, determine the characteristic dimensions corresponding to the virtual characters in the first area, the second area, and the third area according to the number of types and the number of slots, and determine the first character characteristic and the second character characteristic according to the characteristic dimensions; Alternatively, determine the number of types of virtual characters in the game session, determine the maximum number of characters that each participating object controls for each virtual character, determine the characteristic dimensions of the virtual characters in the first area, the second area, and the third area according to the number of types and the maximum number of characters, and determine the first character characteristic and the second character characteristic according to the characteristic dimensions.
14. A training device for an action decision-making model, characterized in that It includes: A sample feature acquisition module, configured to acquire first sample features of multiple participating objects participating in a game session, and acquire second sample features of a target object among the multiple participating objects. The first sample features are used to characterize the public situation information of the participating objects in the game session, and the second sample features are used to characterize the private situation information of the target object in the game session. In the virtual scene corresponding to the game session, there are a first area, a second area, and a third area. The first sample features include first character characteristics of the virtual characters of the participating objects in the first area and the second area, and the second sample features include second character characteristics of the virtual characters of the target object in the third area; A second model processing module, configured to input the first sample features and the second sample features into an action decision model, perform splicing processing on the first sample features and the second sample features to obtain spliced sample features, and obtain a third label and a fourth label according to the spliced sample features. The third label is used to characterize the action type of the target object, and the fourth label is used to characterize the action object corresponding to the action type; A second instruction generation module, configured to generate a training action instruction for the target object according to the third label and the fourth label; A parameter correction module, configured to acquire target result information after the target object executes the training action instruction, determine a target reward value of a target function according to the target result information, and correct the parameters of the action decision model according to the target reward value; The acquiring the first sample features of multiple participating objects participating in a game session and acquiring the second sample features of a target object among the multiple participating objects includes: Determine the number of types of virtual characters in the game session, determine the number of slots for configuring the virtual characters in the first area, the second area, and the third area, determine the characteristic dimensions corresponding to the virtual characters in the first area, the second area, and the third area according to the number of types and the number of slots, and determine the first character characteristic and the second character characteristic according to the characteristic dimensions; Alternatively, determine the number of types of virtual characters in the game session, determine the maximum number of characters that the participating object can control for each virtual character, determine the characteristic dimensions of the virtual characters in the first area, the second area, and the third area according to the number of types and the maximum number of characters, and determine the first character characteristic and the second character characteristic according to the characteristic dimensions.
15. An electronic device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the action instruction generation method described in any one of claims 1 to 7, or implements the training method described in any one of claims 8 to 12.
16. A computer-readable storage medium storing a program, characterized in that, When the program is executed by the processor, it implements the action instruction generation method described in any one of claims 1 to 7, or implements the training method described in any one of claims 8 to 12.
17. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the action instruction generation method described in any one of claims 1 to 7, or implements the training method described in any one of claims 8 to 12.
Citation Information
Patent Citations
Fictitious self-play-based multi-person incomplete information game policy resolving method, device and system as well as storage medium
CN110404264A
Interaction model training method and device, computer equipment and storage medium
CN111111204A