Model training method and device, electronic equipment and computer readable storage medium
By obtaining sample data in game applications and training preset models, and generating twin models to simulate real game environments, the problem of slow sample data production speed and high cost is solved, and the efficiency of game robot development is improved.
Patent Information
- Application Number
- CN202311594472.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-27
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-11-27
AI Technical Summary
In the prior art, the efficiency of game robot development is limited by the slow production speed of sample data, and increasing machines running the game environment to collect data in parallel requires huge costs.
By obtaining the sample data set generated in the running game application, using the preset model to train the sample data, and generate a twin model to simulate the real game environment and predict the status information of the target virtual character.
Improve the production efficiency of sample data, reduce acquisition costs, and shorten the development cycle of game robots based on reinforcement learning.
Smart Images

Figure CN120037662A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a model training method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] With the rise of the artificial intelligence industry, more and more game companies use reinforcement learning technology to develop game robots, i.e., non-player control virtual characters (Non-Player Character, NPC). During the process of developing game robots based on reinforcement learning, a large amount of game behavior data, i.e., sample data, is required. Currently, the slow production speed of sample data is a major bottleneck restricting the development efficiency of game robots.
[0003] In related technologies, by increasing the machines running the game environment and starting multiple game environments to run game robots in parallel, the sample data generated in each game environment is collected to improve the production volume of sample data per unit time, i.e., the production speed of sample data.
[0004] However, since most current game environments are large in scale and require a large amount of system resources when running games, one device can usually only run a limited number of game environments. Related technologies need more machines to effectively improve the production efficiency of sample data. Therefore, related technologies require huge costs. Summary of the Invention
[0005] This application provides a model training method, apparatus, electronic device, and computer-readable storage medium to improve the production efficiency of sample data at low cost.
[0006] In a first aspect, an embodiment of this application provides a model training method, and the method includes:
[0007] Obtain a sample data set, where the sample data set includes multiple pieces of sample data generated during the running of a first game application. Each piece of sample data includes first state information and second state information of a target virtual character in the first game application, and first action information of a first virtual character; wherein, the target virtual character includes the first virtual character and a second virtual character, the second virtual character refers to a virtual character within the field of view of the first virtual character in the game scene of the first game application, and the second state information is the state information of the target virtual character after the first virtual character executes the action corresponding to the first action information;
[0008] Input the first state information and the first action information in the sample data into a preset model to obtain the predicted state information output by the preset model. The predicted state information is the state information of the target virtual character predicted by the preset model according to the first virtual character performing the action corresponding to the first action information.
[0009] Train the preset model with the goal of minimizing a preset loss function, adjust the parameters of the preset model, and use the trained preset model as the twin model corresponding to the first game application. The preset loss function is used to calculate the difference between the predicted state information and the second state information. The twin model is used to predict the next state information of the target virtual character in the first game application according to the current state information of the target virtual character and the first action information performed by the first virtual character.
[0010] In a second aspect, an embodiment of the present application provides a model training device, which includes:
[0011] An acquisition module, configured to acquire a sample data set, where the sample data set includes multiple pieces of sample data generated during the running of the first game application. Each piece of sample data includes the first state information and the second state information of the target virtual character in the first game application, and the first action information of the first virtual character. The target virtual character includes the first virtual character and the second virtual character, and the second virtual character refers to a virtual character within the field of view of the first virtual character in the game scene of the first game application. The second state information is the state information of the target virtual character after the first virtual character performs the action corresponding to the first action information.
[0012] A processing module, configured to input the first state information and the first action information in the sample data into a preset model to obtain the predicted state information output by the preset model. The predicted state information is the state information of the target virtual character predicted by the preset model according to the first virtual character performing the action corresponding to the first action information.
[0013] A control module, configured to train the preset model with the goal of minimizing a preset loss function, adjust the parameters of the preset model, and use the trained preset model as the twin model corresponding to the first game application. The preset loss function is used to calculate the difference between the predicted state information and the second state information. The twin model is used to predict the next state information of the target virtual character in the first game application according to the current state information of the target virtual character and the first action information performed by the first virtual character.
[0014] In a third aspect, an embodiment of the present application provides an electronic device, which includes:
[0015] a memory and a processor, the memory being coupled to the processor;
[0016] the memory is used to store one or more computer instructions;
[0017] the processor is used to execute the one or more computer instructions to implement the model training method according to any one of the above first aspects.
[0018] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which one or more computer instructions are stored, and characterized in that the instructions are executed by a processor to implement the model training method according to any one of the above first aspects.
[0019] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the model training method according to any one of the above first aspects.
[0020] Compared with the prior art, the present application has the following advantages:
[0021] The model training method provided by the present application obtains a sample data set, which includes multiple sample data generated during the operation of a first game application. Each sample data includes first state information and second state information of a target virtual character in the first game application, and first action information of a first virtual character. Among them, the target virtual character includes a first virtual character and a second virtual character, and the second virtual character refers to a virtual character within the field of view of the first virtual character in the game scene of the first game application. The first state information / second state information is the state information of the target virtual character before / after the first virtual character executes the action corresponding to the first action information. The first state information and the first action information in the sample data are input into a preset model to obtain predicted state information output by the preset model. The predicted state information is the state information of the target virtual character predicted by the preset model according to the action corresponding to the first virtual character executing the first action information. The preset model is trained with the goal of minimizing a preset loss function, and the parameters of the preset model are adjusted. The preset model after training is used as a twin model corresponding to the first game application. The preset loss function is used to calculate the difference between the predicted state information and the second state information, and the twin model is used to predict the next state information of the target virtual character in the first game application according to the current state information of the target virtual character and the first action information executed by the first virtual character.
[0022] Compared with the prior art, in the embodiment of the present application, a preset model is trained according to sample data obtained based on running a first game application. The trained preset model can be regarded as a twin model corresponding to the first game application, and this twin model can be as close as possible to the real game environment of the first game application. In this way, the twin model can simulate the real game environment to output states. For example, the current state information of a target virtual character in the real game environment of the first game application is the first state information s 1 , control the first virtual character to execute the action a in the first game application 1 , and the state information of the target virtual character in the real game environment is converted into the second state information s 1 '. Correspondingly, input the first state information s 1 and the first action information a 1 into the twin model corresponding to the first game application. The predicted state information output by the twin model is infinitely close to the second state information s 1 ' of the target virtual character in the real game environment of the first game application. Since the inference and prediction of the second state information of the target virtual character by the twin model do not require image rendering like the real game environment, the production of sample data by the twin model trained based on the model of the present application can greatly improve the production rate of sample data. The amount of sample data produced per unit time far exceeds the production amount of sample data in the real game environment, thereby reducing the acquisition cost of sample data. Further, using the twin model corresponding to the first game application as the game environment for reinforcement learning helps to greatly shorten the development cycle of game robots based on reinforcement learning and solves the problem that the game environment cannot be accelerated during the reinforcement learning training process. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0024] Figure 1 is a schematic flowchart of a model training method provided by one embodiment of the present application;
[0025] Figure 2 is a schematic flowchart of a sample data collection method provided by one embodiment of the present application;
[0026] Figure 3 is a schematic flowchart of performing reinforcement learning training on a target behavior policy network provided by one embodiment of the present application;
[0027] Figure 4 is a schematic structural diagram of a model training device provided by one embodiment of the present application;
[0028] Figure 5 The schematic diagram of the hardware structure of the electronic device provided by one embodiment of the present application.
[0029] Through the above drawings, the specific embodiments of the present application have been shown, and more detailed descriptions will be given in the following text. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed embodiments
[0030] To make the objectives, advantages and features of the present application clearer, the present application will be clearly and completely described below in conjunction with the accompanying drawings and specific embodiments. In the following description, many specific details are set forth to facilitate a thorough understanding of the present application. However, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.
[0031] It should be noted that in the description of the present application, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance, as well as a specific order or sequence. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. In addition, in the description of the present application, unless otherwise specified, the term "plurality" means two or more. The term "and / or" describes the associated relationship of the associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] To facilitate the understanding of the technical solution of the present application, the relevant concepts involved in the present application will be introduced first.
[0033] The basic idea of reinforcement learning is to learn through trial and error. Trial-and-error learning is divided into two steps: First, in the current state of a specific environment, the agent interacts with this specific environment and observes the interaction results. For example, after the interaction, the state information of this specific environment changes from the above-mentioned current state to the next state, and the reward value obtained after the interaction; Then, according to the reward value obtained after the interaction, the corresponding behavior strategy of the agent is optimized. The goal of optimizing the behavior strategy is to maximize the reward, that is, to obtain as many rewards as possible. Therefore, reinforcement learning is a goal-oriented learning method that learns by interacting with the environment.
[0034] The agent, in a specific environment (such as a game application), is controlled by computer artificial intelligence to interact in the game environment corresponding to the game application. Among them, artificial intelligence is a theory, method, technology and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. That is to say, for the non-player controlled virtual character in the first game application of this application, it can be used as an agent, and by controlling this non-player controlled virtual character to simulate the way of a real player controlling a virtual character to play the game, it can cooperate or fight with a real player controlling a virtual character.
[0035] The state information of the virtual character is a description of the current state of the virtual character in the first game application. For example, the state information of the virtual character may include, but is not limited to, information such as the virtual character's health, equipment, position, and affiliated team. Only by way of example, this application does not make specific limitations on the content included in the state information.
[0036] Action. During the process of the agent interacting with a specific environment, the agent has to perform corresponding actions. For example, in the embodiment of this application, the first virtual character can be regarded as an agent, and the first virtual character performs corresponding actions in the game scene of the first game application.
[0037] Reward is a value feedback from the environment to the agent, and the size of the reward value is used to reflect to the agent the quality of the action it performs.
[0038] Next, the prior art involved in this application and the problems existing in the prior art will be described.
[0039] Game robots (i.e., NPCs, non-player controlled virtual characters) are very important roles in a game. They can either accompany players as teammates to complete game tasks or act as opponents to fight against players, enhancing the players' gaming experience. The earliest game robots were programmed by game developers to perform actions according to certain rules, which was a time-consuming and laborious process. With the rise of the artificial intelligence industry, in order to reduce the human input in the production of game robots, more and more game companies use reinforcement learning technology to develop game robots. Reward values for different endings need to be set for the game robots. By continuously inputting the current state of the environment to the robots, the game robots will perform corresponding actions. According to the state changes brought about by the actions in the environment, corresponding reward values are given to the actions performed by the game robots, so that the game robots can know the quality of the actions according to the magnitude of the reward values. In this process of continuous trial and error, the game robots continuously explore the optimal actions to be taken in the current state.
[0040] Reinforcement learning is an online learning method, that is, online learning is carried out during the interaction between the agent and the environment. Therefore, reinforcement learning needs to optimize the behavior strategy of the agent based on the sample data generated during the operation of the first game application. At present, most game operations cannot be accelerated. A single game may take at least 5 minutes to run, or up to 1 hour. And the training process of reinforcement learning requires a large amount of game sample data. Due to the slow generation speed of game sample data, the training of game robots based on reinforcement learning often takes several weeks or even several months. Once there are slight errors in the training code of the robot during the training process or some unexpected behaviors occur during the training process, it may be necessary to modify the training logic and restart the training, which will to a certain extent affect the game development progress. Therefore, how to improve the training speed of game robots is a major bottleneck in the current training and development of game robots using reinforcement learning technology.
[0041] To solve the problem that the game environment cannot be accelerated during the reinforcement learning training process, the prior art provides various technical solutions. The following briefly describes various technical solutions and their respective existing problems.
[0042] Solution 1: Increase the machines for running games. Specifically, multiple machines are used to simultaneously run multiple game environments in parallel to collect the sample data generated in each game environment, improving the generation rate of sample data per unit time and thus increasing the amount of training sample data, and to a certain extent improving the learning efficiency of the reinforcement learning algorithm. Although this solution can improve the collection speed of reinforcement learning sample data, currently most games occupy a large amount of resources during operation, and usually only one game environment can be run on one device. Therefore, this solution requires purchasing more machines, undoubtedly increasing the training cost of game robots.
[0043] Solution 2: Reusing a piece of data multiple times. Specifically, a data cache pool is established to record the sample data that has appeared before. In this way, whenever the environmental generation rate is too slow to meet the training requirements, sample data can be read from the cache pool to train the game robot. That is to say, each training process no longer solely depends on the sample data obtained from the current interaction with the environment, but also draws on some historical sample data to optimize the reinforcement learning algorithm. Although this solution can increase the amount of sample data during the training of the game robot and, to a certain extent, improve the learning efficiency of the reinforcement learning algorithm, since the samples in the cache pool are all samples that have been learned and these samples are not generated by the latest reinforcement policy, the improvement in learning efficiency is very limited.
[0044] Solution 3: Using the model obtained through previous learning as the base model, believing that this base model already has certain experience. Learning new tasks / new game robots on this base model can effectively improve the training efficiency. This method is called transfer learning. Although this solution can use existing knowledge to guide the training direction of the reinforcement policy and improve the training efficiency, the prerequisite is that the existing knowledge must have certain reference significance for the current training target (i.e., the downstream task target). Otherwise, it may have the opposite effect. In order to improve the playability of the game and attract more players, games will set a variety of different game targets and game characters, and it is very difficult to find the commonalities between different targets. Therefore, the application effectiveness of the transfer learning solution in the training process of game robots is very limited.
[0045] To solve at least one of the above problems, the present application provides a model training method, a model training device corresponding to this method, an electronic device capable of implementing this model training method, and a computer-readable storage medium. The following provides embodiments to elaborate on the above method, device, electronic device, and computer-readable storage medium in detail.
[0046] To make the purpose and technical solution of the present application more clearly intuitive, the method provided by the embodiments of the present application will be described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. It can be understood that the following several embodiments can exist independently, and the embodiments and the features in the embodiments can be combined with each other without conflict under the condition that they do not conflict with each other in the present application. For the same or similar content, it will not be repeated in different embodiments. In addition, the step timings in the following method embodiments are only examples and are not strictly limited. In some cases, the steps shown or described can be executed in a different order.
[0047] The present application provides a model training method, apparatus, electronic device, and computer-readable storage medium. Specifically, the model training method according to an embodiment of the present application can be executed by a computer device, where the computer device can be a terminal or a server, etc. The terminal can be a terminal device such as a smart phone, a tablet computer, a laptop computer, a touch screen, a game console, etc., and the terminal can also include a client. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, and big data and artificial intelligence platforms.
[0048] Next, in conjunction with Figure 1 , the model training method provided by an embodiment of the present application will be described. Figure 1 FIG. is a schematic flowchart of the model training method provided by an embodiment of the present application.
[0049] As Figure 1 shown, the model training method includes steps S10 to S30:
[0050] S10. Obtain a sample data set, where the sample data set includes multiple pieces of sample data generated during the running of a first game application. Each piece of sample data includes first state information and second state information of a target virtual character in the first game application, and first action information of a first virtual character. Among them, the target virtual character includes a first virtual character and a second virtual character, and the second virtual character refers to a virtual character within the field of view of the first virtual character in the game scene of the first game application. The second state information is the state information of the target virtual character after the first virtual character executes the action corresponding to the first action information.
[0051] The above-mentioned first game application includes, but is not limited to, any one of game application programs such as first-person game application programs, third-person game application programs, single-player game application programs, and multiplayer online battle arena (MOBA) application programs. The types of the above games can include, but are not limited to, at least one of the following: two-dimensional (2D) game applications, three-dimensional (3D) game applications, virtual reality (VR) game applications, augmented reality (AR) game applications, and mixed reality (MR) game applications. The above is only an example, and the embodiments of the present application do not make any limitations thereto.
[0052] As described above, the first virtual character can be a non-player controlled virtual character (i.e., NPC) or a player controlled virtual character, and the embodiments of the present application do not make any limitations in this regard.
[0053] As described above, the sample data is generated in a real game scenario during the operation of the first game application, that is, during the operation of the first game application on a computer device, by controlling the interaction between the first virtual character and the game environment to obtain the sample data. Each piece of sample data includes the first state information and the second state information of the target virtual character in the first game application, and the first action information of the first virtual character. The first state information is the state information of the target virtual character before the first virtual character executes the action corresponding to the first action information. The second state information is the state information of the target virtual character after the first virtual character executes the action corresponding to the first action information. The target virtual character includes the first virtual character and the second virtual character, and the second virtual character refers to the virtual character within the field of view of the first virtual character in the game scenario of the first game application. It can be understood that usually during the game process of a real player, the player always makes behavioral decisions based on the state of the virtual character he controls and the state information of other virtual characters that the player can observe on the game interface. Among them, the other virtual characters that the player can observe include teammate characters, enemy virtual characters, and other non-player controlled virtual characters. Then, setting the target virtual character as the first virtual character and the second virtual character within the field of view of the first virtual character is to let the first virtual character (i.e., the non-player controlled virtual object) simulate a real player, so as to make the first virtual character simulate a real player to think as much as possible, that is, to let the first virtual character simulate a real player to make behavioral decisions based on its own state information and the state information of the virtual characters within its field of view in the game scenario.
[0054] It should be noted that the first state information of the target virtual character in the sample data is the state information of the target virtual character before the first virtual character executes the action corresponding to the first action information in the game environment. After the first virtual character executes the action corresponding to the first action information, the state information of the target virtual character changes to obtain the next state information of the target virtual character, that is, the second state information.
[0055] Exemplarily, in the game scenario of the first game application, there are virtual characters 1, 2, 3, 4, 5, and 6. Among them, virtual characters 1 - 3 belong to the first camp, and virtual characters 4 - 6 belong to the second camp. The first camp and the second camp are in a hostile relationship with each other. In this first game application, assume that virtual character 1 is a non-player controlled virtual character, and among the virtual characters that virtual character 1 can observe in the game scenario are virtual characters 4 and 6. Then, in this case, the first virtual character is virtual character 1, and the target virtual characters include virtual character 1, virtual character 4, and virtual character 6. This is only an example, and the embodiments of the present application do not make any limitations in this regard.
[0056] Taking the target virtual characters including virtual character 1, virtual character 4, and virtual character 6 as an example, the following is an exemplary description of the content included in the first state information of the above target virtual characters. The first state information of the target virtual characters includes, but is not limited to, the first state information of virtual character 1, virtual character 4, and virtual character 6 respectively. Among them, the state information of each virtual character includes, but is not limited to, the blood volume situation, position information, detailed information of skills (such as cooldown situation, damage value size, etc.). This is only an example, and the embodiments of the present application do not make any limitations on the content included in the state information, and it can be specifically set according to the actual situation.
[0057] S20. Input the first state information and the first action information in the sample data into a preset model, and obtain the predicted state information output by the preset model. The predicted state information is the state information of the target virtual characters predicted by the preset model according to the actions corresponding to the first action information executed by the first virtual character.
[0058] An optional implementation manner is that the network structure of the preset model includes, but is not limited to, any one of the following: feedforward neural network, convolutional neural network, recurrent neural network, and Transform network.
[0059] As described above, the preset model is used to predict the state information of the target virtual character after the first virtual character executes the action corresponding to the first action information based on the first state information of the target virtual character and the first action information of the first character in the sample data, that is, the predicted state information. That is to say, taking any sample data as an example, the sample data includes the first state information and the second state information of the target virtual character in the first game application, and the first action information of the first virtual character. The first state information of the target virtual character and the first action information of the first virtual character are input into the preset model, and the prediction model predicts the state information of the target virtual character after the first virtual character executes the action corresponding to the first action information based on the first state information of the target virtual character and the first action information executed by the first virtual character, when the current state information of the target virtual character is the first state information.
[0060] Then, the first state information of the target virtual character and the first action information of the first virtual character in the sample data are input into the preset model, and the preset model outputs the state information of the target virtual character predicted by it after the first virtual character executes the action corresponding to the first action information.
[0061] Exemplarily, the sample data set includes n pieces of sample data, and the sample data set is {(s 1 , s 1 ', a 1 ), (s 2 = s 1 ', s 2 ', a 2 ), (s 3 = s 2 ', s 3 ', a 3 ),..., (s n = s n-1 ', s n ', a n )}. Among them, taking (s 1 , s 1 ', a 1 ) as an example to exemplarily illustrate the elements in the sample data, in (s 1 , s 1 ', a 1 ), s 1 is the state information of the target virtual character in the first game application at time t1, a 1 is the action executed by the first virtual character in the first game application at time t1, and s 1 ' is the state of the first virtual character in the first game application after executing the action a 1Subsequently, the second state information of the target virtual character at time t2. During the training process, the preset model is trained successively according to each sample data. Specifically, taking (s 1 , s 1 ', a 1 ) in the sample dataset as an example to train the preset model, s 1 and a 1 are input into the preset model, and the preset model outputs the predicted state information of the target virtual character at time t2, that is, the predicted state information is obtained. That is to say, the prediction model actually predicts the second state information of the target virtual character based on the first state information of the target virtual character and the first action information executed by the first virtual character in the sample data.
[0062] S30. The preset model is trained with the goal of minimizing the preset loss function, and the parameters of the preset model are adjusted. The trained preset model is used as the twin model corresponding to the first game application. Among them, the preset loss function is used to calculate the difference between the predicted state information and the second state information, and the twin model is used to predict the next state information of the target virtual character in the first game application based on the first state information of the target virtual character and the first action information executed by the first virtual character.
[0063] As described above, the preset loss function is used to calculate the difference between the second state information of the target virtual character predicted by the preset model output in the first game application and the second state information in the sample data. Among them, the second state information is the state information of the target virtual character after the first virtual character executes the action corresponding to the first action information predicted by the preset model.
[0064] In the embodiment of the present application, according to the difference (i.e., the loss value) calculated by the preset loss function, the parameters of the preset model are adjusted so that the loss value calculated by the trained preset model is infinitely close to zero, that is, the preset model can make more and more accurate predictions on the state information of the target virtual character after the first virtual character executes the action corresponding to the first action information based on the first state information and the first action information in the sample data. After training, the trained preset model is the twin model corresponding to the first game application. Taking any sample data as an example, it can simulate the first game application based on the first state information and the first action information in the same sample data and output predicted state information that is the same as or similar to the second state information of the target virtual character in the sample data.
[0065] Exemplarily, when taking any one of the sample data (s 1 , s 1 ', a 1 ) in the sample dataset as an example to train the preset model, s 1 and a1 Input it into a preset model, and the preset model outputs the state information w(s of the target virtual character in the first game application at time t2 predicted by it 1 ,a 1 ), that is, the second state information. Calculate the loss value loss = |w(s 1 ,a 1 ) - s 1 '| according to a preset loss function. Adjust the parameters of the preset model according to the loss value loss.
[0066] An optional implementation manner. When the training of the preset model is completed based on all the sample data in the sample data set, the trained preset model is obtained.
[0067] Another optional implementation manner. When the number of the sample data on which the above preset model has been trained exceeds a preset sample quantity threshold, the trained preset model is obtained.
[0068] Still another optional implementation manner. When the number of training times in the training process of the above preset model when the loss value is less than a preset loss threshold exceeds a preset number threshold, the trained preset model is obtained.
[0069] In the model training method provided in the embodiments of the present application, a sample data set is obtained. The sample data set includes multiple pieces of sample data generated during the running of the first game application. Each piece of sample data includes the first state information and the second state information of the target virtual character in the first game application, and the first action information of the first virtual character. Among them, the target virtual character includes the first virtual character and the second virtual character. The second virtual character refers to the virtual character within the field of view of the first virtual character in the game scene of the first game application. The first state information / second state information is respectively the state information of the target virtual character before / after the first virtual character executes the action corresponding to the first action information. Input the first state information and the first action information in the sample data into the preset model, and obtain the predicted state information output by the preset model. The predicted state information is the state information of the target virtual character predicted by the preset model according to the action corresponding to the first virtual character executing the first action information. Train the preset model with the goal of minimizing the preset loss function, adjust the parameters of the preset model, and use the trained preset model as the twin model corresponding to the first game application. Among them, the preset loss function is used to calculate the difference between the predicted state information and the second state information, and the twin model is used to predict the next state information of the target virtual character in the first game application according to the current state information of the target virtual character and the first action information executed by the first virtual character.
[0070] Compared with the prior art, in the embodiment of the present application, a preset model is trained according to sample data obtained based on running a first game application. The trained preset model can be regarded as a twin model corresponding to the first game application. This twin model can be as close as possible to the real game environment of the first game application, so that the twin model can simulate the real game environment to output states. For example, the current state information of the target virtual character in the real game environment of the first game application is the first state information s 1 , control the first virtual character to execute the action a in the first game application 1 , and the state information of the target virtual character in the real game environment is converted into the second state information s 1 '. Correspondingly, input the first state information s 1 and the first action information a 1 into the twin model corresponding to the first game application. The predicted state information output by the twin model is infinitely close to the second state information s 1 ' of the target virtual character in the real game environment of the first game application. Since the inference and prediction of the second state information of the target virtual character by the twin model do not require image rendering like the real game environment, the production of sample data by the twin model trained based on the model of the present application can greatly improve the production rate of sample data. The amount of sample data produced per unit time far exceeds the production volume of sample data in the real game environment, thus reducing the acquisition cost of sample data. Further, using the twin model corresponding to the first game application as the game environment for reinforcement learning helps to greatly shorten the development cycle of game robots based on reinforcement learning and solves the problem that the game environment cannot be accelerated during the reinforcement learning training process.
[0071] Based on the above embodiments, the model training method provided by the embodiments of the present application is further described below.
[0072] An alternative implementation manner, the implementation manner of step S10 can also be step S101:
[0073] S101. Loop and execute the following first steps until the number of sample data in the sample data set reaches a preset number threshold.
[0074] The first steps include steps S1011 - S1014:
[0075] S1011. Input the current state information and historical state information of the target virtual character in the first game application into a preset behavior policy network to obtain the second action information output by the preset behavior policy network. The preset behavior policy network is used to predict the action information that the first virtual character will execute next according to the current state information and historical state information of the target virtual character.
[0076] S1012. In the first game application, control the first virtual character to perform the action corresponding to the second action information.
[0077] S1013. After the first virtual character finishes performing the action corresponding to the second action information, obtain the third state information of the target virtual character in the first game application.
[0078] S1014. Use the current state information of the target virtual character as the first state information, the second action information as the first action information, and the third state information as the second state information to construct a sample data and save the sample data to the sample data set.
[0079] As described above, the preset behavior policy network is a network model in reinforcement learning used to learn and determine what action the agent will take next in a specific environment. Among them, the preset behavior policy network can be regarded as a mapping function that maps the current state information of a specific environment to the corresponding action probability distribution. Input the current state information and / or historical state information into the preset behavior policy network, so that the preset behavior policy network outputs a specific action, or outputs the probability value of each possible action, and determines the action with the highest probability value as the action to be executed next.
[0080] In the embodiment of the present application, when the first virtual character is a non-player controlled virtual character, the first game application controls the first virtual character to perform interactive operations in the real game environment of the first game application. Since the first virtual character (i.e., the game robot) that is not controlled by the player has no subjective consciousness, it cannot make a behavioral decision based on the current state information and historical state information of the target virtual character in the first game application, that is, it cannot decide what action to execute. Therefore, for this situation, a preset behavior policy network is introduced in the embodiment of the present application to complete the behavioral decision for the first virtual character. The preset behavior policy network is used to output the action information corresponding to the action that the first virtual character will execute next, that is, the second action information, based on the current state information and historical state information of the target virtual character in the first game application. Subsequently, in the first game application, the game system controls the first virtual character to perform the action corresponding to the second action information, and the state information of the target virtual character changes accordingly, obtaining the state information of the target virtual character after the first virtual character in the first game application finishes performing the action corresponding to the second action information, that is, obtaining the third state information of the target virtual character.
[0081] An optional implementation manner is that the preset behavior policy network can be a neural network model, and the preset behavior policy network can be any one of the following: feedforward neural network, convolutional neural network, recurrent neural network, and Transform network.
[0082] Exemplarily, during the operation of the first game application, the current state information s of the target virtual character is 1 and the historical state information are input into the preset behavior policy network. The preset behavior policy network outputs multiple different actions and the probabilities corresponding to each action. The output information of the preset behavior policy network can be referred to Table 1 as shown.
[0083] Table 1
[0084]
[0085]
[0086] As shown in Table 1, the preset behavior policy network outputs four actions and the probability values of each action respectively. Exemplarily, the action corresponding to the second action information is determined as the action a with the largest probability value 1 . As shown in Table 1, if the probability value of Action 3 is the largest, then Action 3 is determined as the action a 1 . Then, control the first virtual character to execute a in the first game application (i.e., the real game environment) 1 = Action 3, and the state information of the target virtual character changes accordingly. Subsequently, the next state information s 1 ' of the target virtual character in the first game application can be obtained, that is, the third state information of the virtual character. According to the second action information a 1 = Action 3, the current state information s 1 of the target virtual character, and the next state information s 1 ', the constructed sample data is (s 1 , s 1 ', a 1 ). That is to say, the current state information of the target virtual character is used as the first state data s 1 in the sample data, the second action information a 1 is used as the first action information in the sample data, and the third state information s 1 ' of the target virtual character is used as the second state information in the sample data, and a piece of sample data (s 1 , s 1 ', a 1 ) is constructed.
[0087] It can be understood that after the first virtual character executes the action corresponding to the second action information, the current state information of the target virtual character is s 2 = s 1 ', the preset behavior policy network outputs the second action information a 2 , and after controlling the first virtual character to execute a 2, the status information of the target virtual character changes accordingly, and then the next status information s of the target virtual character in the first game application can be obtained 2 ', that is, the third status information of the virtual character. According to the current status information s of the target virtual character before the first virtual character executes the second action information a 2 corresponding to the action 2 = s 1 ', the second action information a 2 , and the third status information s of the target virtual character 2 ', the constructed sample data is (s 2 = s 1 ', s 2 ', a 2 ). And so on, repeating the above steps S1011 - S1014, other sample data such as (s 3 = s 2 ', s 3 ', a 3 ), (s 4 = s 3 ', s 4 ', a 4 ),..., (s n = s n-1 ', s n ', a n ) can be obtained.
[0088] In the embodiment of the present application, in the process of obtaining sample data based on the real game environment of the first game application, a preset behavior strategy network is used to help the first virtual character make an intelligent decision on the next action to be executed based on the current status information and historical status information of the target virtual character, so that the first virtual character can simulate a real player to play the game, which improves the game behavior simulation of the game robot (i.e., the first virtual character), and thus the quality of the obtained sample data is higher. The twin model trained based on the high-quality sample data can also more realistically simulate the first game application for state output.
[0089] An alternative implementation manner, the above first step further includes the following steps S1015 - S1017:
[0090] S1015. Input the third status information into the first network and the second network to obtain a first predicted value output by the first network and a second predicted value output by the second network. The network structures of the first network and the second network are the same.
[0091] S1016. Adjust the parameters of the second network according to the difference between the first predicted value and the second predicted value, and determine the difference as the reward information corresponding to the second action information.
[0092] S1017. Construct a piece of training data based on the current state information, third state information, second action information, and reward information of the target virtual character, and perform reinforcement learning training on the preset behavior policy network based on the training data.
[0093] As described above, the network structures of the first network and the second network are the same, but their initial parameter values are different.
[0094] An optional implementation manner is that the network structures of the first network and the second network include any one of the following: feedforward neural network, convolutional neural network, recurrent neural network, support vector machine, decision tree, random forest, etc. For example only, the embodiments of the present application do not impose any restrictions on this.
[0095] In the embodiments of the present application, the parameters of the first network are frozen so that its parameters remain fixed. The first network is equivalent to a random network whose output has no fixed pattern or rule. Input the third state information of the target virtual character into the first network and the second network to obtain the first predicted value output by the first network and the second predicted value output by the second network. Among them, the explanation of the third state information of the target virtual character can be seen in the explanation of the third state information in the above step S1013, with the same meaning, and will not be repeated here.
[0096] Exemplarily, taking the sample data as (s 1 , s 1 ', a 1 ) as an example, in the sample data, the state information of the target virtual character before and after the first virtual character executes the action a 1 are s 1 and the third state information s 1 ' respectively. When the first network is network f and the second network is network g, input the third state information s 1 ' of the target virtual character into network f, and input the third state information and s 1 ' of the target virtual character into network g. Network f outputs the first predicted value f(s 1 '), and network g outputs the first predicted value g(s 1 ').
[0097] In the embodiments of the present application, according to the difference between the first predicted value f(s 1 ') output by the first network and the second predicted value g(s 1 ') output by the second network (for example, |f(s 1 ') - g(s 1'), the parameters of the second network are adjusted. That is, the difference between the first prediction value and the second prediction value is used as the loss value to adjust the parameters of the second network so that the second network learns the output of the first network. Then, as long as after the first virtual character executes the corresponding action, the third state information (such as s 1 ') of the target virtual character explored is input into the first network and the second network, this indicates that both the first network and the second network have seen this third state information s 1 '. And at this time, the second network has completed the parameter adjustment according to the difference between the first prediction value output by the first network and the second prediction value output by the second network, obtaining the second network with adjusted parameters. Input the third state information s 1 ' into the second network with adjusted parameters, and the difference between the second prediction value output by the second network and the first prediction value output by the first network is very small (or close to zero).
[0098] It can be understood that the difference between the first prediction value and the second prediction value (for example, |f(s 1 ') - g(s 1 ')|) can reflect whether this third state information s 1 ' has been explored by the first virtual character (or understood as the novelty level of this third state information s 1 ' for the first virtual character). That is, if the difference is larger, it indicates that the probability that this state information s 1 ' has not been explored by the first virtual character is greater (that is, the novelty of this third state information s 1 ' for the first virtual character is higher). On the contrary, if the difference is larger, it indicates that the probability that this state information s 1 ' has been explored by the first virtual character is greater (that is, the novelty of this state information s 1 ' for the first virtual character is lower). The difference between the first prediction value output by the first network and the second prediction value output by the second network is determined as the reward information r (i.e., reward) corresponding to the second action information.
[0099] In the embodiment of the present application, subsequently, the current state information, the third state information, the second action information, and the reward information of the target virtual character are constructed into a piece of training data, and the preset behavior policy network is trained by reinforcement learning based on the training data. In this way, during the process of reinforcement learning, according to the numerical value of the reward information, the quality of the action corresponding to the second action information is judged to optimize the preset behavior policy network, so that the first virtual character will continuously explore towards the states it has not seen before (that is, it has a strong curiosity) until it covers as many state spaces as possible. In this way, the finally obtained sample data set has more diverse state information.
[0100] It should be noted that during the training process of reinforcement learning, the preset behavior policy network continuously learns and optimizes by interacting with the real game environment of the first game application. For example, using Monte Carlo sampling or temporal difference learning methods, the parameters of the preset behavior policy network are updated by collecting a series of state information, action information, and corresponding reward information, enabling the first virtual character to better adapt to the environment and generate better actions. In short, by optimizing the preset behavior policy network with training data, the first virtual character will continuously explore towards state information it has never seen before, making the sample data set have more diverse state information.
[0101] Next, in combination with Figure 2 and One specific examples, the above steps S1011 - S1017 will be further described. Figure 2 It is a schematic flowchart of the sample data collection method provided by one embodiment of this application.
[0102] As Figure 2 shown, it includes steps S201 - S209:
[0103] S201. Initialize the first network and the second network with different parameters, where the network structures of the first network and the second network are the same.
[0104] S202. Obtain the preset behavior policy network, which is used to output actions to explore more states.
[0105] S203. Input the current state information of the target virtual character into the preset behavior policy network. The preset behavior policy network outputs the second action information, controls the first virtual character to execute the action corresponding to the second action information, and after the first virtual character finishes executing the action corresponding to the second action information, obtain the third state information of the target virtual character.
[0106] Among them, the target virtual character includes the first virtual character and the virtual characters within the field of vision of the first virtual character in the first game application.
[0107] S204. Input the third state information of the target virtual character into the first network and the second network, and use the difference between the output value of the first network and the output value of the second network as the loss value to optimize the parameters of the second network. At the same time, this loss value will be used as the reward information for reinforcement learning.
[0108] S205. Construct a piece of training data according to the current state information, third state information, second action information, and reward information of the target virtual character in the above steps S203 - S204.
[0109] S206. Perform reinforcement learning training on the preset behavior policy network according to the above training data.
[0110] S207. Use the current state information of the target virtual character as the first state information, the second action information as the first action information, and the third state information as the second state information to construct a sample data and save the sample data to the sample data set.
[0111] S208. Determine whether the number of sample data in the sample data set reaches a preset quantity threshold. If not, execute steps S203 - S208; if so, execute step S209.
[0112] S209. Output the sample data set.
[0113] Next, a detailed description of the implementation method for obtaining sample data for training the twin model corresponding to the first game application by running the first game application is provided.
[0114] An alternative implementation method, the implementation method of step S10 can also be step S102:
[0115] S102. Loop and execute the following second steps until the current round of the first game application ends, to obtain a sample data set including multiple sample data.
[0116] The second steps include S1021 - S1025:
[0117] S1021. Obtain the current state information of the target virtual character in the first game application.
[0118] S1022. Input the current state information of the target virtual character into the target behavior policy network to obtain the third action information output by the target behavior policy network.
[0119] Among them, the target behavior policy network is a network designed to achieve a preset game goal and make action decisions.
[0120] S1023. In the first game application, control the first virtual character to execute the action corresponding to the third action information.
[0121] S1024. After the first virtual character finishes executing the action corresponding to the third action information, obtain the fourth state information of the target virtual character in the first game application.
[0122] S1025. Use the current state information of the target virtual character as the first state information, the third action information as the first action information, and the fourth action state information as the second state information to construct a sample data and save the sample data to the sample data set.
[0123] As described above, the target behavior policy network is a network model that is oriented towards a preset game goal and is used to learn and make decisions on what actions the first virtual character should perform in the real game environment corresponding to the first game application. It should be emphasized that the goals of the above-mentioned preset behavior policy network and the target behavior policy network here are inconsistent. Among them, the preset behavior policy network aims to diversify the state information of the target virtual character that the first virtual character can explore, and thus make action decisions for the first virtual character. The target behavior policy network aims to enable the first virtual character to achieve a certain preset game goal (such as winning the current game round and other game goals), and thus make action decisions for the first virtual character. It can be understood that in practical applications, the behavior policy network is trained and optimized based on reinforcement learning to enable the game robot to complete training tasks or downstream tasks. The behavior policy network after training is the above-mentioned target behavior policy network.
[0124] In the embodiment of the present application, at the start of the first game application, the current state information of the target virtual character is obtained, and the current state information of the target virtual character is input into the target behavior policy network to obtain the third action information output by the target behavior policy network. In the game scenario of the first game application, the first virtual character is controlled to perform the action corresponding to the third action information, and the state information of the target virtual character changes accordingly, obtaining the changed state information of the target virtual character, that is, the fourth state information of the target virtual character. The current state information of the target virtual character is used as the first state information, the third action information is used as the first action information, and the fourth action state information is used as the second state information to construct a sample data and save the sample data to the sample data set. The above steps are repeatedly executed until the current round of the first game application ends, obtaining a sample data set including multiple sample data. In this way, in the real game environment of the first game application, during the process of training the target behavior policy network based on reinforcement learning with the training task or downstream task as the goal, the sample data obtained can be used to train the twin model (or preset model) corresponding to the first game application, which can further improve the similarity between the twin model and the first game application, so that the twin model can more accurately simulate the real game environment of the first game application to output states.
[0125] An optional implementation manner, the model training method provided in the embodiment of the present application may further include the following step S40:
[0126] S40. Repeatedly execute the following third step until the current round of the first game application ends, obtaining multiple target sample data. The target sample data is used for reinforcement learning training of the target behavior policy network. The target behavior policy network is a network that aims to achieve a preset game goal and make action decisions.
[0127] The third step includes S4011 - S4015:
[0128] S4011. Obtain the current state information of the target virtual character in the twin model.
[0129] S4012. Input the current state information into the target behavior policy network to obtain the fourth action information output by the target behavior policy network.
[0130] S4013. Input the current state information and the fourth action information into the twin model to obtain the fifth state information output by the twin model. The fifth state information is the state information of the target virtual character predicted by the twin model according to the action corresponding to the fourth action information executed by the first virtual character.
[0131] S4014. Determine the target reward value according to the preset reward function, the current state information of the target virtual character, and the fifth state information.
[0132] S4015. Construct a target sample data according to the current state information of the target virtual character, the fifth state information, the fourth action information, and the target reward value.
[0133] As described above, the target behavior policy network in step S4015 has the same meaning as the target behavior policy network in step S1022. For the description of the target behavior policy network, reference can be made to the above, and details are not repeated here.
[0134] In the embodiment of the present application, the above step S40 is used to generate subsequent sample data for training and optimizing the target behavior policy network based on reinforcement learning after obtaining the twin model corresponding to the first game application.
[0135] In the embodiment of the present application, the current state information of the target virtual character is obtained from the twin model, and the obtained current state information of the target virtual character is input into the target behavior policy network. Subsequently, the target behavior policy network outputs the fourth action information. The current state information of the target virtual character and the fourth action information are input into the twin model corresponding to the first game application together to obtain the fifth state information output by the twin model. Among them, the fifth state information is the state information of the target virtual character predicted by the twin model after the first virtual character executes the action corresponding to the fourth action information. The target reward value is determined according to the preset reward function, the current state information of the target virtual character, and the fifth state information. A target sample data is constructed according to the current state information of the target virtual character, the fifth state information, the fourth action information, and the target reward value.
[0136] Exemplarily, when the current state information of the target virtual character is s 5 and the fifth state information is s 5', the fourth action information is a 5 In the case of, the target reward value r is obtained based on the preset reward function R 5 = R(s 5 , s 5 '). According to the current state information s 5 , the first state information is s 5 ', the fourth action information is a 5 , the target reward value r 5 = R(s 5 , s 5 '), a target sample data is constructed, that is, (s 5 , s 5 ', a 5 , r 5 ) = (s 5 , s 5 ', a 5 , R(s 5 , s 5 ')).
[0137] In the embodiments of the present application, generating target sample data based on the trained twin model can greatly improve the production rate of target sample data. The amount of target sample data produced per unit time far exceeds the production volume of target sample data in the real game environment, thereby reducing the acquisition cost of target sample data. Further, using the twin model corresponding to the first game application as the game environment for reinforcement learning helps to greatly shorten the development cycle of game robots based on reinforcement learning and solves the problem that the game environment cannot be accelerated during the reinforcement learning training process.
[0138] An optional implementation manner is that the preset reward function is used to calculate the difference between the current state information of the target virtual character and the fifth state information, and determine the difference as the target reward value.
[0139] Exemplarily, assuming that the state information includes the blood volume and position information of the virtual character, and the target virtual characters include virtual character 1 (i.e., the first virtual character), virtual character 3, and virtual character 5. Then, assuming that the current state information and the fifth state information of the target virtual character are as shown in Table 2.
[0140] Table 2
[0141]
[0142] As shown in Table 2, the current state information of the target virtual characters: the health of virtual character 1 is 50, and the position information is (x11, y11); the health of virtual character 3 is 60, and the position information is (x13, y13); the health of virtual character 5 is 40, and the position information is (x15, y15). The fifth state information of the target virtual characters is: the health of virtual character 1 is 30, and the position information is (x21, y21); the health of virtual character 3 is 20, and the position information is (x23, y23); the health of virtual character 5 is 33, and the position information is (x25, y25).
[0143] Then, combining the current state information and the fifth state information of the target virtual characters shown in Table 2, the difference between the current state information and the fifth state information of the target virtual characters calculated can be referred to Table 3 shown below.
[0144] Table 3
[0145]
[0146]
[0147] It should be noted that when the state information includes position information, the difference between two position information can be the distance between two positions, such as Euclidean distance, Manhattan distance, etc. Only for example, the embodiments of the present application do not make any restrictions on this.
[0148] Next, in combination with Figure 3 , the process of obtaining the target behavior policy network by performing reinforcement learning on the twin model corresponding to the first game application obtained according to the model training method provided in the embodiments of the present application will be described. Figure 3 It is a schematic flowchart of the reinforcement learning training for the target behavior policy network provided in one embodiment of the present application.
[0149] First, the process of performing reinforcement learning on the target behavior policy network based on the real game environment, i.e., in the first game application, will be described.
[0150] S301. Obtain the initial target behavior policy network P.
[0151] S302. Obtain the current state information s1 of the target virtual character in the first game application.
[0152] S303. Input the current state information s1 of the target virtual character into the target behavior policy network P, and the target behavior policy network P outputs the action information a1 = P(s1).
[0153] S304. Control the first virtual character to perform the action corresponding to the action information a1 in the first game application, and obtain the next state information s1' and the reward value r1 of the target virtual character in the first game application.
[0154] S305. Optimize the target behavior policy network P using (s1, a1, s1', r1).
[0155] S306. Let s = s1', and determine whether the target behavior policy network P converges. If so, execute step S307; if not, execute steps S302 - S306.
[0156] S307. Obtain the trained target behavior policy network P.
[0157] Next, the process of performing reinforcement learning on the target behavior policy network based on the twin environment corresponding to the first game application, that is, the twin model corresponding to the first game application, will be described.
[0158] S301. Obtain the initial target behavior policy network P.
[0159] S308. Obtain the current state information s2 of the target virtual character in the first game application.
[0160] S309. Input the current state information s2 of the target virtual character into the target behavior policy network P, and the target behavior policy network P outputs the action information a2 = P(s2).
[0161] S310. Input the current state information s2 of the target virtual character and the first action information a2 into the twin model W corresponding to the first game application, and obtain the next state information s2' = W(s2, a2) of the target virtual character and the reward value r2 output by the twin model.
[0162] S311. Optimize the target behavior policy network P using (s2, a2, s2', r2).
[0163] S312. Let s = s2', and determine whether the target behavior policy network P converges. If so, execute step S307; if not, execute steps S308 - S312.
[0164] S307. Obtain the trained target behavior policy network P.
[0165] Next, the model training device provided by the present application will be described. The model training device described below can be mutually corresponding and referred to with the model training method described above.
[0166] Figure 4 It is a schematic structural diagram of the model training device provided in one embodiment of the present application. As Figure 4As shown in the figure, the model training device includes: an acquisition module 401, a processing module 402, and a control module 403.
[0167] The acquisition module is configured to acquire a sample data set, where the sample data set includes multiple pieces of sample data generated during the running of a first game application. Each piece of sample data includes first state information and second state information of a target virtual character in the first game application, and first action information of a first virtual character; wherein, the target virtual character includes the first virtual character and a second virtual character, and the second virtual character refers to a virtual character within the field of view of the first virtual character in the game scene of the first game application. The second state information is the state information of the target virtual character after the first virtual character executes the action corresponding to the first action information.
[0168] The processing module is configured to input the first state information and the first action information in the sample data into a preset model to obtain predicted state information output by the preset model. The predicted state information is the state information of the target virtual character predicted by the preset model according to the action corresponding to the first action information executed by the first virtual character.
[0169] The control module is configured to train the preset model with the goal of minimizing a preset loss function, adjust the parameters of the preset model, and use the trained preset model as the twin model corresponding to the first game application; wherein, the preset loss function is used to calculate the difference between the predicted state information and the second state information; the twin model is used to predict the next state information of the target virtual character in the first game application according to the current state information of the target virtual character and the first action information executed by the first virtual character.
[0170] Optionally, the acquisition module is specifically configured to:
[0171] Loop through the following first step until the number of sample data in the sample data set reaches a preset number threshold;
[0172] The first step includes:
[0173] Input the current state information and historical state information of the target virtual character in the first game application into a preset behavior policy network to obtain second action information output by the preset behavior policy network. The preset behavior policy network is used to predict the action information that the first virtual character will execute next according to the current state information and historical state information of the target virtual character.
[0174] In the first game application, control the first virtual character to execute the action corresponding to the second action information.
[0175] After the first virtual character finishes executing the action corresponding to the second action information, obtain third state information of the target virtual character in the first game application;
[0176] Use the current state information of the target virtual character as the first state information, the second action information as the first action information, and the third state information as the second state information, construct a sample data and save the sample data to the sample data set.
[0177] Optionally, the first step further includes:
[0178] Input the third state information into a first network and a second network, obtain a first prediction value output by the first network and a second prediction value output by the second network, where the network structures of the first network and the second network are the same;
[0179] According to the difference between the first prediction value and the second prediction value, adjust the parameters of the second network, and determine the difference as the reward information corresponding to the second action information;
[0180] Construct a training data according to the current state information, the third state information, the second action information, and the reward information of the target virtual character, and perform reinforcement learning training on the preset behavior policy network based on the training data.
[0181] Optionally, the obtaining module is further configured to:
[0182] Loop to execute the following second step until the current round of the first game application ends, to obtain a sample data set including multiple sample data;
[0183] The second step includes:
[0184] Obtain the current state information of the target virtual character in the first game application;
[0185] Input the current state information of the target virtual character into a target behavior policy network, obtain third action information output by the target behavior policy network, where the target behavior policy network is a network designed to achieve a preset game goal and make action decisions;
[0186] In the first game application, control the first virtual character to execute the action corresponding to the third action information;
[0187] After the first virtual character finishes executing the action corresponding to the third action information, obtain fourth state information of the target virtual character in the first game application;
[0188] Use the current state information of the target virtual character as the first state information, use the third action information as the first action information, and use the fourth action state information as the second state information to construct a sample data and save the sample data to the sample data set.
[0189] Optionally, the device further includes a training module, and the training module is specifically configured to:
[0190] Loop through the following third step until the current round of the first game application ends, to obtain multiple target sample data, which are used for reinforcement learning training of the target behavior policy network. The target behavior policy network is a network that aims to achieve a preset game goal and make action decisions;
[0191] The third step includes:
[0192] Obtain the current state information of the target virtual character in the twin model;
[0193] Input the current state information into the target behavior policy network to obtain the fourth action information output by the target behavior policy network;
[0194] Input the current state information and the fourth action information into the twin model to obtain the fifth state information output by the twin model. The fifth state information is the state information of the target virtual character predicted by the twin model according to the action corresponding to the fourth action information executed by the first virtual character;
[0195] Determine the target reward value according to the preset reward function, the current state information of the target virtual character, and the fifth state information;
[0196] Construct a target sample data according to the current state information of the target virtual character, the fifth state information, the fourth action information, and the target reward value.
[0197] Optionally, the preset reward function is used to calculate the difference between the current state information of the target virtual character and the fifth state information, and determine the difference as the target reward value.
[0198] Optionally, the network structure of the preset model includes any one of the following: feedforward neural network, convolutional neural network, recurrent neural network, Transform network.
[0199] The model training device provided in this embodiment can be used to execute the technical solution of the above-mentioned model training method embodiment. The implementation principle and technical effect are similar, and will not be elaborated here in this embodiment.
[0200] Figure 5 It is a schematic diagram of the hardware structure of an electronic device provided in one embodiment of the present application. As Figure 5 shown, the electronic device 500 in this embodiment includes: a processor 501 and a memory 502; among them
[0201] The memory 502 is used to store computer execution instructions;
[0202] The processor 501 is used to execute the computer execution instructions stored in the memory to implement each step executed by the model training method in the above embodiment. Specifically, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0203] Optionally, the memory 502 can be either independent or integrated with the processor 501.
[0204] When the memory 502 is independently provided, the electronic device further includes a bus 503 for connecting the memory 502 and the processor 501.
[0205] One embodiment of the present application further provides a computer-readable storage medium, in which computer execution instructions are stored. When the processor executes the computer execution instructions, the technical solution corresponding to the model training method executed by the above electronic device is implemented.
[0206] One embodiment of the present application further provides a computer program product. The program product includes: a computer program. The computer program is stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program to enable the electronic device to execute the technical solution corresponding to the model training method in any of the above embodiments.
[0207] Although the present application is disclosed above with preferred embodiments, it is not used to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be defined by the scope of the claims of the present application.
[0208] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be indirect couplings or communication connections through some interfaces, devices or modules, and can be in electrical, mechanical or other forms.
[0209] The integrated modules implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium and include several instructions to enable an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in various embodiments of the present application.
[0210] It should be understood that the above processor can be a central processing module (English: Central Processing Unit, abbreviated as: CPU), and can also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application-specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor.
[0211] The memory may include a high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk, or an optical disc, etc.
[0212] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.
[0213] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0214] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks or optical disks that can store program codes.
[0215] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A model training method, characterized in that, the method includes: Obtain a sample data set, where the sample data set includes multiple pieces of sample data generated during the operation of a first game application. Each piece of sample data includes first state information and second state information of a target virtual character in the first game application, and first action information of a first virtual character; wherein, the target virtual character includes the first virtual character and a second virtual character, and the second virtual character refers to a virtual character within the field of view of the first virtual character in the game scene of the first game application, and the second state information is the state information of the target virtual character after the first virtual character executes the action corresponding to the first action information; Input the first state information and the first action information in the sample data into a preset model to obtain predicted state information output by the preset model. The predicted state information is the state information of the target virtual character predicted by the preset model according to the action corresponding to the first action information executed by the first virtual character; Train the preset model with the goal of minimizing a preset loss function, adjust the parameters of the preset model, and use the trained preset model as the twin model corresponding to the first game application; wherein, the preset loss function is used to calculate the difference between the predicted state information and the second state information, and the twin model is used to predict the next state information of the target virtual character in the first game application according to the current state information of the target virtual character and the first action information executed by the first virtual character.
2. The method according to claim 1, characterized in that, the obtaining of the sample data set includes: Loop and execute the following first step until the number of sample data in the sample data set reaches a preset number threshold; The first step includes: Input the current state information and historical state information of the target virtual character in the first game application into a preset behavior policy network to obtain second action information output by the preset behavior policy network. The preset behavior policy network is used to predict the action information that the first virtual character will execute next according to the current state information and historical state information of the target virtual character; In the first game application, control the first virtual character to execute the action corresponding to the second action information; After the first virtual character finishes executing the action corresponding to the second action information, obtain the third state information of the target virtual character in the first game application; Use the current state information of the target virtual character as the first state information, use the second action information as the first action information, and use the third state information as the second state information to construct a piece of sample data and save the sample data to the sample data set.
3. The method according to claim 2, characterized in that, the first step further includes: Input the third state information into the first network and the second network to obtain a first predicted value output by the first network and a second predicted value output by the second network. The network structures of the first network and the second network are the same; Adjust the parameters of the second network according to the difference between the first predicted value and the second predicted value, and determine the difference as the reward information corresponding to the second action information; Construct a piece of training data based on the current state information of the target virtual character, the third state information, the second action information, and the reward information, and perform reinforcement learning training on the preset behavior policy network based on the training data.
4. The method according to claim 1, wherein, the obtaining of the sample data set includes: Repeatedly execute the following second step until the current round of the first game application ends, to obtain a sample data set including multiple pieces of sample data; The second step includes: Obtain the current state information of the target virtual character in the first game application; Input the current state information of the target virtual character into the target behavior policy network to obtain a third action information output by the target behavior policy network. The target behavior policy network is a network designed to achieve a preset game goal and make action decisions; In the first game application, control the first virtual character to execute the action corresponding to the third action information; After the first virtual character finishes executing the action corresponding to the third action information, obtain the fourth state information of the target virtual character in the first game application; Use the current state information of the target virtual character as the first state information, the third action information as the first action information, and the fourth action state information as the second state information, construct a piece of sample data and save the sample data to the sample data set.
5. The method according to claim 1, wherein, the method further includes: Repeatedly execute the following third step until the current round of the first game application ends, to obtain multiple pieces of target sample data for performing reinforcement learning training on the target behavior policy network. The target behavior policy network is a network designed to achieve a preset game goal and make action decisions; The third step includes: Obtain the current state information of the target virtual character in the twin model; Input the current state information into the target behavior policy network to obtain a fourth action information output by the target behavior policy network; Input the current state information and the fourth action information into the twin model to obtain a fifth state information output by the twin model. The fifth state information is the state information of the target virtual character predicted by the twin model according to the action executed by the first virtual character corresponding to the fourth action information; Determine a target reward value according to a preset reward function, the current state information of the target virtual character, and the fifth state information. Construct a target sample data based on the current state information, the fifth state information, the fourth action information of the target virtual character, and the target reward value.
6. The method according to claim 5, wherein, the preset reward function is used to calculate the difference between the current state information of the target virtual character and the fifth state information, and determine the difference as the target reward value.
7. The method according to claim 1, wherein, the network structure of the preset model includes any one of the following: feedforward neural network, convolutional neural network, recurrent neural network, and Transform network.
8. A model training device, wherein, the device includes: An acquisition module, configured to acquire a sample data set, where the sample data set includes multiple sample data generated during the operation of a first game application, and each sample data includes first state information and second state information of a target virtual character in the first game application, and first action information of a first virtual character; wherein, the target virtual character includes the first virtual character and a second virtual character, and the second virtual character refers to a virtual character within the field of view of the first virtual character in the game scene of the first game application, and the second state information is the state information of the target virtual character after the first virtual character executes the action corresponding to the first action information; A processing module, configured to input the first state information and the first action information in the sample data into a preset model, and obtain predicted state information output by the preset model, where the predicted state information is the state information of the target virtual character predicted by the preset model according to the action corresponding to the first virtual character executing the first action information; A control module, configured to train the preset model with the goal of minimizing a preset loss function, adjust the parameters of the preset model, and use the trained preset model as the twin model corresponding to the first game application; wherein, the preset loss function is used to calculate the difference between the predicted state information and the second state information, and the twin model is used to predict the next state information of the target virtual character in the first game application according to the current state information of the target virtual character and the first action information executed by the first virtual character.
9. An electronic device, wherein, the electronic device includes: A processor; and A memory, configured to store a data processing program, and after the electronic device is powered on and runs the program through the processor, execute the model training method according to any one of claims 1-7.
10. A computer-readable storage medium, wherein, stores a data processing program, and the program is run by a processor to execute the model training method according to any one of claims 1-7.
Citation Information
Patent Citations
Model training method and device, storage medium and electronic equipment
CN113384875A
Motion capture method and device, electronic equipment and storage medium
CN115758157A
Decision-making method for multi-unmanned aerial vehicle network based on imperfect digital twinning
CN116828507A
Class recognition model training method and device, equipment and storage medium
CN117009508A
Data augmentation and batch balancing methods to enhance negation and fairness
US20230153528A1