Model training method and device, electronic equipment and readable storage medium
Patent Information
- Application Number
- CN202510357563.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2026-09-25
AI Technical Summary
然而,游戏场景通常比较复杂且多变,如何控制战斗机器人在复杂多变的游戏场景中模拟人类玩家的决策方式,在游戏对战中表现得更加智能和灵活,是目前亟需解决的问题
[0020]本申请提供的模型训练方法,针对受控虚拟角色获取历史游戏状态数据、与历史游戏状态数据的产生时刻间隔预设时长后受控虚拟角色到达的历史目的位置和历史游戏状态数据对应的历史动作控制数据;根据历史游戏状态数据和受控虚拟角色到达的历史目的位置,生成包括多个状态维度的游戏状态子样本的样本;根据历史动作控制数据,生成样本对应的标签,其中,标签包括多个控制维度的动作控制子标签;根据样本和标签,对待训练的动作预测模型进行训练,得到目标动作预测模型;针对目标动作预测模型配置用于确定目的位置的策略。如此,使得目标动作预测模型输出的动作更加接近人类水平。
Smart Images

Figure CN122806066A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of gaming, specifically to a model training method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] In the gaming industry, combat robots can be controlled via models to perform actions such as movement and shooting, providing players with diverse gaming experiences. However, game scenarios are usually complex and varied. How to control combat robots to simulate the decision-making methods of human players in complex and varied game scenarios, and to make them perform more intelligently and flexibly in game battles, is a problem that urgently needs to be solved. Summary of the Invention
[0003] In view of this, this application provides a model training method, apparatus, electronic device and readable storage medium, so that the action output by the trained target action prediction model is closer to human level.
[0004] In a first aspect, embodiments of this application provide a model training method, the method comprising:
[0005] The controlled virtual character acquires historical game state data, the historical target location reached by the controlled virtual character after a preset time interval between the generation time of the historical game state data and the historical action control data corresponding to the historical game state data. The controlled virtual character is a virtual character controlled by the player object.
[0006] Based on historical game state data and the historical destinations reached by the controlled virtual character, a sample of game state sub-samples including multiple state dimensions is generated.
[0007] Based on historical motion control data, generate labels corresponding to the samples. The labels include motion control sub-labels for multiple control dimensions.
[0008] Based on the samples and labels, the action prediction model to be trained is trained to obtain the target action prediction model;
[0009] Configure a strategy for the target action prediction model to determine the target location.
[0010] Secondly, embodiments of this application provide a model training apparatus, the apparatus comprising:
[0011] The acquisition module is used to acquire historical game state data of the controlled virtual character, the historical destination location reached by the controlled virtual character after a preset time interval from the time when the historical game state data was generated, and the historical action control data corresponding to the historical game state data. The controlled virtual character is a virtual character controlled by the player object.
[0012] The first generation module is used to generate a sample of game state sub-samples including multiple state dimensions based on historical game state data and the historical destination locations reached by the controlled virtual character.
[0013] The second generation module is used to generate labels corresponding to samples based on historical motion control data. The labels include motion control sub-labels for multiple control dimensions.
[0014] The training module is used to train the action prediction model to be trained based on samples and labels, so as to obtain the target action prediction model.
[0015] The configuration module is used to configure the strategy for determining the target location for the target action prediction model.
[0016] Thirdly, embodiments of this application provide an electronic device, including:
[0017] Processor; and
[0018] The memory is used to store data processing programs, which, when the electronic device is powered on and run by the processor, execute the method as described in the first aspect.
[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing a data processing program that is executed by a processor to perform the method as described in the first aspect.
[0020] The model training method provided in this application involves acquiring historical game state data of a controlled virtual character, the historical destination location reached by the controlled virtual character after a preset time interval from the generation time of the historical game state data, and historical action control data corresponding to the historical game state data. Based on the historical game state data and the historical destination location reached by the controlled virtual character, a sample of game state sub-samples including multiple state dimensions is generated. Based on the historical action control data, labels corresponding to the samples are generated, where the labels include action control sub-labels with multiple control dimensions. Based on the samples and labels, the action prediction model to be trained is trained to obtain the target action prediction model. A strategy for determining the destination location is configured for the target action prediction model. This makes the actions output by the target action prediction model more closely resemble human-level performance. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1A flowchart illustrating an example of a model training method provided in this application embodiment;
[0023] Figure 2 A schematic diagram of the structure of the model training device provided in the embodiments of this application;
[0024] Figure 3 A structural block diagram of an electronic device for implementing a model training method is provided for embodiments of this application. Detailed Implementation
[0025] Many specific details are set forth in the following description to provide a full understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this application; therefore, this application is not limited to the specific embodiments disclosed below.
[0026] It should be noted that the terms "first," "second," "third," etc., in the claims, specification, and drawings of this application are used to distinguish similar objects and are not used to describe a specific order or sequence. Such data are interchangeable where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown or described herein. Furthermore, the terms "comprising," "having," and their variations are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.
[0027] It should be understood that in the embodiments of this application, "at least one" means one or more, and "more than one" means two or more. "And / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "Contains A, B and / or C" means containing any one, two, or three of A, B, and C.
[0028] It should be understood that in the embodiments of this application, "B corresponding to A", "B corresponding to A", "A corresponds to B" or "B corresponds to A" means that B is associated with A, and B can be determined based on A. Determining B based on A does not mean that B is determined solely based on A; B can also be determined based on A and / or other information.
[0029] The model training method provided in this application can be executed by an electronic device, such as a terminal or a server. The terminal can be a smartphone, tablet, laptop, or other similar device. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. It is understood that this application does not limit the specific entity executing the model training method.
[0030] Before detailing the technical solution of this application, the relevant technologies will be further introduced.
[0031] In the gaming industry, combat robots can be controlled via models to perform actions such as movement and shooting, providing players with diverse gaming experiences. However, game scenarios are usually complex and varied. How to control combat robots to simulate the decision-making methods of human players in complex and varied game scenarios, and to make them perform more intelligently and flexibly in game battles, is a problem that urgently needs to be solved.
[0032] This application addresses the problem of training a model for controlling the movement of a combat robot, as follows:
[0033] 1. In the existing technology, the model trained in this way can only control combat robots to perform a few actions that are close to human level.
[0034] 2. The samples obtained in the existing technology may have an uneven distribution of actions, which may cause the trained model to learn only a single action.
[0035] 3. In the existing technology, the model trained in this way cannot control the robot to make accurate movement strategies based on game state data.
[0036] It should be noted that the information described above is only for enhancing the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art.
[0037] The model training method provided in this application can train a target action prediction model that can control an intelligent agent character to perform more actions close to human level; by preprocessing the samples, the situation where the trained target action prediction model only learns a single action is avoided; by configuring rules for determining the target position for the target action prediction model, it is possible to control the intelligent agent character to move to the target position through the target action prediction model, and then control the intelligent agent character to complete the underlying action control, thereby realizing macroscopic perception of the game.
[0038] The technical solution of this application will be described in detail below through specific embodiments. It should be noted that the following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments described below are used to explain the technical solution of this application and are not intended to limit actual use.
[0039] Figure 1 This application illustrates a model training method provided in one embodiment, such as... Figure 1 As shown, the method may include the following steps S101 to S105.
[0040] Step S101: For the controlled virtual character, obtain historical game state data, the historical target position reached by the controlled virtual character after a preset time interval between the generation time of the historical game state data and the historical action control data corresponding to the historical game state data. The controlled virtual character is a virtual character controlled by the player object.
[0041] It should be noted that the player object refers to a real player or the device or program acting as a player, that is, the controller of a controlled virtual object. In the following text, the player object will sometimes be simply referred to as the player. Historical game state data refers to the collection of all relevant information recorded at a specific point in time during the game. This data captures the state of the game at a particular moment in its lifecycle, allowing developers, researchers, or players to review and analyze past game events.
[0042] In one optional embodiment, the historical game state data can be classified according to the dimension and category to which it belongs. For example, the historical game state data can first be divided into global state sub-data, item state sub-data, controlled virtual character state sub-data, environmental awareness state sub-data, and character configuration state sub-data, and then each sub-data can be further divided according to its category.
[0043] In one optional embodiment, the global state sub-data includes at least one of the following: game progress time, the location of the virtual object, whether the virtual object has been deployed, and the remaining startup time of the virtual object. Here, game progress time refers to the elapsed time within the game from the start of the game to the current moment.
[0044] In one alternative embodiment, the virtual object may be a virtual explosive or a virtual machine, etc., and this application does not limit it in this regard.
[0045] In one alternative embodiment, the virtual explosive can be a virtual bomb, and the location of the virtual object can be the coordinates of the virtual object in the game scene.
[0046] In one alternative embodiment, the item status sub-data includes at least one of the following: the type of item held by the controlled virtual character, the quantity of the item, and the time interval since the controlled virtual character last used the item.
[0047] In one alternative embodiment, the props can be rifles, machine guns, pistols, bullets, flashbangs, smoke grenades, and incendiary grenades.
[0048] In one optional embodiment, the controlled virtual character state sub-data includes at least one of the following: the location of the controlled virtual character, whether the controlled virtual character belongs to the attacker or defender, and the health value of the controlled virtual character. It should be noted that a game scene may include at least one controlled virtual character.
[0049] In one optional embodiment, the environmental awareness state sub-data includes at least one of the following: whether a ray emitted by the controlled virtual character's body hits an object, and the distance of the hit object from the controlled virtual character. It should be noted that the environmental awareness state sub-data is used to analyze the surrounding environment and avoid obstacles.
[0050] In one optional embodiment, the character configuration status sub-data includes at least one of the following: character configuration information of the controlled virtual character and skill configuration information of the controlled virtual character. The character configuration information indicates the hero category used by the controlled virtual character, and the skill configuration information indicates the skill category used by the controlled virtual character.
[0051] In one alternative embodiment, the state sub-data of multiple dimensions can be represented in the form of vectors and matrices. For example, s = [a, b, c, d, e], where a represents global state sub-data, b represents item state sub-data, c represents controlled virtual character state sub-data, d represents environment awareness state sub-data, and e represents character configuration state sub-data.
[0052] Where 'a' can be a 20-dimensional vector, representing the information included in the global state sub-data. 'b' can be a 22-dimensional vector. Considering that a game scene may include at least one controlled virtual character, a multi-dimensional matrix is used to represent the state sub-data of the controlled virtual character. For example, if the game scene includes 10 controlled virtual characters, then 'c' can be a 10*30 matrix. 'd' can be a 130-dimensional vector, and 'e' can be a 17-dimensional vector.
[0053] The moment the historical game state data is generated refers to the moment when the game state represented by the historical game state data is captured. The historical destination location reached by the controlled virtual character after a preset time interval from the moment the historical game state data is generated refers to the position coordinates of the controlled virtual character in the game scene after the preset time interval. For example, the historical destination location can be the location reached by the controlled virtual character 5 seconds after the moment the historical game state data is generated.
[0054] Historical action control data corresponding to historical game state data refers to the data corresponding to the actions performed by the controlled virtual character in the game state represented by the historical game state data.
[0055] In one optional embodiment, historical motion control data can be classified according to its dimension and category. For example, the historical motion control data can first be divided into movement direction sub-data, posture sub-data, horizontal view sub-data, and vertical view sub-data, and then each sub-data can be further divided according to its category.
[0056] In one alternative embodiment, the movement direction data includes at least one of the following: stationary, moving forward, moving backward, moving up, moving down, moving to the upper left, moving to the lower left, moving to the upper right, and moving to the lower right.
[0057] In one alternative embodiment, the posture sub-data includes at least one of the following: stationary, sprinting, crouching, jumping, aiming, crouching while aiming down sights.
[0058] In one alternative embodiment, the horizontal viewpoint sub-data includes at least one of the following: stationary, leftward movement speed of the first viewpoint, leftward movement speed of the second viewpoint, leftward movement speed of the third viewpoint, leftward movement speed of the fourth viewpoint, leftward movement speed of the fifth viewpoint, leftward movement speed of the sixth viewpoint, rightward movement speed of the first viewpoint, rightward movement speed of the second viewpoint, rightward movement speed of the third viewpoint, rightward movement speed of the fourth viewpoint, rightward movement speed of the fifth viewpoint, and rightward movement speed of the sixth viewpoint.
[0059] In one alternative embodiment, the vertical viewpoint sub-data includes at least one of the following: stationary, upward movement speed of the first viewpoint, upward movement speed of the second viewpoint, upward movement speed of the third viewpoint, upward movement speed of the fourth viewpoint, upward movement speed of the fifth viewpoint, downward movement speed of the first viewpoint, downward movement speed of the second viewpoint, downward movement speed of the third viewpoint, downward movement speed of the fourth viewpoint, and downward movement speed of the fifth viewpoint.
[0060] In one alternative embodiment, the motion control sub-data of multiple dimensions can be represented in the form of vectors. For example, m = [f, g, h, k], where f represents the movement direction sub-data, g represents the attitude sub-data, h represents the horizontal view sub-data, and k represents the vertical view sub-data.
[0061] Where f can be a 9-dimensional vector, g can be a 6-dimensional vector, h can be a 13-dimensional vector, and k can be an 11-dimensional vector.
[0062] In one optional embodiment, historical game state data, historical destination locations reached by the controlled virtual character after a preset time interval from the time the historical game state data was generated, and historical action control data corresponding to the historical game state data can be obtained from multiple sources such as game logs and user behavior records.
[0063] In this step, historical game state data, the historical target position reached by the controlled virtual character after a preset time interval from the time the historical game state data was generated, and the historical motion control data corresponding to the historical game state data are acquired. This allows for the generation of samples based on the acquired data, and the motion prediction model to be trained is then trained using these samples to obtain the target motion prediction model. Furthermore, by acquiring historical motion control data including at least 9 movement directions (including stationary), at least 6 self-poses (including stationary), at least 13 horizontal viewpoint movement speeds (including stationary), and at least 11 vertical viewpoint movement speeds (including stationary), the motion space is expanded, enabling the trained motion prediction model to control the combat robot to perform more actions.
[0064] Step S102: Based on historical game state data and the historical destination locations reached by the controlled virtual character, generate a sample of game state sub-samples including multiple state dimensions.
[0065] In one optional embodiment, generating a sample that includes multiple state dimensions of game state sub-samples based on historical game state data and the historical destination locations reached by the controlled virtual character refers to extracting features from historical game state data and the historical destination locations reached by the controlled virtual character to characterize the historical game state data and the historical destination locations reached by the controlled virtual character, first constructing sub-samples corresponding to the state dimensions, and then constructing a sample based on the sub-samples of each corresponding state.
[0066] In one optional embodiment, the samples include sub-samples of the global state dimension, sub-samples of the item state dimension, sub-samples of the controlled virtual character state dimension, sub-samples of the environment-aware state dimension, and sub-samples of the character configuration state dimension. The sub-samples of the global state dimension include global state sub-data and historical destination location; the sub-samples of the item state dimension include item state sub-data; the sub-samples of the controlled virtual character state dimension include controlled virtual character state sub-data; the sub-samples of the environment-aware state dimension include environment-aware state sub-data; and the sub-samples of the character configuration state dimension include character configuration state sub-data.
[0067] In one optional embodiment, a sample of game state sub-samples including multiple state dimensions is generated based on historical game state data and the historical destination locations reached by the controlled virtual character. This includes: generating a sub-sample of the global state dimension based on the historical destination locations and global state sub-data included in the historical game state data; generating a sub-sample of the item state dimension based on the item state sub-data included in the historical game state data; generating a sub-sample of the controlled virtual character state dimension based on the controlled virtual character state sub-data included in the historical game state data; generating a sub-sample of the environment perception state dimension based on the environment perception state sub-data included in the historical game state data; and generating a sub-sample of the character configuration state dimension based on the character configuration state sub-data included in the historical game state data.
[0068] Step S103: Generate labels corresponding to the samples based on historical motion control data. The labels include motion control sub-labels for multiple control dimensions.
[0069] In one optional embodiment, generating motion control sub-labels that include multiple control dimensions based on historical motion control data can be achieved by extracting features from historical motion control data to characterize the historical motion control data, defining and generating motion control sub-labels corresponding to the control dimensions based on the features.
[0070] In one optional embodiment, the historical motion control data includes movement direction sub-data, posture sub-data, horizontal viewpoint sub-data, and vertical viewpoint sub-data. Based on the historical motion control data, a tag including multiple control dimensions of motion control sub-tags is generated, including: generating motion control sub-tags including the movement direction dimension based on the movement direction sub-data; generating motion control sub-tags including the posture dimension based on the posture sub-data; generating motion control sub-tags including the horizontal viewpoint dimension based on the horizontal viewpoint sub-data; and generating motion control sub-tags including the vertical viewpoint dimension based on the vertical viewpoint sub-data.
[0071] In one alternative embodiment, the motion control sub-label of the movement direction dimension includes at least one of the following: stationary, moving forward, moving backward, moving up, moving down, moving to the upper left, moving to the lower left, moving to the upper right, and moving to the lower right.
[0072] In one alternative embodiment, the motion control sub-label of the posture dimension includes at least one of the following: stationary, sprinting, crouching, jumping, aiming, crouching to scope.
[0073] In one optional embodiment, the motion control sub-label in the horizontal view dimension includes at least one of the following: stationary, leftward movement speed of the first view, leftward movement speed of the second view, leftward movement speed of the third view, leftward movement speed of the fourth view, leftward movement speed of the fifth view, leftward movement speed of the sixth view, rightward movement speed of the first view, rightward movement speed of the second view, rightward movement speed of the third view, rightward movement speed of the fourth view, rightward movement speed of the fifth view, and rightward movement speed of the sixth view.
[0074] In one alternative embodiment, the motion control sub-label in the vertical view dimension includes at least one of the following: stationary, first view upward movement speed, second view upward movement speed, third view upward movement speed, fourth view upward movement speed, fifth view upward movement speed, first view downward movement speed, second view downward movement speed, third view downward movement speed, fourth view downward movement speed, and fifth view downward movement speed.
[0075] In an optional embodiment, considering that the constructed samples may contain low-quality samples and the sample distribution may be uneven, the samples can be screened before training the action prediction model based on the samples and labels. The specific process is as follows:
[0076] In one possible implementation, samples corresponding to labels where the action control sub-labels in multiple control dimensions are all static are deleted. This avoids situations where the trained target action prediction model can only output static actions.
[0077] In one possible implementation, samples corresponding to labels where the action control sub-label is stationary in the horizontal view dimension are deleted. For example, 40% of samples corresponding to labels where the action control sub-label is stationary in the horizontal view dimension are deleted. By deleting samples corresponding to labels where the action control sub-label is stationary in the horizontal view dimension, it is ensured that the trained target action prediction model outputs the correct horizontal view orientation, avoiding the situation where the trained target action prediction model outputs an incorrect horizontal view orientation.
[0078] In one possible implementation, samples corresponding to labels where the motion control sub-label in the movement direction dimension is stationary are deleted. It's easy to understand that if the controlled virtual character's current position in the historical state data differs from its historical destination position (the controlled virtual character's position changes from point A to point B), the movement direction dimension sub-data in the historical motion control data corresponding to the historical state data should not be stationary. That is, samples corresponding to labels where the motion control sub-label in the movement direction dimension is stationary will have a negative impact on model training.
[0079] In one possible implementation, historical destination locations reached by the controlled virtual character are removed from a portion of the samples. For example, 10% of the historical destination locations reached by the controlled virtual character are randomly removed. This allows the trained target action prediction model to output accurate actions even without configured destination locations.
[0080] By removing samples that negatively impact model training, the model can learn more realistic patterns, improving its prediction accuracy, stability, and generalization ability. Furthermore, reducing the number of samples allows the model to converge faster.
[0081] Step S104: Train the action prediction model to be trained based on the samples and labels to obtain the target action prediction model.
[0082] It should be noted that the action prediction model to be trained can be a deep learning model architecture for processing sequence data, such as RNN (Recurrent Neural Networks), Transformer, GRU (Gated Recurrent Unit), TCN (Temporal Convolutional Network), and LSTM (Long Short-Term Memory), etc. This application does not limit it.
[0083] In one optional embodiment, training the action prediction model to be trained based on samples and labels includes: dividing the samples into a training set and a test set; and training the action prediction model to be trained based on the samples in the training set and the labels corresponding to the samples in the training set.
[0084] In one optional embodiment, the action prediction model to be trained is trained based on samples in the training set and the labels corresponding to the samples in the training set, including:
[0085] The samples in the training set are input into the action prediction model to be trained to obtain the predicted actions; the parameters of the action prediction model to be trained are adjusted according to the loss function between the predicted actions and the labels corresponding to the samples in the training set.
[0086] After training the action prediction model based on the samples and labels in the training set, the model can be tested using samples and labels in the test set to prevent overfitting.
[0087] In one optional embodiment, training the action prediction model to be trained based on samples and labels further includes the following steps:
[0088] Step 1: Test the action prediction model trained on the training set based on the samples in the test set and the corresponding labels of the samples in the test set, and obtain the predicted action control data corresponding to the samples in the test set.
[0089] Step 2: Statistically analyze the distribution of actions across different control dimensions in the predicted action control data corresponding to the samples in the test set to obtain the first action distribution;
[0090] Step 3: Statistically analyze the action distribution across different control dimensions of the labels corresponding to the samples in the test set to obtain the second action distribution;
[0091] Step 4: Adjust the sample weights of the samples in the training set based on the difference between the first action distribution and the second action distribution;
[0092] Step 5: Based on the updated sample weights and labels, train the action prediction model trained on the training set until the difference between the first action distribution and the second action distribution meets the preset conditions.
[0093] In step one, the action prediction model trained on the training set is tested based on the samples in the test set and the labels corresponding to the samples in the test set. This includes: inputting the samples in the test set into the action prediction model trained on the training set to obtain the predicted action control data corresponding to the samples in the test set.
[0094] In step two, the motion distribution in different control dimensions of the predicted motion control data corresponding to the samples in the test set can be the data distribution in the movement direction dimension, the data distribution in the posture dimension, the data distribution in the horizontal view dimension, and the data distribution in the vertical view dimension of the predicted motion control data corresponding to the samples in the test set.
[0095] In one optional embodiment, the data distribution along the control dimension in the predicted action control data corresponding to the samples in the test set can be the proportion of the number of actions of each category along that control dimension to the total number of actions along that control dimension. Of course, the data distribution can also take other forms, and this application does not limit it in this way.
[0096] In one optional embodiment, the data distribution in the movement direction dimension of the predicted action control data corresponding to the samples in the test set can be the proportion of the number of actions in the nine categories of stationary, forward, backward, upward, downward, upper left, lower left, upper right, and lower right to the total number of actions in the movement direction dimension. For example, if the test set includes 360 samples, the total number of actions in the movement direction dimension is 360. Among them, the number of stationary actions in the predicted action control data corresponding to the samples in the test set is 36, the number of forward movements is 72, the number of backward movements is 36, the number of upward movements is 12, the number of downward movements is 48, the number of upper left movements is 60, the number of lower left movements is 24, the number of upper right movements is 48, and the number of lower right movements is 24. The data distribution in the direction of movement of the predicted motion control data corresponding to the samples in the test set can be [1 / 10, 1 / 5, 1 / 10, 1 / 30, 2 / 15, 1 / 6, 1 / 15, 2 / 15, 1 / 15].
[0097] In one optional embodiment, the data distribution in the posture dimension of the predicted motion control data corresponding to the samples in the test set can be the proportion of the number of six types of actions—stationary, sprinting, crouching, aiming, and crouching while aiming down sights—to the total number of actions in the posture dimension.
[0098] In one optional embodiment, the data distribution in the horizontal view dimension of the predicted motion control data corresponding to the samples in the test set can be the proportion of the number of actions in the 12 categories—stationary, first view left movement speed, second view left movement speed, third view left movement speed, fourth view left movement speed, fifth view left movement speed, sixth view left movement speed, first view right movement speed, second view right movement speed, third view right movement speed, fourth view right movement speed, fifth view right movement speed, and sixth view right movement speed—to the total number of actions in the horizontal view dimension.
[0099] In one optional embodiment, the data distribution in the vertical view dimension of the predicted motion control data corresponding to the samples in the test set can be the proportion of the number of actions in the vertical view dimension of 10 categories: stationary, first view upward movement speed, second view upward movement speed, third view upward movement speed, fourth view upward movement speed, fifth view upward movement speed, first view downward movement speed, second view downward movement speed, third view downward movement speed, fourth view downward movement speed, and fifth view downward movement speed.
[0100] In step three, the action distribution in different control dimensions of the labels corresponding to the samples in the test set can be the data distribution in the movement direction dimension of the labels corresponding to the samples in the test set, the data distribution in the posture dimension of the labels corresponding to the samples in the test set, the data distribution in the horizontal view dimension of the labels corresponding to the samples in the test set, and the data distribution in the vertical view dimension of the labels corresponding to the samples in the test set.
[0101] In one optional embodiment, the data distribution along the control dimension in the labels corresponding to the samples in the test set can be the proportion of the number of actions of each category along that control dimension to the total number of actions along that control dimension. Of course, the data distribution can also take other forms, and this application does not limit it to any particular form.
[0102] In one optional embodiment, the data distribution in the movement direction dimension of the labels corresponding to the samples in the test set can be the percentage of the number of actions in the nine categories of stationary, forward, backward, upward, downward, upper left, lower left, upper right, and lower right in the total number of actions in the movement direction dimension. For example, if the test set includes 360 samples, the total number of actions in the movement direction dimension is 360. Among them, the number of stationary actions in the labels corresponding to the samples in the test set is 12, the number of forward actions is 72, the number of backward actions is 36, the number of upward actions is 36, the number of downward actions is 24, the number of upper left actions is 60, the number of lower left actions is 24, the number of upper right actions is 48, and the number of lower right actions is 48. Then, the data distribution in the movement direction dimension of the labels corresponding to the samples in the test set can be [1 / 30, 1 / 5, 1 / 10, 1 / 10, 1 / 15, 1 / 6, 1 / 15, 2 / 15, 2 / 15].
[0103] In one optional embodiment, the data distribution in the posture dimension of the labels corresponding to the samples in the test set can be the proportion of the number of actions in the six categories of stillness, sprinting, crouching, aiming, and crouching to aim down sights to the total number of actions in the posture dimension.
[0104] In one optional embodiment, the data distribution in the horizontal perspective dimension of the labels corresponding to the samples in the test set can be the proportion of the number of actions in the 12 categories—stationary, first-view leftward movement speed, second-view leftward movement speed, third-view leftward movement speed, fourth-view leftward movement speed, fifth-view leftward movement speed, sixth-view leftward movement speed, first-view rightward movement speed, second-view rightward movement speed, third-view rightward movement speed, fourth-view rightward movement speed, fifth-view rightward movement speed, and sixth-view rightward movement speed—to the total number of actions in the horizontal perspective dimension.
[0105] In one optional embodiment, the data distribution in the vertical view dimension of the labels corresponding to the samples in the test set can be the proportion of the number of actions in the vertical view dimension of 10 categories: stationary, first view upward movement speed, second view upward movement speed, third view upward movement speed, fourth view upward movement speed, fifth view upward movement speed, first view downward movement speed, second view downward movement speed, third view downward movement speed, fourth view downward movement speed, and fifth view downward movement speed.
[0106] In step four, it's important to note that sample weights are numerical values assigned to samples during the training of a machine learning model, representing the importance of that sample. By adjusting the sample weights, the model can focus more on certain specific samples, thereby changing the overall performance of the model.
[0107] In one optional embodiment, adjusting the sample weights of samples in the training set based on the difference between the first action distribution and the second action distribution includes:
[0108] If the action distribution under the category in the first action distribution is greater than the action distribution under the corresponding category in the second action distribution, reduce the sample weight of the sample corresponding to the action under the category in the training set.
[0109] And / or,
[0110] If the action distribution under the category in the first action distribution is smaller than the action distribution under the corresponding category in the second action distribution, increase the sample weight of the sample corresponding to the action under the category in the training set.
[0111] In one optional embodiment, the predicted motion control data corresponding to the samples in the test set contains 36 stationary actions, 72 forward movement actions, 36 backward movement actions, 12 upward movement actions, 48 downward movement actions, 60 upper-left movement actions, 24 lower-left movement actions, 48 upper-right movement actions, and 24 lower-right movement actions. The data distribution in the movement direction dimension of the predicted motion control data corresponding to the samples in the test set can be [1 / 10, 1 / 5, 1 / 10, 1 / 30, 2 / 15, 1 / 6, 1 / 15, 2 / 15, 1 / 15].
[0112] In the test set, the number of stationary actions corresponding to the samples is 12, the number of forward movement actions is 72, the number of backward movement actions is 36, the number of upward movement actions is 36, the number of downward movement actions is 24, the number of upward-left movement actions is 60, the number of downward-left movement actions is 24, the number of upward-right movement actions is 48, and the number of downward-right movement actions is 48. The data distribution along the movement direction dimension in the labels corresponding to the samples in the test set can be [1 / 30, 1 / 5, 1 / 10, 1 / 10, 1 / 15, 1 / 6, 1 / 15, 2 / 15, 2 / 15]. It can be seen that the distribution of stationary actions in the first action distribution is greater than that in the second action distribution, therefore the sample weight of the samples corresponding to stationary actions in the training set is decreased. The distribution of upward movement actions in the first action distribution is less than that in the second action distribution, therefore the sample weight of the samples corresponding to upward movement actions in the training set is increased. The distribution of downward movement actions in the first action distribution is greater than that in the second action distribution, therefore the sample weight of the samples corresponding to downward movement actions in the training set is decreased. If the distribution of the action to move to the lower right in the first action distribution is smaller than that in the second action distribution, then the sample weight of the sample corresponding to the action to move to the lower right in the training set is increased.
[0113] In step five, the preset condition can be that the difference between the first action distribution and the second action distribution is less than a preset difference threshold.
[0114] When the number of samples in different categories differs greatly, the impact of each category on the loss function can be balanced by increasing the weight of minority class samples and decreasing the weight of majority class samples, thereby helping the model to better learn the features of the minority class.
[0115] Step S105: Configure a strategy for the target action prediction model to determine the target location.
[0116] It should be noted that the strategy configured for the target action prediction model to determine the target location is to enable the model to subsequently determine the corresponding action control data based on the current target location and the current game state data, thereby causing the predicted object to execute the action corresponding to the action control data. Here, the predicted object is an intelligent agent character, which is a virtual character controlled by the game system through multiple models, including the target action prediction model. For example, in the field of shooting games, this intelligent agent character could be a combat robot.
[0117] In one alternative embodiment, configuring a strategy for determining the target location for the target action prediction model includes at least one of the following steps:
[0118] The target action prediction model is configured to use the deployment location of the virtual object as the first sub-strategy for the target location;
[0119] The target action prediction model is configured with a second sub-strategy that uses the fall location of the virtual object as the target location;
[0120] A third sub-strategy is configured for the target action prediction model to prevent the target position from changing when the predicted object reaches a preset range of the target position, until the position of the virtual object changes.
[0121] It should be noted that the virtual object can be a virtual explosive or a virtual machine, etc., and this application does not limit it in this regard.
[0122] In one optional embodiment, the virtual explosive can be a virtual bomb. When the virtual explosive is a virtual bomb, the deployment location of the virtual object is the deployment location of the virtual bomb. For example, a deployment location where a virtual bomb can be deployed is randomly selected as the target location. The drop location of the virtual object can be the drop location of the virtual bomb. It should be noted that the drop location of the virtual bomb refers to the location where the virtual bomb falls after the bearer of the virtual bomb dies.
[0123] By configuring the current destination address based on the state information of virtual objects, the target action prediction model, which is configured with the destination location, can more accurately predict the behavior patterns of human players and control the predicted objects to make more reasonable responses, thereby enhancing the realism and immersion of the game.
[0124] In one optional embodiment, the target action prediction model is used to determine the current target location according to the strategy and obtain the current game state data for the prediction object. Then, the target action prediction model determines the current game state data and the action control data corresponding to the current target location based on the current game state data and the current target location.
[0125] It should be noted that the current destination location is the location the predicted object needs to go to. The target action prediction model determines the corresponding action control data based on the current game state data and the current destination location. Subsequently, the target action prediction model controls the predicted object to execute the actions represented by the action control data corresponding to the current game state data and the current destination location.
[0126] In this embodiment, the controlled virtual character acquires historical game state data, the historical destination location reached by the controlled virtual character after a preset time interval from the generation time of the historical game state data, and historical action control data corresponding to the historical game state data. Based on the historical game state data and the historical destination location reached by the controlled virtual character, a sample of game state sub-samples including multiple state dimensions is generated. Based on the historical action control data, labels corresponding to the samples are generated, wherein the labels include action control sub-labels of multiple control dimensions. Based on the samples and labels, the action prediction model to be trained is trained to obtain the target action prediction model. A strategy for determining the destination location is configured for the target action prediction model. In this way, the action output by the target action prediction model is closer to human level.
[0127] As shown in Table 1, the target action prediction model trained to control the predicted object to reach the target location has an arrival rate of 98.1%, and the predicted object's win rate against the player is 48%. Thus, it can be seen that the target action prediction model trained by the model training method provided in this application has output actions that are quite close to human level.
[0128] Predict the arrival rate of the object to the destination location 98.1% Predict the win rate of the target in a match against a player. 48%
[0129] Table 1
[0130] Corresponding to the data processing method provided in the embodiments of this application, the embodiments of this application also provide a model training device 200, such as... Figure 2 As shown, the device includes:
[0131] The acquisition module 201 is used to acquire historical game state data, historical destination location reached by the controlled virtual character after a preset time interval between the generation time of the historical game state data and the historical action control data corresponding to the historical game state data for the controlled virtual character. The controlled virtual character is a virtual character controlled by the player object.
[0132] The first generation module 202 is used to generate a sample of game state sub-samples including multiple state dimensions based on historical game state data and the historical destination locations reached by the controlled virtual character.
[0133] The second generation module 203 is used to generate labels corresponding to samples based on historical motion control data, wherein the labels include motion control sub-labels of multiple control dimensions;
[0134] Training module 204 is used to train the action prediction model to be trained based on samples and labels to obtain the target action prediction model;
[0135] Configuration module 205 is used to configure a strategy for determining the target location for the target action prediction model.
[0136] In one alternative embodiment, the configuration module is configured to:
[0137] Configure the target action prediction model to use the deployment location of the virtual object as the first sub-strategy for the target location;
[0138] Configure a second sub-strategy for the target action prediction model to use the fall location of the virtual object as the target location;
[0139] A third sub-strategy is configured for the target action prediction model to prevent the target position from changing when the predicted object reaches the target position within a preset range, until the position of the virtual object changes.
[0140] In one alternative embodiment, the training module is used for:
[0141] The samples are divided into training and test sets;
[0142] The action prediction model to be trained is trained based on the samples in the training set and the corresponding labels of the samples in the training set.
[0143] In an optional embodiment, the training module is further configured to:
[0144] Based on the samples in the test set and the labels corresponding to the samples in the test set, the action prediction model trained on the training set is tested to obtain the predicted action control data corresponding to the samples in the test set.
[0145] The first action distribution is obtained by analyzing the action distribution across different control dimensions in the predicted action control data corresponding to the samples in the statistical test set.
[0146] The second action distribution is obtained by analyzing the action distribution across different control dimensions of the labels corresponding to the samples in the test set.
[0147] Adjust the sample weights of the samples in the training set according to the difference between the first action distribution and the second action distribution;
[0148] Based on the updated sample weights and labels, the action prediction model trained on the training set is trained until the difference between the first action distribution and the second action distribution meets the preset conditions.
[0149] In one alternative embodiment, the training module is used for:
[0150] If the action distribution under the category in the first action distribution is greater than the action distribution under the corresponding category in the second action distribution, reduce the sample weight of the sample corresponding to the action under the category in the training set.
[0151] And / or,
[0152] If the action distribution under the category in the first action distribution is smaller than the action distribution under the corresponding category in the second action distribution, increase the sample weight of the sample corresponding to the action under the category in the training set.
[0153] In an optional embodiment, the apparatus further includes a module that performs the following operations:
[0154] Delete samples corresponding to labels where the action control sub-labels of multiple control dimensions are all static.
[0155] In an optional embodiment, the apparatus further includes a module that performs the following operations:
[0156] Remove samples corresponding to labels where the motion control sub-labels in the horizontal view dimension are static.
[0157] In an optional embodiment, the apparatus further includes a module that performs the following operations:
[0158] Remove samples whose motion control sub-labels in the movement direction dimension are stationary.
[0159] In an optional embodiment, the apparatus further includes a module that performs the following operations:
[0160] Remove the historical destination locations reached by the controlled virtual characters in some samples.
[0161] In one alternative embodiment, the first generation module is configured to:
[0162] Based on the global state sub-data included in the historical objective location and historical game state data, generate sub-samples for the global state dimension;
[0163] Based on the item status sub-data included in the historical game status data, generate sub-samples for the item status dimension;
[0164] Based on the controlled virtual character state sub-data included in the historical game state data, generate sub-samples of the controlled virtual character state dimension;
[0165] Based on the environmental perception state sub-data included in the historical game state data, generate sub-samples for the environmental perception state dimension;
[0166] Based on the character configuration status sub-data included in the historical game status data, generate sub-samples for the character configuration status dimension.
[0167] In one optional embodiment, the global state sub-data includes at least one of the following: game progress time, the location of the virtual object, whether the virtual object has been deployed, and the remaining startup time of the virtual object;
[0168] The item status sub-data includes at least one of the following: the type of item held by the controlled virtual character, the quantity of the item, and the time interval since the controlled virtual character last used the item;
[0169] The controlled virtual character status sub-data includes at least one of the following: the location of the controlled virtual character, whether the controlled virtual character belongs to the attacker or the defender, and the health value of the controlled virtual character;
[0170] The environmental perception state sub-data includes at least one of the following: whether the rays emitted by the controlled virtual character's body hit an object, and the distance between the hit object and the controlled virtual character;
[0171] The character configuration status sub-data includes at least one of the following: the character configuration information of the controlled virtual character and the skill configuration information of the controlled virtual character.
[0172] In one alternative embodiment, the second generation module is configured to:
[0173] Based on the movement direction sub-data, generate motion control sub-labels that include the movement direction dimension;
[0174] Based on the posture sub-data, generate motion control sub-labels in the posture dimension;
[0175] Based on the horizontal viewpoint sub-data, generate motion control sub-labels for the horizontal viewpoint dimension;
[0176] Based on the vertical viewpoint sub-data, generate motion control sub-labels in the vertical viewpoint dimension.
[0177] In one optional embodiment, the motion control sub-labels in the movement direction dimension include at least: stationary, moving forward, moving backward, moving up, moving down, moving to the upper left, moving to the lower left, moving to the upper right, and moving to the lower right;
[0178] The action control sub-label of the posture dimension includes at least one of the following: stationary, sprinting, crouching, aiming, and crouching while aiming down sights;
[0179] The motion control sub-label in the horizontal view dimension includes at least one of the following information: stationary, leftward movement speed of the first view, leftward movement speed of the second view, leftward movement speed of the third view, leftward movement speed of the fourth view, leftward movement speed of the fifth view, leftward movement speed of the sixth view, rightward movement speed of the first view, rightward movement speed of the second view, rightward movement speed of the third view, rightward movement speed of the fourth view, rightward movement speed of the fifth view, and rightward movement speed of the sixth view;
[0180] The motion control sub-label in the vertical view dimension includes at least one of the following information: stationary, first-view upward movement speed, second-view upward movement speed, third-view upward movement speed, fourth-view upward movement speed, fifth-view upward movement speed, first-view downward movement speed, second-view downward movement speed, third-view downward movement speed, fourth-view downward movement speed, and fifth-view downward movement speed.
[0181] In one optional embodiment, the target action prediction model is used to determine the current target location according to the strategy and obtain the current game state data for the prediction object. Then, the target action prediction model determines the current game state data and the action control data corresponding to the current target location based on the current game state data and the current target location. The prediction object is an intelligent agent character.
[0182] Corresponding to the model training method provided in the embodiments of this application, the embodiments of this application also provide an electronic device for implementing the model training method, such as... Figure 3 As shown, the electronic device includes:
[0183] The device includes a processor 301 and a memory 302 for storing a program for the display control method. After the device is powered on and the program for the display control method is run by the processor, the following steps are performed:
[0184] The controlled virtual character acquires historical game state data, the historical target location reached by the controlled virtual character after a preset time interval between the generation time of the historical game state data and the historical action control data corresponding to the historical game state data. The controlled virtual character is a virtual character controlled by the player object.
[0185] Based on historical game state data and the historical destinations reached by the controlled virtual character, a sample of game state sub-samples including multiple state dimensions is generated.
[0186] Based on historical motion control data, generate labels corresponding to the samples. The labels include motion control sub-labels for multiple control dimensions.
[0187] Based on the samples and labels, the action prediction model to be trained is trained to obtain the target action prediction model;
[0188] Configure a strategy for the target action prediction model to determine the target location.
[0189] In this embodiment, the controlled virtual character acquires historical game state data, the historical destination location reached by the controlled virtual character after a preset time interval from the generation time of the historical game state data, and historical action control data corresponding to the historical game state data. Based on the historical game state data and the historical destination location reached by the controlled virtual character, a sample of game state sub-samples including multiple state dimensions is generated. Based on the historical action control data, labels corresponding to the samples are generated, wherein the labels include action control sub-labels of multiple control dimensions. Based on the samples and labels, the action prediction model to be trained is trained to obtain the target action prediction model. A strategy for determining the destination location is configured for the target action prediction model. In this way, the action output by the target action prediction model is closer to human level.
[0190] Corresponding to the model training method provided in the embodiments of this application, the embodiments of this application also provide a computer-readable storage medium storing a program that implements the model training method. This program is executed by a processor to perform the following steps:
[0191] The controlled virtual character acquires historical game state data, the historical target location reached by the controlled virtual character after a preset time interval between the generation time of the historical game state data and the historical action control data corresponding to the historical game state data. The controlled virtual character is a virtual character controlled by the player object.
[0192] Based on historical game state data and the historical destinations reached by the controlled virtual character, a sample of game state sub-samples including multiple state dimensions is generated.
[0193] Based on historical motion control data, generate labels corresponding to the samples. The labels include motion control sub-labels for multiple control dimensions.
[0194] Based on the samples and labels, the action prediction model to be trained is trained to obtain the target action prediction model;
[0195] Configure a strategy for the target action prediction model to determine the target location.
[0196] In this embodiment, the controlled virtual character acquires historical game state data, the historical destination location reached by the controlled virtual character after a preset time interval from the generation time of the historical game state data, and historical action control data corresponding to the historical game state data. Based on the historical game state data and the historical destination location reached by the controlled virtual character, a sample of game state sub-samples including multiple state dimensions is generated. Based on the historical action control data, labels corresponding to the samples are generated, wherein the labels include action control sub-labels of multiple control dimensions. Based on the samples and labels, the action prediction model to be trained is trained to obtain the target action prediction model. A strategy for determining the destination location is configured for the target action prediction model. In this way, the action output by the target action prediction model is closer to human level.
[0197] It should be noted that for a detailed description of the model training apparatus, electronic device and computer-readable storage medium provided in the embodiments of this application, please refer to the relevant description of the model training method embodiments provided in the embodiments of this application, which will not be repeated here.
[0198] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
[0199] In a typical configuration, an electronic device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0200] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0201] 1. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable operations, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.
[0202] 2. Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0203] Although this application discloses preferred embodiments as described above, it is not intended to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the scope of protection of this application should be determined by the scope defined in the claims of this application.
Claims
1. A model training method, characterized in that, The method includes: The controlled virtual character acquires historical game state data, the historical destination location reached by the controlled virtual character after a preset time interval between the generation time of the historical game state data and the historical action control data corresponding to the historical game state data, wherein the controlled virtual character is a virtual character controlled by a player object; Based on the historical game state data and the historical destination locations reached by the controlled virtual character, a sample of game state sub-samples including multiple state dimensions is generated; Based on the historical motion control data, a label corresponding to the sample is generated, wherein the label includes motion control sub-labels of multiple control dimensions; Based on the samples and the labels, the action prediction model to be trained is trained to obtain the target action prediction model; Configure a strategy for determining the target location for the target action prediction model.
2. The method according to claim 1, characterized in that, Configuring a strategy for determining the target location for the target action prediction model includes at least one of the following steps: The target action prediction model is configured to use the deployment location of the virtual object as the first sub-strategy for the target location; The target action prediction model is configured with a second sub-strategy that uses the fall location of the virtual object as the target location; A third sub-strategy is configured for the target action prediction model to prevent the target position from changing when the predicted object reaches a preset range of the target position, until the position of the virtual object changes.
3. The method according to claim 1, characterized in that, The step of training the action prediction model to be trained based on the samples and the labels includes: The samples are divided into a training set and a test set; The action prediction model to be trained is trained based on the samples in the training set and the labels corresponding to the samples in the training set.
4. The method according to claim 3, characterized in that, The step of training the action prediction model to be trained based on the samples and the labels further includes: Based on the samples in the test set and the labels corresponding to the samples in the test set, the action prediction model trained by the training set is tested to obtain the predicted action control data corresponding to the samples in the test set. The first action distribution is obtained by statistically analyzing the action distribution across different control dimensions in the predicted action control data corresponding to the samples in the test set. The second action distribution is obtained by statistically analyzing the action distribution across different control dimensions in the labels corresponding to the samples in the test set. The sample weights of the samples in the training set are adjusted according to the difference between the first action distribution and the second action distribution; Based on the updated sample weights and the labels, the action prediction model trained on the training set is trained until the difference between the first action distribution and the second action distribution meets a preset condition.
5. The method according to claim 4, characterized in that, The step of adjusting the sample weights of the samples in the training set according to the difference between the first action distribution and the second action distribution includes: If the action distribution under the category in the first action distribution is greater than the action distribution under the corresponding category in the second action distribution, the sample weight of the sample corresponding to the action under the category in the training set is reduced; And / or, If the action distribution under the category in the first action distribution is smaller than the action distribution under the corresponding category in the second action distribution, increase the sample weight of the sample corresponding to the action under that category in the training set.
6. The method according to claim 1, characterized in that, The method further includes: Delete the samples corresponding to labels where the action control sub-labels of the multiple control dimensions are all static.
7. The method according to claim 1, characterized in that, The method further includes: Remove samples corresponding to labels where the motion control sub-labels in the horizontal view dimension are static.
8. The method according to claim 1, characterized in that, The method further includes: Remove samples whose motion control sub-labels in the movement direction dimension are stationary.
9. The method according to claim 1, characterized in that, The method further includes: The historical destination locations reached by the controlled virtual characters in some of the samples were deleted.
10. The method according to claim 1, characterized in that, The step of generating a sample of game state sub-samples including multiple state dimensions based on the historical game state data and the historical destination locations reached by the controlled virtual character includes: Based on the historical destination location and the global state sub-data included in the historical game state data, a sub-sample of the global state dimension is generated; Based on the item status sub-data included in the historical game status data, generate sub-samples for the item status dimension; Based on the controlled virtual character state sub-data included in the historical game state data, generate sub-samples of the controlled virtual character state dimension; Based on the environmental perception state sub-data included in the historical game state data, generate sub-samples for the environmental perception state dimension; Based on the character configuration status sub-data included in the historical game status data, generate sub-samples for the character configuration status dimension.
11. The method according to claim 10, characterized in that, The global state sub-data includes at least one of the following: game process time, location of the virtual object, whether the virtual object has been deployed, and remaining startup time of the virtual object; The item status sub-data includes at least one of the following: the type of item held by the controlled virtual character, the quantity of the item, and the time interval since the controlled virtual character last used the item; The controlled virtual character status sub-data includes at least one of the following: the location of the controlled virtual character, whether the controlled virtual character belongs to the attacker or the defender, and the health value of the controlled virtual character; The environmental perception state sub-data includes at least one of the following: whether the rays emitted by the controlled virtual character's body hit an object, and the distance between the hit object and the controlled virtual character; The character configuration status sub-data includes at least one of the following: the character configuration information of the controlled virtual character and the skill configuration information of the controlled virtual character.
12. The method according to claim 1, characterized in that, The historical motion control data includes movement direction sub-data, posture sub-data, horizontal viewpoint sub-data, and vertical viewpoint sub-data. The step of generating motion control sub-labels with multiple control dimensions based on the historical motion control data includes: Based on the movement direction sub-data, generate motion control sub-labels that include the movement direction dimension; Based on the posture sub-data, generate motion control sub-labels in the posture dimension; Based on the horizontal viewpoint sub-data, generate motion control sub-labels for the horizontal viewpoint dimension; Based on the vertical viewpoint sub-data, generate motion control sub-labels for the vertical viewpoint dimension.
13. The method according to claim 12, characterized in that, The motion control sub-labels for the movement direction dimension include at least: stationary, moving forward, moving backward, moving up, moving down, moving to the upper left, moving to the lower left, moving to the upper right, and moving to the lower right; The action control sub-label of the posture dimension includes at least one of the following: stationary, sprinting, crouching, aiming, and crouching while aiming down sights; The motion control sub-label of the horizontal view dimension includes at least one of the following information: stationary, leftward movement speed of the first view, leftward movement speed of the second view, leftward movement speed of the third view, leftward movement speed of the fourth view, leftward movement speed of the fifth view, leftward movement speed of the sixth view, rightward movement speed of the first view, rightward movement speed of the second view, rightward movement speed of the third view, rightward movement speed of the fourth view, rightward movement speed of the fifth view, and rightward movement speed of the sixth view; The motion control sub-label of the vertical view dimension includes at least one of the following information: stationary, first view upward movement speed, second view upward movement speed, third view upward movement speed, fourth view upward movement speed, fifth view upward movement speed, first view downward movement speed, second view downward movement speed, third view downward movement speed, fourth view downward movement speed, and fifth view downward movement speed.
14. The method according to claim 1, characterized in that, The target action prediction model is used to determine the current target position and obtain the current game state data for the prediction object according to the strategy. Then, the target action prediction model determines the current game state data and the action control data corresponding to the current target position based on the current game state data and the current target position. The prediction object is an intelligent agent character.
15. A model training device, characterized in that, The device includes: The acquisition module is used to acquire historical game state data of the controlled virtual character, the historical destination position reached by the controlled virtual character after a preset time interval from the time when the historical game state data was generated, and the historical action control data corresponding to the historical game state data. The controlled virtual character is a virtual character controlled by a player object. The first generation module is used to generate a sample of game state sub-samples including multiple state dimensions based on the historical game state data and the historical destination location reached by the controlled virtual character. The second generation module is used to generate a label corresponding to the sample based on the historical motion control data, wherein the label includes motion control sub-labels of multiple control dimensions; The training module is used to train the action prediction model to be trained based on the samples and the labels, so as to obtain the target action prediction model. The configuration module is used to configure a strategy for determining the target location for the target action prediction model.
16. An electronic device, characterized in that, include: processor; as well as A memory for storing a data processing program, which, when powered on and run by the processor, executes the method as described in any one of claims 1 to 14.
17. A computer-readable storage medium, characterized in that, The system contains a data processing program that is executed by a processor to perform the method as described in any one of claims 1 to 14.