Game character behavior control method and device, and electronic device

By transforming and training the parameters of the game's AI model, multiple behavior control models are generated. The target model is randomly selected to control the behavior of the game character, which solves the problem of players recognizing AI control and improves the gaming experience.

CN116510301BActive Publication Date: 2026-02-27NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310310902.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-22
Publication Date
2026-02-27
Estimated Expiration
2043-03-22

AI Technical Summary

Technical Problem

The fixed behavioral output of existing game AI models makes it easy for players to identify and observe the behavioral patterns of game characters, which affects the gaming experience.

Method used

By transforming the parameters of the initial model, multiple intermediate models are generated and trained to obtain multiple behavior control models. The target model is randomly selected to control the behavior of the game character, ensuring the randomness and variability of the behavior.

Benefits of technology

It increases the randomness of game character behavior, preventing players from identifying that game characters are controlled by AI, thus improving the gaming experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116510301B_ABST
    Figure CN116510301B_ABST
Patent Text Reader

Abstract

The application provides a game character behavior control method and device and electronic equipment; the method comprises the following steps: obtaining current state data of a target game; determining a target model from a plurality of behavior control models which are pre-trained, inputting the current state data into the target model, and obtaining an output result; the plurality of behavior control models are obtained through the following training method: transforming model parameters of an initial model based on preset transformation parameters to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models are different; the plurality of intermediate models are trained to obtain the plurality of behavior control models; based on the output result, a target behavior operation is determined, and the game character is controlled to perform the target behavior operation. This method can improve the randomness of the behavior of the game character, avoid the player identifying that the game character is controlled by game AI, and avoid the player observing the behavior rule of the game character, thereby improving the game experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of games, and in particular to a game character behavior control method and device and electronic equipment. BACKGROUND

[0002] Game AI (Artificial Intelligence) is also called a virtual player. A game AI model is trained through machine learning technology, and a game character is controlled through the game AI model, so that the behavior of the game character is similar to that of a real player-controlled game character. When the model parameters of the game AI model are determined and the state data input into the model is the same, the behavior executed by the game character is relatively fixed, which results in a low degree of intelligence of the game character, real players can easily identify which game characters are controlled by the game AI, and can observe the behavior rules of the game characters, which affects the game experience. SUMMARY

[0003] Therefore, the present application aims to provide a game character behavior control method, device and electronic equipment to improve the randomness of the behavior of the game character, avoid players identifying that the game character is controlled by the game AI, and avoid players observing the behavior rules of the game character, thereby improving the game experience.

[0004] In a first aspect, the present application provides a game character behavior control method, which comprises: acquiring current state data of a target game; determining a target model from a plurality of behavior control models that are pre-trained, inputting the current state data into the target model, and obtaining an output result; wherein the plurality of behavior control models are obtained through the following training: based on preset transformation parameters, transforming the model parameters of an initial model to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models are different; the plurality of intermediate models are trained to obtain the plurality of behavior control models; based on the output result, determining a target behavior operation, and controlling the game character to execute the target behavior operation.

[0005] In a second aspect, an embodiment of the present application provides a behavior control apparatus for a game character, the apparatus comprising: a data acquisition module configured to acquire current state data of a target game; an input module configured to determine a target model from a plurality of behavior control models that have been pre-trained, input the current state data into the target model, and obtain an output result; wherein the plurality of behavior control models are obtained by the following method: based on preset transformation parameters, transforming model parameters of an initial model to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models are different; based on a preset loss value threshold, training the plurality of intermediate models to obtain the plurality of behavior control models; and a control module configured to determine a target behavior operation based on the output result, and control the game character to perform the target behavior operation.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, comprising a processor and a memory, the memory storing machine executable instructions capable of being executed by the processor, and the processor executes the machine executable instructions to implement the behavior control method for the game character.

[0007] In a fourth aspect, an embodiment of the present application provides a machine readable storage medium, the machine readable storage medium storing machine executable instructions, and the machine executable instructions, when invoked and executed by a processor, cause the processor to implement the behavior control method for the game character.

[0008] The embodiments of the present application have the following beneficial effects:

[0009] The behavior control method, apparatus and electronic device for the game character acquire current state data of a target game; determine a target model from a plurality of behavior control models that have been pre-trained, input the current state data into the target model, and obtain an output result; wherein the plurality of behavior control models are obtained by the following method: based on preset transformation parameters, transforming model parameters of an initial model to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models are different; training the plurality of intermediate models to obtain the plurality of behavior control models; based on the output result, determining a target behavior operation, and controlling the game character to perform the target behavior operation. In this way, the model parameters of the initial model are transformed by the transformation parameters to obtain a plurality of intermediate models with different model parameters, and the plurality of behavior control models are then obtained by training, the target model is then determined from the plurality of behavior control models, and the behavior operation of the game character is then controlled. This method can improve the randomness of the behavior of the game character, avoid the player from identifying that the game character is controlled by the game AI, and at the same time avoid the player from observing the behavior rule of the game character, thereby improving the game experience.

[0010] Other features and advantages of the present application will be set forth in the descriptions that follow, and in part will be apparent from the description, or can be learned by practice of the application. The purposes and other advantages of the present application will be realized and attained by the structures particularly pointed out in the description, claims and drawings.

[0011] In order to make the above objectives, features and advantages of the present application more apparent, the following will describe a preferred embodiment in detail, and make a detailed description with the attached drawings. BRIEF DESCRIPTION OF DRAWINGS

[0012] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or the prior art description. Obviously, the drawings described below are some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0013] Figure 1 A flow chart of a behavior control method of a game character provided by an embodiment of the present application;

[0014] Figure 2 A flow chart of a training method of a behavior control model provided by an embodiment of the present application;

[0015] Figure 3 A schematic diagram of another training method of a behavior control model provided by an embodiment of the present application;

[0016] Figure 4 A structural schematic diagram of a behavior control device of a game character provided by an embodiment of the present application;

[0017] Figure 5 A structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0018] In order to make the objectives, technical solutions and advantages of the embodiments of the present application more apparent, the technical solutions of the present application will be described clearly and completely below with reference to the drawings. Obviously, the described embodiments are some embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the present application.

[0019] Game AI is an essential element in many games, and a game AI with intelligent performance can bring a better game experience to players. For example, in a Moba (Multiplayer Online Battle Arena) game, there are usually game AIs that control heroes as players do, and the game characters controlled by the game AIs fight against or cooperate with the game characters controlled by the players. More and more games use machine learning techniques to train game AIs with higher strength and more intelligent performance. The most common game AI is based on a neural network model, which includes a convolutional neural network, a recurrent neural network, a long short-term memory network, and a self-attention network, etc.

[0020] For a game AI based on a neural network model, because the parameters of the neural network model are determined once, the behavior output of the game AI is fixed for the state under similar scenarios, and real players can easily identify which game characters are controlled by the game AI, and the behavior template of the game character is captured by the players in a short time, which affects the game experience.

[0021] Based on this, the embodiment of the application provides a game character behavior control method, device and electronic equipment, which can be applied to the control of game AI in various games, for example, the control of game AI in a Moba game.

[0022] In order to facilitate the understanding of the present embodiment, first, the game character behavior control method disclosed by the present embodiment is introduced in detail, as shown in the following Figure 1 The method can be applied to a server, a cloud server or a terminal device, etc.; the method includes the following steps:

[0023] Step S102, obtaining current state data of a target game;

[0024] The target game usually includes multiple game characters, part of which are controlled by real players, and the target game character in the present embodiment is a game AI, that is, the target behavior operation is output by the game character behavior control method provided by the present embodiment, and then the target game character is controlled to execute the target behavior operation.

[0025] The current state data of the target game may, for example, include environment state data of a game match, such as match environment, match progress, number of characters of each party, state, etc.; the current state data may also include character state data of each game character in the game match, such as position, blood volume, attack power, resistance ability, etc. In one way, the current state data may also be the above-mentioned character state data of the target game character to be controlled.

[0026] In a specific implementation, the current state data includes one or more of position data, life value data, physical attack strength data, spell attack strength data, physical defense data, and spell defense data of the target game character.

[0027] The position data can be the position of the target game character in a game scene, which can be expressed using three-dimensional space coordinates of the game scene. The life value data is the blood volume of the target game character, which can include the total blood volume and the current blood volume of the target game character. The physical attack is a direct attack without using skills. The physical attack strength data indicates the ability of the target game character to physically attack. The spell attack is an attack using skills. The spell attack strength data indicates the ability of the target game character to spell attack. The physical defense is a defense against physical attacks, thereby reducing the damage to the target game character caused by physical attacks. The physical defense strength data indicates the defense ability of the target game character against physical attacks. The spell defense is a defense against spell attacks, thereby reducing the damage to the target game character caused by spell attacks. The spell defense strength data indicates the defense ability of the target game character against spell attacks.

[0028] In step S104, a target model is determined from the plurality of behavior control models trained in advance, and the current state data is input into the target model to obtain an output result. The plurality of behavior control models are trained in the following manner: based on the preset transformation parameters, the model parameters of an initial model are transformed to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models are different; and the plurality of intermediate models are trained to obtain the plurality of behavior control models.

[0029] In this embodiment, a plurality of behavior control models are provided, and the model structures of each behavior control model can be the same or different. When the model structures of the behavior control models are the same, the model parameters of different behavior control models are usually different, so that different output results can be obtained when the same current state data is input into different behavior control models, and different behavior operations can be obtained. The behavior control model can be a neural network model, such as a convolutional neural network, a recurrent neural network, a long short-term memory network, and a self-attention network, or other model structures.

[0030] In the target game, there can be multiple game characters that need to be controlled by behavior control models. For each game character, a target model can be randomly determined from the plurality of behavior control models, and the behavior control models can be assigned to the game characters in a certain order. In another way, for each game character, a target model can be determined from the plurality of behavior control models every certain period of time, so that the game character is controlled by different behavior control models, and the behavior of the game character is more variable.

[0031] In order to make different behavior control models obtain different output results under the same state data, and at the same time, make the output results match the behavior operation of the real player, the plurality of behavior control models need a specific training manner. In the embodiment, first, the model parameters of the initial model are transformed by transformation parameters to obtain a plurality of intermediate models, and at least part of the model parameters of different intermediate models are different. The initial model usually includes a plurality of model parameters, and the transformation parameters can be used to control which model parameters are transformed and how to transform these parameters that need to be transformed.

[0032] For example, the transformation parameters can include a transformation threshold, based on the size relationship between the model parameters and the transformation threshold, it is determined whether to transform the model parameters, or a new parameter is generated based on a certain model parameter, and based on the size relationship between the new parameter and the transformation threshold, it is determined whether to transform the model parameters. Further, the transformation parameters can also include a transformation function, and the model parameters are input into the transformation function, and the transformed model parameters are output.

[0033] Then, the plurality of intermediate models are trained, and different intermediate models can use the same training data or different training data. Since the model parameters of the intermediate models are different, the model parameters of the finally trained behavior control models are also different, that is, even if the same state data is input, the output results of different behavior control models will be different, and therefore, the behavior operations of the game characters controlled by different behavior control models will also be different.

[0034] In addition, in order to ensure that the output results of the plurality of behavior control models match the behavior operation of the real player, the plurality of intermediate models can be screened during the training process, for example, part of the intermediate models with smaller loss values are selected, and the selected intermediate models are used as the final plurality of behavior control models.

[0035] Step S106, based on the output result, determining the target behavior operation and controlling the game character to perform the target behavior operation.

[0036] The executable preset behavior operation of the game character usually includes a movement operation, a movement in a specific direction, an attack operation, a use of a specified attack skill, a use of a specified defense skill, etc. In the above output result, the probability value of each preset behavior operation is usually included, based on which the target behavior operation can be determined from a plurality of preset behavior operations, and then the game character is controlled to perform the target behavior operation.

[0037] The method for controlling the behavior of the game character comprises the following steps: obtaining current state data of a target game; determining a target model from a plurality of behavior control models that have been pre-trained, inputting the current state data into the target model, and obtaining an output result; wherein the plurality of behavior control models are obtained through the following method: based on preset transformation parameters, transforming model parameters of an initial model to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models are different; training the plurality of intermediate models to obtain the plurality of behavior control models; based on the output result, determining a target behavior operation, and controlling the game character to perform the target behavior operation. In this way, the model parameters of the initial model are transformed through the transformation parameters to obtain a plurality of intermediate models with different model parameters, and then the plurality of behavior control models are obtained through training. The target model is determined from the plurality of behavior control models, and then the behavior operation of the game character is controlled. This method can improve the randomness of the behavior of the game character, avoid the player from identifying that the game character is controlled by the game AI, and at the same time avoid the player from observing the behavior rule of the game character. This method improves the game experience.

[0038] In a specific implementation, for a game character, a target model is randomly determined from a plurality of behavior control models that have been pre-trained; wherein the target model is used to control the behavior operation of the game character in a target game. In the target game, or in a game session of the target game, for each game character, a target model is randomly determined from the plurality of behavior control models, the current state data is input into the target model, and then the behavior operation of the game character is controlled through the output result of the target model. In other ways, after the current state data is obtained, a target model can be randomly determined, and the behavior operation of the game character is controlled through the output result of the target model. In this way, the same game character can be controlled by different models, which can make the behavior operation of the game character more random.

[0039] The following embodiments describe a specific training method for the plurality of behavior control models.

[0040] First, a plurality of initial models are generated; wherein the initial model is pre-provided with initial model parameters, and the initial model parameters are obtained by respectively initializing the parameters of the plurality of initial models; for at least part of the model parameters of the initial model, a comparison parameter is generated, and based on the size relationship between the comparison parameter and a preset parameter threshold, it is determined whether the model parameter is transformed; if the model parameter is transformed, based on a preset transformation strength parameter, the model parameter is transformed to obtain a plurality of intermediate models after transformation.

[0041] The model structures of the plurality of initial models can be the same or different. For each initial model, the initial model is separately initialized in parameters to obtain initial model parameters. Generally, the initial model includes a plurality of model parameters. In actual implementation, for part of the model parameters, a comparison parameter can be generated to determine whether the model parameter is transformed; or for all model parameters in the model, a comparison parameter can be generated for each model parameter to determine whether the model parameter is transformed.

[0042] For a certain model parameter, the comparison parameter can be randomly generated or generated according to a preset rule. The above preset parameter threshold value can be preset, and the size relationship between the comparison parameter corresponding to each model parameter and the preset parameter threshold value is used to determine whether the model parameter is transformed; for example, when the comparison parameter is less than the preset parameter threshold value, the model parameter needs to be transformed, or when the comparison parameter is greater than the preset parameter threshold value, the model parameter needs to be transformed, and the like.

[0043] When it is determined that the model parameter needs to be transformed, a preset transformation strength parameter is used to change the model parameter. The preset transformation strength parameter can control the amplitude of the increase or decrease of the model parameter. In one way, the preset transformation strength parameter can be added, multiplied or operated with the model parameter to obtain the transformed model parameter.

[0044] In determining whether the model parameter is transformed, in one way, a preset parameter threshold value is set; a comparison parameter is randomly generated for at least part of the model parameters in the initial model; if the comparison parameter is less than the preset parameter threshold value, it is determined that the model parameter is transformed; and if the comparison parameter is greater than or equal to the preset parameter threshold value, it is determined that the model parameter is not transformed.

[0045] The above preset parameter threshold value can also be referred to as a mutation probability, denoted as P. The value range of P is [0, 1]. For a certain model parameter, a comparison parameter p is randomly generated, p = rand(0, 1); that is, the range of p is (0, 1). If p is less than P, the model parameter needs to be transformed, and if p is greater than or equal to P, the model parameter does not need to be transformed. In this example, the greater the value of P, the greater the probability that the model parameter needs to be transformed. For the model, the more model parameters need to be transformed, and after transformation, the greater the difference between the model parameters of different intermediate models.

[0046] In transforming the model parameter, in one way, if the model parameter is transformed, an initial value is determined from a preset value range; the size of the initial value is adjusted by a preset transformation strength parameter to obtain an intermediate value; and the parameter value of the model parameter is adjusted based on the intermediate value to obtain the transformed model parameter.

[0047] In an example, the initial value can be randomly determined from a preset value range, for example, rand(-1, 1), and the initial value is randomly determined from the range of (-1, 1); then the initial value is multiplied by the preset transformation strength parameter to obtain the intermediate value, and the intermediate value is added to the parameter value of the model parameter to obtain the transformed model parameter. In another way, the transformed model parameter can be calculated by the following formula: w1 = w0 + (rand(-1, 1) - 1) * S; wherein w0 is the model parameter, S is the preset transformation strength parameter, and the size of S can be set as needed; w1 is the transformed model parameter.

[0048] The following embodiments describe an implementation of training multiple intermediate models to obtain multiple behavior control models.

[0049] obtain first training data; wherein the first training data includes: historical state data of a specified game role in a target game, and historical behavior operations of the specified game role corresponding to the historical state data; train and process multiple intermediate models based on the first training data until the multiple intermediate models converge to determine loss values corresponding to the intermediate models; determine multiple candidate models from the multiple intermediate models based on the loss values corresponding to the intermediate models; and generate multiple behavior control models based on a preset loss value threshold and the candidate models.

[0050] The specified game role is usually a game role controlled by a real player; as the target game progresses, the game frames are arranged in time sequence. For each game frame, the state data of the specified game role in the game frame, i.e., the historical state data in the first training data, can be collected; for each game frame, the behavior operation of the specified game role under each state data, i.e., the historical behavior operation, also needs to be collected; these historical behavior operations are triggered by the real player through the terminal device. The behavior operation is usually one or more of the aforementioned multiple preset behavior operations.

[0051] The first training data can be used to train and process each intermediate model until each intermediate model converges, and then the loss value of each converged intermediate model is calculated by a preset loss function. In order to make the finally obtained behavior control model have a high degree of intelligence, in this embodiment, multiple candidate models are determined from the multiple intermediate models based on the loss values corresponding to the intermediate models; for example, the intermediate models can be arranged in order from low to high loss value, and then a specified number of intermediate models arranged at the front are determined as candidate models, for example, the top two intermediate models are selected as candidate models.

[0052] The loss value threshold is also set in the embodiment. Based on the loss value threshold, it is determined whether the alternative model needs to be retrained or the training manner of the alternative model is determined to obtain the final plurality of behavior control models.

[0053] The preset loss value threshold is determined by the following manner: obtaining second training data; the second training data includes historical state data of a specified game role in a target game and historical behavior operations of the specified game role corresponding to the historical state data; training the first model based on the second training data until the first model converges to obtain a trained first model; determining the loss value of the trained first model by a preset loss function, and determining the loss value threshold based on the loss value.

[0054] It should be noted that the second training data can be the same as or different from the first training data, but the collection manner of the second training data is usually the same as that of the first training data. The first model can be the same as or different from the initial model; for example, the first model has the same structure as the initial model, but the model parameters can be different.

[0055] The first model is trained by the second training data until the first model converges, and then the loss value of the converged first model is calculated by the loss function, which is represented as Lm. In an implementation manner, the loss value of the first model can be directly determined as the loss value threshold, or the loss value can be calculated to obtain the loss value threshold, for example, the loss value of the first model is multiplied by a preset weight coefficient to obtain the loss value threshold; the weight coefficient is represented as a, and the loss value threshold is a*Lm.

[0056] Based on the preset loss value threshold and the alternative model, in one case, the maximum loss value is determined from the loss value corresponding to the alternative model; if the maximum loss value is less than the loss value threshold, the plurality of alternative models is determined as the plurality of behavior control models. In this case, the maximum loss value of the alternative model is less than the loss value threshold, which indicates that the training result of the alternative model is good, and the plurality of alternative models can be directly used as the plurality of behavior control models without retraining.

[0057] In another case, if the maximum loss value is greater than or equal to the loss value threshold, an evolution parameter is generated, and an initial value is set for the evolution parameter; an updated plurality of initial models is generated based on the alternative model, and the evolution parameter is updated; the following steps are continuously executed until the evolution parameter reaches a preset parameter threshold: the model parameters of the initial model are transformed based on a preset transformation parameter to obtain a plurality of intermediate models.

[0058] In this case, the evolution parameter needs to be set to control the number of cycles, which can be denoted as g, and the initial value of g can be set to 0. When the maximum loss value is greater than or equal to the loss value threshold, it indicates that the training result of the candidate model is poor, and the candidate model needs to be trained again. At this time, the updated multiple initial models are generated based on the candidate model. Since the candidate model is selected from the intermediate model, and the number of intermediate models is the same as the number of initial models, the number of candidate models is less than the number of initial models. Therefore, before entering the next round of cycle training, the updated initial models need to be generated based on the candidate models. For example, the updated initial models are obtained by copying the candidate models, performing parameter transformation and structure transformation on the candidate models, etc. The number of updated initial models is the same as the number of the aforementioned initial models.

[0059] In an implementation, the candidate models are copied to obtain copy models, until the total number of candidate models and copy models is the same as the number of initial models; and the candidate models and copy models are determined as the updated multiple initial models. For example, there are 2 candidate models, and the number of the aforementioned initial models is 6. At this time, the 4 copy models are obtained by copying the candidate models, and the 6 models including the 2 candidate models and the 4 copy models are determined as the updated initial models.

[0060] After the updated multiple initial models are generated, the evolution parameter g is updated, and the updating method is g=g+1, that is, the value of g is increased by 1 each time. In other methods, the evolution parameter can be set when the initial model is generated for the first time, and the initial value of the evolution parameter is set. Then, the value of g is updated after the candidate model is determined.

[0061] After the updated multiple initial models are generated, the aforementioned step of performing transformation processing on the model parameters of the initial models based on the preset transformation parameters to obtain multiple intermediate models can be started to be executed. Then, the subsequent steps are sequentially executed until the value of g reaches the preset parameter threshold, and the cycle training is stopped.

[0062] In the third case, if the maximum loss value is greater than or equal to the loss value threshold, and the evolution parameter reaches the preset parameter threshold, the to-be-trained model with a loss value greater than or equal to the loss value threshold is determined from the candidate model. The third training data is obtained, and the to-be-trained model is trained based on the third training data until the loss value of the to-be-trained model is less than the loss value threshold. The trained to-be-trained model and the candidate model except the to-be-trained model are determined as the multiple behavior control models.

[0063] In this case, if the maximum loss value is greater than or equal to the loss value threshold, and the evolution parameter reaches the preset parameter threshold, it means that the loss value of the model cannot be reduced by transforming the parameters and training, at this time, another model training method needs to be used, that is, a to-be-trained model with a loss value greater than or equal to the loss value threshold is determined from the alternative models, and the loss value threshold is represented as a*Lm; the loss value of the to-be-trained model is greater than or equal to the loss value threshold, that is, the training effect of the to-be-trained model in the foregoing training method is not good, at this time, the third training data is obtained, and the to-be-trained model is trained based on the third training data until the loss value of the to-be-trained model is less than the loss value threshold; the third training data can be the same as the foregoing first training data, or can be different, and the to-be-trained model is directly trained by the third training data, without the need to use the foregoing transformed parameters to transform the model parameters first and then train.

[0064] After the loss value of the to-be-trained model is less than the loss value threshold, the training effect of the to-be-trained model meets the requirement, at this time, the to-be-trained model and the alternative model that has not been trained using the third training data are determined as the behavior control model.

[0065] In order to facilitate understanding, the following Figure 2 The training method of the plurality of behavior control models is shown, including the following steps:

[0066] Step S202, obtaining training data, training a first model based on the training data until the first model converges, obtaining a trained first model; determining a loss value of the trained first model by a preset loss function, and determining a loss value threshold based on the loss value;

[0067] The training data includes historical state data of a specified game role in a target game and historical behavior operations of the specified game role corresponding to the historical state data; the historical state data can be represented as feature data X, and the operation of the game player, that is, the historical behavior operation, is converted into label data Y. The training data can be the first training data, the second training data or the third training data in the foregoing embodiments; or the first training data, the second training data or the third training data in the foregoing embodiments are the same training data, that is, the training data in the foregoing step S202.

[0068] The first model is a neural network model, the first model is identified as M, the loss value of the first model is Lm after being trained by the training data; the loss value threshold is represented as Lm*a, and a is a weight parameter.

[0069] Step S204, generating a plurality of initial models, generating an evolution parameter, and setting an initial value for the evolution parameter;

[0070] The initial model is preset with initial model parameters, and the initial model parameters are obtained by respectively initializing parameters of the plurality of initial models; the number of initial models is represented as N, the parameters of each initial model are independently and randomly initialized to form a new 0th generation, and the evolution parameter is represented as g, and the initial value of g is equal to 0.

[0071] In step S206, a preset parameter threshold and a preset transformation strength parameter are set; a comparison parameter is randomly generated for at least part of the model parameters in the initial model; if the comparison parameter is less than the preset parameter threshold, it is determined that the model parameters are transformed; and if the comparison parameter is greater than or equal to the preset parameter threshold, it is determined that the model parameters are not transformed.

[0072] In step S208, if the model parameters are transformed, an initial value is determined from a preset numerical range; the size of the initial value is adjusted by the preset transformation strength parameter to obtain an intermediate value; the parameter value of the model parameter is adjusted based on the intermediate value to obtain a transformed model parameter, and then a plurality of intermediate models are obtained.

[0073] The preset parameter threshold is also referred to as a mutation probability, and is represented as P; and the preset transformation strength parameter is also referred to as a mutation strength, and is represented as S, wherein P and S are set according to actual application conditions, P is between 0 and 1, and S has no special requirements.

[0074] For each parameter of each initial model in the gth generation, taking a parameter w as an example, a number p = rand(0, 1) is randomly taken from 0 to 1, wherein rand is a function of randomly generating a number, if p < P, the parameter is mutated, and the mutated parameter w = w + (rand(-1, 1) - 1) * S.

[0075] In step S210, the plurality of intermediate models are trained by using the training data; a plurality of candidate models are determined from the plurality of intermediate models based on the loss values corresponding to the intermediate models; and the evolution parameter is updated.

[0076] Specifically, according to the foregoing feature data X and label data Y, the loss value of each intermediate model in the gth generation is calculated, the first n models with the smallest loss function in the gth generation are selected as candidate models, wherein n is less than or equal to N, and the value of g is updated to g + 1.

[0077] In step S212, the maximum loss value is determined from the loss values corresponding to the candidate models; if the maximum loss value is less than a loss value threshold, step S218 is performed; if the maximum loss value is greater than or equal to the loss value threshold, and the evolution parameter is less than a preset parameter threshold, step S216 is performed; if the maximum loss value is greater than or equal to the loss value threshold, and the evolution parameter reaches the preset parameter threshold, step S214 is performed.

[0078] Step S214, determining a to-be-trained model with a loss value greater than or equal to a loss value threshold from the candidate models; obtaining third training data, training the to-be-trained model based on the third training data until the loss value of the to-be-trained model is less than the loss value threshold, determining the to-be-trained model and the candidate models other than the to-be-trained model as updated candidate models, and performing step S218.

[0079] Step S216, copying the candidate models to obtain copied models until the total number of the candidate models and the copied models is the same as the number of the initial models; determining the candidate models and the copied models as updated initial models; and performing step S206.

[0080] Step S218, determining the plurality of candidate models as a plurality of behavior control models.

[0081] Further, Figure 3 A training flow diagram of the model is shown in FIG. 1. Six initial models are shown in Generation 0. After the first mutation, i.e., the first parameter transformation, Generation 1 models are obtained. After training, two models with the lowest loss values are obtained. The two models are copied and parameter transformation is performed again to obtain Generation 2 models. This process is repeated until the final plurality of behavior control models are obtained. Each parameter transformation is a random adjustment of the parameters in the model.

[0082] In the above manner, a specified number of effective convergent models can be finally obtained, and these models have random differences in performance. Only additional genetic algorithm iteration modification is needed on the basis of the original training process, and the development efficiency is high.

[0083] Corresponding to the above method embodiment, refer to Figure 4 A structure diagram of a behavior control device of a game character is shown in FIG. 2. The device includes:

[0084] A data acquisition module 40 is configured to acquire current state data of a target game.

[0085] An input module 42 is configured to determine a target model from a plurality of behavior control models trained in advance, input the current state data into the target model, and obtain an output result. The plurality of behavior control models are trained in the following manner: based on a preset transformation parameter, model parameters of an initial model are transformed to obtain a plurality of intermediate models. At least part of the model parameters of different intermediate models are different. Based on a preset loss value threshold, the plurality of intermediate models are trained to obtain the plurality of behavior control models.

[0086] A control module 44 is configured to determine a target behavior operation based on the output result, and control the game character to perform the target behavior operation.

[0087] The behavior control device of the game character obtains current state data of a target game, determines a target model from a plurality of behavior control models that are pre-trained, inputs the current state data into the target model, and obtains an output result. The plurality of behavior control models are obtained by the following method: based on preset transformation parameters, transforming model parameters of an initial model to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models are different; training the plurality of intermediate models to obtain the plurality of behavior control models; based on the output result, determining a target behavior operation, and controlling the game character to perform the target behavior operation. In this way, the model parameters of the initial model are changed by the transformation parameters to obtain a plurality of intermediate models with different model parameters, and then the plurality of behavior control models are obtained by training. The target model is determined from the plurality of behavior control models, and then the behavior operation of the game character is controlled. This method can improve the randomness of the behavior of the game character, avoid the player identifying that the game character is controlled by the game AI, and avoid the player observing the behavior rule of the game character. This method improves the game experience.

[0088] The device further includes a model training module configured to: for the game character, randomly determine a target model from a plurality of behavior control models that are pre-trained; and wherein the target model is configured to control the behavior operation of the game character in the target game.

[0089] The model training module is further configured to: generate a plurality of initial models; wherein the initial models are preconfigured with initial model parameters, and the initial model parameters are obtained by respectively initializing the parameters of the plurality of initial models; for at least part of the model parameters of the initial models, generate a comparison parameter, and based on the size relationship between the comparison parameter and a preset parameter threshold, determine whether the model parameters are transformed; if the model parameters are transformed, based on a preset transformation strength parameter, transform the model parameters to obtain a plurality of transformed intermediate models.

[0090] The model training module is further configured to: set a preset parameter threshold; for at least part of the model parameters of the initial models, randomly generate a comparison parameter; if the comparison parameter is less than the preset parameter threshold, determine that the model parameters are transformed; and if the comparison parameter is greater than or equal to the preset parameter threshold, determine that the model parameters are not transformed.

[0091] The model training module is further configured to: if the model parameters are transformed, determine an initial value from a preset value range; adjust the size of the initial value by a preset transformation strength parameter to obtain an intermediate value; based on the intermediate value, adjust the parameter value of the model parameters to obtain the transformed model parameters.

[0092] The model training module is further configured to: obtain first training data, wherein the first training data comprises historical state data of a specified game role in a target game and historical behavior operations of the specified game role corresponding to the historical state data; train the plurality of intermediate models based on the first training data until the plurality of intermediate models converge, and determine a loss value corresponding to the intermediate models; determine a plurality of candidate models from the plurality of intermediate models based on the loss value corresponding to the intermediate models; and generate the plurality of behavior control models based on the preset loss value threshold and the candidate models.

[0093] The model training module is further configured to: obtain second training data, wherein the second training data comprises historical state data of a specified game role in a target game and historical behavior operations of the specified game role corresponding to the historical state data; train the first model based on the second training data until the first model converges, and obtain the trained first model; determine a loss value of the trained first model by using a preset loss function, and determine the loss value threshold based on the loss value.

[0094] The model training module is further configured to: determine a maximum loss value from the loss values corresponding to the candidate models; and if the maximum loss value is less than the loss value threshold, determine the plurality of candidate models as the plurality of behavior control models.

[0095] The model training module is further configured to: determine a maximum loss value from the loss values corresponding to the candidate models; and if the maximum loss value is greater than or equal to the loss value threshold, generate an evolution parameter, set an initial value for the evolution parameter, generate an updated plurality of initial models based on the candidate models, and update the evolution parameter; continue to execute from the following step until the evolution parameter reaches a preset parameter threshold: transform the model parameters of the initial models based on a preset transformation parameter to obtain a plurality of intermediate models.

[0096] The model training module is further configured to: copy the candidate models to obtain copied models until the total number of the candidate models and the copied models is the same as the number of the initial models; and determine the candidate models and the copied models as the updated plurality of initial models.

[0097] The model training module is further configured to: if the maximum loss value is greater than or equal to the loss value threshold and the evolution parameter reaches the preset parameter threshold, determine a to-be-trained model from the candidate models, wherein the loss value of the to-be-trained model is greater than or equal to the loss value threshold; obtain third training data, train the to-be-trained model based on the third training data until the loss value of the to-be-trained model is less than the loss value threshold, and determine the trained to-be-trained model and the candidate models other than the to-be-trained model as the plurality of behavior control models.

[0098] The embodiment also provides an electronic device, comprising a processor and a memory, the memory storing machine executable instructions capable of being executed by the processor, and the processor executes the machine executable instructions to implement the behavior control method of the game character. The electronic device can be a server or a terminal device.

[0099] Referring to Figure 5 The electronic device shown in the figure comprises a processor 100 and a memory 101, the memory 101 storing machine executable instructions capable of being executed by the processor 100, and the processor 100 executes the machine executable instructions to implement the behavior control method of the game character.

[0100] Further, Figure 5 The electronic device shown in the figure further comprises a bus 102 and a communication interface 103, and the processor 100, the communication interface 103 and the memory 101 are connected through the bus 102.

[0101] The memory 101 can contain a high-speed random access memory (RAM, Random Access Memory) and can also include a non-volatile memory, for example, at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 103 (which can be wired or wireless), and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus or an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the figure to represent the bus, but it does not mean that there is only one bus or only one type of bus.

[0102] The processor 100 can be an integrated circuit chip with signal processing capability. In the implementation process, each step of the above method can be completed by the integrated logic circuit of hardware in the processor 100 or the instruction in the form of software. The processor 100 described above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Each method, step and logic block disclosed in the embodiment of the present application can be implemented or executed. The general processor can be a microprocessor or the processor can also be any conventional processor. The steps of the method disclosed in combination with the embodiment of the present application can be directly embodied as a hardware code processor for execution, or a combination of hardware and software modules in the code processor for execution. The software module can be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium in the art. The storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101, and combines the hardware to complete the steps of the method of the above embodiment.

[0103] The processor in the above electronic device can realize the following operations in the behavior control method of the game character by executing machine executable instructions:

[0104] Obtain the current state data of the target game; determine a target model from a plurality of behavior control models trained in advance, input the current state data into the target model, and obtain an output result; wherein the plurality of behavior control models are obtained by the following method: based on the preset transformation parameter, the model parameter of the initial model is transformed to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models are different; the plurality of intermediate models are trained to obtain a plurality of behavior control models; based on the output result, determine the target behavior operation, and control the game character to execute the target behavior operation.

[0105] For the game character, a target model is randomly determined from a plurality of behavior control models trained in advance; wherein the target model is used to control the behavior operation of the game character in the target game.

[0106] generate a plurality of initial models; wherein, the initial model is preset with an initial model parameter, and the initial model parameter is obtained by respectively initializing parameters of the plurality of initial models; for at least part of the model parameters in the initial model, a comparison parameter is generated, and based on a size relationship between the comparison parameter and a preset parameter threshold, it is determined whether the model parameter is transformed; if the model parameter is transformed, the model parameter is transformed based on a preset transformation strength parameter, and a plurality of intermediate models after transformation are obtained.

[0107] A preset parameter threshold is set; for at least part of the model parameters in the initial model, a comparison parameter is randomly generated; if the comparison parameter is less than the preset parameter threshold, it is determined that the model parameter is transformed; if the comparison parameter is greater than or equal to the preset parameter threshold, it is determined that the model parameter is not transformed.

[0108] If the model parameter is transformed, an initial value is determined from a preset value range; the size of the initial value is adjusted by a preset transformation strength parameter to obtain an intermediate value; the parameter value of the model parameter is adjusted based on the intermediate value to obtain the transformed model parameter.

[0109] Obtain first training data; wherein, the first training data includes: historical state data of a specified game role in a target game, and historical behavior operation of the specified game role corresponding to the historical state data; based on the first training data, the plurality of intermediate models are trained until the plurality of intermediate models converge, and the loss value corresponding to the intermediate model is determined; based on the preset loss value threshold and the candidate model, a plurality of behavior control models are generated.

[0110] Obtain second training data; wherein, the second training data includes: historical state data of a specified game role in a target game, and historical behavior operation of the specified game role corresponding to the historical state data; based on the second training data, the first model is trained until the first model converges, and the trained first model is obtained; the loss value of the trained first model is determined by a preset loss function, and the loss value threshold is determined based on the loss value.

[0111] Determine the maximum loss value from the loss value corresponding to the candidate model; if the maximum loss value is less than the loss value threshold, the plurality of candidate models are determined as the plurality of behavior control models.

[0112] Determine the maximum loss value from the loss value corresponding to the candidate model; if the maximum loss value is greater than or equal to the loss value threshold, generate an evolution parameter, set an initial value for the evolution parameter; based on the candidate model, generate an updated plurality of initial models, update the evolution parameter; continue to execute from the following steps until the evolution parameter reaches the preset parameter threshold: transform the model parameter of the initial model based on the preset transformation parameter to obtain a plurality of intermediate models.

[0113] copying the candidate model to obtain a copy model, until the total number of the candidate model and the copy model is same as the number of the initial model; and determining the candidate model and the copy model as the plurality of initial models.

[0114] If the maximum loss value is greater than or equal to the loss value threshold, and the evolution parameter reaches the preset parameter threshold, determining a to-be-trained model with a loss value greater than or equal to the loss value threshold from the candidate model; obtaining third training data, training the to-be-trained model based on the third training data until the loss value of the to-be-trained model is less than the loss value threshold, and determining the to-be-trained model and the candidate model except the to-be-trained model as the plurality of behavior control models.

[0115] In the above manner, the model parameters of the initial model are changed by the transformation parameter to obtain a plurality of intermediate models with different model parameters, and the plurality of behavior control models are obtained by training, the target model is determined from the plurality of behavior control models, and the behavior operation of the game character is controlled, which can improve the randomness of the behavior of the game character, avoid the player from identifying that the game character is controlled by the game AI, and avoid the player from observing the behavior rule of the game character, thereby improving the game experience.

[0116] The embodiment also provides a machine readable storage medium, which stores machine executable instructions, and the machine executable instructions cause the processor to implement the above-mentioned behavior control method of the game character when the machine executable instructions are called and executed by the processor.

[0117] The machine executable instructions stored in the above-mentioned machine readable storage medium can implement the following operations in the above-mentioned behavior control method of the game character by executing the machine executable instructions:

[0118] obtaining current state data of a target game; determining a target model from a plurality of behavior control models pre-trained, inputting the current state data into the target model to obtain an output result; wherein the plurality of behavior control models are obtained by training in the following manner: transforming the model parameters of an initial model based on a preset transformation parameter to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models are different; training the plurality of intermediate models to obtain the plurality of behavior control models; determining a target behavior operation based on the output result, and controlling the game character to perform the target behavior operation.

[0119] For a game character, a target model is randomly determined from a plurality of behavior control models pre-trained; wherein the target model is used to control the behavior operation of the game character in a target game.

[0120] Generate a plurality of initial models; wherein, the initial model is preset with an initial model parameter, and the initial model parameter is obtained by respectively initializing the parameters of the plurality of initial models; for at least part of the model parameters in the initial model, generate a comparison parameter, and determine whether the model parameter is transformed based on the size relationship between the comparison parameter and the preset parameter threshold; if the model parameter is transformed, transform the model parameter based on the preset transformation strength parameter to obtain a plurality of intermediate models after transformation.

[0121] Set a preset parameter threshold; for at least part of the model parameters in the initial model, randomly generate a comparison parameter; if the comparison parameter is less than the preset parameter threshold, determine that the model parameter is transformed; if the comparison parameter is greater than or equal to the preset parameter threshold, determine that the model parameter is not transformed.

[0122] If the model parameter is transformed, determine an initial value from a preset numerical range; adjust the size of the initial value through a preset transformation strength parameter to obtain an intermediate value; adjust the parameter value of the model parameter based on the intermediate value to obtain the transformed model parameter.

[0123] Obtain first training data; wherein, the first training data includes: historical state data of a specified game role in a target game, and historical behavior operations of the specified game role corresponding to the historical state data; train the plurality of intermediate models based on the first training data until the plurality of intermediate models converge to determine the loss value corresponding to the intermediate model; determine a plurality of candidate models from the plurality of intermediate models based on the loss value corresponding to the intermediate model; generate a plurality of behavior control models based on the preset loss value threshold and the candidate models.

[0124] Obtain second training data; wherein, the second training data includes: historical state data of a specified game role in a target game, and historical behavior operations of the specified game role corresponding to the historical state data; train the first model based on the second training data until the first model converges to obtain a trained first model; determine the loss value of the trained first model through a preset loss function, and determine the loss value threshold based on the loss value.

[0125] Determine the maximum loss value from the loss values corresponding to the candidate models; if the maximum loss value is less than the loss value threshold, determine the plurality of candidate models as the plurality of behavior control models.

[0126] Determine the maximum loss value from the loss values corresponding to the candidate models; if the maximum loss value is greater than or equal to the loss value threshold, generate an evolution parameter and set an initial value for the evolution parameter; generate an updated plurality of initial models based on the candidate models and update the evolution parameter; continue to execute from the following steps until the evolution parameter reaches the preset parameter threshold: transform the model parameters of the initial model based on the preset transformation parameter to obtain a plurality of intermediate models.

[0127] copying the candidate model to obtain a copy model, until the total number of the candidate model and the copy model is same as the number of the initial model; and determining the candidate model and the copy model as the plurality of initial models.

[0128] If the maximum loss value is greater than or equal to the loss value threshold, and the evolution parameter reaches the preset parameter threshold, determining a to-be-trained model with a loss value greater than or equal to the loss value threshold from the candidate model; obtaining third training data, training the to-be-trained model based on the third training data until the loss value of the to-be-trained model is less than the loss value threshold, and determining the to-be-trained model and the candidate model except the to-be-trained model as the plurality of behavior control models.

[0129] In the above manner, the model parameters of the initial model are changed by changing the parameters to obtain a plurality of intermediate models with different model parameters, and the plurality of behavior control models are obtained by training, the target model is determined from the plurality of behavior control models, and the behavior operation of the game character is controlled, which can improve the randomness of the behavior of the game character, avoid the player from identifying that the game character is controlled by the game AI, and avoid the player from observing the behavior rule of the game character, thereby improving the game experience.

[0130] The game character behavior control method and device and the computer program product of the electronic equipment provided in the embodiments of the present application include a machine-readable storage medium storing program codes, the program codes include instructions for executing the method described in the foregoing method embodiments, and the specific implementation can be referred to the method embodiments, which will not be described here.

[0131] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described system and device can refer to the corresponding process in the foregoing method embodiments, which will not be described here.

[0132] In addition, in the description of the embodiments of the present application, unless otherwise explicitly specified and limited, the terms "mounting", "connecting", "connecting" should be understood in a broad sense, for example, can be fixedly connected, can also be detachably connected, or integrally connected; can be mechanically connected, can also be electrically connected; can be directly connected, can also be indirectly connected through an intermediate medium, can be the communication inside two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood according to the specific circumstances.

[0133] If the functions are realized in the form of software function units and sold or used as independent products, they can be stored in a machine readable storage medium. Based on this understanding, the technical solutions of the present application or the parts of the present application that essentially contribute to the prior art or the parts of the technical solutions can be embodied in the form of software products. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in the various embodiments of the present application. The aforementioned storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0134] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", and the like indicate the orientation or positional relationship shown in the drawings, and are only for the purpose of facilitating the description of the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" are only for the purpose of description, and cannot be understood as indicating or implying relative importance.

[0135] Finally, it should be noted that the above embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, and are not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily think of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed by the present application, or make equivalent replacements to some of the technical features; and these modifications, changes or replacements do not cause the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A behavior control method of a game character, characterized by, The method comprises: acquiring current state data of a target game; determining a target model from a plurality of behavior control models trained in advance, inputting the current state data into the target model, and obtaining an output result; wherein the plurality of behavior control models are obtained by the following method: based on preset transformation parameters, transforming model parameters of an initial model to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models are different; training the plurality of intermediate models to obtain the plurality of behavior control models; The step of transforming the model parameters of the initial model based on the preset transformation parameters to obtain a plurality of intermediate models comprises: generating a plurality of initial models; wherein the initial model is preset with initial model parameters, and the initial model parameters are obtained by respectively initializing the parameters of the plurality of initial models; for at least part of the model parameters of the initial model, a comparison parameter is generated, and based on the size relationship between the comparison parameter and a preset parameter threshold, it is determined whether the model parameter is transformed; if the model parameter is transformed, the model parameter is transformed based on a preset transformation strength parameter to obtain a plurality of transformed intermediate models; Based on the output result, a target behavior operation is determined, and the game character is controlled to perform the target behavior operation.

2. The method of claim 1, wherein, The step of determining a target model from a plurality of behavior control models trained in advance comprises: For the game character, a target model is randomly determined from a plurality of behavior control models trained in advance; wherein the target model is used to control the behavior operation of the game character in the target game.

3. The method of claim 1, wherein, The step of generating a comparison parameter for at least part of the model parameters of the initial model, and determining whether the model parameter is transformed based on the size relationship between the comparison parameter and a preset parameter threshold comprises: setting the preset parameter threshold; randomly generating a comparison parameter for at least part of the model parameters of the initial model; if the comparison parameter is less than the preset parameter threshold, it is determined that the model parameter is transformed; if the comparison parameter is greater than or equal to the preset parameter threshold, it is determined that the model parameter is not transformed.

4. The method of claim 1, wherein, If the model parameter is transformed, the step of transforming the model parameter based on a preset transformation strength parameter comprises: if the model parameter is transformed, an initial value is determined from a preset value range; adjusting the size of the initial value through the preset transformation strength parameter to obtain an intermediate value; adjusting the parameter value of the model parameter based on the intermediate value to obtain the transformed model parameter.

5. The method of claim 1, wherein, The step of training the plurality of intermediate models to obtain the plurality of behavior control models comprises: acquiring first training data; wherein the first training data comprises: historical state data of a specified game character in the target game, and historical behavior operations of the specified game character corresponding to the historical state data; training the plurality of intermediate models based on the first training data until the plurality of intermediate models converge to determine the loss value corresponding to the intermediate model; determine a plurality of candidate models from the plurality of intermediate models based on loss values corresponding to the intermediate models; generate the plurality of behavior control models based on a preset loss value threshold and the candidate models.

6. The method of claim 5, wherein, Before the step of generating the plurality of behavior control models based on the preset loss value threshold and the candidate models, the method further comprises: obtain second training data, wherein the second training data comprises historical state data of a specified game role in the target game and historical behavior operations of the specified game role corresponding to the historical state data; train a first model based on the second training data until the first model converges to obtain a trained first model; the first model is the same as the initial model, or the first model has the same structure as the initial model but different model parameters; determine a loss value of the trained first model by a preset loss function, and determine the loss value threshold based on the loss value.

7. The method of claim 5, wherein, The step of generating the plurality of behavior control models based on the preset loss value threshold and the candidate models comprises: determine a maximum loss value from loss values corresponding to the candidate models; if the maximum loss value is less than the loss value threshold, determine the plurality of candidate models as the plurality of behavior control models.

8. The method of claim 5, wherein, The step of generating the plurality of behavior control models based on the preset loss value threshold and the candidate models comprises: determine a maximum loss value from loss values corresponding to the candidate models; if the maximum loss value is greater than or equal to the loss value threshold, generate an evolution parameter and set an initial value for the evolution parameter; generate updated plurality of initial models based on the candidate models and update the evolution parameter; continue to perform the following steps until the evolution parameter reaches a preset parameter threshold: transform model parameters of an initial model based on a preset transformation parameter to obtain a plurality of intermediate models.

9. The method of claim 8, wherein, The step of generating updated plurality of initial models based on the candidate models comprises: copy the candidate model to obtain a copied model until the total number of the candidate model and the copied model is the same as the number of the initial model; determine the candidate model and the copied model as the updated plurality of initial models.

10. The method of claim 8, wherein, The method further comprises: if the maximum loss value is greater than or equal to the loss value threshold and the evolution parameter reaches the preset parameter threshold, determine a to-be-trained model with a loss value greater than or equal to the loss value threshold from the candidate models; obtain third training data, train the to-be-trained model based on the third training data until the loss value of the to-be-trained model is less than the loss value threshold, and determine the trained to-be-trained model and candidate models other than the to-be-trained model from the candidate models as the plurality of behavior control models; wherein the third training data comprises historical state data of a specified game role in the target game and historical behavior operations of the specified game role corresponding to the historical state data.

11. An apparatus for controlling the behavior of a game character, characterized by: The device comprises: a data acquisition module configured to obtain current state data of a target game; The input module is configured to determine a target model from a plurality of behavior control models that have been pre-trained, input the current state data into the target model, and obtain an output result. The plurality of behavior control models are obtained by: performing transformation processing on model parameters of an initial model based on preset transformation parameters to obtain a plurality of intermediate models; at least part of the model parameters of different intermediate models being different; and performing training processing on the plurality of intermediate models based on a preset loss value threshold to obtain the plurality of behavior control models. The input module is further configured to generate a plurality of initial models, wherein the initial models are preconfigured with initial model parameters, the initial model parameters are obtained by performing parameter initialization on the plurality of initial models respectively, for at least part of the model parameters in the initial models, generate a comparison parameter, and determine whether the model parameters are transformed based on a size relationship between the comparison parameter and a preset parameter threshold. If the model parameters are transformed, perform transformation processing on the model parameters based on a preset transformation strength parameter to obtain a plurality of transformed intermediate models. The control module is configured to determine a target behavior operation based on the output result, and control the game character to perform the target behavior operation.

12. An electronic device, comprising: A processor and a memory are included, the memory stores machine executable instructions that can be executed by the processor, and the processor executes the machine executable instructions to implement the game character behavior control method of any one of claims 1-10.

13. A machine-readable storage medium, characterized in that, The machine readable storage medium stores machine executable instructions, and when the machine executable instructions are called and executed by the processor, the machine executable instructions cause the processor to implement the game character behavior control method of any one of claims 1-10.

Citation Information

Patent Citations

  • Player imitation method and device and readable storage medium

    CN110052031A

  • Game role model determination method and device and server

    CN111450537A