Role processing method and device in game, electronic equipment and readable storage medium
By loading the agent character in the virtual scene customized by the player and controlling its behavior according to the virtual object attributes, the problem that the agent character cannot adapt to the customized scene is solved, and the good performance and game experience of the agent character in the customized scene is achieved.
Patent Information
- Application Number
- CN202510329338.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-07-25
AI Technical Summary
The existing agent characters are difficult to adapt to the virtual scenes customized by players and cannot be played normally in the virtual scenes customized by players, resulting in the inability to improve the player's gaming experience.
By determining the customized virtual scenes and target tasks set by the player, loading the agent role, and controlling the agent role to perform virtual behavior in the custom virtual scene according to the properties of the virtual object to complete the target task, and obtaining the target agent role suitable for the custom virtual scene.
Make the agent character better perform in specific custom virtual scenes, improve the player's gaming experience in custom virtual scenes, and reduce the development and maintenance costs of the agent character.
Smart Images

Figure CN120361547A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a method for processing characters in a game, a device for processing characters in a game, an electronic device, and a computer-readable storage medium. Background Art
[0002] In a multiplayer game, the presence of an agent character (usually also referred to as a robot character or an AI character) can often improve the player experience well. For example, an agent character can be added during the practice process for practice through human-machine battles. Another example is that when the number of players participating in a game session is insufficient, an agent character can be used to fill in the position to ensure the normal start of the game session.
[0003] However, existing agent characters are usually specially customized by game planners, and any modification to the game virtual scene requires simultaneous update of the behavior logic of the agent character. However, some current games allow players to customize the virtual scene, and the specially customized agent characters will be difficult to adapt to the virtual scene customized by players, and the agent characters cannot play normally in the virtual scene customized by players. Thus, it is difficult to improve the player game experience through agent characters. Summary of the Invention
[0004] The present disclosure provides a method for processing characters in a game, a device for processing characters in a game, an electronic device, and a computer-readable storage medium to solve or at least partially solve the above problems, specifically as follows.
[0005] In a first aspect, the present disclosure provides a method for processing characters in a game, the method including:
[0006] Determining a custom virtual scene built by a first player and a target task for the custom virtual scene; wherein the custom virtual scene includes virtual objects;
[0007] Loading an agent character in the custom virtual scene;
[0008] Controlling the agent character to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual objects so that the agent character completes the target task and obtains a target agent character applicable to the custom virtual scene.
[0009] In a second aspect, the present disclosure further provides a device for processing characters in a game, the device including:
[0010] A determination module, configured to determine a custom virtual scene built by a first player and a target task for the custom virtual scene; wherein the custom virtual scene includes virtual objects;
[0011] A loading module for loading an agent character in the custom virtual scene;
[0012] A control module for controlling the agent character to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual object, so that the agent character completes the target task and obtains a target agent character suitable for the custom virtual scene.
[0013] In a third aspect, the present disclosure also provides an electronic device, including: a processor, a memory, and computer program instructions stored on the memory and executable on the processor;
[0014] When the processor executes the computer program instructions, the method for processing characters in a game described in the first aspect above is implemented.
[0015] In a fourth aspect, the present disclosure also provides a computer-readable storage medium, in which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method for processing characters in a game described in the first aspect above is implemented.
[0016] The exemplary embodiments of the present disclosure have the following beneficial effects:
[0017] The method for processing characters in a game provided by the present disclosure first determines a custom virtual scene built by a first player and a target task for the custom virtual scene, where the custom virtual scene includes virtual objects; then, an agent character is loaded in the custom virtual scene; afterwards, according to the object attributes of the virtual object, the agent character is controlled to perform virtual behaviors in the custom virtual scene, so that the agent character completes the target task and obtains a target agent character suitable for the custom virtual scene. In the present disclosure, since the target tasks to be completed are different in different custom virtual scenes built by players, and there are also certain differences in the ability requirements for agent characters, therefore, training the agent character using a specific custom virtual scene can enable the agent character to have better performance in the specific custom virtual scene, thereby improving the adaptability of the agent character to the specific custom virtual scene. When a player needs to add an agent character to a specific custom virtual scene for a game, the game experience of the player in the custom virtual scene can be improved through the trained agent character. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A flowchart of a method for processing characters in a game provided by one embodiment of the present disclosure;
[0019] Figure 2 A block diagram of a device for processing characters in a game provided by one embodiment of the present disclosure;
[0020] Figure 3 It is a schematic diagram of the logical structure of an electronic device for implementing character processing in a game provided by one embodiment of the present disclosure. Detailed implementation manners
[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some of the embodiments of the present disclosure, rather than all of the embodiments. Components of the embodiments of the present disclosure usually described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, every other embodiment obtained by a person skilled in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0022] In this specification, the terms "a", "an", "the", and "said" are used to indicate the existence of one or more elements / components / etc.; the terms "comprising" and "having" are used to mean an open inclusion and mean that there may be additional elements / components / etc. in addition to the listed elements / components / etc.; the terms "first" and "second", etc. are only used as labels and do not limit the quantity of their objects.
[0023] It should be understood that in the embodiments of the present disclosure, "at least one" means one or more, and "a plurality" means two or more. "And / or" is merely a description of the association relationship of associated objects, indicating that there may be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally means that the associated objects before and after are in an "or" relationship. "Including A, B, and / or C" means including any one or any two or all three of A, B, and C.
[0024] It should be understood that in the embodiments of the present disclosure, "B corresponding to A", "B corresponding to A relatively", "A corresponding to B relatively", or "B corresponding to A relatively" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.
[0025] In one embodiment of the present disclosure, the character processing method in the game can be run on a local terminal device or a server. When the character processing method in the game is run on a server, the method can be implemented and executed based on a cloud interaction system, wherein the cloud interaction system includes a server and a client device.
[0026] In an optional implementation, various cloud applications can be run under the cloud interaction system, such as cloud games. Taking cloud games as an example, cloud games refer to a game mode based on cloud computing. In the operation mode of cloud games, the operating body of the game program and the main body of the game screen presentation are separated. The storage and operation of the character processing method in the game are completed on the cloud game server. The role of the client device is used for receiving and sending data and presenting the game screen. For example, the client device can be a display device with data transmission function close to the user side, such as a mobile terminal, a TV, a computer, a handheld computer, etc.; but the cloud game server in the cloud is used for information processing. When playing the game, the player operates the client device to send an operation instruction to the cloud game server. The cloud game server runs the game according to the operation instruction, encodes and compresses the game screen and other data, and returns it to the client device through the network. Finally, the client device decodes and outputs the game screen.
[0027] In an optional embodiment, taking a game as an example, a local terminal device stores a game program and is used to present a game screen. The local terminal device is used to interact with the player through a graphical user interface, that is, the game program is downloaded and installed by an electronic device and run conventionally. The local terminal device may provide the graphical user interface to the player in a variety of ways, for example, it may be rendered and displayed on a display screen of the terminal, or provided to the player through a holographic projection. For example, the local terminal device may include a display screen and a processor, the display screen is used to present a graphical user interface, the graphical user interface includes a game screen, and the processor is used to run the game, generate a graphical user interface, and control the display of the graphical user interface on the display screen.
[0028] In a possible implementation, an embodiment of the present invention provides a method for processing a character in a game, providing a graphical user interface through a terminal device, wherein the terminal device can be the local terminal device mentioned above, or can be a client device in the cloud interaction system mentioned above.
[0029] Figure 1 A flowchart of a method for processing a role in a game provided by one embodiment of the present disclosure is shown. Figure 1 As shown, the method includes the following steps S101 to S103.
[0030] Step S101: Determine the custom virtual scene built by the first player and the target task for this custom virtual scene, where the custom virtual scene includes virtual objects.
[0031] Optionally, the game can include racing games, exploration games, etc.
[0032] The first player is one of the players in the game. The first player can build a virtual scene in the game, that is, a custom virtual scene, so as to obtain a UGC (User Generated Content) game map. The custom virtual scene built by the first player can be used in the game rounds participated by the first player himself, or provided to other players for use in game rounds where the first player does not participate. The present disclosure does not limit this.
[0033] In an optional implementation manner, the game can provide scene modules for players to build virtual scenes. The scene module refers to a virtual object or a component of a virtual object. Players can build a virtual scene from scratch by means of splicing, stacking, etc.
[0034] In another optional implementation manner, the game can provide a template scene, which includes an initial virtual scene. Players can add or delete scene modules in this initial virtual scene, or adjust the position and posture of the scene modules therein, so as to build the required virtual scene with a small amount of adjustment on the basis of the template scene.
[0035] In order to make the intelligent agent character applicable to the custom virtual scene of the first player, the intelligent agent character can be trained in the custom virtual scene of the first player.
[0036] In the present disclosure, the target task for this custom virtual scene, that is, the task required to win in this custom virtual scene. If a virtual character completes the target task in this custom virtual scene, it can be determined that the virtual character wins.
[0037] Optionally, the target task can include at least one of the following:
[0038] Reach a preset area in the custom virtual scene;
[0039] Reach a preset area in the custom virtual scene within the first preset time;
[0040] Not be eliminated within the second preset time;
[0041] The number of target virtual objects collected in the custom virtual scene exceeds the first quantity threshold.
[0042] For the first type of task described above, if a virtual character starts moving from an initial position in a custom virtual scene and finally reaches a preset area in the custom virtual scene, it can be determined that the virtual character has won. Optionally, the preset area can be specified by the player.
[0043] For the second type of task described above, if a virtual character starts moving from an initial position in a custom virtual scene and finally reaches a preset area in the custom virtual scene within a specified time, it can be determined that the virtual character has won. Optionally, both the first preset time and the preset area can be specified by the player.
[0044] For the third type of task described above, if a virtual character is not eliminated within a specified time, it can be determined that the virtual character has won. Optionally, the second preset time can be specified by the player.
[0045] For the fourth type of task described above, if the number of target virtual objects collected by a virtual character in a custom virtual scene exceeds a specified quantity threshold, it can be determined that the virtual character has won. Among them, the object attribute of the target virtual object can be allowing the character to collect. Optionally, both the target virtual object and the first quantity threshold can be specified by the player.
[0046] Step S102: Load an agent character in the custom virtual scene.
[0047] After the first player builds a custom virtual scene, the training process of the agent character can be started. First, an agent character can be loaded in the custom virtual scene.
[0048] In one embodiment of the present disclosure, this step can be implemented in the following manner: Load a pre-trained agent character in the custom virtual scene.
[0049] In this embodiment, the agent character that needs to be trained in the custom virtual scene can be a pre-trained agent character. That is, an agent character with at least one general ability such as collision, movement, attack judgment, obstacle avoidance, collection, and confrontation obtained through pre-training can be used as a basis, and then the playing ability in a specific virtual scene (the custom virtual scene built by the player in the present disclosure) can be obtained through fine-tuning, so that the pre-trained agent character can learn the behavior patterns related to the specific virtual scene through fine-tuning. In this way, the development cost of the agent character can be reduced.
[0050] In an alternative embodiment, the pre-training process of the agent character can be executed on the game server.
[0051] Step S103: According to the object attributes of the virtual object, control the intelligent agent character to perform virtual behaviors in the custom virtual scene, so that the intelligent agent character can complete the target task and obtain the target intelligent agent character applicable to the custom virtual scene.
[0052] In different custom virtual scenes built by players, the scene structures will be different. Therefore, when training the intelligent agent character, it is necessary to use the object attributes of the virtual objects in the custom virtual scene to control the character's behaviors, so that the intelligent agent character can make virtual behaviors suitable for the custom virtual scene.
[0053] In different custom virtual scenes built by players, in addition to the different scene structures, the tasks required to win will also be different. Therefore, the target task needs to be used as one of the judgment indicators for whether to stop training the intelligent agent character, that is, one of the judgment indicators for whether the training of the intelligent agent character converges. According to whether the intelligent agent character can complete the target task after training, it is judged whether to stop training the intelligent agent character. If the intelligent agent character can complete the target task after training, it means that the intelligent agent character already has the ability to complete the target task in this custom virtual scene, and the training of the intelligent agent character can be stopped. If the intelligent agent character cannot complete the target task after training, it means that the intelligent agent character does not yet have the ability to complete the target task in this custom virtual scene, and the training of the intelligent agent character can be continued until the intelligent agent character can complete the target task after training, so as to have the ability to complete the target task in this custom virtual scene.
[0054] In summary, in different custom virtual scenes built by players, the target tasks are different, so the expected behaviors of the intelligent agent character will also have certain differences. Therefore, the intelligent agent character trained for a specific custom virtual scene will have better performance in the specific custom virtual scene, thus improving the adaptability of the intelligent agent character to the specific custom virtual scene and giving players a better gaming experience. In addition, there will also be certain differences in the scene structures, prop types, and executable operations of different custom virtual scenes. For example, some custom virtual scenes support attack behaviors while some do not. Therefore, there are also significant differences in the ability requirements for the intelligent agent character in different custom virtual scenes.
[0055] In one implementation, after the intelligent agent role is loaded, the intelligent agent role can be controlled to perform M training sessions in the custom virtual scene according to the object attributes of the virtual objects in the custom virtual scene. Each training session includes controlling the intelligent agent role to perform N virtual behaviors in the custom virtual scene M and N are both positive integers. If the intelligent agent role completes the target task N times in the last training session of the M training sessions, the training of the intelligent agent role can be stopped to obtain a target intelligent agent role suitable for the custom virtual scene.
[0056] If the intelligent agent role can complete the target task N times in the last training session of the M training sessions, it indicates that the intelligent agent role has converged for the custom virtual scene and can complete the target task in the custom virtual scene. At this time, the training of the intelligent agent role can be stopped to obtain a target intelligent agent role suitable for the custom virtual scene. The target intelligent agent role can play normally in the virtual scene customized by the first player. When a game needs to be played in this custom virtual scene, the player can choose to add this target intelligent agent role to enhance the gaming experience.
[0057] The method provided by the present disclosure does not require a large amount of manual development and adjustment for the intelligent agent role. Therefore, it can be applied to a large number of and complex UGC virtual scenes, saving the development and maintenance costs for the intelligent agent role.
[0058] In one implementation, after controlling the intelligent agent role to perform a virtual behavior in the custom virtual scene once, the model parameters of the intelligent agent role can be adjusted according to whether the intelligent agent role completes the target task this time.
[0059] During the training process, the intelligent agent role needs to perform a total of M×N actions of performing virtual behaviors in the custom virtual scene. After each action, the model parameters of the intelligent agent role can be adjusted once according to whether the intelligent agent role completes the target task this time. In this way, each action of the intelligent agent role during the training process can change the model parameters of the intelligent agent role, enabling the decision-making ability of the intelligent agent role to be improved to a certain extent, so that the intelligent agent role is closer to completing the target task in the next action. The intelligent agent role gradually gains the ability to complete the target task in this custom virtual scene through parameter adjustment after each action. In this implementation, the intelligent agent role performs M×N actions during the training process. Correspondingly, the model parameters of the intelligent agent role need to be adjusted M×N times.
[0060] Optionally, the step of adjusting the model parameters of the intelligent agent role according to whether the intelligent agent role completes the target task this time can be implemented in the following way:
[0061] If the intelligent agent role completes the target task this time, a positive incentive is given to the model parameters of the intelligent agent role;
[0062] If the intelligent agent role fails to complete the target task this time, a negative incentive or zero incentive is given to the model parameters of the intelligent agent role.
[0063] If the intelligent agent role completes the target task after an action, a positive incentive can be given to the model parameters of the intelligent agent role. If the intelligent agent role fails to complete the target task after an action, a negative incentive or zero incentive can be given to the model parameters of the intelligent agent role. Among them, positive incentives and negative incentives are opposite incentives. For example, if the positive incentive is to add 1 (+0.5, +2, +3, etc.) to a certain model parameter of the intelligent agent role, the negative incentive can be to subtract 1 (-0.5, -2, -3, etc.) from this model parameter, and the zero incentive can be to keep this model parameter unchanged.
[0064] That is to say, if the intelligent agent role completes the target task after an action, a positive feedback can be given to the model parameters of the intelligent agent role, so that the intelligent agent role can learn some behavior patterns from this action, which is convenient for gradually approaching the goal of completing the target task in the next action.
[0065] In one embodiment of the present disclosure, the step of controlling the intelligent agent role to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual object may include: controlling the intelligent agent role to move on the surface of the first virtual object, where the object attribute of the first virtual object is that the role is allowed to pass through.
[0066] As described above, the custom virtual scene has its unique scene structure. Therefore, when training the intelligent agent role, it is necessary to use the object attributes of the virtual objects in the custom virtual scene to control the role's behaviors. In this embodiment, the virtual objects in the custom virtual scene may include virtual objects whose object attributes are that the role is allowed to pass through. When the intelligent agent role plays in this custom virtual scene, if it encounters such virtual objects, it can move on the surface of this virtual object. The intelligent agent role can move to avoid attacks, go to a specified location or area, etc.
[0067] Exemplarily, the first virtual object may be a virtual plot (including land plot, wood plot, grass plot, etc.), a virtual road surface, a virtual geometric body, etc.
[0068] In one embodiment of the present disclosure, the steps of controlling an agent character to perform virtual behaviors in a custom virtual scene according to the object attributes of virtual objects may further include: in response to the location of the agent character in the custom virtual scene and the location of a second virtual object satisfying a preset condition, controlling the agent character to collect the second virtual object, where the object attribute of the second virtual object is that it allows the character to collect it.
[0069] In this embodiment, the virtual objects in the custom virtual scene may further include virtual objects whose object attribute is that they allow the character to collect them. When the agent character plays in this custom virtual scene, if it encounters such virtual objects, it can collect the virtual objects. The collected virtual objects can be used as one of the judgment indicators for whether the agent character has completed the target task (for example, the target task is a collection task), and can also be used as an additional reward or additional penalty during settlement.
[0070] Optionally, the situation where the location of the agent character in the custom virtual scene and the location of the second virtual object satisfy the preset condition may include the following cases:
[0071] The location of the agent character in the custom virtual scene overlaps with the location of the second virtual object;
[0072] Or,
[0073] The distance between the location of the agent character in the custom virtual scene and the location of the second virtual object does not exceed a distance threshold.
[0074] In this embodiment, if the location of the agent character overlaps with the location of the second virtual object that allows the character to collect it, or the distance from the location of the agent character to the location of the second virtual object that allows the character to collect it does not exceed the distance threshold, it can be considered that the agent character is currently within the position range where it can collect the second virtual object, and then the agent character can be controlled to collect the second virtual object.
[0075] Exemplarily, the second virtual object may be virtual coins, virtual gems, virtual plants, virtual food, etc.
[0076] In one embodiment of the present disclosure, the steps of controlling an agent character to perform virtual behaviors in a custom virtual scene according to the object attributes of virtual objects may further include: in response to the agent character performing an interaction trigger behavior on a third virtual object, controlling the third virtual object to interact with the agent character, where the object attribute of the third virtual object is that it allows interaction with the character.
[0077] In this embodiment, the virtual objects in the customized virtual scene may further include virtual objects whose object attributes allow interaction with the character. When the intelligent agent character plays in this customized virtual scene, if it encounters such virtual objects, it can interact with them. Among them, the interaction method may be related to factors such as the type and structure of the virtual objects.
[0078] Optionally, controlling the third virtual object to interact with the intelligent agent character may include at least one of the following:
[0079] Controlling the third virtual object to attack the intelligent agent character;
[0080] Controlling the third virtual object to carry the intelligent agent character to move;
[0081] Controlling the third virtual object to make the intelligent agent character fly;
[0082] Controlling the third virtual object to change the combination method so that the third virtual object provides an interaction space with the intelligent agent character;
[0083] Controlling the third virtual object to change the combination method so that the third virtual object provides an interaction component with the intelligent agent character.
[0084] For the first above-mentioned interaction method, the third virtual object may have an attack ability, such as shooting arrows, dropping fruits, etc. Exemplarily, if the intelligent agent character is within the attack range of the third virtual object, or performs a virtual behavior that triggers the third virtual object to attack, the third virtual object can attack the intelligent agent character.
[0085] For the second above-mentioned interaction method, the third virtual object may have the ability to carry the character to move. For example, when the intelligent agent character is on the surface of the third virtual object, it can trigger the third virtual object to move, thereby driving the intelligent agent character to move. Another example is that when the intelligent agent character is on the surface of the third virtual object, it can trigger the third virtual object to tilt, thereby rolling the intelligent agent character from one end of the third virtual object to the other end. Exemplarily, if the intelligent agent character is on the surface of the third virtual object or connected to the third virtual object, the third virtual object can carry the intelligent agent character to move.
[0086] For the above-mentioned third interaction method, the third virtual object may have the ability to enable the character to fly. For example, when the agent character is located at a specific position / area (such as an ejection point, a launch point) on the surface of the third virtual object, or when a virtual behavior that triggers the third virtual object to exert a force on the agent character that can make the agent character fly is executed, the agent character can fly under the action of the third virtual object. Exemplarily, if the agent character is located at the ejection point on the surface of the third virtual object, the third virtual object can eject the agent character so that the agent character can fly.
[0087] For the above-mentioned fourth interaction method, the third virtual object may have the ability to change its own combination method. If the agent character is located at a specific position / area on the surface of the third virtual object or is not more than a certain distance threshold away from the third virtual object, the third virtual object can change its own combination method so that the third virtual object provides an interaction space with the agent character. For example, by opening a door, moving an obstacle, etc., the space inside the third virtual object is revealed, or by recombining its own components to form an interaction space. This interaction space can allow the character to pass through or allow the character to stand, etc. Exemplarily, the third virtual object can open a passage or door that a character can pass through by changing the combination method.
[0088] For the above-mentioned fifth interaction method, the third virtual object may have the ability to change its own combination method. If the agent character is located at a specific position / area on the surface of the third virtual object or is not more than a certain distance threshold away from the third virtual object, the third virtual object can change its own combination method so that the third virtual object provides interaction components with the agent character. For example, a step for a character to stand on is extended, or an obstacle that hinders the character is extended.
[0089] Exemplarily, the third virtual object can be a virtual mechanism device, virtual doors and windows, virtual weapons, virtual plants, etc.
[0090] In one embodiment of the present disclosure, the character processing method in the game may further include the following steps: in response to a reward object determination operation performed by the first player, a reward object is selected from the virtual objects included in the custom virtual scene or a virtual object is added to the custom virtual scene as a reward object.
[0091] In this embodiment, the first player can select one or more virtual objects from the virtual objects included in the custom virtual scene built by himself / herself as reward objects through the reward object determination operation, or add one or more virtual objects to the custom virtual scene built by himself / herself as reward objects. Optionally, the reward object determination operation can include at least one of the following: the operation of selecting a virtual object as a reward object from the custom virtual scene; the operation of adding a virtual object as a reward object to the custom virtual scene.
[0092] Among them, the reward object is used in the training process of the agent role. The reward object can be one of the judgment indicators for whether to stop the training of the agent role, that is, one of the judgment indicators for whether the training of the agent role converges. In this embodiment, both the target task and the reward object are judgment indicators for whether to stop the training of the agent role, and a comprehensive judgment needs to be made by combining the two.
[0093] Among them, the nature of the reward object is also a virtual object. The reward object can be a virtual object selected by the first player from the virtual objects included in the custom virtual scene, or a virtual object additionally added by the first player on the basis of the custom virtual scene. If the reward object can be a virtual object selected by the first player from the virtual objects included in the custom virtual scene, then the reward object is valid in the game session, that is, it is allowed to be displayed in the game session. If the reward object is a virtual object additionally added by the first player on the basis of the custom virtual scene, then the reward object is invalid in the game session, that is, it is not allowed to be displayed in the game session.
[0094] Optionally, the object attribute of the reward object can be allowing the character to pass through or allowing the character to collect. In one implementation, the agent role can collect the reward objects during the training process, and the number of collected reward objects will be used to evaluate the training effect of the agent role to determine whether to stop the training of the agent role. In another implementation, the agent role can pass through the reward objects during the training process, and the number of passed reward objects will be used to evaluate the training effect of the agent role to determine whether to stop the training of the agent role.
[0095] Optionally, the step of stopping the training of the agent role and obtaining the target agent role applicable to the custom virtual scene if the agent role completes the target task N times in the last training of M times can be implemented in the following way: if the agent role completes the target task N times in the last training of M times, and the number of reward objects obtained or passed by the agent role in the last training meets the first preset condition, stop the training of the agent role and obtain the target agent role applicable to the custom virtual scene.
[0096] The position of the reward object can reflect the first player's understanding of the custom virtual scene. The position of the reward object is also the key position that the first player believes the character needs to pass through to complete the target task in the custom virtual scene. If the intelligent agent character can complete the target task N times in the last training of M trainings, and the number of reward objects obtained or passed by the intelligent agent character in the last training meets the first preset condition, it indicates that the intelligent agent character has converged for this custom virtual scene and can generally complete the target task in this custom virtual scene according to the first player's understanding of the custom virtual scene. That is, the intelligent agent character has the ability to complete the target task by means of the key positions understood by the first player through training. At this time, the training of the intelligent agent character can be stopped, so as to obtain the target intelligent agent character applicable to this custom virtual scene. The target intelligent agent character can play normally in the virtual scene customized by the first player. When a game needs to be played in this custom virtual scene, the player can choose to join this target intelligent agent character to improve the game experience.
[0097] Determining whether the training of the intelligent agent character converges only based on whether the target task is completed will result in the problem of sparse rewards, and sparse rewards will cause the training of the intelligent agent character not to converge or not to converge quickly. In an exemplary embodiment of the present disclosure, a reward object can be determined, and on the basis of the target task, the reward object is added as one of the judgment indicators for whether the training of the intelligent agent character converges, so that every time a reward object is obtained or passed by the intelligent agent character during the training process, it can be regarded as an intermediate goal for the intelligent agent character to achieve the target task. In this way, the problem of sparse rewards caused by determining whether the training of the intelligent agent character converges only based on whether the target task is completed can be solved.
[0098] In one implementation, during the period of controlling the intelligent agent character to perform a virtual behavior in the custom virtual scene once, every time the intelligent agent character obtains or passes a reward object, the model parameters of the intelligent agent character can be adjusted; after controlling the intelligent agent character to perform a virtual behavior in the custom virtual scene once, according to whether the number of reward objects obtained or passed by the intelligent agent character this time meets the first preset condition, and whether the intelligent agent character completes the target task this time, the model parameters of the intelligent agent character can be adjusted.
[0099] During the training process, the intelligent agent role needs to perform a total of M×N actions of executing virtual behaviors in this custom virtual scenario. In each action, whenever the intelligent agent role obtains or passes by a reward object, the model parameters of the intelligent agent role can be adjusted once. In addition, after each action, it is also necessary to adjust the model parameters of the intelligent agent role according to whether the intelligent agent role has completed the target task in this action. In this implementation, the intelligent agent role performs M×N actions during the training process. Correspondingly, the model parameters of the intelligent agent role need to be adjusted M×N+K times, where K is the total number of reward objects obtained or passed by the intelligent agent role during M×N actions.
[0100] Optionally, the reward objects can have an order, which can be set by the first player or determined according to a preset strategy, such as determining according to the specified time order or reverse order of the reward objects. In each action, whenever the intelligent agent role obtains or passes by a reward object in the order of the reward objects, the model parameters of the intelligent agent role can be adjusted. For example, if the first reward object obtained by the intelligent agent role in an action is the reward object ranked first in the order of the reward objects, the model parameters of the intelligent agent role can be adjusted for the first time; if the second reward object obtained by the intelligent agent role in an action is the reward object ranked second in the order of the reward objects, the model parameters of the intelligent agent role can be adjusted for the second time. However, if the second reward object obtained by the intelligent agent role in an action is not the reward object ranked second in the order of the reward objects, such as the reward object ranked first, third, or fourth, the model parameters of the intelligent agent role will not be adjusted. In this way, the intelligent agent role can learn the behavior pattern of completing the target task according to the specified route (that is, the route of passing by each reward object in the order of the reward objects), so that the intelligent agent role learns the route of passing by key positions in a certain order.
[0101] Optionally, in each action, whenever the intelligent agent role obtains or passes by a reward object and the reward object is currently obtained or passed by the intelligent agent role for the first time, the model parameters of the intelligent agent role can be adjusted. In this implementation, the reward object only triggers the adjustment of the model parameters of the intelligent agent role when it is obtained or passed by the intelligent agent role for the first time. If the intelligent agent role obtains or passes by the reward object again (such as the second time, the third time), the adjustment of the model parameters of the intelligent agent role will not be triggered. In this way, it can be avoided that the intelligent agent role learns the route of repeatedly passing by key positions.
[0102] Optionally, in each action, whenever the agent character obtains or passes by a reward object in the order of the reward objects and the reward object is currently obtained or passed by the agent character for the first time, the model parameters of the agent character can be adjusted. In this embodiment, it is possible to enable the agent character to learn the route of passing through key positions in a certain order and avoid the agent character learning the route of repeatedly passing through key positions.
[0103] Optionally, whenever the agent character obtains or passes by a reward object, the step of adjusting the model parameters of the agent character can be implemented in the following manner:
[0104] Whenever the agent character obtains or passes by a reward object, a positive incentive is given to the model parameters of the agent character.
[0105] If the agent character can receive a positive feedback on its model parameters every time it obtains or passes by a reward object in an action, the agent character can learn the behavior pattern of reaching the position of the reward object from this action, which is convenient for gradually approaching the goal of completing the target task in the next action.
[0106] Optionally, according to whether the number of reward objects obtained or passed by the agent character this time meets the first preset condition and whether the agent character completes the target task this time, the step of adjusting the model parameters of the agent character can be implemented in the following manner:
[0107] If the number of reward objects obtained or passed by the agent character this time meets the first preset condition, a positive incentive is given to the model parameters of the agent character;
[0108] If the number of reward objects obtained or passed by the agent character this time does not meet the first preset condition, a negative incentive or zero incentive is given to the model parameters of the agent character;
[0109] If the agent character completes the target task this time, a positive incentive is given to the model parameters of the agent character;
[0110] If the agent character does not complete the target task this time, a negative incentive or zero incentive is given to the model parameters of the agent character.
[0111] In this embodiment, during an action in the process of training the agent role, both the completion of the target task and the number of reward objects obtained or passed through meeting the first preset condition will positively motivate the model parameters of the agent role. Among them, the degree of positive motivation can be the same or different. During an action in the process of training the agent role, both the non-completion of the target task and the number of reward objects obtained or passed through not meeting the first preset condition will negatively motivate or zero-motivate the model parameters of the agent role. Optionally, the non-completion of the target task can be negatively motivated, and the number of reward objects obtained or passed through not meeting the first preset condition can be negatively motivated; Optionally, the non-completion of the target task can be negatively motivated, and the number of reward objects obtained or passed through not meeting the first preset condition can be zero-motivated; Optionally, the non-completion of the target task can be zero-motivated, and the number of reward objects obtained or passed through not meeting the first preset condition can be negatively motivated; Optionally, the non-completion of the target task can be zero-motivated, and the number of reward objects obtained or passed through not meeting the first preset condition can be zero-motivated. Among them, if both the non-completion of the target task and the number of reward objects obtained or passed through not meeting the first preset condition are negatively motivated, the degree of negative motivation can be the same or different.
[0112] If the situation of completing the target task (assuming positive motivation) and the number of reward objects obtained or passed through not meeting the first preset condition (assuming negative motivation) occurs, then after combining the motivations generated by these two situations, whether the model parameters of the agent role are finally positively motivated or negatively motivated depends on the degree of motivation for the above two situations. If the degree of motivation for completing the target task is greater, then it may finally be positively motivated. If the degree of motivation for the number of reward objects obtained or passed through not meeting the first preset condition is greater, then it may finally be negatively motivated.
[0113] Similarly, if the situation of not completing the target task (assuming negative motivation) and the number of reward objects obtained or passed through meeting the first preset condition (assuming positive motivation) occurs, then after combining the motivations generated by these two situations, whether the model parameters of the agent role are finally positively motivated or negatively motivated depends on the degree of motivation for the above two situations. If the degree of motivation for not completing the target task is greater, then it may finally be negatively motivated. If the degree of motivation for the number of reward objects obtained or passed through meeting the first preset condition is greater, then it may finally be positively motivated.
[0114] In one embodiment of the present disclosure, the reward objects include one or more of the following: marked points, virtual items. Among them, the virtual items can include virtual gold coins, virtual gems, virtual food, etc., and the present disclosure does not make specific limitations on this.
[0115] In one implementation of the present disclosure, it is possible to determine whether the number of marked points passed by the agent role in the last training of M trainings meets the first preset condition. In another implementation of the present disclosure, it is possible to determine whether the number of virtual items obtained by the agent role in the last training of M trainings meets the first preset condition.
[0116] In one embodiment of the present disclosure, that the number of reward objects obtained or passed by the agent role in the last training meets the first preset condition may include any one of the following two cases.
[0117] The first case: The average number of reward objects obtained or passed by the agent role after performing virtual behaviors N times in a custom virtual scenario in the last training exceeds a second quantity threshold.
[0118] For example, in the last training of M trainings, if the total number of virtual items obtained by the agent role in N actions is K1, then the average number of virtual items obtained by the agent role in these N actions is K1 / N. If K1 / N > T1, where T1 is the second quantity threshold, it means that the number of virtual items obtained by the agent role in N actions in the last training of M trainings meets the first preset condition.
[0119] Another example, in the last training of M trainings, if the total number of marked points passed by the agent role in N actions is K1, then the average number of marked points passed by the agent role in these N actions is K1 / N. If K1 / N > T1, where T1 is the second quantity threshold, it means that the number of marked points passed by the agent role in N actions in the last training of M trainings meets the first preset condition.
[0120] The second case: The average total number of reward objects obtained or passed by the agent role after performing virtual behaviors N times in a custom virtual scenario in the last training exceeds a third quantity threshold.
[0121] For example, in the last training of M trainings, if the total number of virtual items obtained by the agent role in N actions is K2, if K2 > T2, where T2 is the third quantity threshold, it means that the number of virtual items obtained by the agent role in N actions in the last training of M trainings meets the first preset condition. Optionally, T2 = NT1.
[0122] Another example, in the last training of M trainings, if the total number of marked points passed by the agent role in N actions is K2, if K2 > T2, where T2 is the second quantity threshold, it means that the number of marked points passed by the agent role in N actions in the last training of M trainings meets the first preset condition. Optionally, T2 = NT1.
[0123] Exemplarily, virtual items may have virtual values. For example, one virtual gold coin may be worth 1000 in-game coins or game points. In an alternative embodiment, it is also possible to determine whether to stop training the agent character by judging whether the virtual value of the obtained virtual item meets certain value conditions. For example, if the virtual value of the reward object obtained or passed by the agent character in the last training meets the second preset condition, the training of the agent character is stopped.
[0124] In one embodiment, the virtual value of the virtual item obtained by the agent character in the last training meeting the second preset condition may include:
[0125] The average virtual value of the reward objects obtained or passed by the agent character after performing virtual behaviors N times in the custom virtual scene in the last training exceeds the first value threshold;
[0126] Or,
[0127] The total virtual value of the reward objects obtained or passed by the agent character after performing virtual behaviors N times in the custom virtual scene in the last training exceeds the second value threshold.
[0128] In one embodiment of the present disclosure, the character processing method in the game may further include the following steps S01 to step S02.
[0129] Step S01: After one training of the agent character, display the travel route of the agent character in the custom virtual scene during this training in the graphical user interface provided by the terminal controlled by the first player, and display the total number of reward objects obtained or passed by the agent character in each historical training process or the change trend of this total number with the number of training times in the graphical user interface.
[0130] In this embodiment, the travel route of the agent character in the custom virtual scene can be recorded in each game of each training process. After each training is completed, the travel route of the agent character in the custom virtual scene during this training can be displayed in the graphical user interface provided by the terminal controlled by the first player. Optionally, the travel route of the agent character in the custom virtual scene in the last game of this training process can be displayed. In this way, it is convenient for the first player to observe the specific behavior pattern of the agent character and judge whether the agent character adapts to the custom virtual scene.
[0131] In addition, after each training session is completed, the total number of reward objects obtained or passed by the agent character during each historical training process, or the trend of change of this total number with the number of training sessions, can be displayed in the graphical user interface provided by the terminal controlled by the first player. In this way, it is convenient for the first player to determine whether the training process of the agent character is gradually converging.
[0132] Optionally, the travel route of the agent character in the custom virtual scene during the training process, and the total number of reward objects obtained or passed by the agent character during each historical training process, or the trend of change of this total number with the number of training sessions, can be visually displayed by means such as charts. For example, the travel route can be displayed in the top-down plane map of the custom virtual scene. The specific display method is not limited in this disclosure.
[0133] Step S02: After displaying the travel route of the agent character in the custom virtual scene during the current training process, and the total number of reward objects obtained or passed by the agent character during each historical training process, or the trend of change of this total number with the number of training sessions, in response to a continue training instruction for the agent character, control the agent character to perform the next training in the custom virtual scene.
[0134] After the first player observes and analyzes the travel route of the agent character during the current training process, and the total number of reward objects obtained or passed by the agent character during each historical training process, or the trend of change of this total number with the number of training sessions, if it is determined that the agent character has not yet adapted to the custom virtual scene and the training process of the agent character has not converged, the first player can perform a continue training operation for the agent character to issue a continue training instruction for the agent character. For example, a control for confirming whether to continue training can be displayed in the graphical user interface. This control provides two options: "Yes" and "No". The first player can click the "Yes" option in this control, and then can control the agent character to continue the next training in the custom virtual scene until the training process converges and then stop the training.
[0135] In one embodiment of the present disclosure, the character processing method in this game may further include the following step S03.
[0136] Step S03: If the number of reward objects obtained or passed by the agent character in the last training does not meet the first preset condition, display prompt information for the reward objects in the custom virtual scene in the graphical user interface provided by the terminal controlled by the first player, where the prompt information is used to prompt the first player to adjust the reward objects in the custom virtual scene.
[0137] If the position or quantity of the reward objects is set inappropriately, it will be difficult for the training of the agent character to converge. Therefore, in this embodiment, the training process can be monitored. When it is found that the training does not converge, the first player can be prompted to adjust the reward objects so that the training process can converge as soon as possible. Specifically, in this embodiment, after the agent character is trained multiple times, if the quantity of the reward objects obtained or passed by the agent character in the last training does not meet the first preset condition, prompt information for the reward objects in the custom virtual scene can be displayed in the graphical user interface provided by the terminal controlled by the first player, so as to prompt the first player to adjust the reward objects in the custom virtual scene.
[0138] Optionally, the adjustment of the reward objects includes adding reward objects. Correspondingly, the character processing method in this game can further include the following steps: after the above prompt information is displayed, in response to the addition operation for the reward objects, add reward objects in the custom virtual scene.
[0139] Exemplarily, the first player can perform a double-click operation of the left mouse button, a touch double-click operation, or a right mouse button click to display an option list and click the option to add a reward object in the option list at the scene position where the reward object is desired to be set in the custom virtual scene, etc., to add a reward object at this scene position.
[0140] Optionally, the adjustment of the reward objects includes adjusting the position of the reward objects in the custom virtual scene. Correspondingly, the character processing method in this game can further include the following steps: after the above prompt information is displayed, in response to the position adjustment operation for the reward objects, adjust the position of the reward objects in the custom virtual scene.
[0141] Exemplarily, the first player can drag the reward object whose position is desired to be moved to the target scene position in the custom virtual scene, or perform an operation such as clicking at the target scene position after selecting the reward object, so as to move the reward object from the current scene position to the target position.
[0142] Optionally, the adjustment of the reward objects includes deleting the reward objects. Correspondingly, the character processing method in this game can further include the following steps: after the above prompt information is displayed, in response to the deletion operation for the reward objects, delete the reward objects in the custom virtual scene.
[0143] Exemplarily, the first player can perform a double-click operation of the left mouse button, a touch double-click operation, a right mouse button click to display an option list and click the option to delete the reward object in the option list for the reward object to be deleted in the custom virtual scene, etc., so as to delete the reward object.
[0144] In one implementation, adjusting the position of the rewarded object in the custom virtual scene can be achieved by deleting the rewarded object at the first position in the custom virtual scene and adding the rewarded object at the second position in the custom virtual scene.
[0145] When the number or position of the rewarded objects is set improperly, the training process may be difficult to converge. By adding rewarded objects, deleting rewarded objects, adjusting the positions of rewarded objects, etc., the convergence speed can be accelerated and the training process of the agent role can be shortened.
[0146] In one embodiment of the present disclosure, the character processing method in the game may further include the following steps S04 to step S06.
[0147] Step S04: After obtaining the target agent role applicable to the custom virtual scene, in response to the game start operation for the custom virtual scene, determine the player virtual characters participating in the game round, where the player virtual characters are controlled by the players.
[0148] The custom virtual scene built by the first player can be provided for the players in the game. After obtaining the target agent role applicable to the custom virtual scene of the first player, the player can select the custom virtual scene as the virtual scene of the game round. Exemplarily, after the first player selects the custom virtual scene, the player can start the game round by clicking the "Start Game" control or other ways that can start the game round. In response to the above game start operation for the custom virtual scene, first determine which player virtual characters participate in the game round and how many there are. Among them, the player virtual character refers to the virtual character controlled by the player.
[0149] Step S05: If the number of player virtual characters is less than the character number threshold corresponding to the round mode of the game round, load the target agent role in the custom virtual scene so that the number of characters participating in the game round is equal to the character number threshold.
[0150] In an alternative implementation, the player can enable the function of automatically filling positions by the agent role. This function can automatically fill the number of characters required for the round mode by the agent role when the number of player virtual characters is less than the character number threshold corresponding to the round mode of the game round at the start of the game.
[0151] Exemplarily, for example, if the game mode currently selected by the first player is 3V3, that is, in the game match, the virtual characters participating in the game need to be divided into two camps, and each camp includes 3 virtual characters respectively. Then, starting the game match requires 6 virtual characters. If there are only 4 virtual characters of the player participating in the game match currently, 2 target intelligent agent characters can be loaded in this custom virtual scene, so that the number of characters participating in the game match reaches 6, thus making up for the number of characters required for the current game mode.
[0152] Among them, the present disclosure does not make specific limitations on how to allocate the camps for the virtual characters of the players participating in the game match and the target intelligent agent characters for filling in the positions. Exemplarily, the player controlling the virtual character of the player can select the camp, and the target intelligent agent character can be automatically supplemented into the camp with insufficient number of characters. Another example is that both the virtual characters of the player and the target intelligent agent characters can be randomly allocated to the camps. Optionally, when allocating the camps, the number of target intelligent agent characters included in each camp can be the same, or the difference does not exceed a third quantity threshold.
[0153] Step S06: Control the virtual characters of the players and the target intelligent agent characters to play a game in this custom virtual scene.
[0154] After making up for the number of characters required for the current game mode, the virtual characters of the participating players and the loaded target intelligent agent characters can be controlled to play a game in this custom virtual scene, thereby enhancing the gaming experience of the players.
[0155] In one embodiment of the present disclosure, the method for processing characters in the game provided by the present disclosure can be executed on the game client. Optionally, the problem of excessive waiting time for players can be avoided by executing it in the background.
[0156] The character processing method in the game provided by the present disclosure first determines a custom virtual scene built by a first player and a target task for the custom virtual scene, where the custom virtual scene includes virtual objects; then, an intelligent agent character is loaded in the custom virtual scene; after that, according to the object attributes of the virtual objects, the intelligent agent character is controlled to perform virtual behaviors in the custom virtual scene so that the intelligent agent character completes the target task and obtains a target intelligent agent character suitable for the custom virtual scene. In the present disclosure, since the target tasks to be completed are different in different custom virtual scenes built by players, and there are also certain differences in the ability requirements for intelligent agent characters, therefore, training the intelligent agent character using a specific custom virtual scene can enable the intelligent agent character to have better performance in the specific custom virtual scene, thereby improving the adaptability of the intelligent agent character to the specific custom virtual scene. When a player needs to add an intelligent agent character to play in a specific custom virtual scene, the game experience of the player in the custom virtual scene can be improved through the trained intelligent agent character.
[0157] Corresponding to the character processing method in the game provided by the embodiments of the present disclosure, the embodiments of the present disclosure also provide a character processing device in the game. As Figure 2 shown, the device 700 includes:
[0158] A determination module 701, configured to determine a custom virtual scene built by a first player and a target task for the custom virtual scene; where the custom virtual scene includes virtual objects;
[0159] A loading module 702, configured to load an intelligent agent character in the custom virtual scene;
[0160] A control module 703, configured to control the intelligent agent character to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual objects, so that the intelligent agent character completes the target task and obtains a target intelligent agent character suitable for the custom virtual scene.
[0161] Optionally, the controlling the intelligent agent character to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual objects includes:
[0162] Controlling the intelligent agent character to move on the surface of a first virtual object, where the object attribute of the first virtual object is that a character is allowed to pass through.
[0163] Optionally, the controlling the intelligent agent character to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual objects further includes:
[0164] In response to the fact that the location of the agent role in the custom virtual scene and the location of the second virtual object meet a preset condition, control the agent role to collect the second virtual object, where the object attribute of the second virtual object is that it allows the role to collect it.
[0165] Optionally, the situation that the location of the agent role in the custom virtual scene and the location of the second virtual object meet a preset condition includes:
[0166] The location of the agent role in the custom virtual scene overlaps with the location of the second virtual object;
[0167] Or,
[0168] The distance between the location of the agent role in the custom virtual scene and the location of the second virtual object does not exceed a distance threshold.
[0169] Optionally, the step of controlling the agent role to perform virtual behaviors in the custom virtual scene according to the object attribute of the virtual object further includes:
[0170] In response to the agent role performing an interaction trigger behavior on a third virtual object, control the third virtual object to interact with the agent role, where the object attribute of the third virtual object is that it allows interaction with the role.
[0171] Optionally, the step of controlling the third virtual object to interact with the agent role includes at least one of the following:
[0172] Control the third virtual object to attack the agent role;
[0173] Control the third virtual object to carry the agent role to move;
[0174] Control the third virtual object to make the agent role fly;
[0175] Control the third virtual object to change the combination method so that the third virtual object provides an interaction space with the agent role;
[0176] Control the third virtual object to change the combination method so that the third virtual object provides an interaction component with the agent role.
[0177] Optionally, the step of controlling the agent role to perform virtual behaviors in the custom virtual scene according to the object attribute of the virtual object so that the agent role completes the target task and obtains a target agent role applicable to the custom virtual scene includes:
[0178] According to the object attributes of the virtual object, control the intelligent agent to perform M times of training in the custom virtual scene, where each training includes controlling the intelligent agent to perform virtual behaviors N times in the custom virtual scene, and both M and N are positive integers;
[0179] If the intelligent agent completes the target task N times in the last training of the M times of training, stop training the intelligent agent to obtain a target intelligent agent applicable to the custom virtual scene.
[0180] Optionally, the device is further configured to:
[0181] After controlling the intelligent agent to perform a virtual behavior in the custom virtual scene once, adjust the model parameters of the intelligent agent according to whether the intelligent agent completes the target task this time.
[0182] Optionally, the device is further configured to:
[0183] In response to the reward object determination operation performed by the first player, select a reward object from the virtual objects included in the custom virtual scene or add a virtual object in the custom virtual scene as a reward object;
[0184] The step of stopping training the intelligent agent to obtain a target intelligent agent applicable to the custom virtual scene if the intelligent agent completes the target task N times in the last training of the M times of training includes:
[0185] If the intelligent agent completes the target task N times in the last training of the M times of training, and the number of the reward objects obtained or passed by the intelligent agent in the last training meets the first preset condition, stop training the intelligent agent to obtain a target intelligent agent applicable to the custom virtual scene.
[0186] Optionally, the device is further configured to:
[0187] During the period of controlling the intelligent agent to perform a virtual behavior in the custom virtual scene once, whenever the intelligent agent obtains or passes a reward object, adjust the model parameters of the intelligent agent;
[0188] After controlling the intelligent agent to perform a virtual behavior in the custom virtual scene once, adjust the model parameters of the intelligent agent according to whether the number of the reward objects obtained or passed by the intelligent agent this time meets the first preset condition and whether the intelligent agent completes the target task this time.
[0189] Optionally, when the intelligent agent character obtains or passes by a reward object, adjusting the model parameters of the intelligent agent character includes:
[0190] When the intelligent agent character obtains or passes by a reward object, giving positive incentives to the model parameters of the intelligent agent character.
[0191] Optionally, according to whether the number of reward objects obtained or passed by the intelligent agent character this time meets the first preset condition, and whether the intelligent agent character completes the target task this time, adjusting the model parameters of the intelligent agent character includes:
[0192] If the number of reward objects obtained or passed by the intelligent agent character this time meets the first preset condition, giving positive incentives to the model parameters of the intelligent agent character;
[0193] If the number of reward objects obtained or passed by the intelligent agent character this time does not meet the first preset condition, giving negative incentives or zero incentives to the model parameters of the intelligent agent character;
[0194] If the intelligent agent character completes the target task this time, giving positive incentives to the model parameters of the intelligent agent character;
[0195] If the intelligent agent character does not complete the target task this time, giving negative incentives or zero incentives to the model parameters of the intelligent agent character.
[0196] The character processing device in the game provided by the present disclosure first determines the custom virtual scene built by the first player and the target task for the custom virtual scene, where the custom virtual scene includes virtual objects; then, loads the intelligent agent character in the custom virtual scene; afterwards, according to the object attributes of the virtual objects, controls the intelligent agent character to perform virtual behaviors in the custom virtual scene so that the intelligent agent character completes the target task and obtains the target intelligent agent character suitable for the custom virtual scene. In the present disclosure, since in different custom virtual scenes built by players, the target tasks to be completed are different and there are certain differences in the ability requirements for the intelligent agent character, therefore, training the intelligent agent character using a specific custom virtual scene can enable the intelligent agent character to have better performance in the specific custom virtual scene, thereby improving the adaptability of the intelligent agent character to the specific custom virtual scene. When the player needs to add an intelligent agent character to play in a specific custom virtual scene, the game experience of the player in the custom virtual scene can be improved through the trained intelligent agent character.
[0197] Next, an electronic device provided by an embodiment of the present disclosure will be introduced. Please refer toFigure 3 , Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Among them, a character processing device in the game described in the embodiment of the present disclosure can be deployed on the electronic device 800 to implement the functions in the embodiment of the present disclosure. Specifically, the electronic device 800 includes: a receiver 801, a transmitter 802, a processor 803, and a memory 804 (where the number of processors 803 in the electronic device 800 can be one or more, Figure 3 taking one processor as an example), where the processor 803 may include an application processor 8031 and a communication processor 8032. In some embodiments of the present disclosure, the receiver 801, the transmitter 802, the processor 803, and the memory 804 can be connected through a bus or other means.
[0198] The memory 804 may include a read-only memory and a random access memory, and provide instructions and data to the processor 803. A part of the memory 804 may also include a non-volatile random access memory (NVRAM). The memory 804 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, where the operation instructions may include various operation instructions for implementing various operations.
[0199] The processor 803 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, where the bus system may include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clear illustration, all kinds of buses are referred to as the bus system in the figure.
[0200] The method disclosed in the above embodiments of the present disclosure can be applied to or implemented by the processor 803. The processor 803 can be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed through the integrated logic circuit in the hardware of the processor 803 or instructions in software form. The above-mentioned processor 803 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 803 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 804, and the processor 803 reads the information in the memory 804 and combines its hardware to complete the steps of the above method.
[0201] The receiver 801 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 802 can be used to output digital or character information through the first interface; the transmitter 802 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 802 can also include a display device such as a display screen.
[0202] In the embodiments of the present disclosure, the application processor 8031 in the processor 803 is used to execute the character processing method in the game in the embodiments of the present disclosure. It should be noted that the specific manner in which the application processor 8031 executes each step is based on the same concept as each method embodiment in the present disclosure, and the technical effects brought by it are the same as those of each method embodiment in the present disclosure. For specific content, reference can be made to the description in the method embodiments shown above in the present disclosure, and details will not be repeated here.
[0203] The embodiments of the present disclosure also provide a chip for running instructions, and this chip is used to execute the technical solution of the character processing method in the game in the above embodiments.
[0204] An embodiment of the present disclosure also provides a computer-readable storage medium storing computer instructions, which, when running on a processor, cause the processor to execute the technical solution of the method for processing a character in the game in the above embodiment.
[0205] An embodiment of the present disclosure also provides a computer program product including a computer program, which is used to execute the technical solution of the method for processing a character in the game in the above embodiment when executed by a processor.
[0206] The above computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or dedicated server.
[0207] It should be understood that the present disclosure is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
[0208] Although the present disclosure is disclosed above in preferred embodiments, it is not intended to limit the present disclosure. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the scope defined by the claims of the present disclosure.
Claims
1. A method for processing characters in a game, characterized in that The method includes: Determining a custom virtual scene built by a first player and a target task for the custom virtual scene; wherein, the custom virtual scene includes virtual objects; Loading an agent character in the custom virtual scene; Controlling the agent character to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual objects, so that the agent character completes the target task and obtains a target agent character applicable to the custom virtual scene.
2. The method according to claim 1, wherein The controlling the agent character to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual objects includes: Controlling the agent character to move on the surface of a first virtual object, wherein the object attribute of the first virtual object is that the character is allowed to pass through.
3. The method according to claim 2, wherein The controlling the agent character to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual objects further includes: In response to the position of the agent character in the custom virtual scene and the position of a second virtual object satisfying a preset condition, controlling the agent character to collect the second virtual object, wherein the object attribute of the second virtual object is that the character is allowed to collect.
4. The method according to claim 3, wherein The position of the agent character in the custom virtual scene and the position of the second virtual object satisfying the preset condition includes: The position of the agent character in the custom virtual scene overlapping with the position of the second virtual object; Or, The distance between the position of the agent character in the custom virtual scene and the position of the second virtual object not exceeding a distance threshold.
5. The method according to claim 2, wherein The controlling the agent character to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual objects further includes: In response to the agent character performing an interaction trigger behavior on a third virtual object, controlling the third virtual object to interact with the agent character, wherein the object attribute of the third virtual object is that it is allowed to interact with the character.
6. The method according to claim 5, wherein The controlling the third virtual object to interact with the agent character includes at least one of the following: Controlling the third virtual object to attack the agent character; Controlling the third virtual object to carry the agent character to move; Controlling the third virtual object to make the agent character fly; Controlling the third virtual object to change the combination method so that the third virtual object provides an interaction space with the agent character; Controlling the third virtual object to change the combination method so that the third virtual object provides an interaction component with the agent character.
7. The method according to claim 1, characterized in that, The controlling the agent character to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual objects, so that the agent character completes the target task and obtains a target agent character applicable to the custom virtual scene, includes: Control the intelligent agent to perform M times of training in the custom virtual scene according to the object attributes of the virtual object, where each training includes controlling the intelligent agent to perform N times of virtual behaviors in the custom virtual scene, and both M and N are positive integers; If the intelligent agent completes the target task N times in the last training of the M times of training, stop training the intelligent agent to obtain a target intelligent agent applicable to the custom virtual scene.
8. The method according to claim 7, wherein The method further includes: After controlling the intelligent agent to perform a virtual behavior in the custom virtual scene once, adjust the model parameters of the intelligent agent according to whether the intelligent agent completes the target task this time.
9. The method according to claim 7, characterized in that, The method further includes: In response to the reward object determination operation performed by the first player, select a reward object from the virtual objects included in the custom virtual scene or add a virtual object in the custom virtual scene as a reward object; The step of, if the intelligent agent completes the target task N times in the last training of the M times of training, stop training the intelligent agent to obtain a target intelligent agent applicable to the custom virtual scene, includes: If the intelligent agent completes the target task N times in the last training of the M times of training, and the number of the reward objects obtained or passed by the intelligent agent in the last training meets the first preset condition, stop training the intelligent agent to obtain a target intelligent agent applicable to the custom virtual scene.
10. The method according to claim 9, characterized in that, The method further includes: During the process of controlling the intelligent agent to perform a virtual behavior in the custom virtual scene once, whenever the intelligent agent obtains or passes by a reward object, adjust the model parameters of the intelligent agent; After controlling the intelligent agent to perform a virtual behavior in the custom virtual scene once, adjust the model parameters of the intelligent agent according to whether the number of the reward objects obtained or passed by the intelligent agent this time meets the first preset condition and whether the intelligent agent completes the target task this time.
11. The method according to claim 10, wherein The step of, whenever the intelligent agent obtains or passes by a reward object, adjust the model parameters of the intelligent agent, includes: Whenever the intelligent agent obtains or passes by a reward object, perform positive excitation on the model parameters of the intelligent agent.
12. The method according to claim 10, characterized in that, The step of, adjust the model parameters of the intelligent agent according to whether the number of the reward objects obtained or passed by the intelligent agent this time meets the first preset condition and whether the intelligent agent completes the target task this time, includes: If the number of the reward objects obtained or passed by the intelligent agent this time meets the first preset condition, perform positive excitation on the model parameters of the intelligent agent; If the number of the reward objects obtained or passed by the intelligent agent this time does not meet the first preset condition, perform negative excitation or zero excitation on the model parameters of the intelligent agent; If the agent role completes the target task this time, a positive incentive is given to the model parameters of the agent role; If the agent role fails to complete the target task this time, a negative incentive or zero incentive is given to the model parameters of the agent role.
13. A character processing device in a game, characterized in that The device includes: A determination module, configured to determine a custom virtual scene built by a first player, and a target task for the custom virtual scene; wherein the custom virtual scene includes virtual objects; A loading module, configured to load an agent role in the custom virtual scene; A control module, configured to control the agent role to perform virtual behaviors in the custom virtual scene according to the object attributes of the virtual objects, so that the agent role completes the target task and obtains a target agent role applicable to the custom virtual scene.
14. An electronic device, characterized in that, It includes: A processor, a memory, and computer program instructions stored on the memory and executable on the processor; When the processor executes the computer program instructions, it implements the character processing method in the game according to any one of claims 1 to 12 above.
15. A computer-readable storage medium, characterized in that, Computer program instructions are stored in the computer-readable storage medium, and when the computer program instructions are executed by a processor, they are used to implement the character processing method in the game according to any one of claims 1 to 12 above.