Virtual object control method and apparatus, storage medium, and electronic device

By acquiring and processing environmental perception data, masking invisible information, and predicting and combining the actions of virtual objects, the problem of low control accuracy caused by the monotonous representation of virtual objects is solved, resulting in higher control accuracy and a richer gaming experience.

WO2025227991A1PCT designated stage Publication Date: 2025-11-06TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Patent Information

Application Number
PCT/CN2025/084085
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-30
Filing Date
2025-03-21
Publication Date
2025-11-06

AI Technical Summary

Technical Problem

The limited representation of virtual objects in existing technologies results in low control accuracy and a lack of flexibility, restricting their application in diverse and complex tasks.

Method used

By acquiring environmental perception data of the target virtual scene, using preset parameters to mask invisible data, predicting different types of sub-actions based on the environmental perception data, and combining them to determine the target action, the invisible effect is simulated, improving anthropomorphism and control accuracy.

Benefits of technology

It enables flexible and natural representation of virtual objects, improves control accuracy and player engagement, and enhances the richness and intelligence of the gaming experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025084085_06112025_PF_FP_ABST
    Figure CN2025084085_06112025_PF_FP_ABST
Patent Text Reader

Abstract

The present application discloses a virtual object control method and apparatus, a storage medium, and an electronic device. The method comprises: displaying a target virtual scene; acquiring environment perception data corresponding to a first virtual object in the target virtual scene, the first virtual object representing an object controlled by a system, and a portion of data associated with a target virtual object in the environment perception data being determined by whether the target virtual object is visible to the first virtual object; determining action packet data on the basis of the environment perception data, and controlling the first virtual object to execute a target action according to the action packet data, the action packet data comprising different types of sub-actions predicted on the basis of the environment perception data, and the target action being determined by combining different types of sub-actions. The present application solves the technical problem of the accuracy of control of a virtual object being low due to the presentation of the virtual object being singular. The embodiments of the present application can be applied to various scenarios, such as cloud technology, artificial intelligence and smart traffic.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for controlling virtual object, storage medium and electronic device

[0001] The present application claims priority from the Chinese patent application No. 2024105445062 filed on April 30, 2024, and entitled "Method and device for controlling virtual object, storage medium and electronic device", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present disclosure relates to the field of computers, and more specifically to a method for controlling virtual objects. BACKGROUND

[0003] At present, in the related art, virtual objects are often controlled by pre-set rules. For example, a plurality of rule conditions are pre-set, the state of a virtual object is determined mechanically based on the rule conditions, and then the virtual object performs corresponding actions. In short, virtual objects often exhibit single behavior when processing tasks, lack flexibility, and are limited in application in diversified and complex tasks, resulting in the technical problem of low control accuracy of virtual objects.

[0004] At present, there is no effective solution to the above problems. SUMMARY

[0005] The embodiments of the present application provide a method and device for controlling virtual objects, a storage medium and an electronic device to at least solve the technical problem of low control accuracy of virtual objects due to single behavior of virtual objects.

[0006] According to an aspect of an embodiment of the present application, a method for controlling virtual objects is provided, comprising: displaying a target virtual scene, wherein the target virtual scene represents a game scene in which a target virtual object in a target application is located, and the target virtual object represents an object controlled by a target account logged into the target application; obtaining environment perception data corresponding to a first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by a system, and part of the data associated with the target virtual object in the environment perception data is determined by whether the target virtual object is visible to the first virtual object, and in the case that the target virtual object is not visible to the first virtual object, the part of the data is set to be masked by a pre-set parameter; determining action package data according to the environment perception data, and controlling the first virtual object to perform a target action according to the action package data, wherein the action package data includes different types of sub-actions predicted based on the environment perception data, and the target action is determined by combining the different types of sub-actions.

[0007] According to a further aspect of the embodiments of the present application, a device for controlling a virtual object is also provided, comprising: a display module configured to display a target virtual scene, wherein the target virtual scene represents a game scene in which a target virtual object in a target application is located, and the target virtual object represents an object controlled by a target account logged into the target application; an acquisition module configured to acquire environment perception data corresponding to a first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by a system, and part of data associated with the target virtual object in the environment perception data is determined by whether the target virtual object is visible to the first virtual object, and the part of data is set to be masked by a preset parameter in a case that the target virtual object is not visible to the first virtual object; and a control module configured to determine action package data according to the environment perception data, and control the first virtual object to perform a target action according to the action package data, wherein the action package data comprises different types of sub-actions predicted based on the environment perception data, and the target action is determined by combination of the different types of sub-actions.

[0008] According to a further aspect of the embodiments of the present application, a computer readable storage medium is also provided, and the computer readable storage medium stores a computer program, wherein the computer program is configured to execute the above-mentioned method for controlling a virtual object when running.

[0009] According to a further aspect of the embodiments of the present application, a computer program product is provided, comprising a computer program, and the computer program is executed by a processor to implement the above-mentioned method for controlling a virtual object.

[0010] According to a further aspect of the embodiments of the present application, an electronic device is also provided, comprising a memory and a processor, wherein the memory stores a computer program and transmits the computer program to the processor; and the processor is configured to execute the above-mentioned method for controlling a virtual object by using the computer program.

[0011] In the embodiments of the present application, a target virtual scene is determined, wherein the target virtual scene represents a picture scene in which a target virtual object in a target application is located, and the target virtual object represents an object controlled by a target account logged in the target application; environment perception data corresponding to a first virtual object in the target virtual scene is acquired, wherein the first virtual object represents an object controlled by the system, and part of data associated with the target virtual object in the environment perception data is determined according to whether the target virtual object is visible to the first virtual object, and in the case that the target virtual object is not visible to the first virtual object, the part of data is set to be masked by a preset parameter; action package data is determined according to the environment perception data, and the first virtual object is controlled to perform a target action according to the action package data, wherein the action package data includes different types of sub-actions predicted based on the environment perception data, and the target action is determined in a manner of combination of the different types of sub-actions. In other words, in the target virtual scene, the environment perception data is acquired based on the state information of whether the target virtual object is visible to the first virtual object. If the target virtual object is not visible to the first virtual object, part of the data in the environment perception data needs to be masked by a preset parameter, so as to simulate the effect of invisibility, so that the first virtual object can simulate the perception of a human user who does not cheat, thereby improving the anthropomorphism, i.e. the degree of intelligence, of the first virtual object controlled by the system from the source of environmental perception. Further, the action package data is determined according to the environment perception data, and the first virtual object performs the target action according to the action package data, which is more flexible and natural, more expressive, and more accurate in control, thereby improving the stickiness between the user and the target application. That is, the technical effect of improving the accuracy of the target action based on the environment perception data is achieved, and the technical problem of low control accuracy of the virtual object due to the single performance of the virtual object is solved.

[0012] On the other hand, in the target virtual scene, the above-mentioned environment perception data is determined based on not only the state information of whether the target virtual object is visible to the first virtual object, but also the state information of the first virtual object itself. Further, the environment perception data is acquired in combination with the state information of the first virtual object itself and the state information of whether the target virtual object is visible to the first virtual object, so as to ensure the accuracy and timeliness of the environment perception data, so that the environment perception data can fully reflect the current state of the picture scene in which the first virtual object is located, facilitate the generation of the action package data based on the environment perception data, and make the target action represented by the action package data more consistent with the current state of the first virtual object, thereby improving the accuracy and the degree of intelligence of the target action, and enabling the first virtual object to provide players with more flexible and rich gaming experience through the execution of the target action, and improving the stickiness between the players and the game.

[0013] In addition, by deploying the target action inference model in an inference server, a distinction from an application server is achieved, and action package data corresponding to a control mode of a virtual object is generated in the inference server, so as to achieve the technical purpose of reducing the computing overhead of the application server for controlling the anthropomorphic virtual object.

[0014] In another aspect, by masking part of the environmental perception data by using preset parameters, the data transmission overhead can be reduced on the basis of improving the anthropomorphic effect of the first virtual object. BRIEF DESCRIPTION OF DRAWINGS

[0015] FIG. 1 is a schematic diagram of an application environment of an optional virtual object control method according to an embodiment of the present application;

[0016] FIG. 2 is a flowchart of an optional virtual object control method according to an embodiment of the present application;

[0017] FIG. 3 is a schematic diagram of an optional virtual object control method according to an embodiment of the present application;

[0018] FIG. 4 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application;

[0019] FIG. 5 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application;

[0020] FIG. 6 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application;

[0021] FIG. 7 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application;

[0022] FIG. 8 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application;

[0023] FIG. 9 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application;

[0024] FIG. 10 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application;

[0025] FIG. 11 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application;

[0026] FIG. 12 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application;

[0027] FIG. 13 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application;

[0028] FIG. 14 is a schematic diagram of another optional method for controlling a virtual object according to an embodiment of the present application;

[0029] FIG. 15 is a schematic diagram of another optional method for controlling a virtual object according to an embodiment of the present application;

[0030] FIG. 16 is a schematic diagram of an optional device for controlling a virtual object according to an embodiment of the present application;

[0031] FIG. 17 is a schematic diagram of an optional product for controlling a virtual object according to an embodiment of the present application;

[0032] FIG. 18 is a schematic diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to make the personnel in the technical field better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should belong to the scope of protection of the present application.

[0034] It should be noted that the terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to the process, method, product or device.

[0035] First, some of the nouns or terms that appear in the description of the embodiments of the present application are applicable to the following explanations:

[0036] Reinforcement learning: Reinforcement learning (RL) is one of the paradigms of machine learning, also known as reinforcement learning, which studies the intelligent agent making decisions in a sequence of decision sequences to obtain rewards, and constantly improves its decision-making to obtain the maximum reward value. After repeated iterations, the intelligent agent obtains an optimal mapping from perception to decision.

[0037] Supervised learning: Supervised learning (SL) is one of the paradigms of machine learning, which is a data analysis method that trains an iterative learning algorithm from a labeled data set, so that the model can accurately classify data or predict results.

[0038] Depth map: A depth map is an image that takes the distance (depth) from an image collector to each point in a scene as a pixel value, which can accurately depict the pure depth information in front of the scene. The value closer to 1 indicates that the front is brighter and the field of view is more transparent.

[0039] Annular ray: A number of rays are shot outward from the center to the periphery of the center to describe the obstacle terrain information around the periphery of the center, which is a way of environmental perception.

[0040] Action mask: In deep reinforcement learning training, the action mask (AM) temporarily shields "unreasonable or unexecutable action dimensions" when the network infers an action distribution, and then samples actions from "reasonable or executable actions" according to the original probability. This can ensure that the inferred action can obtain valid data samples in the game environment, and avoid wasting computing resources due to returning invalid actions.

[0041] Strong human-likeness: A non-player character (NPC) or game character in a game whose behavior and actions are very similar to those of a real player, with diverse and intelligent characteristics, making it difficult for human players to distinguish the authenticity of the character.

[0042] Behavior tree: A tree structure used to control the decision-making of NPCs in a game.

[0043] Proximal policy optimization algorithm: The proximal policy optimization (PPO) algorithm is an online deep reinforcement learning algorithm based on policy gradient optimization for continuous or discrete action spaces.

[0044] Auxiliary reward: In reinforcement learning algorithms, the game agent can achieve various differentiated performances according to the differentiated goals, and the rewards for achieving the auxiliary goals of virtual objects are designed.

[0045] Game Agent: In the game scene, the machine that executes reinforcement learning is called an agent, which is a game artificial intelligence (AI) driven by a neural network. The environment with which it interacts is called the environment. The game agent makes reasonable decisions by analyzing the game environment, and can cooperate or compete with other agents or players, making it behave like a human being.

[0046] Atomic Action: In the game, the most basic and fine-grained primitive action operated by the game character. Such actions cannot be further disassembled and refined.

[0047] Fully Connected Layer: The fully connected layer (FC) is one of the components of the multilayer perceptron (MLP) applied in the deep neural network, which maps the features of the previous layer to the next layer in the form of vector calculation.

[0048] Convolutional Layer: Convolutional layer is a set of parallel feature maps that slide different convolution kernels over the input image and perform certain matrix operations. There are one-dimensional convolution, two-dimensional convolution and three-dimensional convolution.

[0049] Full Perception: There are many types of maps in shooting games, and the terrain scene is complex and variable. Limited by the limited visible angle range, a lot of information cannot be observed. If all information is assumed to be visible to the game agent, it is called full perception. In essence, full perception is a way of cheating to obtain information about the opponent player.

[0050] Non-full perception: In shooting games, in order to train a strong human-like AI to make various human-like operations and make it difficult for human players to distinguish the authenticity of the AI, it is necessary to align the input information of human players and ensure the human-like nature of the environment perception input. The way of perceiving environmental input by using real human players without cheating is called non-full perception.

[0051] Shooting Confidence Region: In shooting games, the shooting confidence region defines the range of valid shooting. Shooting outside the confidence region is invalid and will not cause damage to the enemy. Only shooting within the confidence region is valid.

[0052] Autoregression: In statistics, the method for processing time series is to predict the performance of the current time x using the same variable x at each time. This regression analysis method that uses the same variable x to predict its own performance is called autoregression, which has time series correlation.

[0053] Overfitting: In statistics, overfitting is the problem of constructing a model that appears to fit a dataset very closely, but does not generalize well to other datasets or predict future observations.

[0054] The present application will be described in conjunction with the embodiments below:

[0055] According to an aspect of the embodiments of the present application, a virtual object control method is provided. Optionally, in the present embodiment, the virtual object control method can be applied to a hardware environment composed of a server 101 and a terminal device 103 as shown in FIG. 1. As shown in FIG. 1, the server 101 is connected to the terminal device 103 through a network, and can be used to provide services such as the virtual object control method for the terminal device or an application 107 installed on the terminal device. The application can be a video application, an instant messaging application, a browser application, an educational application, a game application, etc. A database 105 can be set on the server or independently of the server, and can be used to provide data storage services for the server 101, for example, a game data storage server. The network can include but is not limited to a wired network and a wireless network. The wired network includes a local area network, a metropolitan area network, and a wide area network. The wireless network includes Bluetooth, wireless network communication technology (WIFI), and other wireless communication networks. The terminal device 103 can be a terminal device configured with an application, and can include but is not limited to at least one of a mobile phone (such as an Android phone, an iOS phone, etc.), a notebook computer, a tablet computer (Portable Android Device, PAD), a palm computer, a MID (Mobile Internet Devices), a desktop computer, a smart television, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, a virtual reality (Virtual Reality, VR) terminal, an augmented reality (Augmented Reality, AR) terminal, a mixed reality (Mixed Reality, MR) terminal, etc. The server can be a single server, a server cluster composed of multiple servers, or a cloud server.

[0056] In combination with FIG. 1, the virtual object control method can be executed by an electronic device, which can be a terminal device or a server. The virtual object control method can be implemented by the terminal device or the server respectively, or by the terminal device and the server together.

[0057] In an exemplary embodiment, the following can be included but are not limited to:

[0058] S1, display a target virtual scene on the terminal device 103, wherein the target virtual scene represents a game scene in which a target virtual object in a target application (application program 107 shown in FIG. 1) is located, and the target virtual object represents an object controlled by a target account logged into the target application;

[0059] S2, obtain, by the server 101, environmental perception data corresponding to a first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by the system, and part of the environmental perception data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object, and in the case that the target virtual object is not visible to the first virtual object, the part of the data is set to be masked by a preset parameter;

[0060] S3, the server 101 or the terminal device 103 determines action package data according to the environmental perception data, and controls the first virtual object to perform a target action according to the action package data, wherein the action package data includes different types of sub-actions predicted based on the environmental perception data, and the target action is determined by combination of the different types of sub-actions.

[0061] The above is only an example, and the present embodiment is not specifically limited.

[0062] Optionally, as an optional implementation, as shown in FIG. 2, the control method of the virtual object includes:

[0063] S202, determine a target virtual scene, wherein the target virtual scene represents a picture scene in which a target virtual object in a target application is located, and the target virtual object represents an object controlled by a target account logged into the target application.

[0064] Optionally, in the embodiments of the present application, the target virtual scene can be understood as a picture scene in a target application such as a game. The target application can be, but is not limited to, a game application, a social media application, a fitness application, an education application, a business application, a music application, etc. These applications can run on a mobile phone, a tablet computer or a computer, and provide functions such as user interaction, entertainment, learning or business activities. The target virtual object can include, but is not limited to, an image displayed in a target virtual picture of the target application by a user through a target account. The target virtual object is an object controlled by the user through the target account. The target virtual object can be understood as a virtual entity representing the user. The user can perform various activities and interactions in the target virtual object through the target virtual object, such as interacting with other users on a virtual social platform, playing games in a virtual game, experiencing various scenes in a virtual reality environment, etc. It can be understood that the target virtual object can include, but is not limited to, a virtual character, a virtual pet, a virtual item, etc. The target virtual object can represent the user's presence in the virtual world and provide the user with a rich virtual experience.

[0065] Taking a game application as an example, a player enters the game using a target account on a terminal device. The game picture displayed on the terminal device is the target virtual scene. Assuming that the game application is a shooting type game, the target virtual scene can include, but is not limited to, enemy characters, virtual equipment, city streets, jungles, deserts and different terrains, virtual props such as medical kits and ammunition kits, etc. The character controlled by the player in the target virtual scene through the target account is the target virtual object.

[0066] Illustratively, FIG. 3 is a schematic diagram of an optional control method of a virtual object according to an embodiment of the present application. The target virtual scene can be as shown in FIG. 3, which includes eight different terrain areas, namely area A, area B, area C, area D, area E, area F, area G and area H. The player can fight in various terrains, and can also use the terrain to avoid and take strategic action.

[0067] S204, acquire environment perception data corresponding to the first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by the system, and part of the data associated with the target virtual object in the environment perception data is determined by whether the target virtual object is visible to the first virtual object, and in the case that the target virtual object is not visible to the first virtual object, the part of the data is set to be masked by a preset parameter. Optionally, in the embodiments of the present application, the first virtual object can be understood as a game agent, and the first virtual object and the target virtual object can interact, including but not limited to fighting, talking, executing game tasks, etc., and the environment perception data refers to the data that the first virtual object as the target virtual object can perceive in the target virtual scene, or in other words, the data that the first virtual object can perceive in the target virtual scene if the user controls it through the target account, and the environment perception data can include but is not limited to the current position, health value, whether injured, depth map, surrounding ray, whether firing, and whether visible between the first virtual object and the target virtual object of the virtual object (such as the first virtual object or the target virtual object, etc.), wherein whether visible between the first virtual object and the target virtual object can be determined by ray detection. For example, as shown in FIG. 3, the area E in the target virtual scene includes the first virtual object and the target virtual object.

[0068] Further, FIG. 4 is a schematic diagram of an optional virtual object control method according to an embodiment of the present application, and the first virtual object is taken as a game agent in a game application for example. The first virtual object and the target virtual object are in the same building, the first virtual object is in a hall on the first floor of the building, and the target virtual object is in a bedroom on the second floor. At this time, the first virtual object and the target virtual object are not visible to each other, for example, the target virtual object does not exist in the game field of view of the first virtual object, that is, the first virtual object cannot determine the position of the target virtual object in the game field of view, and at this time, the information that the target virtual object cannot be seen needs to be masked in the environment perception data to deeply restore the game perception of the human player based on the non-full perception mode. For example, the basic information (such as orientation, position, etc.) of the target virtual object and whether the target virtual object can see the first virtual object all need to be masked. At this time, for the first target virtual object, the estimated position of the target virtual object can be used instead of the accurate position of the target virtual object, and then the relative position and relative angle between the two are calculated based on this.

[0069] Exemplarily, if the target virtual object is invisible to the first virtual object, part of the environmental perception data will be masked, which can include but is not limited to being determined by setting a preset parameter, which is a pre-set value. For example, the preset parameter is set to a special value (-1, 0), and the corresponding part of the data is filled with (-1, 0) to achieve depth restoration of the game perception of the human player, ensuring the anthropomorphism from the source of environmental perception, and providing data support and potential for subsequent model training anthropomorphism strategy, wherein the value of the preset parameter can be flexibly set, which is not limited in the present application. By using the preset parameter to mask part of the environmental perception data, the data transmission overhead can be reduced on the basis of improving the anthropomorphism effect of the first virtual object. Similarly, as shown in FIG. 4, if the target virtual object moves after a period of time and comes to the first floor hall, at this time, there is the first virtual object in the game field of view of the target virtual object, that is, the target virtual object and the first virtual object are visible. At this time, part of the environmental perception data associated with the target virtual object is determined by the target virtual object being visible to the first virtual object, such as the basic information of the target virtual object: orientation, position, etc., but it is not necessary to obtain the life value, bullet number and virtual prop range in the basic information, because even if the player is in the real world, the real-time bullet number and accurate blood volume information of the visible target enemy are unknown, and the elimination of these almost cheating information not only further simulates the alignment of human player perception input, but also avoids the model from learning speculative behavior by using these accurate information, thereby polluting the subsequent target action reasoning model strategy training and convergence.

[0070] It should be noted that FIG. 5 is a schematic diagram of an optional virtual object control method according to an embodiment of the present application, and the environmental perception data can include but is not limited to as shown in FIG. 5, the target virtual object being invisible to the first virtual object corresponds to part of the information 502 shown in FIG. 5, and the target virtual object being visible to the first virtual object corresponds to part of the information 504 shown in FIG. 5.

[0071] S206, determining action package data according to the environmental perception data, and controlling the first virtual object to perform a target action according to the action package data, wherein the action package data includes different types of sub-actions predicted based on the environmental perception data, and the target action is determined by combination of the different types of sub-actions.

[0072] Optionally, in the embodiments of the present application, the action package data can include but is not limited to the action data to be performed by the first virtual object, and the action package data includes data of different action types, for example, the moving direction is eastward movement, the moving manner is fast walking, etc. The target action can include a combination of a plurality of sub-actions indicated by the action package data, for example, the target action includes eastward movement and fast walking, and further, the first virtual object will move eastward in the manner of fast walking.

[0073] FIG. 6 is a schematic diagram of an optional control method of a virtual object according to an embodiment of the present application. The action package data can include but is not limited to several types of data as shown in FIG. 6, which can include but is not limited to the sum value of any one value in [0°, 30°, 60°, 90°, 120°, 150°, 180°, 210°, 240°, 270°, 300°, 330°] and the horizontal angle of the first virtual object, the horizontal angle of the horizontal aiming target virtual object, and the horizontal angle between the target virtual object and the target virtual object, wherein the environment perception data includes the angle between the target virtual object and the target virtual object; the sum value of any one value in [-8°, -5°, -2°, 0°, 2°, 5°, 8°] and the vertical angle of the vertical aiming target virtual object, the vertical angle of the horizontal aiming target virtual object, and the vertical angle between the target virtual object and the target virtual object, wherein the environment perception data includes the angle between the target virtual object and the target virtual object; whether to fire: 0 represents not to fire, and 1 represents to fire; moving manner: 0 represents static step, 1 represents fast walking, 2 represents sprint, and 3 represents staying in place; pathfinding manner: 0 represents atomic movement, 1 represents behavior tree-based auxiliary movement, 2 represents maintaining the previous moving manner, and 3 represents stopping auxiliary movement; moving direction: 16-dimensional atomic movement with an interval of 22.5° and a placeholder with no specific meaning (representing auxiliary movement); posture type: 0 represents standing, 1 represents crouching, 2 represents lying down, and 3 represents jumping; side type: 0 represents no side, 1 represents left side, and 2 represents right side; special type: 0 represents opening the scope, 1 represents changing the ammunition, and 2 represents neither opening the scope nor changing the ammunition, etc.

[0074] Further, FIG. 7 is a schematic diagram of an optional method for controlling a virtual object according to an embodiment of the present application, which can include but is not limited to periodically or at a preset time node obtaining environment perception data to determine action package data, and can also obtain environment perception data in real time, and determine action package data according to environment perception data at a periodic or preset time node. Moreover, the target action indicated by the action package data can include the same or different sub-actions at different time frames. In other words, the determined game of the sub-actions can be the same or different. As shown in FIG. 7, the gun barrel horizontal aiming orientation sub-action is updated every 2 time frames, the gun barrel vertical aiming orientation sub-action is updated every 2 time frames, the whether to fire sub-action is updated every 1 time frame, the moving manner sub-action is updated every 3 time frames, the special type sub-action is updated every 2 time frames, the moving direction sub-action is updated every 4 time frames, the pathfinding direction sub-action is updated every 5 time frames, the posture type sub-action is updated every 3 time frames, and the side type sub-action is updated every 2 time frames.

[0075] Taking a shooting game application as an example, in the case of obtaining environment perception data by means including but not limited to ray detection, the action package data can be determined based on the environment perception data. Specifically, the environment perception data can be taken as an input of a pre-trained target action inference model, and the action package data can be output by the target action inference model. For example, the environment perception data is obtained at a first time frame, and after 2 time frames, the action package data is obtained at a second time frame. The action package data is obtained by the target action inference model based on the environment perception data, so that the first virtual object performs the target action indicated by the action package data.

[0076] It should be noted that the training of the target action inference model and the deployment position are not in the background server (i.e., application server) of the target application. In other words, taking the game application as an example, the background server such as the dedicated server (DS) where the game service logic requested by the player to play the game is located, and the AI service module (i.e., inference server) where the target action inference model is located are different. The dedicated server only provides game service logic and cannot output action package data according to environmental perception data, wherein the game service logic can include but is not limited to game rules, game systems, game processes, and game balance, etc. Among them, the game rules: the basic framework of the game, which defines the interaction mode of the player in the game, including the goal of the game, the winning condition, the losing condition, the ability and characteristics of the player role, etc.; the game system: various mechanisms and functions in the game, including the attributes of the player role, the skill system, the equipment system, the battle system, the economic system, the social system, etc.; the game process: the whole experience process of the player in the game, including the start of the game, the tasks and challenges in the process, the settlement and rewards at the end, etc.; the game balance: the balance relationship between various elements in the game, including the balance of the role ability, the balance of the game system, the balance of the game difficulty, etc. The AI service module where the target action inference model is located can output action package data according to environmental perception data.

[0077] Exemplarily, FIG. 8 is a schematic diagram of an optional virtual object control method according to an embodiment of the present application. As shown in FIG. 8, taking the game application as an example, the player receives a game picture in the game client, the game picture displays a target virtual scene, the game service logic in the target virtual scene is obtained from the dedicated server, the environmental perception data corresponding to the first virtual object in the target virtual scene is transmitted to the distributed reinforcement learning network, the action package data is output by the target action inference model, and the first target virtual object performs the corresponding target action in the game picture by analyzing the action package data, for example, running forward, or first running forward 50 meters, and then lying down and keeping still.

[0078] In an exemplary embodiment, the virtual object control method proposed in the present application can be applied to the application scenario of determining the action of the game agent in the game. Through the embodiments of the present application, the current action of the game agent can be quickly determined in a short time, thereby improving the development efficiency of the game agent. Specifically, FIG. 9 is a schematic diagram of an optional virtual object control method according to an embodiment of the present application. As shown in FIG. 9:

[0079] S902, obtaining the game service logic of the current game frame from the dedicated server, including but not limited to game rules, game processes, game systems, etc.

[0080] S904-1, construct the state data of the i-th frame corresponding to the first virtual object as the environment perception data, and send the environment perception data to the online inference service, and synchronously execute S906-1 and S906-2;

[0081] S906-1, the online inference service calculates the action package data through static ray query and target action inference model, and returns the action package data to the game client or the special server;

[0082] S906-2, the client continues to execute the game business logic;

[0083] S908, the game client or the special server obtains the target action by analyzing the action package data, so that the first virtual object performs the target action, for example, the target action is to run forward for 50 meters first, and then lie down and keep still;

[0084] S910, repeat steps S902 to S908 until the game ends and the player exits the game application.

[0085] In another exemplary embodiment, the virtual object control method proposed in the present application can be applied to the application scenario of intelligent traffic. Through the embodiment of the present application, the vehicle driving path of the virtual vehicle in the intelligent traffic system can be quickly determined in a short time, and the path and driving operation of the virtual vehicle are ensured to be correct based on the environment perception data, which can include but is not limited to:

[0086] S1, obtain the intelligent traffic business logic of the current frame from the special server, including but not limited to traffic rules, traffic processes, traffic systems, etc.;

[0087] S2, construct the state data of the i-th frame corresponding to the first virtual object as the environment perception data, and send the environment perception data to the online inference service, and synchronously execute S3-1 and S3-2;

[0088] S3-1, the online inference service calculates the action package data through static ray query and target action inference model, and returns the action package data to the client or the special server;

[0089] The static ray query can be a static collision and ray detection query, which is not limited in the present application.

[0090] S3-2, the client continues to execute the intelligent traffic business logic;

[0091] S4, the client or the special server obtains the target action by analyzing the action package data, so that the virtual vehicle (the first virtual object) performs the target action;

[0092] S5, repeat the steps of S1 to S4 until the virtual vehicle reaches the destination.

[0093] It should be noted that the embodiments of the present application can be applied to games, cloud technology, artificial intelligence, intelligent transportation and various scenes.

[0094] Through the embodiments of the present application, the target virtual scene is determined, wherein the target virtual scene represents a picture scene in which a target virtual object in a target application is located, and the target virtual object represents an object controlled by a target account logged in the target application. The environment perception data corresponding to a first virtual object in the target virtual scene is obtained, wherein the first virtual object represents an object controlled by the system, and the part of the data associated with the target virtual object in the environment perception data is determined by whether the target virtual object is visible to the first virtual object. In the case that the target virtual object is not visible to the first virtual object, the part of the data is set to be masked by a preset parameter. The action package data is determined according to the environment perception data, and the first virtual object is controlled to perform the target action according to the action package data, wherein the action package data includes different types of sub-actions predicted based on the environment perception data, and the target action is determined in a manner of combination of different types of sub-actions. In other words, in the target virtual scene, the environment perception data is obtained based on the state information of whether the target virtual object is visible to the first virtual object. If the target virtual object is not visible to the first virtual object, part of the data in the environment perception data needs to be masked by a preset parameter, so as to simulate the effect of invisibility, so that the first virtual object can simulate the perception of a non-cheating human user, thereby improving the anthropomorphism of the system-controlled first virtual object from the source of environmental perception, i.e. the degree of intelligence, and further, by determining the action package data from the environmental perception data, the first virtual object performs the target action according to the action package data, which is more flexible and natural, rich in performance, and higher in control accuracy, thereby improving the stickiness between the user and the target application. That is, the technical effect of improving the accuracy of the target action based on the environment perception data is achieved, and the technical problem of low control accuracy of the virtual object due to the single performance of the virtual object is solved.

[0095] In another aspect, in the target virtual scene, the aforementioned environmental perception data includes state information of the first virtual object itself in addition to the state information based on whether the target virtual object and the first virtual object are visible to each other. Further, the environmental perception data is obtained in combination with the state information of the first virtual object itself and the state information based on whether the target virtual object and the first virtual object are visible to each other, so as to ensure the accuracy and timeliness of the environmental perception data, so that the environmental perception data can fully reflect the current state of the first virtual object in the scene, and facilitate subsequent generation of action package data based on the environmental perception data, so that the target action represented by the action package data can be more consistent with the current state of the first virtual object, improve the accuracy and intelligent degree of the target action, and enable the first virtual object to provide players with more flexible and rich game experience by executing the target action, thereby improving the stickiness between the players and the game.

[0096] As an optional solution, the aforementioned determining the action package data based on the environmental perception data and controlling the first virtual object to perform the target action according to the action package data includes: predicting the first sub-action based on the environmental perception data to obtain a first set of sub-action confidences, wherein one sub-action confidence in the first set of sub-action confidences corresponds to one action parameter of the first sub-action; predicting the second sub-action based on the environmental perception data to obtain a second set of sub-action confidences, wherein one sub-action confidence in the second set of sub-action confidences corresponds to one action parameter of the second sub-action, and the first sub-action and the second sub-action represent mutually independent sub-actions; determining a sub-action confidence in the first set of sub-action confidences that meets a preset condition as a target first sub-action confidence, and determining an action parameter corresponding to the target first sub-action confidence as first sub-action data corresponding to the first sub-action; determining a sub-action confidence in the second set of sub-action confidences that meets the preset condition as a target second sub-action confidence, and determining an action parameter corresponding to the target second sub-action confidence as second sub-action data corresponding to the second sub-action; and generating the action package data based on the first sub-action data and the second sub-action data, and controlling the first virtual object to perform the target action according to the action package data.

[0097] Optionally, in the embodiments of the present application, the first sub-action and the second sub-action can include, but are not limited to, firing, squatting, reloading, turning sideways, standing still, opening the scope, adjusting the horizontal aiming angle of the muzzle to 30°, adjusting the vertical aiming angle of the muzzle to 8°, etc., and the first sub-action and the second sub-action do not affect each other, for example, the first sub-action is squatting, and the second sub-action is reloading, then the first sub-action and the second sub-action do not affect each other, the first sub-action and the second sub-action can be executed synchronously, the first sub-action and the second sub-action are orthogonal, and for another example, the first sub-action is sprinting, and the second sub-action is turning sideways, then the first sub-action and the second sub-action affect each other, and when the first sub-action and the second sub-action are executed synchronously, the first virtual object will appear model twitching phenomenon, affecting the game experience of the player.

[0098] Exemplarily, the first sub-action and the second sub-action can be predicted based on the environment perception data, wherein the first sub-action has a first group of sub-action confidence, and the first group of sub-action confidence is used to determine the action parameter of the first sub-action, for example, the type of the first sub-action is whether to fire, then the optional action parameter of the first sub-action has two, respectively corresponding to firing and not firing, which can include, but are not limited to, setting the action parameter to 0, the first sub-action represents not firing, and the corresponding sub-action confidence is 30; setting the action parameter to 1, the first sub-action represents firing, and the corresponding sub-action confidence is 70.

[0099] Similarly, the second sub-action has a second group of sub-action confidence, and the second group of sub-action confidence is used to determine the action parameter of the second sub-action, and the type of the second sub-action is the moving way, then the optional action parameter of the second sub-action has four, respectively corresponding to standing, squatting, crawling, and jumping, which can include, but are not limited to, setting the action parameter to 0, the second sub-action represents standing, and the corresponding sub-action confidence is 30; setting the action parameter to 1, the second sub-action represents squatting, and the corresponding sub-action confidence is 35; setting the action parameter to 2, the second sub-action represents crawling, and the corresponding sub-action confidence is 15; setting the action parameter to 3, the second sub-action represents jumping, and the corresponding sub-action confidence is 20.

[0100] Further, the preset condition can be flexibly set, and the application does not make any limitation thereon. Taking whether to fire as the type of the first sub-action and the moving mode as the type of the second sub-action as an example, the preset condition is set as follows: the action parameter corresponding to the sub-action confidence with the maximum value in the first group of sub-action confidences is taken as the action parameter of the first sub-action, and the first sub-action data is generated; and the action parameter corresponding to the sub-action confidence with the maximum value in the second group of sub-action confidences is taken as the action parameter of the second sub-action, and the second sub-action data is generated. At this time, the action parameter of the first sub-action is 1, and the first sub-action indicates firing; and the action parameter of the second sub-action is 1, and the second sub-action indicates crouching.

[0101] For another example, the preset condition is set as follows: the action parameter corresponding to the sub-action confidence with the maximum value in the first group of sub-action confidences is taken as the action parameter of the first sub-action, and the first sub-action data is generated; and the action parameter corresponding to the sub-action confidence with the maximum value in the second group of sub-action confidences is taken as the action parameter of the second sub-action, and the second sub-action data is generated. If there are multiple same sub-action confidences with the maximum value in the first group of sub-action confidences, then the action parameter corresponding to a sub-action confidence randomly selected from the multiple same sub-action confidences is taken as the action parameter of the first sub-action. Similarly, if there are multiple same sub-action confidences with the maximum value in the second group of sub-action confidences, then the action parameter corresponding to a sub-action confidence randomly selected from the multiple same sub-action confidences is taken as the action parameter of the second sub-action.

[0102] For example, the first sub-action includes: the action parameter is set as 0, the first sub-action indicates not firing, and the corresponding sub-action confidence is 50; the action parameter is set as 1, the first sub-action indicates firing, and the corresponding sub-action confidence is 50. Then, one action parameter can be determined from the action parameter with the value of 0 and the action parameter with the value of 1 corresponding to the first sub-action by means of random sampling, so as to generate the first sub-action data. Similarly, the second sub-action includes: the action parameter is set as 0, the first sub-action indicates standing, and the corresponding sub-action confidence is 25; the action parameter is set as 1, the first sub-action indicates crouching, and the corresponding sub-action confidence is 25; the action parameter is set as 2, the first sub-action indicates lying down, and the corresponding sub-action confidence is 25; and the action parameter is set as 3, the first sub-action indicates jumping, and the corresponding sub-action confidence is 25. Then, one action parameter can be determined from the action parameter with the value of 0, the action parameter with the value of 1, the action parameter with the value of 2, and the action parameter with the value of 3 corresponding to the first sub-action by means of random sampling, so as to generate the second sub-action data.

[0103] In an example embodiment, the action package data can include but is not limited to first sub-action data and second sub-action data, and the target action can include but is not limited to a first sub-action and a second sub-action, for example, the first sub-action data is action data for the first virtual object to perform a shooting action, including but not limited to the position, time, and virtual prop used to perform the shooting, and the second sub-action data is action data for the first virtual object to perform a squatting action, including but not limited to the position and time to perform the squatting, wherein the execution time of the first sub-action indicated by the first sub-action data and the execution time of the second sub-action indicated by the second sub-action data can be the same or different, for example, the corresponding target action includes shooting and squatting, including but not limited to shooting first and then squatting, or squatting first and then shooting, or performing squatting and shooting at the same time, so as to increase the diversity and richness of the target action, so that the target action can better fit the current state of the first virtual object, thereby improving the accuracy of the target action.

[0104] According to the embodiments of the present application, the actions that the first virtual object can perform are divided into a plurality of mutually independent sub-actions, and the first sub-action and the second sub-action can be predicted based on the environment perception data, thereby reducing the difficulty of action prediction. From the first group of sub-action confidences and the second group of sub-action confidences, the target first sub-action confidence that meets the preset condition is selected as the first sub-action data, and the target second sub-action confidence that meets the preset condition is selected as the second sub-action data, and the action package data is obtained based on the first sub-action data and the second sub-action data, so as to control the first virtual object to perform the target action based on the action package data. Therefore, the accuracy of the sub-action data corresponding to each sub-action is higher through the confidences and the preset conditions, the accuracy of the action package data is improved, and the diversity and richness of the target action are increased, so that the target action can better fit the current state of the first virtual object, the first virtual object will be more flexible and natural when performing the action, the performance will be more rich, the control accuracy will be higher, and the stickiness between the user and the target application is improved.

[0105] As an optional solution, the above determining the action package data according to the environment perception data and controlling the first virtual object to perform the target action according to the action package data includes: predicting the first sub-action according to the environment perception data at a first frequency to obtain the first group of sub-action confidences, wherein the first frequency represents a preset execution frequency of the first sub-action; predicting the second sub-action according to the environment perception data at a second frequency to obtain the second group of sub-action confidences, wherein the second frequency represents a preset execution frequency of the second sub-action, and the first frequency and the second frequency are different.

[0106] In an example embodiment, the first sub-action is periodically predicted multiple times according to the environmental perception data at a first frequency, each time obtaining the first set of sub-action confidences, i.e., the first frequency is used as the update frequency of the first sub-action, and the first sub-action data is updated to obtain the corresponding first sub-action. For example, the first sub-action is updated at a first time point, and the first sub-action is updated every 2 time frames. Therefore, the first frequency is 2 time frames. Similarly, the second sub-action is periodically predicted multiple times according to the environmental perception data at a second frequency, each time obtaining the second set of sub-action confidences, i.e., the second sub-action is updated at a second time point, and the second sub-action is updated every 3 time frames. Therefore, the second frequency is 3 time frames. It can be understood that the first frequency and the second frequency can be the same or different.

[0107] For example, as shown in FIG. 7, it is assumed that the first sub-action is a gun muzzle horizontal aiming direction sub-action, and the first frequency is 2 time frames, i.e., the gun muzzle horizontal aiming direction sub-action is updated every 2 time frames. The second sub-action is a moving direction sub-action, and the second frequency is 3 time frames, i.e., the moving direction sub-action is updated every 3 time frames.

[0108] According to the embodiments of the present application, different sub-actions correspond to different update frequencies, e.g., the first sub-action is periodically updated at a first frequency, and the second sub-action is periodically updated at a second frequency. This ensures that different sub-actions can be updated independently of each other, so as to achieve the purpose of parallel processing of multiple sub-actions, save computing resources, and improve the generation efficiency of target actions.

[0109] As an optional solution, the method further includes: in the case where the first sub-action represents an attack operation, the first sub-action is predicted according to the environmental perception data at a first frequency to obtain the first set of sub-action confidences; in the case where the second sub-action represents a sub-action other than the attack operation, the second sub-action is predicted according to the environmental perception data at a second frequency to obtain the second set of sub-action confidences, wherein the first frequency is greater than the second frequency; or, in the case where the first sub-action represents a pathfinding operation, the first sub-action is predicted according to the environmental perception data at a first frequency to obtain the first set of sub-action confidences; in the case where the second sub-action represents a sub-action other than the pathfinding operation, the second sub-action is predicted according to the environmental perception data at a second frequency to obtain the second set of sub-action confidences, wherein the first frequency is less than the second frequency.

[0110] Optionally, in the embodiments of the present application, the attack operation can be understood as that the first virtual object will launch an attack on the target virtual object, which may cause damage to the target virtual object; and the pathfinding operation refers to that the first virtual object moves in the target scene, which may include but is not limited to atomic movement, auxiliary movement based on a behavior tree, maintaining the previous movement mode, stopping auxiliary movement, etc.

[0111] Exemplarily, taking a shooting game as an example, the attack action is firing, and then, if the first sub-action indicates firing, at this time, the first sub-action can be predicted using the first frequency, and the first group of sub-action confidences corresponding to the first sub-action are as follows: the action parameter is set to 0, the first sub-action indicates not firing, and the corresponding sub-action confidence is 50; the action parameter is set to 1, the first sub-action indicates firing, and the corresponding sub-action confidence is 50; as shown in FIG. 7, the second sub-action can be any action other than firing, for example, a movement mode, at this time, the second sub-action can be predicted using the second frequency, and the first frequency is greater than the second frequency, that is, the prediction frequency of the first sub-action is higher than that of the second sub-action, for example, the first sub-action is predicted once every 2 time frames; and the second sub-action is predicted once every 5 time frames.

[0112] In addition, if the first sub-action indicates a pathfinding operation, at this time, the first sub-action can be predicted using the first frequency, and the first group of sub-action confidences corresponding to the first sub-action are as follows: the action parameter is set to 0, the first sub-action indicates atomic movement, and the corresponding sub-action confidence is 50; the action parameter is set to 1, the first sub-action indicates auxiliary movement based on a behavior tree, and the corresponding sub-action confidence is 10; the action parameter is set to 2, the first sub-action indicates maintaining the previous movement mode, and the corresponding sub-action confidence is 20; the action parameter is set to 3, the first sub-action indicates stopping auxiliary movement, and the corresponding sub-action confidence is 20; as shown in FIG. 7, the second sub-action can be any action other than pathfinding, for example, a movement mode, at this time, the second sub-action can be predicted using the second frequency, and the first frequency is less than the second frequency, that is, the prediction frequency of the first sub-action is lower than that of the second sub-action, for example, the first sub-action is predicted once every 5 time frames; and the second sub-action is predicted once every 2 time frames.

[0113] By the embodiments of the present application, the prediction frequency of the sub-action indicating the attack operation is set to the maximum, and the prediction frequency of the sub-action indicating the pathfinding operation is set to the minimum, so that the target action is more humanized and more suitable for the actual game scene, and the interaction between the first virtual object and the target virtual object is more flexible and intelligent.

[0114] As an optional solution, the above generating the action package data according to the first sub-action data and the second sub-action data, and controlling the first virtual object to perform the target action according to the action package data, comprises: obtaining a pre-determined invalid action combination; in a case where the action combination composed of the first sub-action data and the second sub-action data does not belong to the invalid action combination, generating the action package data according to the first sub-action data and the second sub-action data, and controlling the first virtual object to perform the target action according to the action package data; in a case where the action combination composed of the first sub-action data and the second sub-action data belongs to the invalid action combination, the action package data is no longer generated based on the first sub-action data and the second sub-action data.

[0115] Optionally, in the embodiment of the present application, the invalid action combination refers to an action combination in which a plurality of sub-actions are combined together to cause the first virtual object to perform an abnormal action. The invalid action combination can be one or multiple, and one invalid action combination can include but is not limited to a plurality of sub-actions. The invalid action combination can be pre-set by a person, which is not limited in the present application. It can be understood that if the first virtual object synchronously performs a plurality of sub-actions included in one invalid action combination, for example, the invalid action combination includes a fast walking sub-action and a lying down sub-action, then the first virtual object will have a twitching phenomenon, which affects the game experience of the player.

[0116] In an exemplary embodiment, if the action combination composed of the first sub-action indicated by the first sub-action data and the second sub-action indicated by the second sub-action data does not belong to the invalid action combination, the first sub-action and the second sub-action can be taken as the target action, and the first sub-action and the second sub-action do not affect each other. Conversely, if the action combination composed of the first sub-action indicated by the first sub-action data and the second sub-action indicated by the second sub-action data belongs to the invalid action combination, the action package data is no longer generated based on the first sub-action data and the second sub-action data, which is equivalent to deleting the first sub-action data and the second sub-action data. This ensures the steering flexibility of the first virtual object and effectively alleviates the twitching phenomenon of the model, and enhances the human-like performance of the first virtual object.

[0117] By the embodiments of the present application, it can be determined in advance that the combination of a plurality of sub-actions will cause an abnormal motion combination of the first virtual object, i.e., an invalid motion combination. If the combination of the first sub-action and the second sub-action belongs to the invalid motion combination, in order to avoid the first virtual object from appearing twitching and shaking and other phenomena, affecting the game experience of the player, the first sub-action data and the second sub-action data can be deleted, and the action package data is no longer generated based on the first sub-action data and the second sub-action data, which is equivalent to pruning a plurality of sub-actions, reducing the complexity of the action. If the combination of the first sub-action and the second sub-action does not belong to the invalid motion combination, the action package data is generated according to the first sub-action data and the second sub-action data, which effectively alleviates the model twitching and shaking phenomenon while ensuring the steering flexibility of the first virtual object, and enhances the human-like performance of the first virtual object.

[0118] As an optional solution, the above-mentioned obtaining the environmental perception data corresponding to the first virtual object in the target virtual scene comprises: in the case that the target virtual object is visible to the first virtual object, obtaining full-quantity attribute information corresponding to the first virtual object, the partial data, relative position information between the target virtual object and the first virtual object, and voiceprint perception information of the first virtual object; the full-quantity attribute information, the partial data, the relative position information, and the voiceprint perception information are collectively determined as the environmental perception data; in the case that the target virtual object is invisible to the first virtual object, obtaining full-quantity attribute information corresponding to the first virtual object, relative position information between the target virtual object and the first virtual object, and voiceprint perception information of the first virtual object; the full-quantity attribute information, the relative position information, and the voiceprint perception information are collectively determined as the environmental perception data, wherein the partial data is set to be shielded according to the preset parameter.

[0119] Exemplarily, taking a shooting game as an example for illustration:

[0120] Continuing to refer to FIG. 5, the full attribute information can be the base information of the first virtual object, which can include but is not limited to 68-dimensional scalar values and 56-dimensional Boolean values. The 68-dimensional scalar values include the ray lengths of the first virtual object’s eyes and muzzle to each part of the target virtual object, the (muzzle) orientation, the speed, the (muzzle and each part) position, the horizontal ray length of each part, the overhead ray length, the number of bullets, the virtual prop range, the blood volume of each part, and the vegetation blocking condition. The 56-dimensional Boolean values include whether the first virtual object (each part) is injured, whether the first virtual object (fires, aims, reloads, folds, sprints, walks, turns left and right, stands, crouches, lies, jumps, and smokes), whether the target virtual object is within the first virtual object’s field of view (e.g., the game field of view), whether the target virtual object (including each part) is exposed to the first virtual object’s eyes and muzzle, and whether the target virtual object (including each part) can be seen and shot by the first virtual object.

[0121] The relative position information can include but is not limited to 77-dimensional scalar values, including the Euclidean three-dimensional distance, the two-dimensional horizontal distance, the absolute deflection angle, the (including each part) relative deflection angle, the (including each part) ΔX and ΔY in the polar coordinate system, the ΔX, ΔY, and ΔZ in the Cartesian coordinate system, whether the target virtual object is within the first virtual object’s field of view, whether the target virtual object (including each part) is exposed to the first virtual object’s eyes and muzzle, and whether the target virtual object (including each part) can be seen and shot by the first virtual object.

[0122] The voiceprint perception information can include but is not limited to 24-dimensional scalar values, including the sound of the gun, the sound of the footsteps, the sound of the grenade in the air, the sound of the landing grenade, the distance of each voiceprint point from the first virtual object and the corresponding ΔX, ΔY, and ΔZ, and the distance of the last voiceprint position of the target virtual object from the first virtual object.

[0123] In addition, the previous decision information can include but is not limited to the various action information made by the first virtual object in the last decision game frame, and the first state request at the start of each round of the game is filled with 0.0.

[0124] In an exemplary embodiment, if the target virtual object is visible to the first virtual object, the environment perception data can include but is not limited to the full attribute information, the partial data, the relative position information between the target virtual object and the first virtual object, the voiceprint perception information of the first virtual object, and the previous decision information of the first virtual object. The partial data can refer to the partial information 504 in the head 5, which includes but is not limited to the ray lengths of the target virtual object’s eyes and muzzle to each part of the first virtual object, the (muzzle) orientation, the speed, the (muzzle and each part) position, the horizontal ray length of each part, and the overhead ray length.

[0125] In yet another exemplary embodiment, if the target virtual object is not visible to the first virtual object, the environmental perception data can include, but is not limited to, full attribute information, partially occluded data by preset parameters, relative position information between the target virtual object and the first virtual object, voiceprint perception information of the first virtual object, and previous decision information of the first virtual object. The occluded partial data can refer to the partial information 502 of FIG. 5.

[0126] Through the embodiments of the present application, the environmental perception data is refined based on whether the target virtual object is visible to the first virtual object, so that different environmental perception data is determined for different situations, so that the environmental perception data is based on the data obtained by personification, which is equivalent to the first virtual object being able to simulate the perception of a non-cheating human user, thereby improving the accuracy of the action package data by improving the accuracy of the environmental perception data, thereby increasing the diversity and richness of the target action, so that the target action can be more in line with the current state of the first virtual object, and the first virtual object will be more flexible and natural when performing the action, the performance will be more rich, and the control accuracy will be higher, thereby improving the stickiness between the user and the target application.

[0127] In addition, the environmental perception data also includes the state information of the first virtual object itself, and further, the environmental perception data is obtained in combination with the state information of the first virtual object itself and whether the first virtual object and the target virtual object are visible, so as to ensure the accuracy and timeliness of the environmental perception data, so that the environmental perception data can fully reflect the current state of the first virtual object in the scene, facilitate subsequent generation of action package data based on the environmental perception data, so that the target action represented by the action package data can be more in line with the current state of the first virtual object, improve the accuracy and intelligent degree of the target action, so that the first virtual object can provide players with more flexible and rich game experience by performing the target action, thereby improving the stickiness between the players and the game.

[0128] As an optional solution, the above determining the action package data based on the environmental perception data and controlling the first virtual object to perform the target action according to the action package data includes: sending the environmental perception data to a pre-trained target action inference model through the target application, wherein the target action inference model is deployed on an inference server, and the target application is deployed on an application server; receiving a return data packet, wherein the action package data returned by the target action inference model in response to the environmental perception data is the return data packet, and the sending speed of the environmental perception data to the target action inference model is less than or equal to the return speed of the return data packet returned by the target action inference model; and controlling the first virtual object to perform the target action according to the action package data.

[0129] Optionally, in the embodiments of the present application, the target action inference model can include but is not limited to a neural network model, the inference server and the application server can include but are not limited to a cloud server, a distributed server, etc., and the inference server and the application server are two different servers. The application server can deploy a target application, but cannot deploy the target action inference model. The target action inference model needs to be deployed on the inference server to realize the predicted output of the action package data.

[0130] Exemplarily, the environment perception data is taken as the input of the target action inference model, the output of the target action inference model is taken as the action package data, and the sending speed of the environment perception data is less than or equal to the returning speed of the returned data package, so as to ensure the timeliness of the target action, so that the first virtual object can execute the target action in time, and the dynamic balance of the sending and receiving of the target package data is achieved.

[0131] Through the embodiments of the present application, the target action inference model is deployed on the inference server to realize the differentiation from the application server, and the action package data corresponding to the control mode of the virtual object is generated on the inference server to achieve the technical purpose of reducing the computing overhead of the application server for controlling the anthropomorphic virtual object. Moreover, the sending speed of the environment perception data is less than or equal to the returning speed of the returned data package, so as to ensure the timeliness of the target action, so that the first virtual object can execute the target action in time, and the dynamic balance of the sending and receiving of the target package data is achieved.

[0132] As an optional solution, the method further includes: in the case that the sending speed exceeds the returning speed, reducing the sending frequency of the environment perception data to reduce the sending speed; or in the case that the number of accumulated action package data of the target action inference model exceeds a preset number threshold, controlling the first virtual object to execute a behavior tree action through a behavior tree logic, wherein the accumulated action package data represents the action package data sent to the application server by the inference server, and the action package data that has not realized the control of the first virtual object to execute the corresponding target action according to the action package data sent by the inference server. The behavior tree is a tree structure used to control the decision of an NPC in a game. The behavior tree logic is the structure of the behavior tree, and the behavior tree action is an action included in the behavior tree.

[0133] Optionally, in the embodiments of the present application, the above-mentioned accumulated action package data can be understood as follows: after the application server receives the action package data sent by the inference server, the application server does not timely control the first virtual object to perform the target action corresponding to the action package data. For example, at the first time point, the action package data A is received, and the action package data A indicates the target action A. At the second time point, the action package data B and the action package data C are received, the action package data B indicates the target action B, and the action package data C indicates the target action C. However, at the third time point, the first virtual object does not perform the target action A, the target action B, and the target action C. Therefore, the number of the above-mentioned accumulated action package data is 3. Assuming that the above-mentioned preset number threshold is 2 (wherein the value of the preset number threshold can be flexibly set), at this time, the above-mentioned first virtual object will be controlled to perform the behavior tree action through the behavior tree logic.

[0134] Exemplarily, in the embodiments of the present application, the packet sending speed of the environmental perception data can be reduced by increasing the request interval frame number, or the inference time of the target action inference model can be reduced by optimizing the structure of the target action inference model to accelerate the packet sending speed of the action package data. When the number of the accumulated action package data exceeds the above-mentioned preset number threshold, the behavior tree logic such as the behavior tree AI can be selected through the business logic. The behavior tree AI (Behavior Tree AI) is an artificial intelligence technology used for game development and robot control. It describes and organizes the behaviors of characters or entities through a tree structure. The behavior tree AI can automatically select appropriate behaviors to respond to environmental changes and player instructions according to different conditions. The behavior tree AI is usually composed of a series of nodes, including condition nodes, behavior nodes, and combination nodes. The condition nodes are used to judge whether the current environmental conditions meet the execution of a certain behavior, the behavior nodes describe specific behavior actions, and the combination nodes are used to organize and manage the relationships and execution orders between different nodes. By designing and configuring different nodes and the relationships between the nodes, developers can flexibly build complex behavior logic, so that characters or entities can exhibit more intelligent and natural behaviors. In order to ensure the stable and normal operation of the game logic, improve the usability, flexibility, and security of the action package data.

[0135] Through the embodiments of the present application, although the speed of the inference server response packet is affected by the inference calculation time consumption and network fluctuation delay and other factors, the response packet speed may not keep up with the sending packet speed in time, and as the number of response packets increases, the accumulated action packet data also increases, and the action timeliness of the target action inference model becomes worse due to the non-timely response packet, resulting in that the inference service loses significance, so it is necessary to ensure that the response packet speed is greater than or equal to the sending packet speed, and at least to achieve the dynamic balance of the sending and receiving packets. Specifically, if the response packet speed of the inference server is greater than or equal to the sending packet speed of the application server, the sending packet speed of the application server can be reduced, or the model structure can be optimized to reduce the inference time consumption to speed up the response packet speed, and if the number of response packet accumulations is too large, the bottom line is rolled back to the behavior tree AI to avoid the phenomenon that the player feels that the game frame is stuck, and to improve the user experience.

[0136] As an optional solution, before the method of determining the action package data according to the environmental perception data, controlling the first virtual object to perform the target action according to the action package data, the method further includes: performing distributed training on the initial action reasoning model to obtain the target action reasoning model, wherein, in each game session construction process of the distributed training, the game session construction refers to the configuration constructed for a game session: dividing the target virtual scene into a plurality of key areas, the key area being a part of the virtual scene in the target virtual scene, wherein each key area in the plurality of key areas is pre-set with a division probability; sampling a target key area from the plurality of key areas according to the division probability, the target key area being one of the plurality of key areas and being sampled based on the division probability; taking a target position (such as a center of the target key area, a shooting muzzle of a virtual prop of the first sample virtual object, etc.) in the target key area as a first area center, sampling a first sample position of a first sample virtual object in a first area with a first pathfinding distance as a radius of the first area, wherein the first pathfinding distance represents a maximum pathfinding distance pre-set for the first sample virtual object; taking the first sample position as a center, a second pathfinding distance as an outer circle radius, and a third pathfinding distance as an inner circle radius, sampling a second sample position of a second sample virtual object in a circular ring, wherein the circular ring is an area obtained by removing the inner circle from the outer circle, the second pathfinding distance represents a maximum pathfinding distance pre-set for the second sample virtual object, and the third pathfinding distance represents a minimum pathfinding distance pre-set for the second sample virtual object; in a case where there is a virtual obstacle between the first sample position and the second sample position, generating the first sample virtual object at the first sample position and generating the second sample virtual object at the second sample position, and based on first sample environmental perception data corresponding to the first sample virtual object and second sample environmental perception data corresponding to the second sample virtual object, controlling the first sample virtual object and the second sample virtual object to play a game session until the game session of the first sample virtual object and the second sample virtual object ends.

[0137] Optionally, in the embodiments of the present application, the initial action reasoning model can include but is not limited to a neural network model, and the initial action reasoning model can be trained in a distributed manner to obtain the target action reasoning model, for example, a plurality of game sessions are set, the initial action reasoning model is used to determine the target action of the first virtual object in each game session, and one of the key areas is selected as the target key area with equal probability to improve the probability of game session occurrence.

[0138] Exemplarily, FIG. 10 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application. As shown in FIG. 10, a key region is randomly selected as a target key region, and then a first pathfinding distance is set with the shooting muzzle of the virtual prop of the first sample virtual object in the target key region as the center. The first pathfinding distance can be determined by the distance between the muzzle position of the virtual prop of the first sample virtual object and the enemy position, but is not limited thereto. Then, the second pathfinding radius and the third pathfinding radius are determined based on the first pathfinding distance, the first sample virtual object is generated at the first sample position, and the second sample virtual object is generated at the second sample position. Specifically,

[0139] FIG. 11 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application. As shown in FIG. 11, it is assumed that the distance between the muzzle position of the virtual prop of the first sample virtual object and the enemy position is d, and the included angle between the straight line from the shooting impact point of the virtual prop of the first sample virtual object to the first sample virtual object and the horizontal straight line of the shooting impact point of the virtual prop of the first sample virtual object is θ. θ represents the maximum angle of the deviation of the muzzle when a normal human player shoots, and the shooting beyond the angle θ will become meaningless shooting. θ can be defined as the inverse tangent function of 40 / 1500, but the present application does not limit θ.

[0140] Further, as shown in FIG. 11, the shooting impact point of the virtual prop is taken as the center, the second pathfinding distance is taken as the outer radius, and the third pathfinding distance is taken as the inner radius. The second sample position of the second sample virtual object is on the annulus with the inner radius of the second pathfinding distance and the outer radius of the third pathfinding distance. The maximum distance between the second sample position of the second sample virtual object and the first sample position is the second pathfinding distance, and the minimum distance between the second sample position of the second sample virtual object and the first sample position is the third pathfinding distance.

[0141] Further, if there is a virtual obstacle between the first sample position and the second sample position, the virtual obstacle can include but is not limited to buildings, vehicles and the like in the game picture. In other words, the first sample virtual object can hide behind the virtual obstacle, or the first sample virtual object can hide behind the virtual obstacle. At this time, the first sample virtual object is generated at the first sample position in the target key region, the second sample virtual object is generated at the second sample position, and the game is started. The first sample virtual object and the second sample virtual object can include but are not limited to fighting, talking, performing tasks and the like, so as to quickly determine the positions of the first sample virtual object and the second sample virtual object and improve the training efficiency of the initial action reasoning model.

[0142] According to the embodiments of the present application, each key area is completely symmetrical to the two parties (such as the first virtual object and the target virtual object) in the key area, and the birth points of the two parties are selected according to a specific logic. In order to avoid the situation that one party is killed at the beginning of the game, it is required that there is an obstacle between the birth points of the two parties to block each other, so as to ensure the sample diversity. During training, a key area is randomly selected first, and then an internal birth point combination is further randomly selected for the selected key area. When one party dies or the game is over, a single game is ended, so as to improve the training efficiency of the initial action reasoning model.

[0143] As an optional solution, the above-mentioned distributed training of the initial action reasoning model to obtain the target action reasoning model includes: performing non-full-amount processing on the first sample environment perception data and the second sample environment perception data to obtain first sample data and second sample data; in the case that the first sample virtual object is visible to the second sample virtual object, performing feature extraction on the first sample data to obtain first sample features, wherein the first sample features include part of the data of the second sample virtual object; in the case that the second sample virtual object is invisible to the first sample virtual object, performing feature extraction on the second sample data to obtain second sample features, wherein the second sample features include part of the data of the first sample virtual object after being masked by the preset parameters; inputting the first sample features and the second sample features into the initial action reasoning model respectively to determine the preset index values corresponding to the first sample virtual object and the second sample virtual object respectively, and updating the initial action reasoning model according to the preset index values to determine the target action reasoning model, wherein the output of the feature network layer associated with the part of the data is uniformly masked as a target value during the updating of the initial action reasoning model.

[0144] Optionally, in the embodiments of the present application, the first sample environment perception data corresponds to a first sample virtual object, and the second sample environment perception data corresponds to a second sample virtual object, and the non-full-amount processing is to mask part of the data in the environment perception data (such as the first sample environment perception data or the second sample environment perception data). For example, part of the data in the first sample environment perception data is masked (i.e., the aforementioned masking), for example, the health value of the second sample virtual object is masked, and the value is set to -1.0 to obtain the first sample data; similarly, part of the data in the second sample environment perception data is masked, for example, the health value of the second sample virtual object is masked, and the value is set to -1.0 to obtain the second sample data. The preset index value can be understood as a personification index of the first sample virtual object and the second sample virtual object, for example, the side rate: the reasonable side can attack the enemy; the squat rate: the reasonable squat can hide the position of itself; the avoidance rate: the time length ratio of not being found in each game; the shooting rate: the reasonable shooting frequency in each game; the static step rate: the frequency of hiding the sound print by using the static step without being found; the attack rate after hiding for 3 seconds: the attack rate after hiding for 2.5 seconds; the dark shooting rate: shooting in a place where the enemy cannot see; the cover utilization rate: reasonably using the cover such as column and box to hide; the walking position dispersion rate: the walking position strategy is flexible and variable, and the virtual object does not stay in the same area for a long time; the head exposure rate when unable to shoot: the virtual object cannot use the virtual prop or is blocked by the terrain, so that the head is exposed but cannot shoot; the retreat rate after shooting: the virtual object retreats to find cover while shooting; the dangerous area stay rate: the virtual object stays in a dangerous open area without cover; and the large turning rate: the virtual object cannot suddenly change direction, which avoids the personification effect being too bad.

[0145] It can be understood that FIG. 12 is a schematic diagram of another optional virtual object control method according to an embodiment of the present application. The above-mentioned non-full-amount processing mode is shown in FIG. 12. Taking a first sample virtual object as an example, a mask is added to a feature extraction network layer related to a second sample virtual object. When the second sample virtual object is visible, normal training logic is used. When the second sample virtual object (i.e., the target for the first sample virtual object) is invisible, part of the data in the second sample environmental perception data corresponding to the second sample virtual object is filled with false data -1.0, which does not have an actual physical meaning. Therefore, the output of the feature network layer connected to the current network layer is uniformly masked to 0, and the gradient is intercepted and returned to the previous network layer. That is, the A network is trained when the second sample virtual object is visible, and the B network is trained when the second sample virtual object is invisible. Part of the network parameters of the A network and part of the network parameters of the B network are frozen and cannot be trained and updated when the second sample virtual object is invisible. This realizes a dynamic neural network structure during the entire training process. By performing non-full-amount processing on the first sample environmental perception data, the first sample data is obtained. By performing non-full-amount processing on the second sample environmental perception data, the second sample data is obtained. This can more realistically simulate the game perception input of a human player and ensure the humanization of the virtual object at the input data level of the initial action reasoning model.

[0146] Further, if the first sample virtual object is visible to the second sample virtual object, the first sample feature can be determined by performing a feature extraction operation on the first sample data, and the first sample feature includes the unmasked data in the second sample environmental perception data. Similarly, if the second sample virtual object is visible to the first sample virtual object, the second sample feature can be determined by performing a feature extraction operation on the second sample data, and the second sample feature includes the unmasked data in the first sample environmental perception data.

[0147] For example, the first sample feature can be used as the input of the initial action reasoning model to obtain a preset index value of the first sample virtual object, and the second sample feature can be used as the input of the initial action reasoning model to obtain a preset index value of the second sample virtual object. The initial action reasoning model is updated using the preset index value until the initial action reasoning model training is completed. For example, the preset index value is a mask utilization rate of 80%. When the mask utilization rate reaches 80%, it is considered that the initial action reasoning model training is successful, and the expected preset index value of the first sample virtual object and the preset index value of the second sample virtual object can be obtained. The initial action reasoning model is determined as the target action reasoning model, and the purpose of accurately predicting the action package data using the target action reasoning model is achieved.

[0148] By the embodiments of the present application, the first sample data is obtained by performing non-full-quantity processing on the first sample environment perception data, and the second sample data is obtained by performing non-full-quantity processing on the second sample environment perception data, which can more truly simulate the game perception input of a human player and ensure the personification of the virtual object at the input data level of the initial action reasoning model. Non-full-quantity processing is also performed at the initial action reasoning model training level to synchronously adapt to the input information subjected to non-full-quantity processing. Specifically, in the non-full-quantity processing manner of the initial action reasoning model, a mask is added to a feature extraction network layer related to target enemy (for example, for the first sample virtual object, the second sample virtual object is the target enemy) information. When the target enemy is visible, normal training logic is used. When the target enemy is invisible, the target enemy information is filled with false data -1.0 at this time, which does not have actual physical meaning. The output of the feature network layer connected thereto is uniformly masked as a target value, and the gradient is intercepted and returned to the previous network layer. That is, the A network is trained when the target enemy is visible, and the B network is trained when the target enemy is invisible. Part of the network parameters of the A network and part of the network parameters of the B network are frozen and cannot be trained and updated when the target enemy is invisible, so as to realize a dynamic neural network structure in the entire training process.

[0149] As an optional solution, the above-mentioned obtaining the environment perception data corresponding to the first virtual object in the target virtual scene comprises: obtaining first environment perception data corresponding to the i-th frame when the service of the target virtual scene proceeds to the i-th frame; obtaining second environment perception data corresponding to the i+n-th frame when the service of the target virtual scene proceeds to the i+n-th frame, wherein the service remains to execute a preset service logic from the i-th frame to the i+n-th frame, and i and n are positive integers; and the above-mentioned determining action package data according to the environment perception data and controlling the first virtual object to perform a target action according to the action package data comprises: determining first action package data according to the first environment perception data and controlling the first virtual object to perform a first target action according to the first action package data at the i+j-th frame; and determining second action package data according to the second environment perception data and controlling the first virtual object to perform a second target action according to the second action package data at the i+n+j-th frame, wherein j is a non-negative integer.

[0150] In an exemplary embodiment, the first environment perception data corresponds to data corresponding to the service logic executed by the first virtual object at the i-th frame, and the second environment perception data corresponds to data corresponding to the service logic executed by the first virtual object at the i+n-th frame, in other words, the generation time of the second environment perception data is later than the generation time of the first environment perception data.

[0151] Specifically, in the process of generating the corresponding first action package data using the first environment perception data, the above-mentioned business will continue to execute the preset business logic, and similarly, in the process of generating the corresponding second action package data using the second environment perception data, the above-mentioned business also continues to execute the preset business logic, in other words, the above-mentioned preset business logic will not change with the first environment perception data and the second environment perception data, for example, the business preset logic is to increase a virtual prop in the target virtual scene every 2 frames, the number of virtual props in the i-th frame is 10, n is 2, and the number of virtual props in the i+2-th frame is 11.

[0152] Further, the first action package data can be determined based on the first environment perception data, so that the first virtual object performs the target action corresponding to the first action package data at the i+j-th frame, wherein the time interval from obtaining the first environment perception data to the first virtual object performing the target action indicated by the first action package data can be represented as the j-th frame, where j can take the value of 0 or other positive integers, if j takes the value of 0, it means that the first virtual object performs the target action corresponding to the first action package data at the i-th frame.

[0153] Similarly, the second action package data can be determined based on the second environment perception data, so that the first virtual object performs the target action corresponding to the second action package data at the i+n+j-th frame, wherein the time interval from obtaining the second environment perception data to the first virtual object performing the target action indicated by the second action package data can be represented as the j-th frame, where j can take the value of 0 or other positive integers, if j takes the value of 0, it means that the first virtual object performs the target action corresponding to the second action package data at the i+n-th frame.

[0154] By the embodiments of the present application, the environment perception data is acquired every fixed time, and in the process of generating the corresponding action package data using the environment data, the preset business logic will be continuously executed, that is, the preset business logic will not change with the environment perception data. That is, the preset business logic and the target action are implemented by an asynchronous inference manner, that is, if there is no target action generated, the preset business logic is executed, and if there is a target action generated, the target action is executed, which not only does not need to consume a large amount of server performance to do a lot of collision and ray detection, reduces the performance consumption of the server, but also avoids the waiting time caused by waiting for the target action prediction, avoids the frame freezing phenomenon, and improves the user experience. In addition, FIG. 13 is a schematic diagram of another optional control method of a virtual object according to an embodiment of the present application, as shown in FIG. 13, the preset business logic is a game business logic, in the process of generating the corresponding first sample action package data using the first sample environment perception data (i.e., the i-th frame state data), that is, the i-th frame state data is sent to the AI training service based on the state request, so that the AI training service generates the action return. The above business will terminate the preset business logic, in other words, during the time of generating the first sample action package data according to the first sample environment perception data and the first sample virtual object executing the sample target action indicated by the first sample action package data (i.e., executing the return action instruction), the above business will be suspended and in synchronous waiting. For example, the preset business logic is that a virtual prop in a target virtual scene increases every 2 frames, the number of virtual props in the i-th frame is 10, n is 2, and the number of virtual props in the i+2-th frame is still 10. At this time, in the case that the first sample virtual object executes the sample target action indicated by the first sample action package data, the preset business logic will be automatically restored, that is, in the next i+3-th frame, the number of virtual props will be 11.

[0155] As an optional solution, the above determining the action package data according to the environment perception data, and controlling the first virtual object to perform the target action according to the action package data, comprises: in a case that the environment perception data acquired at the i-th frame indicates that the first virtual object needs to be controlled to first launch a virtual prop within a preset time length, acquiring a maximum angle of aiming deviation preset for the first virtual object, wherein i is a positive integer; determining a first confidence region according to a distance between the first virtual object and the target virtual object and the maximum angle of aiming deviation, wherein a radius of the first confidence region is r, and r is greater than 0; sampling a first confidence point from the first confidence region, and determining a second confidence region with the first confidence point as a center and r / x as a radius; and sampling a second confidence point from the second confidence region, wherein the second confidence point represents a aiming reticle of the first virtual object corresponding to the environment perception data acquired at the i-th frame.

[0156] Optionally, in the embodiments of the present application, taking a shooting game as an example, the virtual props can include but are not limited to virtual shooting props, as shown in FIG. 11, the first virtual object first fires the virtual prop within a preset time length, the maximum angle of aiming deviation can be represented as θ shown in FIG. 11, which represents the maximum angle of deviation of the muzzle when a normal human player shoots, and shooting beyond the θ angle will become meaningless shooting, θ can include but is not limited to an inverse tangent function defined as 40 / 1500, at this time, the maximum search radius is tan(θ)*d, and the present application does not limit θ.

[0157] Exemplarily, the first confidence region is obtained based on the distance between the first virtual object and the target virtual object and the maximum angle of aiming deviation, which can be understood as a coarse-grained target shooting confidence domain in which the first virtual object can shoot the target virtual object, and the first confidence point is a position in the first execution region, for example, shooting possibility trajectory detection is performed on each part of the target virtual object, and each part is sorted according to the shooting possibility detection result of each part in the order of the hit priority of each part, and the body part with the highest hit priority is taken as the aiming base point.

[0158] Further, the aiming base point can be taken as the first confidence point, assuming that x is 3, then r / 3 represents the radius of the second confidence region, and each point in the second confidence region is sampled to obtain a second confidence point, which can be understood as an aiming reticle when the first virtual object aims at the target virtual object to shoot in the i th frame, and the shooting confidence domain in which the first virtual object can shoot the target virtual object is further narrowed, and if the second confidence point is not in the first confidence region, the firing action of the first virtual object is terminated, and the first environmental perception data is waited for the next time to ensure that the shooting action of the first virtual object is effective shooting in the second confidence region, and the shooting hit rate of the first virtual object is improved.

[0159] Through the embodiments of the present application, when the first virtual object first fires the virtual prop, the first confidence region can be obtained based on the distance between the first virtual object and the target virtual object and the maximum angle of aiming deviation, which is equivalent to a large circle region, and the second confidence region is obtained in the first confidence region, which is equivalent to a small circle. A shooting method of random selection in a small circle and truncation outside a large circle is used to further narrow the shooting confidence domain in which the first virtual object can shoot the target virtual object, and if the second confidence point is not in the first confidence region, the firing action of the first virtual object is terminated, and the first environmental perception data is waited for the next time to ensure that the shooting action of the first virtual object is effective shooting in the second confidence region, and the shooting hit rate of the first virtual object is improved.

[0160] As an optional solution, after the second confidence point is sampled from the second confidence area, the method further includes: when the environment perception data obtained at the i+m frame indicates that the first virtual object needs to control the virtual prop to be fired for the kth time, determining a third confidence area with a third confidence point determined for the (k-1)th time as a center and r / x as a radius, where m is a positive integer and m is less than the preset time length, k is a positive integer greater than or equal to 2, and when k=2, the third confidence point is the second confidence point; sampling a target confidence point from the third confidence area, where the target confidence point represents a sighting reticle of the first virtual object corresponding to the environment perception data obtained at the i+m frame.

[0161] Optionally, in the embodiments of the present application, the first virtual object can fire the virtual prop for multiple times, for example, k is 3, and when the first virtual object fires the virtual prop for the third time, a third confidence area can be determined with a third confidence point determined for the second time as a center and r / x as a radius, so that the corresponding shooting landing point of the first virtual object when firing the virtual prop for the third time is located in the third confidence area, in other words, the shooting confidence area determined at the current time is determined by the confidence point at the previous time, for example, if k is 3, the third confidence point is the second confidence point, and the corresponding shooting landing point of the first virtual object when firing the virtual prop for the second time is located in the second confidence area.

[0162] Further, by sampling each point in the third confidence area, a target confidence point can be obtained, which can be understood as a sighting reticle of the first virtual object when aiming at the second virtual object for shooting at the i+m frame. In the case that the first virtual object continuously fires the virtual prop, the horizontal and vertical offset angles of the virtual prop muzzle when shooting each time can be calculated based on the sighting reticle determined each time, so as to simulate the human player's pressing gun shooting effect. At the same time, since the target confidence point is obtained by random sampling in the third confidence area, the muzzle produces a scattering effect, so that the shooting performance of the first virtual object under the first perspective of the playback lens can be restored to restore the human player's shooting effect, and the shooting anthropomorphism is greatly improved compared with the original behavior tree AI shooting method, thereby enhancing the shooting anthropomorphism of the AI from the bottom layer of the game client.

[0163] Exemplarily, the control method of the virtual object proposed in the application can be applied to the application scenario of general training of non-player characters in games. Specifically, based on a large-scale reinforcement learning training framework, a general training scheme for producing a strong human-like AI of a shooting game can be designed in a complex three-dimensional perception scene and a high-dimensional coupled structured action space by combining a proximal policy optimization algorithm. Human-like auxiliary enhancement processing is performed in three dimensions of perception input, network decision, and bottom execution to realize the construction of a high-intensity and strong human-like AI training scheme in a three-dimensional complex scene and a high-dimensional discrete action space. In other words, based on multi-modal environmental perception input, a non-full perception method of aligning human players is used to process environmental information of different modalities, and a dynamic neural network structure is used in the model training process to synchronize and adapt to non-full input information, thereby ensuring the human-likeness of the perception end. Moreover, in combination with the playing habits of high-level players of shooting games, different action decision frequencies are configured for different action heads without losing the completeness and flexibility of actions, so as to achieve the purpose of pruning the action exploration space at the same time to reduce the difficulty of strategy exploration and training. In addition, it also includes a large and small circle auxiliary shooting method, which combines a realistic atomic firing logic with recoil and scattering, and truly realizes the shaking and scattering of the muzzle of the human player when shooting, thereby meeting the requirements of strong human-like AI at the bottom layer of the game client. These human-like enhancement processes make the AI behavior more close to human operation habits and better help the business team to produce content.

[0164] Specifically, assuming that the application scenario in the game is a farm map, which is the map with the highest player activity, and in the case that the game has already deployed a virtual object with a behavior tree version, the virtual object is used for accompanying play and filling the number of matches. Since the existing AI logic of the behavior tree version is easy to be recognized by players, the players lose the fun of the battle, therefore, through the embodiment of the application, a high-intensity and strong human-like AI (i.e., the first virtual object described above) can be produced by reinforcement learning technology for deployment in high-end matches for accompanying play. The first virtual object can move throughout the map, reasonably use cover battles and pulling, move flexibly and naturally, and show high human-likeness and diversity in shooting performance. The game elimination playback can be aligned with the human player's perspective, and the playing method shows high human-likeness and diversity, which truly achieves the purpose of strong human-like accompanying play, thereby driving the game activity and player participation.

[0165] In addition, since the first virtual object is placed in any scene of the map, and any map area can fight with the player, high requirements are put forward for the generalization ability of the model. Through the detailed optimization of the model training battle logic and scene development in the embodiments of the present application, the training scene covers the whole map, a plurality of key battle areas (i.e. the aforementioned key areas) are delimited, one of the battle areas (i.e. the aforementioned target key area) is selected according to equal probability when each game starts, so as to improve the probability of main event occurrence; the two parties in each battle area are completely symmetrical, and the birth points of the two parties are selected according to a specific logic. In order to avoid that one party is defeated at the moment of starting the game, it is required that there are obstacles between the birth points of each other to ensure sample diversity; during training, a battle area is randomly selected first, and for the selected battle area, an internal birth point combination is further randomly selected, and the single game ends if one party is defeated or the game is timed out, so as to improve the training efficiency.

[0166] Exemplarily, the overall structure of the model training stage in the embodiments of the present application can include but is not limited to: a game client, a game environment integrated with an AI Service SDK (Artificial Intelligence Service Software Development Kit) for interface calling, and a dedicated server (i.e. the aforementioned application server) providing single game service. After the game state (i.e. the aforementioned environmental perception data) is input into the network, the instruction action return packet (i.e. the aforementioned action packet data) can be output to the game client for execution, so as to ensure the model training effect. The game client sends a starting game request to the AI server (i.e. the aforementioned inference server), and the AI server returns corresponding battle configuration data according to the training requirements. After the game client obtains the battle configuration data, the game client constructs the game (i.e. the aforementioned game construction process) and generates a reinforcement learning Agent (an entity capable of perceiving environment and taking action) (i.e. the aforementioned first virtual object). Each request response game frame sample is processed into state data (i.e. the aforementioned first sample environmental perception data and second sample environmental perception data) and sent to the distributed reinforcement learning framework in the AI server for interaction. The AI training service provided by the AI server predicts the action of the current game request frame, splices the action into executable action instructions (i.e. action packet data) and returns them to the client game engine (i.e. the aforementioned application server) for execution. The AI training service continues to wait for the next game frame sample request response. The model training stage will report business side statistical indicators regularly for detecting the training effect. When each game ends, the game client will call the related logic interface to destroy the Agent of the current game, and automatically start the next round of game according to the same configuration. In this way, data is repeatedly collected and provided to the reinforcement learning framework for training.

[0167] In an exemplary embodiment, the model training phase can include but is not limited to adopting a synchronous inference mode, in each game of the training phase, there is no human player, various business behavior tree AI and irrelevant game business logic in the corresponding dedicated server (DS), only two training battle AIs (i.e. the first virtual object and the target virtual object mentioned above) are running in the DS, the game business logic is very clean and pure, and a large amount of server performance is not consumed for a lot of collision and ray detection, the DS is in a synchronous waiting state in the time interval from sending the current frame request response of the game to getting the action return packet, and only after getting the action return packet, the game business logic is continued to run, so the DS of the training phase is a continuous cycle process from state request to action return packet, and the action request is made every n game frames, and the synchronous inference mode of the model training phase can unify the game business logic of the DS itself and the reinforcement learning training under the same time line.

[0168] The model online phase can include but is not limited to adopting an asynchronous inference mode, in each game of the online phase, a large number of online players exist in the corresponding DS, there are a lot of game interaction logics between players and players, and between players and various AIs (such as the first virtual object mentioned above), if the online phase also adopts the synchronous inference mode, a lot of collision and ray detection will be done in the DS for building state data when the DS sends the game frame request response each time, most of the server performance will be consumed in the related technology, and the server resource consumption will be proportional to the increase of the AI deployment quantity in each game after online. In addition, in the process from state request to action return packet each time, the DS is in a synchronous waiting state, and the game logic of the DS itself will not be run in this period, and only after the action return packet is obtained, the game logic of the DS itself is continued to run, so that the players will feel that the game frame is stuck in the synchronous waiting empty window period (which may be only tens of milliseconds), and the frame sticking phenomenon will be further amplified due to the model inference time and the return packet delay caused by network fluctuation. Through the embodiment of the present application, the game business logic of the DS and the inference service logic can be isolated, it is ensured that the inference service will not block the game business logic of the DS itself, and a lot of collision and ray detection work originally performed on the DS side is stripped out of the DS environment and directly migrated to the inference service provided by the inference server for development, the inference service will pre-save a large amount of static collision data, and the basic state request response sent by the DS is queried for ray detection each time, so that a lot of DS performance is not consumed for dynamic ray detection on the DS side each time, and the developer only needs to focus on the inference service logic, so as to improve the development efficiency.

[0169] Exemplarily, the asynchronous inference mode can be adopted in the online stage. The asynchronous inference has two timelines. The DS side sends the basic environment perception data to the inference service at a timing (for example, every n frames, n = 4). The inference service side performs static collision and ray detection query work on the sent environment perception data according to the first-in first-out principle, processes the network state data for inference calculation, and then returns the action instruction package to the DS side in turn according to the first-in first-out principle. After the DS stably sends the state request response every n frames, it does not need to wait for the action return package and continues to run the game business logic itself. At the same time, it detects whether there is an action return package every frame. If there is, it immediately executes the latest action. In addition, the speed of the action return package of the inference service side is affected by the inference calculation time consumption and network fluctuation delay and other factors, which may cause the return package speed to be unable to keep up with the sending speed. As the number of return packages increases, the timeliness of the action predicted by the model (i.e., the target action inference model described above) becomes worse and worse, which causes the inference service to lose its meaning. Through the embodiments of the present application, the return package speed of the inference service is ensured to be greater than or equal to the sending speed, and at least reaches the dynamic balance of sending and receiving packages. Specifically, the sending speed of the environment perception data can be reduced by increasing the request interval frame number n, or the inference time consumption can be reduced by optimizing the model structure to speed up the return speed of the action package data. If the number of return package accumulations is too large, the bottom line can be handled by rolling back to the behavior tree AI. The asynchronous inference can prevent data blocking of the main thread, ensure that the neural network (i.e., the target action inference model described above) inference time consumption and the action return package delay caused by network fluctuation do not affect other game business logic, reduce the server occupation of the business side, and ensure the safety of model inference. If the inference fails due to network fluctuation and other factors, the game logic can be switched to the ordinary behavior tree AI to ensure the stable and normal operation of the game logic. The hybrid reinforcement learning access logic has the characteristics of easy-to-use, flexible, safe and the like.

[0170] For example, in a three-dimensional scene similar to the real world, there are a large number of static, dynamic and environment perception related objects, and the details are very realistic. The game operation panel also includes a large number of rocker actions that can be executed simultaneously. In such a three-dimensional complex and variable game environment, if a virtual object with high strength and strong personification needs to be trained, personification enhancement processing needs to be performed on the environment perception side, the network decision side and the bottom execution side. Therefore, in the training process, the embodiments of the present application can use the following three methods to assist in enhancing the strategy personification.

[0171] (1) Environment perception side: non-full quantity modification.

[0172] Specifically, the expansion from 2D environment to 3D environment makes the environment entity have additional state dimensions such as height, z-axis orientation, etc., greatly increasing the state space. Large maps in 3D environment, real environment scenes such as doors, windows, stairs, buildings, and corridors further increase the difficulty of 3D environment perception. Considering the integrity and computational complexity of 3D environment perception, a general fusion multi-modal environment perception scheme is designed, and the complete 3D environment information is composed of multiple scalar information, three layers of surrounding rays, and 2D depth map in terms of complexity. At the same time, in order to align and restore the game perception of human players, the strategy anthropomorphism is strengthened from the environment perception input level, and the non-full-quantity perception method is adopted to simulate the perception input of real human players to the game environment, as follows:

[0173] a) First virtual object information: full-quantity perception method, because humans in the real world also have full-quantity perception of their own information;

[0174] b) Target virtual object information: non-full-quantity perception method, which needs to be discussed in two cases: target virtual object visible and target virtual object invisible. Because humans in the real world cannot fully acquire the information of visible target enemies, and the information of invisible target enemies is unknown, the specific methods are as follows:

[0175] 1) Target virtual object visible: obtain the basic information of the target virtual object, including orientation, position, etc., but need to exclude information such as blood volume, bullet count, and virtual prop range, because even human players in the real world cannot know the real-time bullet count and accurate blood volume information of visible target enemies. Excluding these almost cheating information not only further simulates the human player's perception input, but also avoids the model from learning speculative behavior, thereby polluting the model strategy training and convergence.

[0176] 2) Target virtual object invisible: the basic information of the target virtual object, whether the target virtual object can see the first virtual object, and other information are all shielded, and only the estimated position of the target enemy is used to replace the accurate position of the target enemy. Based on this, the relative position and relative angle between the two are calculated, because human players cannot know the various basic information of the target enemy and whether the target enemy can see themselves when the target enemy is invisible. The shielded information is all filled with a special value -1.0. This non-full-quantity perception method deeply restores the game perception of human players and ensures anthropomorphism from the source of environment perception, providing data support and potential for subsequent model training of anthropomorphic strategy.

[0177] 3) Multiple ray information: full-sensing mode, including depth map and various body rays, the previous environment interaction module mentioned DS and inference service isolation logic, which can avoid excessive model training from occupying business side performance overhead. Further, the ray detection query in the inference service can be divided into simple collision query and complex collision query. The simple collision query data is mainly used for path perception, and the accuracy requirement is not high. The rays are relatively sparse, such as three-layer surrounding rays, and horizontal rays in each part. The complex collision query data is used for scenes with high detail requirements, such as depth map, overhead ray, visibility judgment ray, and shootability judgment ray. The rays with low accuracy requirements can be replaced by simple collision data to further reduce the performance consumption of the inference service.

[0178] Based on multi-modal environment perception, non-full-quantity processing can more realistically simulate the game perception input of human players and ensure the anthropomorphism at the level of neural network input data. In addition, non-full-quantity processing is also required at the internal training level of the neural network to adapt to non-full-quantity perception input information. The non-full-quantity processing of the neural network specifically adds a mask to the feature extraction network layer related to the target enemy information. When the target enemy is visible, the normal training logic is used. When the target enemy is invisible, the feature network layer output connected to it is uniformly masked to 0, and the gradient is intercepted back to the previous network layer. That is, the A network is trained when the target enemy is visible, and the B network is trained when the target enemy is invisible. Part of the network parameters of the A network and part of the network parameters of the B network are frozen and cannot be trained and updated when the target enemy is invisible, realizing the dynamic neural network structure in the entire training process.

[0179] (2) Network decision side - multi-frequency prediction.

[0180] There are multiple controllable joystick actions in a three-dimensional game scene, and 3D environment entities have more rich actions than 2D environment entities, such as controlling the view angle, squatting, jumping and other actions, and different types of actions can be executed simultaneously, resulting in an explosion of action space combinations, increasing the difficulty of AI exploration and training. In view of the complexity of the action space, a multi-scale corner design based on player data prior is constructed to simulate and approach the habits of player view angle turning and aiming, and an action scale design under different virtual prop states is constructed to simulate the differences in player's field of view and perception under different virtual props and camera states, and finally align the player's behavior in the 3D environment. In terms of reducing action complexity, a three-finger modeling scheme is adopted to prune a large number of illegal and invalid parallel operations that represent the upper limit of human operation, thereby achieving double improvement in training efficiency and anthropomorphism. Based on the combinability of human left and right hands and sub-actions, a number of orthogonal sub-actions (i.e., the aforementioned first sub-action and second sub-action) are defined, and sub-actions can occur simultaneously.

[0181] In combination with the actual performance of the model and the habits of high-level human players, appropriate action decision frequencies are adjusted for different action heads (i.e., the aforementioned sub-actions), and the strategy anthropomorphism is strengthened from the network decision level. Action sampling does not have to be performed simultaneously between different action heads, i.e., multi-frequency prediction, to select the optimal action combination from a large action space combination in a short time. The advantages of multi-frequency prediction are as follows:

[0182] a) Taking into account high-frequency and low-frequency action sampling heads, the model twitching phenomenon is effectively alleviated while ensuring the flexibility of turning, and the model anthropomorphism performance is enhanced;

[0183] b) Pruning the action exploration space at the same time, the strategy exploration and training difficulty is effectively reduced without compromising the completeness and flexibility of the model actions;

[0184] c) The sampling frequency of the action head can be manually adjusted according to the actual effect of the model, facilitating the training of differentiated style models.

[0185] Since shooting is an action that is very sensitive to time, the decision frequency is the highest. A two-frame decision interval is used to capture potential fast-moving target enemies, and then a quick response is made to enter the firing battle. Some action heads do not require such high decision frequency, such as the sprint action in the movement mode. If the sprint is immediately switched to a static step, the model performance will be very twitching and jittering, and the anthropomorphism will be reduced. Therefore, in actual combat, the sprint should be maintained for at least a certain period of time, and the movement mode action head only needs a low prediction frequency during training. In this way, different actions have their own different sampling frequencies, which not only improves the performance of anthropomorphism, but also greatly reduces the difficulty of model training.

[0186] In addition, in addition to multi-frequency prediction between different action heads, the prediction frequency of the same action head will also change dynamically during training, which is reflected in the differences between combat scenes and non-combat scenes. The action prediction frequency associated with combat will change during the model combat process. For non-combat scenes, AI pays more attention to pathfinding, and does not require fast and high-frequency action logic, but in combat scenes, the winner is often determined within a few seconds, so AI pays more attention to micro-operation. At this time, the actions related to the position need a higher decision frequency in order to achieve the human-like playing method of shooting two guns and then quickly pulling back to hide, and then shooting two guns and pulling back to hide again, which can better simulate the playing behavior of advanced players. The dynamic multi-frequency prediction mechanism not only improves the model humanization, but also greatly reduces the difficulty of model training.

[0187] (3) Bottom execution side - shooting humanization.

[0188] For first-person shooting games, shooting feel is a direct intuitive feeling for players in the process of FPS games, and is also an important part of FPS games. Good shooting feel is also the basis for the success of the game. In FPS games, shooting is a process in which players control virtual props to move the crosshair when the crosshair and the enemy model coincide, and cause damage to the enemy under the premise of ignoring bullet speed and drop. Normal human players' aiming operations can be divided into the following four categories:

[0189] 1) Positioning: continuously measure input to keep the crosshair (i.e. the aforementioned aiming crosshair) on the moving target enemy (e.g. for the first virtual object, the target virtual object is the target enemy of the first virtual object); 2) Jumping: through instantaneous reaction and rapid measurement and input, the crosshair is accurately moved to the target point once, and whether to open the mirror (ADS or waist shot) is decided according to the situation; 3) Stabilization: through continuous fine-tuning of measurement input, the crosshair is kept on a relatively fixed target (gun pressing); 4) Prediction: by observing the target movement characteristics, the crosshair is moved to the coincidence point of the predicted trajectory destination and the target enemy.

[0190] Wherein, the aiming operation will be different according to the virtual props, distance, player action and aiming method, and the four operations will be different in different games. The selection of the appropriate aiming method is a concentrated embodiment of the game operation skills of normal human players. Even if the crosshair is perfectly aimed at the target enemy, it can still be found that sometimes the bullet will fly in another direction. This phenomenon of bullet deviation is called scattering. Generally, in the process of continuous shooting, the distribution of bullets is not a point in the middle of the crosshair, but will be randomly arranged within a certain range. The range of this bullet distribution is the scattering degree of the virtual props, and the degree of scattering is related to the type of shooting props, which is more obvious in waist shooting. In FPS games, after a normal human player completes a shot, the crosshair of the shooting prop will not remain in its original position but will shift as a whole within a certain period of time. This is caused by the recoil of the shooting prop. The recoil of the shooting prop can be roughly divided into vertical recoil and horizontal recoil. The former will cause the crosshair to shift upwards, while the latter will cause the crosshair to shift left and right. The shift of the crosshair will cause the deviation of the bullet. The recoil of the shooting prop is one of the key factors that affect whether a normal human player can hit the target. In order to reduce the impact of recoil, normal players need to use the way of pressing the gun (controlling the movement of the gun barrel and the trajectory distribution in the opposite direction) to make the bullets as concentrated as possible.

[0191] In summary, when playing FPS games, normal human players will first choose the appropriate aiming method according to the type of virtual props they hold. Due to the existence of scattering and the influence of virtual prop recoil during the shooting process, players need to use the gun pressing operation to improve the hit rate, making gun pressing shooting a necessary skill for high-level FPS players and an important factor in evaluating behavioral anthropomorphism. The existing behavior tree AI shooting method does not support real virtual prop movement logic and does not require trajectory detection. It only needs to manually configure a hit rate parameter. If the hit rate parameter is configured to 100%, it can hit the enemy steadily. If it is less than 100%, the behavior tree AI shooting will hit the enemy according to the configured hit rate probability. The shooting logic of the behavior tree AI is a result-oriented embodiment, not by real bullets causing damage, but by directly deducting the target enemy's blood volume according to the probability. Result-oriented shooting can cause the following problems:

[0192] a) Result-oriented shooting has no scattering and shooting prop recoil effect, and cannot simulate the bullet deviation distribution when players shoot;

[0193] b) It is impossible to achieve the human player's gun pressing shooting operation mentioned above, which leads to a serious lack of shooting anthropomorphism in the behavior tree first-person perspective in the elimination playback lens;

[0194] c) The strength of the behavior tree is highly dependent on the adjustment of the hit rate parameter. If the hit rate is too high, the AI may exhibit the behavior of headshotting every shot, which is unrealistic. If the hit rate is too low, the strength of the behavior tree AI will be affected, leading to an imbalance between the strength and the humanization of the behavior tree AI, further affecting the player's game experience and enjoyment.

[0195] Therefore, in the embodiments of the present application, in order to truly restore the human player's gun-pressing shooting operation, a shooting logic supporting real trajectory detection with scattering and shooting prop recoil is developed, which truly aligns the human player's shooting method from the game client execution level, ensuring that the shooting humanization is established at the underlying execution logic. At the same time, in order to achieve the human player's gun-pressing shooting operation and the scattering effect of the simulated shooting prop, a shooting method of small circle inside random + large circle outside truncation is used. In FIG. 11, θ is defined as the inverse tangent function of 40 / 1500, and 40 / 1500 is a reference value given by the business team based on game experience. d represents the distance between the AI gun muzzle position and the target enemy position, and θ represents the maximum angle of the gun muzzle deviation when the normal human player shoots. Shooting beyond this angle will become meaningless shooting. The circle with a larger radius in FIG. 11 is the shooting confidence domain M corresponding to different distances d. Shooting within the region is considered normal shooting, and shooting outside the region is considered invalid shooting. If the AI model does not predict a shooting instruction, the neural network freely selects the horizontal and numerical gun muzzle deviation angle. If the model continuously predicts a shooting instruction, manual assistance is required to calculate the horizontal and vertical gun muzzle deviation angles to continuously pull the crosshair back onto the target enemy. The scattering effect of the human player's shooting should be simulated during each crosshair pull-back and deviation process. The gun-pressing effect of the human player's shooting should be simulated during each crosshair deviation and pull-back process. The specific method is as follows:

[0196] As shown in FIG. 11, the target enemy position is taken as the first confidence point, a first pathfinding radius r=(40 / 1500)xd is taken as the first pathfinding distance (i.e., the radius corresponding to the first pathfinding distance), a coarse-grained target shooting confidence region of the target enemy corresponding to the distance d is determined, shooting possibility trajectory detection is performed on each body part of the target enemy, the shooting possibility detection results of each part are sorted according to the hit priority of each part, the body part with the highest hit priority is taken as the aiming base point p0, p0 is taken as the center of a small circle, the radius of the small circle is r / x, the coefficient x can be adjusted according to the actual effect, a fine-grained part shooting confidence region N1 is determined, a point p1 is randomly selected as the aiming point in N1, if p1 is out of the coarse-grained target shooting confidence region M, the selection is truncated, to ensure that the shooting is an effective shooting within M, p1 is taken as the center of a small circle for the next shooting, a fine-grained shooting confidence region N2 is determined with a radius of r / 3, a point p2 is randomly selected as the aiming point in N2, if p2 is out of the coarse-grained target shooting confidence region M, the selection is truncated, to ensure that the shooting is an effective shooting within M, the process is repeated, the aiming position coordinates of each shot are obtained, the horizontal and vertical offset angles of the AI muzzle are calculated based on the aiming coordinates, when continuous prediction shooting is performed, the effect of human player shooting can be simulated, and the shooting performance under the first view of the playback lens AI can be restored to the effect of human player shooting, the shooting personification is greatly improved compared with the original behavior tree AI shooting method, and the shooting personification of the AI is enhanced from the bottom of the game client.

[0197] As shown in FIG. 14, the human target shooting hole distribution diagram and the AI target shooting hole distribution diagram are compared, it can be seen that the AI shot hole landing points are more dispersed than the human player shot hole landing points, the AI can simulate the shooting prop scattering effect and the gun pressing shooting operation when the human player shoots, in order to more truly restore the human shooting effect and exhibit the shooting prop scattering and gun pressing shooting ability like a human player, a continuous shooting target comparison experiment can be performed, and the AI shooting effect is shown in FIG. 15, in which the shot hole landing points of each two shots are relatively close, within the normal player operation deviation range, and also have some random disturbances, the shooting method of random selection in a small circle and truncation outside a large circle realizes the cycle of AI muzzle back and forth, offset and back and forth in a continuous shooting process, and it is just because of the continuity of the disturbance offset that the shooting prop scattering and gun pressing effect when the human player shoots is simulated, and the personified behavior of the AI shooting is indirectly increased.

[0198] Further, after the game starts and training is completed, the DS obtains runtime state data and divides it into different modal perception data according to Table 1. The environment information of different modes is independently extracted by different types of neural networks. In this way, the current game situation state can be more accurately and finely represented. FIG. 15 is a schematic diagram of another optional control method of a virtual object according to an embodiment of the present application. As shown in FIG. 15, the inference network (i.e., the multi-action head decision in FIG. 15) inputs the output hidden state encoding information of the LSTM (Long Short-Term Memory neural network) and global information. The global information mainly includes absolute or relative information about the first virtual object and the target virtual object. The target virtual object information is reacquired in a full perception mode, that is, after the target enemy loses the field of view, the real information of the target enemy is still acquired. This global input feature including “cheating” information can reduce the value estimation bias and variance of the value network, which is conducive to the convergence and stability of the strategy. This method is only used in the model training stage. In fact, only the strategy network is needed in the actual online deployment stage, and the value network is not needed. The self-recurrent embedding is introduced into the model design. The one-hot encoding (a coding method for converting a categorical variable into a binary vector) processing is performed on each layer of action head sampling action to form an embedding vector, which is then stacked with the input vector of the current layer to serve as the input of the next layer of action head. The influence is formed layer by layer to cause a cause-and-effect dependence between the front and back, so that the subsequent action head decision is more reasonable and natural. The one-hot encoding is performed on the action sampled by the first layer of action head, the embedding vector is formed through the fully connected layer, and then the LSTM output information is spliced to form a new embedding vector output to the second layer of action head. The action head sampling is continued to form an embedding vector through one-hot encoding, and then the input vector of the second layer is stacked to form the input of the third layer. This process is repeated until the last action head. The advantages of the target action inference model training structure design are as follows:

[0199] 1) The highly complex neural network fuses multi-modal input information such as lists, images, scalars, and Booleans; 2) The self-recurrent network action head design decouples the dependency between the structured action spaces of the game; 3) The interactive layered action mask accelerates the training efficiency and also improves the strategy personification; 4) Different types of input are activated separately to improve the accuracy of network input perception; 5) When the target enemy is invisible, the enemy non-full perception input is closer to the real world. The dynamic network structure adapts to this non-full perception input, thereby ensuring that the model can learn similar human playing operations;

[0200] In summary, the embodiment of the present application gives the humanization enhancement processing skills of training FPS game AI from three levels of environmental perception, network decision and bottom execution. First, the non-full perception method is used to align the human player's game perception method from the training data source, and the dynamic neural network structure is realized in the training process by combining the non-full processing logic of the neural network. Second, the dynamic multi-frequency decision mechanism not only can relieve the AI twitching and shaking phenomenon, enhance its humanization performance, but also greatly prunes the action space dimension, speeds up the training efficiency and reduces the training difficulty. Discarding the result-based shooting of behavior tree AI, the atomic shooting similar to the real human player with scattering and recoil is developed, and the large and small circle auxiliary shooting method is used, which truly realizes the scattering and gun pressure effect of the human player's shooting, and further enhances the behavior humanization performance from the bottom execution end. Based on large-scale reinforcement learning training and the above three humanization enhancement processing methods, the reinforcement learning AI feedback of the embodiment of the present application is as follows: 1) The recognition of complex obstacles is more flexible. Compared with the behavior tree version, the reinforcement learning AI can recognize the shelter position and go to hide without pre-constructing the shelter point, and can use various unconventional shelters for hiding from the battle; 2) The behavior is flexible and variable, and has different battle performances under the same fighting scene, which can improve the repeat playability; 3) The action space completely covers the player's action space, and completes all the actions that the player can perform, so as to go to places that the behavior tree cannot go and complete operations that the behavior tree cannot do; 4) The firing humanization and behavior humanization are done to the extent that the behavior tree cannot do, and the game elimination playback can also be aligned with the human player's perspective, which truly realizes the purpose of virtual objects.

[0201] Further, (1) because the users on the product line are more active and have large-scale landing landing data, early behavior pattern guidance can be attempted using human data. The initial training stage of the model will have a lot of repeated and invalid explorations, and it is even easy to fall into some unpredictable special states, which will have irreversible long-term effects on strategy learning. The strategy falls into a local optimum point and is difficult to jump out. Using supervised learning, the AI is initially at the starting point of the ability of a human player. Further, using reinforcement learning to further explore the strategy can improve the exploration efficiency and effectively avoid the initial strategy learning from falling into a local optimum. The subsequent personification training can also achieve twice the result with half the effort.(2) In the actual deployment process of game interaction, there will be communication delay problems in asynchronous online reasoning, including but not limited to communication delay between mobile phone mobile terminal and DS, communication delay between DS and online reasoning service. By isolating the DS and the online reasoning service, the DS has run for t time from sending a state request s to the online reasoning service -> the DS gets the action return a in this period of time t. At this time, the a action instruction that is obtained may not be applicable to the latest DS state. In short, the timeliness of the action a is reduced. If the communication delay between the mobile phone mobile terminal and the DS and network fluctuations and other factors are superimposed, some time-sensitive sub-action return packets will not work well when executed on the latest mobile terminal, resulting in a decline in model reasoning effect. This is the disadvantage of asynchronous online reasoning. In this regard, state delay and action delay can be introduced during model training, so that the neural network can adapt to the impact of the delay, thereby learning how to make decisions in the presence of delay. When the model is actually deployed online, the model can automatically adapt to the delay introduced by asynchronous reasoning, so that the model reasoning effect will not decrease too much. Or in the asynchronous reasoning way, after the DS receives the action return packet, the DS side further adjusts and adapts the action instruction returned back to the latest DS state to artificially offset the impact of the decrease in action timeliness caused by the network communication delay between the DS and the online reasoning service. Through these two methods, the impact of the decline in model reasoning effect caused by asynchronous reasoning can be alleviated to some extent.

[0202] Further, through the embodiments of the present application, large-scale distributed reinforcement learning can be used to speed up the collection and training of data for the purpose of strategy learning. Non-full perception processing of multi-modal environment input can meet the humanization requirements of the model input end. Frequency division decision is made for different action heads to make the action flexible and natural, which meets the humanization requirements of the model decision end. The large and small circles assist in shooting to achieve the gun muzzle shaking effect when the player shoots, which meets the humanization requirements of the game execution end. The comprehensive reward and punishment mechanism and the rich humanization indicators are used to evaluate the humanization of the model, which further speeds up the model adaptation and landing.

[0203] It can be understood that in the specific embodiments of the present application, related data such as user information is involved, and when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of countries and regions.

[0204] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.

[0205] According to another aspect of the embodiments of the present application, a virtual object control device for implementing the control method of the virtual object is also provided. As shown in FIG. 16, the device includes:

[0206] The display module 1602 is configured to determine a target virtual scene, wherein the target virtual scene represents a picture scene in which a target virtual object in a target application is located, and the target virtual object represents an object controlled by a target account logged in the target application.

[0207] The acquisition module 1604 is configured to acquire environmental perception data corresponding to a first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by the system, and part of the data associated with the target virtual object in the environmental perception data is determined by whether the target virtual object is visible to the first virtual object. In the case that the target virtual object is not visible to the first virtual object, the part of the data is set to be masked by a preset parameter.

[0208] The control module 1606 is configured to determine action package data according to the environmental perception data, and control the first virtual object to perform a target action according to the action package data, wherein the action package data includes different types of sub-actions predicted based on the environmental perception data, and the target action is determined by combining different types of sub-actions.

[0209] As an optional solution, the device is configured to determine the action package data according to the environment perception data, and control the first virtual object to perform the target action according to the action package data in the following manner: predict the first sub-action according to the environment perception data to obtain a first set of sub-action confidences, wherein one sub-action confidence in the first set of sub-action confidences corresponds to an action parameter of the first sub-action; predict the second sub-action according to the environment perception data to obtain a second set of sub-action confidences, wherein one sub-action confidence in the second set of sub-action confidences corresponds to an action parameter of the second sub-action, and the first sub-action and the second sub-action represent mutually independent sub-actions; determine a sub-action confidence in the first set of sub-action confidences that meets a preset condition as a target first sub-action confidence, and determine an action parameter corresponding to the target first sub-action confidence as first sub-action data corresponding to the first sub-action; determine a sub-action confidence in the second set of sub-action confidences that meets the preset condition as a target second sub-action confidence, and determine an action parameter corresponding to the target second sub-action confidence as second sub-action data corresponding to the second sub-action; and generate the action package data according to the first sub-action data and the second sub-action data, and control the first virtual object to perform the target action according to the action package data.

[0210] As an optional solution, the device is configured to predict the first sub-action according to the environment perception data to obtain a first set of sub-action confidences, and predict the second sub-action according to the environment perception data to obtain a second set of sub-action confidences in the following manner: predict the first sub-action according to the environment perception data at a first frequency to obtain the first set of sub-action confidences, wherein the first frequency represents a preset execution frequency of the first sub-action; and predict the second sub-action according to the environment perception data at a second frequency to obtain the second set of sub-action confidences, wherein the second frequency represents a preset execution frequency of the second sub-action, and the first frequency and the second frequency are different.

[0211] As an optional solution, the device is further configured to: in a case where the first sub-action represents an attack operation, predict the first sub-action according to the environment perception data at a first frequency to obtain the first set of sub-action confidence; in a case where the second sub-action represents an operation other than the attack operation, predict the second sub-action according to the environment perception data at a second frequency to obtain the second set of sub-action confidence, wherein the first frequency is greater than the second frequency; or in a case where the first sub-action represents a pathfinding operation, predict the first sub-action according to the environment perception data at a first frequency to obtain the first set of sub-action confidence; in a case where the second sub-action represents an operation other than the pathfinding operation, predict the second sub-action according to the environment perception data at a second frequency to obtain the second set of sub-action confidence, wherein the first frequency is less than the second frequency.

[0212] As an optional solution, the device is configured to generate the action package data according to the first sub-action data and the second sub-action data by: obtaining a predetermined invalid action combination; in a case where an action combination composed of the first sub-action data and the second sub-action data does not belong to the invalid action combination, generating the action package data according to the first sub-action data and the second sub-action data, and controlling the first virtual object to perform the target action according to the action package data.

[0213] As an optional solution, the device is configured to obtain the environment perception data corresponding to the first virtual object in the target virtual scene by: in a case where the target virtual object is visible to the first virtual object, obtaining full attribute information corresponding to the first virtual object, the partial data, relative position information between the target virtual object and the first virtual object, and voiceprint perception information of the first virtual object; determining the full attribute information, the partial data, the relative position information, and the voiceprint perception information as the environment perception data together; in a case where the target virtual object is not visible to the first virtual object, obtaining full attribute information corresponding to the first virtual object, partial data shielded by the preset parameter, relative position information between the target virtual object and the first virtual object, and voiceprint perception information of the first virtual object; and determining the full attribute information, the relative position information, and the voiceprint perception information as the environment perception data together.

[0214] As an optional solution, the device is further configured to determine the action package data according to the environment perception data by: sending the environment perception data to a pre-trained target action inference model through the target application, wherein the target action inference model is deployed on an inference server, and the target application is deployed on an application server; receiving a return data packet, wherein the return data packet is the action package data returned by the target action inference model in response to the environment perception data, and a sending speed of the environment perception data to the target action inference model is less than or equal to a return speed of the return data packet returned by the target action inference model; and controlling the first virtual object to perform a target action according to the action package data.

[0215] As an optional solution, the device is further configured to: in a case where the sending speed exceeds the return speed, reduce a sending frequency of the environment perception data; or in a case where a number of accumulated action package data of the target action inference model exceeds a preset number threshold, control the first virtual object to perform a behavior tree action through a behavior tree logic, wherein the accumulated action package data indicates action package data that is not implemented to control the first virtual object to perform.

[0216] As an optional solution, the device is further used for: before the device determines the action package data according to the environment perception data and controls the first virtual object to perform the target action according to the action package data, the device performs distributed training on an initial action reasoning model to obtain the target action reasoning model, wherein, in each game construction process of the distributed training, the following steps are performed: dividing the target virtual scene into a plurality of key areas, wherein each key area in the plurality of key areas is pre-configured with a division probability; sampling a target key area from the plurality of key areas according to the division probability; sampling a first sample position of a first sample virtual object in a first area with the target key area as a center and a first path distance as a radius of the first area, wherein the first path distance represents a maximum path distance pre-configured for the first sample virtual object; sampling a second sample position of a second sample virtual object in a circular ring with the first sample position as a center, a second path distance as an outer radius, and a third path distance as an inner radius, wherein the second path distance represents a maximum path distance pre-configured for the second sample virtual object, and the third path distance represents a minimum path distance pre-configured for the second sample virtual object; in a case where there is a virtual obstacle between the first sample position and the second sample position, generating the first sample virtual object at the first sample position and generating the second sample virtual object at the second sample position; and based on first sample environment perception data corresponding to the first sample virtual object and second sample environment perception data corresponding to the second sample virtual object, controlling the first sample virtual object and the second sample virtual object to play a game until the game of the first sample virtual object and the second sample virtual object ends.

[0217] As an optional solution, the device is configured to obtain the target action reasoning model by performing distributed training on the initial action reasoning model in the following manner: performing non-full-volume processing on the first sample environment perception data to obtain first sample data, and performing non-full-volume processing on the second sample environment perception data to obtain second sample data; performing feature extraction on the first sample data to obtain first sample features in a case where the first sample virtual object is visible to the second sample virtual object, wherein the first sample features include partial data of the second sample virtual object; performing feature extraction on the second sample data to obtain second sample features in a case where the second sample virtual object is not visible to the first sample virtual object, wherein the second sample features include partial data of the first sample virtual object after being masked by the preset parameters; inputting the first sample features and the second sample features into the initial action reasoning model respectively, determining preset index values corresponding to the first sample virtual object and the second sample virtual object respectively, and updating the initial action reasoning model according to the preset index values to determine the target action reasoning model, wherein the output of a feature network layer associated with the partial data is uniformly masked as a target value during the updating of the initial action reasoning model.

[0218] As an optional solution, the device is configured to obtain the target action reasoning model by performing distributed training on the initial action reasoning model in the following manner: performing non-full-volume processing on the first sample environment perception data to obtain first sample data, and performing non-full-volume processing on the second sample environment perception data to obtain second sample data; performing feature extraction on the first sample data to obtain first sample features in a case where the first sample virtual object is visible to the second sample virtual object, wherein the first sample features include partial data of the second sample virtual object; performing feature extraction on the second sample data to obtain second sample features in a case where the second sample virtual object is not visible to the first sample virtual object, wherein the second sample features include partial data of the first sample virtual object after being masked by the preset parameters; inputting the first sample features and the second sample features into the initial action reasoning model respectively, determining preset index values corresponding to the first sample virtual object and the second sample virtual object respectively, and updating the initial action reasoning model according to the preset index values to determine the target action reasoning model, wherein the output of a feature network layer associated with the partial data is uniformly masked as a target value during the updating of the initial action reasoning model.

[0219] As an optional solution, the apparatus is further configured to determine the action package data according to the environmental perception data by: when the environmental perception data acquired at the i th frame indicates that the first virtual object needs to be controlled to first launch a virtual prop within a preset time length, acquiring a maximum angle of aiming deviation preset for the first virtual object, where i is a positive integer; determining a first confidence region according to a distance between the first virtual object and the target virtual object and the maximum angle of aiming deviation, where a radius of the first confidence region is r, and r is greater than 0; sampling a first confidence point from the first confidence region, and determining a second confidence region with the first confidence point as a center and r / x as a radius; and sampling a second confidence point from the second confidence region, where the second confidence point represents a reticle of the first virtual object corresponding to the environmental perception data acquired at the i th frame.

[0220] As an optional solution, the apparatus is further configured to: after the second confidence point is sampled from the second confidence region, when the environmental perception data acquired at the i+m th frame indicates that the first virtual object needs to be controlled to launch a virtual prop for the k th time, determining a third confidence region with a third confidence point determined for the k-1 th time as a center and r / x as a radius, where m is a positive integer, m is less than the preset time length, k is a positive integer greater than or equal to 2, and when k=2, the third confidence point is the second confidence point; and sampling a target confidence point from the third confidence region, where the target confidence point represents a reticle of the first virtual object corresponding to the environmental perception data acquired at the i+m th frame.

[0221] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.

[0222] As to the apparatus in the above embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and will not be described in detail here.

[0223] According to an aspect of the present application, a computer program product is provided, which includes a computer program.

[0224] The serial numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0225] FIG. 17 schematically shows a computer system structural diagram of an electronic device for implementing embodiments of the present application.

[0226] It should be noted that the computer system 1700 of the electronic device shown in FIG. 17 is only an example, and should not bring any limitation to the functions and use range of embodiments of the present application.

[0227] As shown in FIG. 17, the computer system 1700 includes a central processing unit 1701 (CPU), which can perform various appropriate actions and processes according to programs stored in a read-only memory 1702 (ROM) or loaded from a storage portion 1708 to a random access memory 1703 (RAM). In the random access memory 1703, various programs and data required for system operation are also stored. The central processing unit 1701, the read-only memory 1702, and the random access memory 1703 are connected to each other through a bus 1704. An input / output interface 1705 (I / O interface) is also connected to the bus 1704.

[0228] The following components are connected to the input / output interface 1705: an input portion 1706 including a keyboard, a mouse, and the like; an output portion 1707 including a cathode ray tube (CRT), a liquid crystal display (LCD), and the like, and a speaker, and the like; a storage portion 1708 including a hard disk, and the like; and a communication portion 1709 including a network interface card such as a local area network card, a modem, and the like. The communication portion 1709 performs communication processing via a network such as the Internet. A drive 1710 is also connected to the input / output interface 1705 as necessary. A removable media 1711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is mounted on the drive 1710 as necessary, so that a computer program read therefrom is installed in the storage portion 1708 as necessary.

[0229] In particular, according to embodiments of the present application, the processes described in each of the method flowcharts can be implemented as a computer software program. For example, embodiments of the present application include a computer program product for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication portion 1709, and / or installed from the removable media 1711. When the computer program is executed by the central processing unit 1701, various functions defined in the system of the present application are performed.

[0230] In such an embodiment, the computer program can be downloaded and installed from the network by the communication section 1709, and / or installed from the detachable medium 1711. When the computer program is executed by the central processing unit 1701, various functions provided by the embodiments of the present application are performed.

[0231] According to still another aspect of the embodiments of the present application, an electronic device for implementing the control method of the virtual object is also provided. The electronic device can be the terminal device or the server shown in FIG. 1. The embodiments of the present application are described taking the electronic device as the terminal device as an example. As shown in FIG. 18, the electronic device includes a memory 1802 and a processor 1804. The memory 1802 stores a computer program and transmits the computer program to the processor 1804. The processor 1804 is configured to execute the steps in any of the method embodiments described above by the computer program.

[0232] Optionally, in the embodiment, the electronic device can be located in at least one of the network devices in the computer network.

[0233] Optionally, in the embodiment, the processor can be configured to execute the method in the embodiments of the present application by the computer program.

[0234] Optionally, those skilled in the art can understand that the structure shown in FIG. 18 is only schematic, and FIG. 18 does not limit the structure of the electronic device. For example, the electronic device can further include more or less components (such as a network interface, etc.) than those shown in FIG. 18, or have a different configuration from that shown in FIG. 18.

[0235] The memory 1802 can be used to store software programs and modules, such as program instructions / modules corresponding to the virtual object control method and device in the embodiments of the present application. The processor 1804 executes various functions and data processing by running the software programs and modules stored in the memory 1802, that is, implements the virtual object control method described above. The memory 1802 can include a high-speed random access memory, and can also include a non-volatile memory such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 1802 can further include a memory remotely arranged with respect to the processor 1804, which can be connected to the terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. Specifically, the memory 1802 can be used to store environmental perception data, action package data, and the like, but is not limited to this. As an example, as shown in FIG. 18, the memory 1802 can include but is not limited to the display module 1602, the acquisition module 1604, and the control module 1606 in the virtual object control device described above. In addition, other module units in the virtual object control device described above can also be included, but are not limited to this, and will not be described in detail in this example.

[0236] Optionally, the transmission device 1806 described above is used to receive or send data via a network. Specific examples of the above-mentioned network can include wired networks and wireless networks. In one example, the transmission device 1806 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable to communicate with the Internet or a local area network. In one example, the transmission device 1806 is a radio frequency (Radio Frequency, RF) module, which is used to communicate with the Internet in a wireless manner.

[0237] In addition, the electronic device described above further includes a display 1808 for displaying the environmental perception data and the action package data, and a connection bus 1810 for connecting various module components in the electronic device.

[0238] In other embodiments, the terminal device or the server described above can be a node in a distributed system, where the distributed system can be a blockchain system, which can be a distributed system formed by the plurality of nodes communicating through a network. Wherein, the nodes can form a peer-to-peer network, and any form of computing device, such as a server, a terminal, and the like, can become a node in the blockchain system by joining the peer-to-peer network.

[0239] According to an aspect of the present application, a computer readable storage medium is provided, and a processor of an electronic device reads the computer instructions from the computer readable storage medium. The processor executes the computer instructions, so that the electronic device executes the virtual object control method provided in various optional implementations of the virtual object control aspect.

[0240] Optionally, in the embodiment, the computer readable storage medium can be configured to store the computer instructions for executing the method in the embodiments of the present application.

[0241] Optionally, in the embodiment, a person skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by programs instructing the hardware related to the terminal device, and the programs can be stored in a computer readable storage medium, and the storage medium can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0242] The serial numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.

[0243] The integrated units in the above embodiments, if realized in the form of software function units and sold or used as independent products, can be stored in the computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or the whole or part of the technical solutions can be embodied in the form of software products, and the computer software products are stored in the storage medium, including a plurality of instructions for causing one or more electronic devices to execute all or part of the steps of the methods described in the embodiments of the present application.

[0244] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0245] In the several embodiments provided by the present application, it should be understood that the disclosed application programs can be implemented in other ways. Of course, the above described device embodiments are only schematic. The division of the units is only a logical function division. There can be another division manner in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, accessors or buses, and can be electrical, mechanical or in other forms.

[0246] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment scheme.

[0247] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0248] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled persons in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which should be considered as the protection scope of the present application.

Claims

1. A method for controlling a virtual object, the method being performed by an electronic device, the method comprising: determining a target virtual scene, wherein the target virtual scene represents a picture scene in which a target virtual object in a target application is located, and the target virtual object represents an object controlled by a target account logged into the target application; obtaining environmental perception data corresponding to a first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by a system, and a part of data associated with the target virtual object in the environmental perception data is determined according to whether the target virtual object is visible to the first virtual object, and the part of data is set to be masked by a preset parameter when the target virtual object is not visible to the first virtual object; determining action package data according to the environmental perception data, and controlling the first virtual object to perform a target action according to the action package data, wherein the action package data comprises different types of sub-actions predicted based on the environmental perception data, and the target action is determined by combining the different types of sub-actions.

2. The method of claim 1, wherein the determining the action package data according to the environmental perception data, and controlling the first virtual object to perform the target action according to the action package data comprises: predicting a first sub-action according to the environmental perception data to obtain a first set of sub-action confidences, wherein one sub-action confidence in the first set of sub-action confidences corresponds to an action parameter of the first sub-action; predicting a second sub-action according to the environmental perception data to obtain a second set of sub-action confidences, wherein one sub-action confidence in the second set of sub-action confidences corresponds to an action parameter of the second sub-action, and the first sub-action and the second sub-action represent mutually independent sub-actions; determining a sub-action confidence in the first set of sub-action confidences that meets a preset condition as a target first sub-action confidence, and determining an action parameter corresponding to the target first sub-action confidence as first sub-action data corresponding to the first sub-action; determining a sub-action confidence in the second set of sub-action confidences that meets the preset condition as a target second sub-action confidence, and determining an action parameter corresponding to the target second sub-action confidence as second sub-action data corresponding to the second sub-action; and generating the action package data according to the first sub-action data and the second sub-action data, and controlling the first virtual object to perform the target action according to the action package data.

3. The method of claim 2, wherein the predicting the first sub-action according to the environmental perception data to obtain the first set of sub-action confidences comprises: predicting the first sub-action according to the environmental perception data at a first frequency to obtain the first set of sub-action confidences, wherein the first frequency represents a preset execution frequency of the first sub-action; and the predicting the second sub-action according to the environmental perception data to obtain the second set of sub-action confidences comprises: predicting the second sub-action according to the environmental perception data at a second frequency to obtain the second set of sub-action confidences, wherein the second frequency represents a preset execution frequency of the second sub-action. ​ ​ ​ ​ ​ ​ ​ ​ ​ ​ According to the environmental perception data, the second sub-action is predicted at a second frequency, to obtain the second set of sub-action confidence, wherein the second frequency represents a preset execution frequency of the second sub-action, and the first frequency and the second frequency are different.

4. The method of claim 3, wherein the first set of sub-action confidence is obtained by predicting the first sub-action at a first frequency according to the environmental perception data, and the second set of sub-action confidence is obtained by predicting a second sub-action according to the environmental perception data, and the predicting the first sub-action at the first frequency according to the environmental perception data to obtain the first set of sub-action confidence, and the predicting the second sub-action at the second frequency according to the environmental perception data to obtain the second set of sub-action confidence, comprise: in the case that the first sub-action represents an attack operation, the first set of sub-action confidence is obtained by predicting the first sub-action at the first frequency according to the environmental perception data; in the case that the second sub-action represents a sub-action other than the attack operation, the second set of sub-action confidence is obtained by predicting the second sub-action at the second frequency according to the environmental perception data, wherein the first frequency is greater than the second frequency; or, in the case that the first sub-action represents a pathfinding operation, the first set of sub-action confidence is obtained by predicting the first sub-action at the first frequency according to the environmental perception data; in the case that the second sub-action represents a sub-action other than the pathfinding operation, the second set of sub-action confidence is obtained by predicting the second sub-action at the second frequency according to the environmental perception data, wherein the first frequency is less than the second frequency.

5. The method of any one of claims 2-4, wherein the action package data is generated according to the first sub-action data and the second sub-action data, and the first virtual object is controlled to perform the target action according to the action package data, comprises: obtaining a predetermined invalid action combination; in the case that the action combination composed of the first sub-action data and the second sub-action data does not belong to the invalid action combination, the action package data is generated according to the first sub-action data and the second sub-action data, and the first virtual object is controlled to perform the target action according to the action package data.

6. The method of any one of claims 1-5, wherein the environmental perception data corresponding to the first virtual object in the target virtual scene is obtained, comprises: in the case that the target virtual object is visible to the first virtual object, obtaining full attribute information corresponding to the first virtual object, the partial data, relative position information between the target virtual object and the first virtual object, and voiceprint perception information of the first virtual object; and determining the full attribute information, the partial data, the relative position information, and the voiceprint perception information as the environmental perception data together.

7. The method of any one of claims 1-6, wherein the target virtual object is controlled to perform the target action according to the action package data, comprises: in the case that the target virtual object is visible to the first virtual object, the target virtual object is controlled to perform the target action according to the action package data; in the case that the target virtual object is not visible to the first virtual object, the target virtual object is controlled to perform the target action according to the action package data and the environmental perception data. In a case where the target virtual object is invisible to the first virtual object, obtaining full-quantity attribute information corresponding to the first virtual object, part of data shielded by the preset parameter, relative position information between the target virtual object and the first virtual object, and voiceprint perception information of the first virtual object; and determining the full-quantity attribute information, the relative position information, and the voiceprint perception information as the environment perception data.

7. The method of any one of claims 1-6, wherein determining the action package data according to the environment perception data and controlling the first virtual object to perform a target action according to the action package data comprises: sending, by the target application, the environment perception data to a pre-trained target action inference model, wherein the target action inference model is deployed on an inference server, and the target application is deployed on an application server; receiving a return data package, wherein the return data package is the action package data returned by the target action inference model in response to the environment perception data, wherein a sending speed of the environment perception data to the target action inference model is less than or equal to a return speed of the return data package returned by the target action inference model, and controlling the first virtual object to perform a target action according to the action package data.

8. The method of claim 7, further comprising: in a case where the sending speed exceeds the return speed, reducing a sending frequency of the environment perception data; or in a case where a number of accumulated action package data of the target action inference model exceeds a preset number threshold, controlling the first virtual object to perform a behavior tree action by a behavior tree logic, wherein the accumulated action package data represents action package data sent from the inference server to the application server.

9. The method of claim 7 or 8, wherein before determining the action package data according to the environment perception data and controlling the first virtual object to perform a target action according to the action package data, the method further comprises: performing distributed training on an initial action inference model to obtain the target action inference model, wherein each game session of the distributed training is constructed as follows: dividing the target virtual scene into a plurality of key regions, wherein each key region in the plurality of key regions is pre-provided with a division probability; sampling a target key region from the plurality of key regions according to the division probability; sampling a first sample position of a first sample virtual object in a first region centered at a target position in the target key region and having a first pathfinding distance as a radius of the first region, wherein the first pathfinding distance represents a maximum pathfinding distance pre-provided for the first sample virtual object; and sampling a second sample position of a second sample virtual object in a second region centered at a target position in the target key region and having a second pathfinding distance as a radius of the second region, wherein the second pathfinding distance represents a maximum pathfinding distance pre-provided for the second sample virtual object. sampling a second sample virtual object at a second sample position, the second sample position being centered on the first sample position, a second pathfinding distance being an outer radius of a circle, and a third pathfinding distance being an inner radius of the circle, the second pathfinding distance representing a maximum pathfinding distance pre-set for the second sample virtual object, and the third pathfinding distance representing a minimum pathfinding distance pre-set for the second sample virtual object; in a case where there is a virtual obstacle between the first sample position and the second sample position, generating the first sample virtual object at the first sample position and generating the second sample virtual object at the second sample position, and controlling the first sample virtual object and the second sample virtual object to play a game based on first sample environment perception data corresponding to the first sample virtual object and second sample environment perception data corresponding to the second sample virtual object until the game of the first sample virtual object and the second sample virtual object ends.

10. The method of claim 9, wherein the distributed training of the initial action inference model to obtain the target action inference model comprises: performing non-full-volume processing on the first sample environment perception data to obtain first sample data, and performing non-full-volume processing on the second sample environment perception data to obtain second sample data; in a case where the first sample virtual object is visible to the second sample virtual object, performing feature extraction on the first sample data to obtain first sample features, wherein the first sample features include partial data of the second sample virtual object; in a case where the second sample virtual object is not visible to the first sample virtual object, performing feature extraction on the second sample data to obtain second sample features, wherein the second sample features include partial data of the first sample virtual object after being masked by the preset parameter; inputting the first sample features and the second sample features into the initial action inference model to determine preset index values corresponding to the first sample virtual object and the second sample virtual object, respectively, and updating the initial action inference model according to the preset index values to determine the target action inference model, wherein, during the updating of the initial action inference model, an output of a feature network layer associated with the partial data is uniformly masked as a target value.

11. The method of any one of claims 1-10, The acquiring the environmental perception data corresponding to the first virtual object in the target virtual scene comprises: when the service of the target virtual scene is performed to an i-th frame, obtaining first environment perception data corresponding to the i-th frame; when the service of the target virtual scene is performed to an i+n-th frame, obtaining second environment perception data corresponding to the i+n-th frame, wherein the service performs a preset service logic from the i-th frame to the i+n-th frame, and i and n are positive integers. The determining the action package data according to the environment perception data and the controlling the first virtual object to perform a target action according to the action package data comprises: determining first action package data according to the first environment perception data, and controlling the first virtual object to perform a first target action according to the first action package data at the i+jth frame; determining second action package data according to the second environment perception data, and controlling the first virtual object to perform a second target action according to the second action package data at the i+n+jth frame, wherein j is a non-negative integer.

12. The method of any one of claims 1-11, wherein the determining the action package data according to the environment perception data and the controlling the first virtual object to perform a target action according to the action package data comprises: when the environment perception data obtained at the ith frame indicates that the first virtual object needs to be controlled to first launch a virtual prop within a preset time length, obtaining a maximum angle of aiming deviation preset for the first virtual object, wherein i is a positive integer; determining a first confidence region according to a distance between the first virtual object and the target virtual object and the maximum angle of aiming deviation, wherein a radius of the first confidence region is r, and r is greater than 0; sampling a first confidence point from the first confidence region, and determining a second confidence region with the first confidence point as a center and r / x as a radius; sampling a second confidence point from the second confidence region, wherein the second confidence point represents a target of the first virtual object corresponding to the environment perception data obtained at the ith frame.

13. The method of claim 12, after the sampling the second confidence point from the second confidence region, the method further comprises: when the environment perception data obtained at the i+mth frame indicates that the first virtual object needs to be controlled to launch a virtual prop for the kth time, determining a third confidence region with a third confidence point determined for the (k-1)th time as a center and r / x as a radius, wherein m is a positive integer and m is less than the preset time length, k is a positive integer greater than or equal to 2, and when k=2, the third confidence point is the second confidence point; sampling a target confidence point from the third confidence region, wherein the target confidence point represents a target of the first virtual object corresponding to the environment perception data obtained at the i+mth frame.

14. A device for controlling a virtual object, the device comprising: a display module configured to determine a target virtual scene, wherein the target virtual scene represents a picture scene in which a target virtual object is located in a target application, and the target virtual object represents an object controlled by a target account logged into the target application. An acquisition module is configured to acquire environmental perception data corresponding to a first virtual object in a target virtual scene, wherein the first virtual object represents an object controlled by a system, and a part of the environmental perception data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object, and the part of the environmental perception data is set to be occluded by a preset parameter when the target virtual object is not visible to the first virtual object. A control module is configured to determine action package data according to the environmental perception data, and control the first virtual object to perform a target action according to the action package data, wherein the action package data includes different types of sub-actions predicted based on the environmental perception data, and the target action is determined by combination of the different types of sub-actions.

15. A computer readable storage medium, the computer readable storage medium comprising a stored computer program, wherein, The computer program is configured to execute the method in any one of claims 1 to 13.

16. A computer program product comprising a computer program which, when executed by a processor, implements the steps of the method in any one of claims 1 to 13.

17. An electronic device comprising a memory and a processor, the memory having stored therein a computer program and transmitting the computer program to the processor, the processor being configured to execute the method in any one of claims 1 to 13 by the computer program.

Citation Information

Patent Citations

  • Virtual object operation control method and device, electronic device and storage medium

    CN109445662A

  • Object control method and device, storage medium and electronic device

    CN109499068A

  • Virtual object control method and device, computer readable storage medium and equipment

    CN111318017A

  • Interaction method and device in game, computer equipment and storage medium

    CN113952723A

  • Game data processing method, device and equipment and computer readable storage medium

    CN114288663A

Cited By

  • Method and system for detecting collision interference in airborne product manufacturing production line

    CN121413278A

  • Micro-service software optimization method based on reinforcement learning

    CN121743170A

  • Robot generalization control method, system and equipment

    CN122210628A