Virtual object control method and device, storage medium and electronic equipment

By acquiring and processing environmental perception data of the target virtual scene, occluding invisible data, predicting and combining sub-actions to generate action package data, the problem of the single representation of virtual objects is solved, achieving higher control accuracy and flexibility, and improving the game experience.

CN120860591APending Publication Date: 2025-10-31TENCENT DIGITAL TIANJIN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410544506.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

The limited representation of virtual objects in existing technologies results in low control accuracy and a lack of flexibility, restricting their application in diverse and complex tasks.

Method used

By acquiring environmental perception data from the target virtual scene, using preset parameters to mask invisible data, combining environmental perception data to predict different types of sub-actions, generating action package data, controlling virtual objects to execute target actions, and using a distributed training model to improve action accuracy.

Benefits of technology

It improves the accuracy and flexibility of virtual object movements, provides a richer gaming experience, reduces data transmission and computational overhead, and enhances the anthropomorphic effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120860591A_ABST
    Figure CN120860591A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual object control method and device, a storage medium and electronic equipment. The method comprises the following steps: displaying a target virtual scene; environment perception data corresponding to a first virtual object in the target virtual scene is obtained, the first virtual object represents an object controlled by a system, and part of data associated with the target virtual object in the environment perception data is determined by whether the target virtual object is visible to the first virtual object or not; action packet data is determined according to the environment perception data, the first virtual object is controlled to execute a target action according to the action packet data, the action packet data comprises different types of sub-actions obtained through prediction based on the environment perception data, and the target action is determined by combining the different types of sub-actions. The technical problem of low control accuracy of the virtual object caused by single performance of the virtual object is solved, and the embodiment of the invention can be applied to various scenes such as cloud technology, artificial intelligence, smart traffic and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computers, and more specifically, to a method and apparatus for controlling virtual objects, a storage medium, and an electronic device. Background Technology

[0002] Currently, in related technologies, virtual objects are often controlled by pre-set rules. That is, several rule conditions are set in advance, and the state of the virtual object is mechanically determined based on the rule conditions, thereby causing the virtual object to perform corresponding actions. In short, virtual objects often exhibit a single performance and lack flexibility when processing tasks, which limits their application in diverse and complex tasks. Therefore, this leads to the technical problem of low control accuracy of virtual objects due to their single performance.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a method and apparatus for controlling virtual objects, a storage medium and an electronic device, to at least solve the technical problem that the control accuracy of virtual objects is low due to the single representation of virtual objects.

[0005] According to one aspect of the embodiments of this application, a method for controlling a virtual object is provided, comprising: displaying a target virtual scene, wherein the target virtual scene represents a game scene in which a target virtual object in a target application is located, and the target virtual object represents an object controlled by a target account logged into the target application; acquiring environmental awareness data corresponding to a first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by a system, and a portion of the environmental awareness data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object, and if the target virtual object is not visible to the first virtual object, the portion of data is set to be occluded by a preset parameter; determining action package data based on the environmental awareness data, and controlling the first virtual object to perform a target action according to the action package data, wherein the action package data includes different types of sub-actions predicted based on the environmental awareness data, and the target action is determined by combining the different types of sub-actions.

[0006] According to another aspect of the embodiments of this application, a control device for a virtual object is also provided, comprising: a display module for displaying a target virtual scene, wherein the target virtual scene represents a game scene in which a target virtual object in a target application is located, and the target virtual object represents an object controlled by a target account logged into the target application; an acquisition module for acquiring environmental perception data corresponding to a first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by the system, and a portion of the environmental perception data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object, and if the target virtual object is not visible to the first virtual object, the portion of the data is set to be occluded by a preset parameter; and a control module for determining action package data based on the environmental perception data and controlling the first virtual object to perform a target action according to the action package data, wherein the action package data includes different types of sub-actions predicted based on the environmental perception data, and the target action is determined by combining the different types of sub-actions.

[0007] Optionally, the device is configured to determine action package data based on the environmental perception data in the following manner, and control the first virtual object to perform a target action according to the action package data: predicting a first sub-action based on the environmental perception data to obtain a first set of sub-action confidence scores, wherein one sub-action confidence score in the first set of sub-action confidence scores corresponds to an action parameter of the first sub-action; predicting a second sub-action based on the environmental perception data to obtain a second set of sub-action confidence scores, wherein one sub-action confidence score in the second set of sub-action confidence scores corresponds to an action parameter of the second sub-action, and the first sub-action and the second sub-action represent sub-actions that do not affect each other; determining the action parameter corresponding to the sub-action confidence score in the first set of sub-action confidence scores that meets a preset condition as the first sub-action data corresponding to the first sub-action; determining the action parameter corresponding to the sub-action confidence score in the second set of sub-action confidence scores that meets the preset condition as the second sub-action data corresponding to the second sub-action; generating the action package data based on the first sub-action data and the second sub-action data, and controlling the first virtual object to perform the target action according to the action package data.

[0008] Optionally, the device is configured to determine action package data based on the environmental perception data in the following manner, and control the first virtual object to perform a target action according to the action package data: predicting the first sub-action according to the environmental perception data at a first frequency to obtain the confidence level of the first set of sub-actions, wherein the first frequency represents a preset execution frequency of the first sub-action; predicting the second sub-action according to the environmental perception data at a second frequency to obtain the confidence level of the second set of sub-actions, wherein the second frequency represents a preset execution frequency of the second sub-action, and the first frequency and the second frequency are different.

[0009] Optionally, the apparatus is further configured to: when the first sub-action represents an attack operation, predict the first sub-action according to the environmental perception data at a first frequency to obtain a first set of sub-action confidence scores; when the second sub-action represents a sub-action other than the attack operation, predict the second sub-action according to the environmental perception data at a second frequency to obtain a second set of sub-action confidence scores, wherein the first frequency is greater than the second frequency; when the first sub-action represents a pathfinding operation, predict the first sub-action according to the environmental perception data at a first frequency to obtain a first set of sub-action confidence scores; when the second sub-action represents a sub-action other than the pathfinding operation, predict the second sub-action according to the environmental perception data at a second frequency to obtain a second set of sub-action confidence scores, wherein the first frequency is less than the second frequency.

[0010] Optionally, the device is configured to generate the action package data based on the first sub-action data and the second sub-action data in the following manner, and control the first virtual object to perform the target action according to the action package data: obtaining a predetermined invalid action combination; if the action combination composed of the first sub-action data and the second sub-action data does not belong to the invalid action combination, generating the action package data based on the first sub-action data and the second sub-action data, and controlling the first virtual object to perform the target action according to the action package data; if the action combination composed of the first sub-action data and the second sub-action data belongs to the invalid action combination, deleting the first sub-action data and the second sub-action data from the action package data.

[0011] Optionally, the device is configured to acquire environmental perception data corresponding to a first virtual object in the target virtual scene in the following manner: when the target virtual object is visible to the first virtual object, acquire the full attribute information of the first virtual object, the partial data, the relative position information between the target virtual object and the first virtual object, and the voiceprint perception information of the first virtual object; and jointly determine the full attribute information, the partial data, the relative position information, and the voiceprint perception information as the environmental perception data; when the target virtual object is not visible to the first virtual object, acquire the full attribute information of the first virtual object, the relative position information between the target virtual object and the first virtual object, and the voiceprint perception information of the first virtual object; and jointly determine the full attribute information, the relative position information, and the voiceprint perception information as the environmental perception data, wherein the partial data is set to be occluded according to the preset parameters.

[0012] Optionally, the device is configured to determine action package data based on the environmental perception data and control the first virtual object to perform a target action according to the action package data in the following manner: sending the environmental perception data to a pre-trained target action inference model through the target application, wherein the target action inference model is deployed on an inference server and the target application is deployed on an application server; receiving the action package data returned by the target action inference model in response to the environmental perception data, wherein the packet sending speed of the environmental perception data to the target action inference model is less than or equal to the packet return speed of the target action inference model; and controlling the first virtual object to perform the target action according to the action package data.

[0013] Optionally, the device is further configured to: adjust the transmission frequency of the environmental perception data to reduce the transmission speed when the packet transmission speed exceeds the packet return speed; and control the first virtual object to execute behavior tree actions through behavior tree logic when the number of accumulated action packet data of the target action inference model exceeds a preset number threshold, wherein the accumulated action packet data represents action packet data for which the control of the first virtual object to execute according to the accumulated action packet data has not been implemented.

[0014] Optionally, the device is further configured to: before determining the action package data based on the environmental perception data and controlling the first virtual object to execute the target action according to the action package data, perform distributed training on the initial action inference model to obtain the target action inference model, wherein the game construction process of each distributed training is as follows: dividing the target virtual scene into multiple key regions, wherein each key region in the multiple key regions has a pre-set division probability; sampling the target key region from the multiple key regions according to the division probability; sampling the first sample position of the first sample virtual object with the target key region as the center and a first pathfinding distance as the radius, wherein the first pathfinding radius represents the maximum pathfinding radius pre-set for the first sample virtual object; and sampling the first sample position with the first sample position as the center and a second pathfinding distance as the outer radius. The inner circle radius is used as the third pathfinding distance. The second sample position of the second sample virtual object is obtained by sampling within the circular ring. The second pathfinding radius represents the maximum pathfinding radius preset for the second sample virtual object, and the third pathfinding radius represents the minimum pathfinding radius preset for the second sample virtual object. When virtual obstacles exist between the first sample position and the second sample position, the first sample virtual object is generated at the first sample position, and the second sample virtual object is generated at the second sample position. Based on the first sample environment perception data corresponding to the first sample virtual object and the second sample environment perception data corresponding to the second sample virtual object, the first sample virtual object and the second sample virtual object are controlled to engage in a game until the game ends.

[0015] Optionally, the device is used to perform distributed training on the initial action inference model to obtain the target action inference model in the following manner: performing non-full processing on the first sample environment perception data and the second sample environment perception data to obtain first sample data and second sample data; when the first sample virtual object is visible to the second sample virtual object, extracting features from the first sample data to obtain first sample features, wherein the first sample features include partial data of the second sample virtual object; when the second sample virtual object is not visible to the first sample virtual object, extracting features from the second sample data to obtain second sample features, wherein the second sample features include partial data of the first sample virtual object after being masked by the preset parameters; inputting the first sample features and the second sample features into the initial action inference model respectively, determining preset index values ​​corresponding to the first sample virtual object and the second sample virtual object respectively, and updating the initial action inference model according to the preset index values ​​to determine the target action inference model, wherein during the update of the initial action inference model, the output mask of the feature network layer associated with the partial data is uniformly masked as the target value.

[0016] Optionally, the device is configured to acquire environmental awareness data corresponding to a first virtual object in the target virtual scene in the following manner: when the service in the target virtual scene reaches the i-th frame, acquire the first environmental awareness data corresponding to the i-th frame; when the service in the target virtual scene reaches the i+n-th frame, acquire the second environmental awareness data corresponding to the i+n-th frame, wherein the service maintains the execution of preset service logic from the i-th frame to the i+n-th frame, and i and n are both positive integers; the step of determining action package data based on the environmental awareness data and controlling the first virtual object to execute a target action according to the action package data includes: determining the first action package data based on the first environmental awareness data and controlling the first virtual object to execute a first target action according to the first action package data at the i+j-th frame; determining the second action package data based on the second environmental awareness data and controlling the first virtual object to execute a second target action according to the second action package data at the i+n+j-th frame, wherein j is a non-negative integer.

[0017] Optionally, the device is configured to determine action package data based on the environmental perception data in the following manner, and control the first virtual object to perform a target action according to the action package data: when the environmental perception data acquired in the i-th frame indicates that the first virtual object needs to be controlled to launch a virtual prop for the first time within a preset time period, the maximum aiming deviation angle preset for the first virtual object is obtained, where i is a positive integer; a first confidence region is determined based on the distance between the first virtual object and the target virtual object and the maximum aiming deviation angle, where the radius of the first confidence region is r, and r is greater than 0; a first confidence point is sampled from the first confidence region, and a second confidence region is determined with the first confidence point as the center and r / x as the radius; a second confidence point is sampled from the second confidence region, where the second confidence point represents the aiming crosshair of the first virtual object corresponding to the environmental perception data acquired in the i-th frame.

[0018] Optionally, the device is further configured to: after obtaining the second confidence point from the second confidence region, if the environmental perception data obtained in the (i+m)th frame indicates that the first virtual object needs to launch a virtual prop for the kth time, determine a third confidence region with the third confidence point determined in the (k-1)th frame as the center and r / x as the radius, where m is a positive integer and m is less than the preset duration, k is a positive integer greater than or equal to 2, and when k = 2, the third confidence point is the second confidence point; obtain a target confidence point from the third confidence region, where the target confidence point represents the aiming crosshair of the first virtual object corresponding to the environmental perception data obtained in the (i+m)th frame.

[0019] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the control method of the virtual object described above when running.

[0020] According to another aspect of the embodiments of this application, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the control method of the virtual object described above.

[0021] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the control method of the virtual object described above through the computer program.

[0022] In this embodiment, a target virtual scene is displayed, where the target virtual scene represents the game scene where the target virtual object in the target application is located, and the target virtual object represents the object controlled by the target account logged into the target application. Environmental awareness data corresponding to the first virtual object in the target virtual scene is obtained, where the first virtual object represents the object controlled by the system. Part of the environmental awareness data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object. If the target virtual object is not visible to the first virtual object, part of the data is set to be occluded by preset parameters. Action package data is determined based on the environmental awareness data, and the first virtual object is controlled to execute the target action according to the action package data. The action package data includes different types of sub-actions predicted based on the environmental awareness data, and the target action is determined by combining different types of sub-actions. In other words, in the target virtual scene, environmental awareness data is obtained based on the visibility information between the target virtual object and the first virtual object. Then, action package data is determined through the environmental awareness data to enable the first virtual object to execute the target action according to the action package data. This achieves the technical effect of improving the accuracy of the target action based on environmental awareness data, solving the technical problem of low control accuracy of virtual objects due to their singular representation.

[0023] On the other hand, in the target game scene, the aforementioned environmental perception data, in addition to being determined based on the visibility information between the target virtual object and the first virtual object, also includes the visibility information of the first virtual object itself. Furthermore, by combining the visibility information of the first virtual object itself with the visibility information between the first virtual object and the target virtual object, environmental perception data is obtained, ensuring the accuracy and timeliness of the environmental perception data. This allows the environmental perception data to fully reflect the current state of the first virtual object in the target game screen, facilitating the subsequent generation of action package data based on the environmental perception data. This makes the target action represented by the action package data more consistent with the current state of the first virtual object, improving the accuracy and intelligence of the target action. This enables the first virtual object to provide players with a more flexible and richer gaming experience by executing the target action, increasing the stickiness between players and the game.

[0024] Furthermore, by deploying the inference model on the inference server to differentiate it from the application server, and by generating the action package data corresponding to the control method of the virtual object on the inference server, the technical objective of reducing the computational overhead of the application server for controlling the anthropomorphic virtual object is achieved.

[0025] On the other hand, by using preset parameters to mask some environmental data, the human-like effect of virtual objects can be improved while reducing data transmission overhead. Attached Figure Description

[0026] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0027] Figure 1 This is a schematic diagram of an application environment for an optional virtual object control method according to an embodiment of this application;

[0028] Figure 2 This is a flowchart illustrating an optional virtual object control method according to an embodiment of this application;

[0029] Figure 3 This is a schematic diagram of an optional virtual object control method according to an embodiment of this application;

[0030] Figure 4 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0031] Figure 5 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0032] Figure 6 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0033] Figure 7 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0034] Figure 8 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0035] Figure 9 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0036] Figure 10 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0037] Figure 11 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0038] Figure 12 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0039] Figure 13This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0040] Figure 14 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0041] Figure 15 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application;

[0042] Figure 16 This is a schematic diagram of the structure of an optional virtual object control device according to an embodiment of this application;

[0043] Figure 17 This is a schematic diagram of the structure of an optional virtual object control product according to an embodiment of this application;

[0044] Figure 18 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of this application. Detailed Implementation

[0045] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0046] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0047] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows:

[0048] Reinforcement Learning: Reinforcement learning (RL) is one of the paradigms of machine learning. It studies how an agent makes decisions about the environment in a series of decision sequences to obtain rewards, and continuously improves its own decisions to obtain the maximum reward value. After repeated iterations, the agent obtains an optimal mapping from perception to decision.

[0049] Supervised learning (SL) is a paradigm of machine learning and a data analysis method that trains an iterative learning algorithm on a labeled dataset to enable the model to accurately classify data or predict results.

[0050] Depth map: also known as a depth map, is an image that uses the distance (depth) from the image sensor to each point in the scene as pixel values. It can depict the pure depth information in front in detail. The closer the value is to 1, the brighter the front is and the clearer the field of view is.

[0051] Ring rays: Starting from itself, several rays are emitted outward around its own perimeter. They are used to describe the terrain information of obstacles around itself and are a way of environmental perception.

[0052] Action Masking: In deep reinforcement learning training, Action Masking (AM) temporarily masks "unreasonable or unexecutable action dimensions" when the network infers the action distribution, and then samples actions from "reasonable or executable actions" according to the original probability. This can ensure that the inferred actions can obtain effective data samples in the game environment and avoid wasting computing resources due to returning invalid actions.

[0053] Strong anthropomorphism: The NPCs or game characters controlled in the game behave and act very similarly to real players. Their behavior is diverse and intelligent, making it difficult for human players to distinguish the characters from real ones.

[0054] Behavior tree: A tree-like structure used to control the decisions of NPCs in a game.

[0055] Proximal Policy Optimization (PPO) is an online deep reinforcement learning algorithm based on policy gradient optimization and oriented towards continuous or discrete action spaces.

[0056] Auxiliary rewards: In reinforcement learning algorithms, game agents are enabled to achieve various differentiated performances. Based on the differentiated goals, rewards are designed to help achieve the target virtual objects.

[0057] Atomic actions: In games, these are the most basic and fine-grained primitive actions performed by game characters. These actions cannot be further broken down or refined.

[0058] Fully Connected Layer: The fully connected layer is one of the many components of the multilayer perceptron (MLP) applied in deep neural networks. It maps the features of the previous layer to the next layer in the form of vector computation.

[0059] Convolutional layer: A convolutional layer is a set of parallel feature maps that are formed by sliding different convolutional kernels on the input image and performing certain matrix operations. There are one-dimensional convolution, two-dimensional convolution and three-dimensional convolution.

[0060] Full perception: Shooting games have a wide variety of maps and complex and varied terrain. Due to the limited range of visible angles, much information is not observable. If we assume that all information is visible to the game's intelligent agent, it is called full perception. In essence, full perception is a perception method that uses cheating to obtain information about the opponent's player.

[0061] Incomplete perception: In shooting games, in order to train a highly human-like AI that can perform various human-like actions and make it difficult for human players to distinguish the AI ​​from the real thing, it is necessary to align the input information of human players and ensure that the anthropomorphism is valid from the perspective of environmental perception input. This is called incomplete perception, which uses the method of a real player without cheating to perceive environmental input.

[0062] Shooting confidence zone: In shooting games, the shooting confidence zone defines the range of valid shots. Shots outside the confidence zone are invalid shots and will not cause damage to enemies. Only shots within the confidence zone are valid shots.

[0063] Autoregression: A statistical method for processing time series data. It uses the previous time points of the same variable x to predict the performance of x at the current time point. This regression analysis method, which uses the same variable x to predict its own performance, is called autoregression and has time series correlation.

[0064] Overfitting: In statistics, overfitting refers to the phenomenon where a particular dataset is matched too closely or precisely, resulting in an inability to fit other data or predict future observations well.

[0065] The present application will be described below with reference to embodiments:

[0066] According to one aspect of the embodiments of this application, a method for controlling virtual objects is provided. Optionally, in this embodiment, the above-mentioned method for controlling virtual objects can be applied to, for example... Figure 1 The hardware environment shown consists of server 101 and terminal device 103. For example... Figure 1As shown, server 101 is connected to terminal device 103 via a network and can be used to provide services to terminal device or application 107 installed on terminal device. The application can be video application, instant messaging application, browser application, educational application, game application, etc. Database 105 can be set up on the server or independently of the server to provide data storage services for server 101, such as a game data storage server. The network mentioned above can include, but is not limited to, wired networks and wireless networks. The wired network includes local area networks, metropolitan area networks, and wide area networks. The wireless network includes Bluetooth, WIFI, and other networks that enable wireless communication. Terminal device 103 can be a terminal configured with an application, and can include, but is not limited to, at least one of the following: mobile phones (such as Android phones, iOS phones, etc.), laptops, tablets, handheld computers, MID (Mobile Internet Devices), PADs, desktop computers, smart TVs, smart voice interaction devices, smart home appliances, vehicle terminals, aircraft, virtual reality (VR) terminals, augmented reality (AR) terminals, mixed reality (MR) terminals, and other computer devices. The server mentioned above can be a single server, a server cluster composed of multiple servers, or a cloud server.

[0067] Combination Figure 1 As shown, the control method of the virtual object can be executed by an electronic device, which can be a terminal device or a server. The control method of the virtual object can be implemented by the terminal device or the server respectively, or by the terminal device and the server together.

[0068] In one exemplary embodiment, the following may be included, but are not limited to:

[0069] S1, the target virtual scene is displayed on the terminal device 103, wherein the target virtual scene represents the target application ( Figure 1 The game scene in the application 107 shown shows the target virtual object, which represents an object controlled by the target account logged into the target application.

[0070] S2, obtain environmental perception data corresponding to the first virtual object in the target virtual scene through server 101. The first virtual object represents an object controlled by the system. The part of the environmental perception data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object. If the target virtual object is not visible to the first virtual object, part of the data is set to be occluded by preset parameters.

[0071] S3, the server 101 or the terminal device 103 determines the action package data based on the environmental perception data, and controls the first virtual object to execute the target action according to the action package data. The action package data includes different types of sub-actions predicted based on the environmental perception data, and the target action is determined by combining different types of sub-actions.

[0072] The above is merely an example, and this embodiment does not impose any specific limitations.

[0073] Alternatively, as an alternative implementation method, such as Figure 2 As shown, the control methods for the aforementioned virtual objects include:

[0074] S202, Display the target virtual scene, where the target virtual scene represents the game scene in which the target virtual object in the target application is located, and the target virtual object represents the object controlled by the target account logged into the target application;

[0075] Optionally, in this embodiment, the aforementioned target virtual scene can be understood as a scene within a target application such as a game. The target application here can be, but is not limited to, game applications, social media applications, fitness applications, educational applications, business applications, music applications, etc. These applications can run on mobile phones, tablets, or computers, providing users with functions such as interaction with other users, entertainment, learning, or business activities. The aforementioned target virtual object can include, but is not limited to, the image of a user logging into the target application with a target account in the target virtual scene. This target virtual object can be understood as a virtual entity representing the user. Through the target virtual object, the user can perform various activities and interactions, such as interacting with other users on a virtual social platform, playing games in a virtual game, and experiencing various scenes in a virtual reality environment. It is understood that the target virtual object can include, but is not limited to, virtual characters, virtual pets, virtual items, etc., representing the user's existence in the virtual world and providing the user with a rich and varied virtual experience.

[0076] Taking a game application as an example, when a player enters the game using the target account on a terminal device, the game screen displayed on the terminal device is the aforementioned target virtual scene. Assuming the game application is a shooting game, then the target virtual scene may include, but is not limited to, enemy characters, virtual equipment, city streets, jungles, deserts and other different terrains, as well as virtual props such as medkits and ammunition packs.

[0077] For example, Figure 3 This is a schematic diagram of an optional virtual object control method according to an embodiment of this application. The target virtual scene can be as follows: Figure 3As shown, there are eight different terrain areas: Area A, Area B, Area C, Area D, Area E, Area F, Area G, and Area H. Players can fight in various terrains and also use the terrain for evasion and strategic actions.

[0078] S204, Obtain environmental perception data corresponding to the first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by the system, and the part of the environmental perception data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object. If the target virtual object is not visible to the first virtual object, part of the data is set to be occluded by preset parameters.

[0079] Optionally, in this embodiment, the first virtual object can be understood as a game agent. The first virtual object and the target virtual object can interact, including but not limited to combat, conversation, and performing game tasks. Here, the game agent can be understood as a machine that performs reinforcement learning in the game scene. It is not a virtual object directly controlled by the player, but a game AI driven by a neural network. The thing that interacts with the game agent is called the environment. The game agent makes reasonable decisions by analyzing the game environment and can cooperate and compete with other agents or players, making it behave like a human. The environmental perception data can include but not limited to the current position, health, whether it is injured, depth map, surrounding ray, whether it is firing, and whether the first virtual object and the target virtual object are visible to each other. The visibility between the first virtual object and the target virtual object can be determined by ray detection.

[0080] For example, such as Figure 3 As shown, region E in the target virtual scene includes the first virtual object and the target virtual object.

[0081] Furthermore, Figure 4 This is a schematic diagram of an optional virtual object control method according to an embodiment of this application, such as... Figure 4As shown, taking the first virtual object as an example of an intelligent game entity in a game application, the first virtual object and the target virtual object are in the same building. The first virtual object is in the lobby on the first floor of the building, and the target virtual object is in the bedroom on the second floor. At this time, the first virtual object and the target virtual object are not visible to each other. The corresponding environmental perception data may include, but is not limited to, the basic information of the target virtual object: orientation, position, etc. In addition, it is not necessary to obtain overly precise information such as health points, number of bullets, and range of virtual items in the basic information. The information on whether the target virtual object can see the first virtual object needs to be masked. Only the estimated position of the target virtual object is used to replace the precise position of the target virtual object. Based on this, the relative position and relative angle between them are calculated. At this time, since the target virtual object is not visible to the first virtual object, it can be understood that the first virtual object does not exist in the target virtual object's game field of view, and the position of the first virtual object cannot be determined in the game field of view. At this time, the part of the environmental perception data associated with the target virtual object indicates that the target virtual object is not visible to the first virtual object.

[0082] For example, if the target virtual object is not visible to the first virtual object, then some data in the environmental perception data will be occluded. This can be determined by setting preset parameters, for example, the preset parameters are set to special values ​​(-1, 0), and the corresponding data values ​​are all filled with (-1, 0) as placeholders, so as to achieve a deep restoration of the game perception of human players, ensuring human-likeness from the source of environmental perception, and providing data support and potential possibilities for subsequent model training of human-like strategies. The value of the preset parameters can be set flexibly, and this application does not impose any limitations on it.

[0083] Similarly, as Figure 4 As shown, if the target virtual object moves to the lobby on the first floor after a period of time, a first virtual object exists in the target virtual object's game field of view. The location of the first virtual object can be determined in the game field of view. At this time, the part of the environmental perception data associated with the target virtual object indicates that the target virtual object is visible to the first virtual object. The basic information of the target virtual object includes its orientation and position. Furthermore, it is not necessary to obtain overly precise information such as health, bullet count, and virtual item range in the basic information. The information on whether the target virtual object can see the first virtual object does not need to be masked. This corresponds to the fact that in a real game scenario, human players cannot obtain the real-time bullet count and precise health information of the currently visible first virtual object. This further aligns the deep simulation with the perceptual input of human players.

[0084] It should be noted that, Figure 5 This is a schematic diagram of an optional virtual object control method according to an embodiment of this application. The environmental perception data may include, but is not limited to, such as... Figure 5As shown, the target virtual object is invisible to the first virtual object, corresponding to... Figure 5 The data shown in part 502, the target virtual object is visible to the first virtual object, corresponding to... Figure 5 The data shown is 504.

[0085] S206, based on the motion package data determined by occlusion through preset parameters, control the first virtual object to execute the target action according to the motion package data, wherein the motion package data includes different types of sub-actions predicted based on environmental perception data, and the target action is determined by combining different types of sub-actions.

[0086] Optionally, in the embodiments of this application, the aforementioned action package data may include, but is not limited to, action data to be performed by the first virtual object. The action package data includes data of different action types, such as moving eastward in the direction of movement and walking quickly in the mode of movement. The aforementioned target action may include a combination of multiple sub-actions indicated by the action package data. For example, the target action includes moving eastward and walking quickly. Further, the first virtual object will move eastward using the mode of walking quickly.

[0087] Figure 6 This is a schematic diagram of an optional virtual object control method according to an embodiment of this application. The action package data may include, but is not limited to, the following: Figure 6 The data types shown may include, but are not limited to, the sum of any one of the following values ​​for the horizontal aiming direction of the first virtual object: [0°, 30°, 60°, 90°, 120°, 150°, 180°, 210°, 240°, 270°, 300°, 330°] and the horizontal angle between the horizontally aiming target virtual object and the target virtual object. The environmental perception data includes the angle between the horizontal angles of the target virtual object and the target virtual object; and the sum of any one of the following values ​​for the vertical aiming direction of the gun muzzle: [-8°, -5°, -2°, 0°, 2°, 5°, 8°] and the vertical angle between the horizontally aiming target virtual object and the target virtual object. The environmental perception data includes... The data includes: the angle between the target virtual object and the target virtual object perpendicular to each other; whether to fire: 0 indicates no fire, 1 indicates fire; movement mode: 0 indicates walking still, 1 indicates walking fast, 2 indicates sprinting, 3 indicates standing still; pathfinding mode: 0 indicates atomic movement, 1 indicates assisted movement based on behavior tree, 2 indicates maintaining the previous movement mode, 3 indicates stopping assisted movement; movement direction: 16-dimensional atomic movement at 22.5° intervals, placeholders with no specific meaning (indicating assisted movement); posture type: 0 indicates standing, 1 indicates crouching, 2 indicates lying down, 3 indicates jumping; side-stepping type: 0 indicates not side-stepping, 1 indicates left side, 2 indicates right side; special types: 0 indicates aiming down sights, 1 indicates reloading, 2 indicates neither aiming down sights nor reloading, etc.

[0088] Furthermore, Figure 7 This is a schematic diagram of an optional virtual object control method according to an embodiment of this application. It may include, but is not limited to, periodically acquiring environmental perception data or determining action package data at preset time points, or acquiring environmental perception data in real time, and determining action package data based on the environmental perception data periodically or at preset time points. Furthermore, in different time frames, the target action indicated by the action package data may include the same or different sub-actions; in other words, the determination of sub-actions may be the same or different, such as... Figure 7 As shown, the sub-action of aiming horizontally is updated every 2 time frames, the sub-action of aiming vertically is updated every 2 time frames, the sub-action of firing is updated every 1 time frame, the sub-action of movement mode is updated every 3 time frames, the sub-action of special type is updated every 2 time frames, the sub-action of movement direction is updated every 4 time frames, the sub-action of pathfinding direction is updated every 5 time frames, the sub-action of posture type is updated every 3 time frames, and the sub-action of sidestepping type is updated every 2 time frames.

[0089] Taking shooting game applications as an example, when environmental perception data is obtained through methods including but not limited to ray detection, action package data can be determined based on the environmental perception data. Specifically, this may include, but is not limited to, using the environmental perception data as input to a pre-trained target action inference model, and having the target action inference model output action package data. For example, if environmental perception data is obtained in the first time frame, and after two time frames, action package data is obtained in the second time frame, the first virtual object will execute the target action indicated by the action package data.

[0090] It's important to note that the training and deployment of the target action reasoning model are not located on the target application's backend server. In other words, taking a game application as an example, the dedicated DS server where the game business logic requested by the player resides and the AI ​​service module where the target action reasoning model resides are different. The dedicated DS server only provides the game business logic and cannot output action package data based on environmental awareness data. The game business logic may include, but is not limited to: game rules: the basic framework of the game, defining how players interact in the game, including game objectives, victory conditions, defeat conditions, player character abilities and characteristics, etc.; game system: various mechanisms and functions in the game, including player character attributes, skill system, equipment system, combat system, economic system, social system, etc.; game flow: the entire experience of the player in the game, including the start of the game, ongoing tasks and challenges, and end-of-game settlement and rewards, etc.; game balance: the balance between various elements in the game, including character ability balance, game system balance, and game difficulty balance, etc. The AI ​​service module where the target action reasoning model resides, on the other hand, can output action package data based on environmental awareness data.

[0091] For example, Figure 8 This is a schematic diagram of an optional virtual object control method according to an embodiment of this application, such as... Figure 8 As shown, taking a game application as an example, the player receives the game screen in the game client, obtains the game business logic through the DS server, transmits the environmental perception data to the distributed reinforcement learning network, and the target action inference model outputs action package data. By parsing the action package data, the first target virtual object performs the corresponding target action in the game screen, such as running forward, or running forward 50 meters and then lying down and remaining still.

[0092] In one exemplary embodiment, the virtual object control method proposed in this application can be applied to application scenarios where the actions of intelligent game entities in a game are determined. Through the embodiments of this application, the current actions of intelligent game entities can be quickly determined in a short time, thereby improving the development efficiency of intelligent game entities. Specifically... Figure 9 This is a schematic diagram of an optional virtual object control method according to an embodiment of this application, such as... Figure 9 As shown:

[0093] S902 retrieves the game business logic of the current game frame from a dedicated server, including but not limited to game rules, game flow, and game system;

[0094] S904-1, construct the state data of the i-th frame corresponding to the first virtual object as the above-mentioned environmental perception data, and send the environmental perception data to the online inference service, and execute S906-1 and S906-2 simultaneously.

[0095] S906-1, the online inference service calculates action package data through static ray query and target action inference model, and sends the action package data back to the game client or dedicated server;

[0096] S906-2, The client continues to execute the game's business logic;

[0097] S908, the game client or dedicated server obtains the target action by parsing the action packet data, so that the first virtual object performs the target action. For example, the target action is to first run forward 50 meters, and then lie down and remain still.

[0098] S910, repeat steps S902 to S908 until the game ends and the player exits the game application.

[0099] In yet another exemplary embodiment, the virtual object control method proposed in this application can be applied to intelligent transportation application scenarios. Through the embodiments of this application, the vehicle driving path of a virtual vehicle in an intelligent transportation system can be quickly determined in a short time. Based on environmental perception data, it can be ensured that the path traveled by the virtual vehicle and the driving operation are correct, which may include, but is not limited to:

[0100] S1, obtains the intelligent transportation business logic of the current frame from the dedicated server, including but not limited to traffic rules, traffic processes, traffic systems, etc.

[0101] S2, construct the state data of the i-th frame corresponding to the first virtual object as the above-mentioned environmental perception data, and send the environmental perception data to the online inference service, and execute S3-1 and S3-2 simultaneously;

[0102] S3-1, the online inference service calculates action package data through static ray query and target action inference model, and sends the action package data back to the client or dedicated server;

[0103] The S3-2 client continues to execute intelligent transportation business logic;

[0104] S4, the client or dedicated server obtains the target action by parsing the action packet data, so that the virtual vehicle (the first virtual object mentioned above) executes the target action;

[0105] S5. Repeat steps S1 to S4 until the virtual vehicle reaches its destination.

[0106] It should be noted that the embodiments of this application can be applied to various scenarios such as games, cloud technology, artificial intelligence, and smart transportation.

[0107] This application's embodiments employ a display target virtual scene, where the target virtual scene represents the game scene where a target virtual object in a target application resides, and the target virtual object represents an object controlled by a target account logged into the target application. Environmental awareness data corresponding to a first virtual object in the target virtual scene is acquired. The first virtual object represents an object controlled by the system. A portion of the environmental awareness data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object. If the target virtual object is not visible to the first virtual object, a portion of the data is occluded using preset parameters. Action package data is determined based on the environmental awareness data, and the first virtual object is controlled to execute a target action according to the action package data. The action package data includes different types of sub-actions predicted based on the environmental awareness data. The target action is determined by combining different types of sub-actions. In other words, in the target virtual scene, environmental awareness data is acquired based on the visibility information between the target virtual object and the first virtual object. Then, action package data is determined using the environmental awareness data to enable the first virtual object to execute the target action according to the action package data. This achieves the technical effect of improving the accuracy of target actions based on environmental awareness data, solving the technical problem of low control accuracy of virtual objects due to their singular representation.

[0108] On the other hand, in the target game scene, the aforementioned environmental perception data, in addition to being determined based on the visibility information between the target virtual object and the first virtual object, also includes the state information of the first virtual object itself. Furthermore, by combining the state information of the first virtual object itself with the visibility information between the first virtual object and the target virtual object, environmental perception data is obtained, ensuring the accuracy and timeliness of the environmental perception data. This allows the environmental perception data to fully reflect the current state of the first virtual object in the target game screen, facilitating the subsequent generation of action package data based on the environmental perception data. This makes the target action represented by the action package data more consistent with the current state of the first virtual object, improving the accuracy of the target action. As a result, the first virtual object can provide players with a more flexible and richer gaming experience by executing the target action, increasing the stickiness between players and the game.

[0109] As an optional approach, the above-mentioned determination of action package data based on the aforementioned environmental perception data, and control of the first virtual object to execute the target action according to the aforementioned action package data, includes: predicting a first sub-action based on the aforementioned environmental perception data to obtain a first set of sub-action confidence scores, wherein one sub-action confidence score in the first set of sub-action confidence scores corresponds to one action parameter of the first sub-action; predicting a second sub-action based on the aforementioned environmental perception data to obtain a second set of sub-action confidence scores, wherein one sub-action confidence score in the second set of sub-action confidence scores corresponds to one action parameter of the second sub-action. The parameters correspond to each other, with the first sub-action and the second sub-action representing sub-actions that do not affect each other; the action parameters corresponding to the confidence values ​​of the sub-actions in the first group of sub-actions that meet the preset conditions are determined as the first sub-action data corresponding to the first sub-action; the action parameters corresponding to the confidence values ​​of the sub-actions in the second group of sub-actions that meet the preset conditions are determined as the second sub-action data corresponding to the second sub-action; the action package data is generated based on the first sub-action data and the second sub-action data, and the first virtual object is controlled to execute the target action according to the action package data.

[0110] Optionally, in the embodiments of this application, the aforementioned first sub-action and second sub-action may include, but are not limited to, firing, crouching, reloading, sidestepping, walking silently, aiming down sights, adjusting the horizontal aiming angle of the gun to 30°, adjusting the vertical aiming angle of the gun to 8°, etc., and the first sub-action and the second sub-action do not affect each other. For example, if the first sub-action is crouching and the second sub-action is reloading, then the first sub-action and the second sub-action do not affect each other, and the first sub-action and the second sub-action can be executed synchronously. The first sub-action and the second sub-action are orthogonal. For another example, if the first sub-action is sprinting and the second sub-action is sidestepping, then the first sub-action and the second sub-action affect each other. Furthermore, when the first sub-action and the second sub-action are executed synchronously, the first virtual object will experience model twitching and shaking, affecting the player's gaming experience.

[0111] For example, based on environmental perception data, the first sub-action and the second sub-action can be predicted separately. The first sub-action has a set of first sub-action confidence scores, which are used to determine the action parameters of the first sub-action. For example, if the type of the first sub-action is whether to fire, then the first sub-action can have two action parameters, corresponding to firing and not firing respectively. This can include, but is not limited to, setting the action parameter to 0, indicating that the first sub-action does not fire, with a corresponding sub-action confidence score of 30; and setting the action parameter to 1, indicating that the first sub-action fires, with a corresponding sub-action confidence score of 70.

[0112] Similarly, the second sub-action has a second set of sub-action confidence scores. This set of sub-action confidence scores is used to determine the action parameters of the second sub-action. The type of the second sub-action is movement mode. Therefore, the second sub-action can have four action parameters, corresponding to standing, squatting, lying down, and jumping, respectively. It can include, but is not limited to, setting the action parameter to 0, which indicates that the first sub-action is standing, and the corresponding sub-action confidence score is 30; setting the action parameter to 1, which indicates that the first sub-action is squatting, and the corresponding sub-action confidence score is 35; setting the action parameter to 2, which indicates that the first sub-action is lying down, and the corresponding sub-action confidence score is 15; and setting the action parameter to 3, which indicates that the first sub-action is jumping, and the corresponding sub-action confidence score is 20.

[0113] Furthermore, the aforementioned preset conditions can be flexibly set, and this application does not impose any limitations on them. Taking the example of the first sub-action being whether to fire and the second sub-action being the movement method, the preset conditions are set as follows: the action parameter corresponding to the sub-action with the highest confidence value in the first group of sub-actions is used as the action parameter of the first sub-action to generate the first sub-action data; the action parameter corresponding to the sub-action with the highest confidence value in the first group of sub-actions is used as the action parameter of the second sub-action to generate the second sub-action data. At this time, the action parameter of the first sub-action is 1, and the first sub-action represents firing; the action parameter of the second sub-action is 1, and the first sub-action represents crouching.

[0114] For example, the preset conditions are set as follows: the action parameter corresponding to the sub-action confidence value with the largest value in the first group of sub-action confidence is used as the action parameter of the first sub-action to generate the first sub-action data; the action parameter corresponding to the sub-action confidence value with the largest value in the first group of sub-action confidence is used as the action parameter of the second sub-action to generate the second sub-action data; if there are multiple identical sub-action confidence values ​​with the largest value in the first group of sub-action confidence, then the action parameter corresponding to one of these identical sub-action confidence values ​​will be randomly selected as the action parameter of the first sub-action; similarly, if there are multiple identical sub-action confidence values ​​with the largest value in the second group of sub-action confidence, then the action parameter corresponding to one of these identical sub-action confidence values ​​will be randomly selected as the action parameter of the second sub-action.

[0115] For example, the first sub-action includes: with the action parameter set to 0, the first sub-action represents not firing, and the corresponding sub-action confidence level is 50; with the action parameter set to 1, the first sub-action represents firing, and the corresponding sub-action confidence level is 50. Then, an action parameter can be randomly selected from the action parameters with values ​​of 0 and 1 corresponding to the first sub-action to generate the first sub-action data. Similarly, the second sub-action includes: with the action parameter set to 0, the first sub-action represents standing, and the corresponding sub-action confidence level is 25; with the action parameter set to 1, the first sub-action represents squatting, and the corresponding sub-action confidence level is 25; with the action parameter set to 2, the first sub-action represents lying down, and the corresponding sub-action confidence level is 25; with the action parameter set to 3, the first sub-action represents jumping, and the corresponding sub-action confidence level is 25. Then, an action parameter can be randomly selected from the action parameters with values ​​of 0, 1, 2, and 3 corresponding to the first sub-action to generate the second sub-action data.

[0116] In an exemplary embodiment, the action package data may include, but is not limited to, first sub-action data and second sub-action data. The target action may include, but is not limited to, the first sub-action and the second sub-action. For example, the first sub-action data is the action data of the first virtual object performing the firing action, including but not limited to the firing position, time, and virtual props used. The second sub-action data is the action data of the first virtual object performing the squatting action, including but not limited to the squatting position and time. The execution time of the first sub-action indicated by the first sub-action data and the execution time of the second sub-action indicated by the second sub-action data may be the same or different. For example, the corresponding target action includes firing and squatting, including but not limited to firing first and then squatting, or squatting first and then firing, or performing squatting and firing simultaneously, so as to increase the diversity and richness of the target actions, thereby making the target actions more consistent with the current state of the first virtual object, so as to improve the accuracy of the target actions.

[0117] As an optional approach, the above-mentioned determination of action package data based on the environmental perception data and control of the first virtual object to perform the target action according to the action package data includes: predicting the first sub-action according to the environmental perception data at a first frequency to obtain the confidence level of the first set of sub-actions, wherein the first frequency represents the preset execution frequency of the first sub-action; predicting the second sub-action according to the environmental perception data at a second frequency to obtain the confidence level of the second set of sub-actions, wherein the second frequency represents the preset execution frequency of the second sub-action, and the first frequency and the second frequency are different.

[0118] In an exemplary embodiment, the first sub-action data can be updated using a first frequency as the update frequency of the first sub-action to obtain the corresponding first sub-action. For example, if the first sub-action is updated at a first time point, the first sub-action is updated once every 2 time frames, then the first frequency is 2 time frames. Similarly, if the second sub-action is updated at a second time point, the second sub-action is updated once every 3 time frames, then the second frequency is 3 time frames. It can be understood that the first frequency and the second frequency can be the same or different.

[0119] For example, such as Figure 7 As shown, assume the first sub-action is the horizontal aiming direction sub-action, with a first frequency of 2 time frames, where the horizontal aiming direction sub-action is updated every 2 time frames; the second sub-action is the movement direction sub-action, with a second frequency of 3 time frames, where the movement direction sub-action is updated every 3 time frames.

[0120] In the embodiments of this application, the first sub-action is updated periodically at a first frequency and the second sub-action is updated periodically at a second frequency, ensuring that different sub-actions can be updated independently of each other, so as to achieve the purpose of parallel processing of multiple sub-actions, saving computing resources and improving the generation efficiency of the target action.

[0121] As an optional approach, the method further includes: when the first sub-action represents an attack operation, predicting the first sub-action according to the environmental perception data at a first frequency to obtain the confidence level of the first set of sub-actions; when the second sub-action represents a sub-action other than the attack operation, predicting the second sub-action according to the environmental perception data at a second frequency to obtain the confidence level of the second set of sub-actions, wherein the first frequency is greater than the second frequency; when the first sub-action represents a pathfinding operation, predicting the first sub-action according to the environmental perception data at a first frequency to obtain the confidence level of the first set of sub-actions; when the second sub-action represents a sub-action other than the pathfinding operation, predicting the second sub-action according to the environmental perception data at a second frequency to obtain the confidence level of the second set of sub-actions, wherein the first frequency is less than the second frequency.

[0122] Optionally, in the embodiments of this application, the above-mentioned attack operation can be understood as the first virtual object launching an attack on the target virtual object, which may cause damage to the target virtual object; the above-mentioned pathfinding operation may include, but is not limited to, atomic movement, behavior tree-based assisted movement, maintaining the previous movement mode, stopping assisted movement, etc.

[0123] For example, taking a shooting game as an example, the above attack action is firing. If the first sub-action indicates firing, then the first frequency can be used to predict the first sub-action. The confidence level of the first group of sub-actions corresponding to the first sub-action is expressed as follows: if the action parameter is set to 0, the first sub-action indicates not firing, and the corresponding sub-action confidence level is 50; if the action parameter is set to 1, the first sub-action indicates firing, and the corresponding sub-action confidence level is 50. Figure 7 As shown, the second sub-action can be any action other than firing, such as movement. In this case, the second frequency can be used to predict the second sub-action, and the first frequency is greater than the second frequency. That is, the prediction frequency of the first sub-action is higher than the prediction frequency of the second sub-action. For example, the first sub-action is predicted once every 2 time frames; the second sub-action is predicted once every 5 time frames.

[0124] Similarly, if the first sub-action represents a pathfinding operation, then the first frequency can be used to predict the first sub-action. The confidence level of the first group of sub-actions corresponding to the first sub-action is expressed as follows: when the action parameter is set to 0, the first sub-action represents an atomic movement, and the corresponding sub-action confidence level is 50; when the action parameter is set to 1, the first sub-action represents an auxiliary movement based on the behavior tree, and the corresponding sub-action confidence level is 10; when the action parameter is set to 2, the first sub-action represents maintaining the previous movement method, and the corresponding sub-action confidence level is 20; when the action parameter is set to 3, the first sub-action represents stopping the auxiliary movement, and the corresponding sub-action confidence level is 20; for example... Figure 7 As shown, the second sub-action can be any action other than pathfinding, such as movement. In this case, the second frequency can be used to predict the second sub-action, and the first frequency is less than the second frequency. That is, the prediction frequency of the first sub-action is lower than the prediction frequency of the second sub-action. For example, the first sub-action is predicted once every 5 time frames, and the second sub-action is predicted once every 2 time frames.

[0125] By using the embodiments of this application, the prediction frequency of the sub-actions representing attack operations is set to the maximum, and the prediction frequency of the sub-actions representing pathfinding operations is set to the minimum, making the target actions more human-like and more in line with the actual game scene, and ensuring that the interaction between the first virtual object and the target virtual object is more flexible and intelligent.

[0126] As an optional approach, the above-mentioned method of generating the action package data based on the first sub-action data and the second sub-action data, and controlling the first virtual object to execute the target action according to the action package data, includes: obtaining a predetermined invalid action combination; if the action combination composed of the first sub-action data and the second sub-action data does not belong to the invalid action combination, generating the action package data based on the first sub-action data and the second sub-action data, and controlling the first virtual object to execute the target action according to the action package data; if the action combination composed of the first sub-action data and the second sub-action data belongs to the invalid action combination, deleting the first sub-action data and the second sub-action data from the action package data.

[0127] Optionally, in the embodiments of this application, the above-mentioned invalid action combination may include, but is not limited to, several sub-actions. The invalid action combination can be preset by human, and this application does not limit it. It is understood that if the first virtual object executes these several sub-actions simultaneously, for example, if the invalid action combination includes the fast walking sub-action and the lying down sub-action, then the first virtual object will exhibit twitching and shaking, affecting the player's game experience.

[0128] In an exemplary embodiment, if the action combination consisting of the first sub-action indicated by the first sub-action data and the second sub-action indicated by the second sub-action data is not an invalid action combination, then the first sub-action and the second sub-action can be used as the target action, and the first sub-action and the second sub-action do not affect each other. Conversely, if the action combination consisting of the first sub-action indicated by the first sub-action data and the second sub-action indicated by the second sub-action data is an invalid action combination, then the first sub-action data and the second sub-action data need to be deleted from the action package data. This ensures the turning flexibility of the first virtual object while effectively alleviating the model's twitching and shaking phenomenon, and enhances the anthropomorphic performance of the first virtual object.

[0129] As an optional approach, the acquisition of environmental perception data corresponding to the first virtual object in the target virtual scene includes: when the target virtual object is visible to the first virtual object, acquiring the full attribute information of the first virtual object, the partial data, the relative position information between the target virtual object and the first virtual object, and the voiceprint perception information of the first virtual object; and determining the full attribute information, the partial data, the relative position information, and the voiceprint perception information together as the environmental perception data; when the target virtual object is not visible to the first virtual object, acquiring the full attribute information of the first virtual object, the relative position information between the target virtual object and the first virtual object, and the voiceprint perception information of the first virtual object; and determining the full attribute information, the relative position information, and the voiceprint perception information together as the environmental perception data, wherein the partial data is set to be occluded according to the preset parameters.

[0130] For example, let's take a shooting game as an example:

[0131] The aforementioned full attribute information may include, but is not limited to, 68-dimensional scalar values: the eyes and muzzle of the first virtual object, specifically including the ray length of each part of the target virtual object, the direction of the muzzle, the speed, the position of the muzzle and each part, the horizontal ray length of each part, the ray length above the head, the number of bullets, the range of virtual props, the health of each part, and the vegetation obstruction status; 56-dimensional Boolean values: whether the first virtual object (each part) is injured, whether the first virtual object (fires, aims, reloads, folds its gun, sprints, walks silently, turns to the left or right, stands, crouches, lies prone, jumps, or is obscured by smoke), whether the target virtual object is within the field of view of the first virtual object, whether the target virtual object (including each part) is exposed to the eyes and muzzle of the first virtual object, and whether the target virtual object (including each part) can be seen and shot by the first virtual object.

[0132] The aforementioned relative position information may include, but is not limited to: 77-dimensional scalar values: Euclidean three-dimensional distance, two-dimensional horizontal distance, absolute deflection angle, (including each part) relative deflection angle, (including each part) ΔX and ΔY in polar coordinate system, ΔX, ΔY, and ΔZ in Cartesian coordinate system, whether the target virtual object is within the field of view of the first virtual object, whether the target virtual object (including each part) is exposed to the eyes and muzzle of the first virtual object, and whether the target virtual object (including each part) can be seen and shot by the first virtual object.

[0133] The aforementioned voiceprint perception information may include, but is not limited to: 24-dimensional scalar values: gunshots, footsteps, grenade sounds in the air, grenade sounds on the ground, the distance between each voiceprint point and the first virtual object and the corresponding ΔX, ΔY, ΔZ, and the distance between the last voiceprint position of the target virtual object and the first virtual object.

[0134] In addition, previous decision information may include, but is not limited to: information on the various actions performed by the first virtual object in the previous decision game frame, which is filled with 0.0 in the first state request at the start of each game round.

[0135] In an exemplary embodiment, if the target virtual object is visible to the first virtual object, then the environmental perception data may include, but is not limited to, full attribute information, partial data, relative position information between the target virtual object and the first virtual object, voiceprint perception information of the first virtual object, and previous decision information of the first virtual object. The partial data here may be masked according to the preset parameters, including but not limited to the ray length from the eyes and muzzle of the target virtual object to each part of the first virtual object, the direction of the (muzzle) and the speed, the position (muzzle and each part), the horizontal ray length of each part, the ray length above the head, and other masked information. Several data values ​​in the partial data may be set to a special value -1.0 for padding.

[0136] In yet another exemplary embodiment, if the target virtual object is visible to the first virtual object, then the environmental perception data may include, but is not limited to, full attribute information, relative position information between the target virtual object and the first virtual object, voiceprint perception information of the first virtual object, and previous decision information of the first virtual object.

[0137] As an optional approach, the above-mentioned determination of action package data based on the aforementioned environmental perception data and control of the first virtual object to execute the target action according to the aforementioned action package data includes: sending the aforementioned environmental perception data to a pre-trained target action inference model through the aforementioned target application, wherein the aforementioned target action inference model is deployed on an inference server and the aforementioned target application is deployed on an application server; receiving the aforementioned action package data returned by the aforementioned target action inference model in response to the aforementioned environmental perception data, wherein the packet sending speed of the aforementioned environmental perception data to the aforementioned target action inference model is less than or equal to the packet return speed of the aforementioned target action inference model returning the aforementioned action package data; and controlling the aforementioned first virtual object to execute the target action according to the aforementioned action package data.

[0138] Optionally, in the embodiments of this application, the target action inference model may include, but is not limited to, a neural network model, and the inference server and application server may include, but are not limited to, cloud servers, distributed servers, etc. Furthermore, the inference server and the application server are two different servers. The application server can deploy the target application, but cannot deploy the target action inference model. The target action inference model needs to be deployed on the inference server to realize the prediction output of the action package data.

[0139] For example, environmental perception data is used as input to the target action reasoning model, and the model output is used as action packet data. Furthermore, the speed at which environmental perception data is sent is less than or equal to the speed at which action packet data is returned, so as to ensure the timeliness of the target action and enable the first virtual object to execute the target action in a timely manner, thereby achieving a dynamic balance between sending and receiving target packet data.

[0140] As an optional solution, the above method further includes: adjusting the transmission frequency of the environmental perception data to reduce the transmission speed when the packet sending speed exceeds the packet return speed; and controlling the first virtual object to execute behavior tree actions through behavior tree logic when the number of accumulated action packet data of the target action inference model exceeds a preset threshold, wherein the accumulated action packet data represents action packet data for which the control of the first virtual object to execute according to the accumulated action packet data has not been implemented.

[0141] Optionally, in this embodiment, the aforementioned stacked action package data can be understood as follows: after receiving the action package data, the application server does not promptly control the first virtual object to execute the target action corresponding to the action package data. For example, at the first time point, action package data A is received, indicating target action A; at the second time point, action package data B and action package data C are received, indicating target action B and target action C respectively. However, at the third time point, the first virtual object does not execute target action A, target action B, or target action C. In this case, the number of stacked action package data is 3. Assuming the preset quantity threshold is 2, where the value of the preset quantity threshold can be flexibly set, the behavior tree logic will control the first virtual object to execute behavior tree actions. Specifically:

[0142] The speed of receiving action package data can be accelerated by increasing the request interval frame rate to reduce the packet sending speed of environmental awareness data, or by optimizing the structure of the target action inference model to reduce inference time. When the number of accumulated action packages exceeds the aforementioned preset threshold, the system can be switched to Behavior Tree AI through business logic. Behavior Tree AI is an artificial intelligence technology used in game development and robot control. It describes and organizes the behavior of characters or entities through a tree structure. Behavior Tree AI can automatically select appropriate behaviors to respond to changes in the environment and player commands based on different situations and conditions. Behavior Tree AI typically consists of a series of nodes, including condition nodes, behavior nodes, and composite nodes. Condition nodes are used to determine whether the current environmental conditions meet the execution of a certain behavior, behavior nodes describe the specific behavior actions, and composite nodes are used to organize and manage the relationships and execution order between different nodes. By designing and configuring different nodes and the relationships between nodes, developers can flexibly construct complex behavioral logic, enabling characters or entities to exhibit more intelligent and natural behaviors. This ensures the stable and normal operation of the game logic and improves the usability, flexibility, and security of action package data.

[0143] As an optional solution, before determining the action package data based on the aforementioned environmental perception data and controlling the first virtual object to execute the target action according to the aforementioned action package data, the method further includes: performing distributed training on the initial action inference model to obtain the aforementioned target action inference model, wherein the game construction process of each of the aforementioned distributed training sessions is as follows: dividing the aforementioned target virtual scene into multiple key regions, wherein each of the aforementioned multiple key regions has a pre-set division probability; sampling the target key region from the aforementioned multiple key regions according to the aforementioned division probability; sampling the first sample position of the first sample virtual object with the aforementioned target key region as the center and a first pathfinding distance as the radius, wherein the aforementioned first pathfinding radius represents the maximum pathfinding radius pre-set for the aforementioned first sample virtual object; and sampling the first sample position of the first sample virtual object with the aforementioned first sample position as the center and a second pathfinding distance as the radius. The outer radius is defined as the distance, and the inner radius is defined as the third pathfinding distance. The second sample position of the second sample virtual object is obtained by sampling within the annulus. The second pathfinding radius represents the maximum pathfinding radius preset for the second sample virtual object, and the third pathfinding radius represents the minimum pathfinding radius preset for the second sample virtual object. If there are virtual obstacles between the first sample position and the second sample position, a first sample virtual object is generated at the first sample position, and a second sample virtual object is generated at the second sample position. Based on the first sample environment perception data corresponding to the first sample virtual object and the second sample environment perception data corresponding to the second sample virtual object, the first sample virtual object and the second sample virtual object are controlled to engage in a game until the game ends.

[0144] Optionally, in the embodiments of this application, the initial action reasoning model may include, but is not limited to, a neural network model, and may include, but is not limited to, training the initial action reasoning model using a distributed training method to obtain the target action reasoning model. For example, multiple game matches are set up, and the initial action reasoning model is used to determine the target action of the first virtual object in each match, and one key region is selected as the target key region with equal probability.

[0145] For example, Figure 10 This is a schematic diagram illustrating another optional virtual object control method according to an embodiment of this application, such as... Figure 10As shown, a key area is randomly selected as the target key area. Then, taking the firing muzzle of the virtual prop of the first sample virtual object in the target key area as the center, a first pathfinding distance is set. The first pathfinding distance can be determined by, but is not limited to, the distance between the muzzle position of the virtual prop of the first sample virtual object and the enemy position. Then, based on the first pathfinding distance, the aforementioned second pathfinding radius and third pathfinding radius are determined. A first sample virtual object is generated at the first sample position, and a second sample virtual object is generated at the second sample position. Specifically:

[0146] Figure 11 This is a schematic diagram illustrating another optional virtual object control method according to an embodiment of this application, such as... Figure 11 As shown, assuming the distance between the muzzle position of the virtual prop of the first sample virtual object and the enemy position is d, the angle between the straight line from the shooting point of the virtual prop of the first sample virtual object to the first sample virtual object and the horizontal straight line of the shooting point of the virtual prop of the first sample virtual object is set to θ. Here, θ represents the maximum angle at which the muzzle will deviate when a normal human player shoots. Shooting beyond the angle θ will become meaningless shooting. θ can be, but is not limited to, being defined as the arctangent function of 40 / 1500. In this case, the maximum pathfinding radius is tan(θ)*d. This application does not impose any restrictions on θ.

[0147] Furthermore, such as Figure 11 As shown, the shooting point of the virtual prop is the center of the circle, the second pathfinding distance is the outer radius, the third pathfinding distance is the inner radius, and the second sample position of the second sample virtual object is located on a ring with an inner radius equal to the second pathfinding distance and an outer radius equal to the third pathfinding distance. The maximum distance between the second sample position and the first sample position of the second sample virtual object is the aforementioned second pathfinding distance, and the minimum distance between the second sample position and the first sample position of the second sample virtual object is the aforementioned third pathfinding distance.

[0148] Furthermore, if there are virtual obstacles between the first sample location and the second sample location, these virtual obstacles may include, but are not limited to, buildings, vehicles, etc. in the game screen. In other words, the first sample virtual object can use the virtual obstacles as cover to avoid obstacles, or the first sample virtual object can use the virtual obstacles as cover to avoid obstacles. In this case, the first sample virtual object will be generated at the first sample location in the target area, and the second sample virtual object will be generated at the second sample location to start the game. The first sample virtual object and the second sample virtual object may engage in combat, conversation, or perform tasks, etc., in order to quickly determine the positions of the first sample virtual object and the second sample virtual object and improve the training efficiency of the initial action reasoning model.

[0149] As an optional approach, the above-mentioned distributed training of the initial action inference model to obtain the target action inference model includes: performing non-full processing on the first sample environment perception data and the second sample environment perception data to obtain first sample data and second sample data; when the first sample virtual object is visible to the second sample virtual object, performing feature extraction on the first sample data to obtain first sample features, wherein the first sample features include partial data of the second sample virtual object; when the second sample virtual object is not visible to the first sample virtual object, performing feature extraction on the second sample data to obtain second sample features, wherein the second sample features include partial data of the first sample virtual object after being masked by the preset parameters; inputting the first sample features and the second sample features into the initial action inference model respectively, determining the preset index values ​​corresponding to the first sample virtual object and the second sample virtual object respectively, and updating the initial action inference model according to the preset index values ​​to determine the target action inference model, wherein during the updating process of the initial action inference model, the output of the feature network layer associated with the partial data is uniformly masked as the target value.

[0150] Optionally, in this embodiment, the first sample environment perception data corresponds to the first sample virtual object, and the second sample environment perception data corresponds to the second sample virtual object. The aforementioned non-full processing can be understood as masking a portion of the data in the first sample environment perception data. For example, masking the health value of the second sample virtual object and setting it to -1.0 to obtain the first sample data; similarly, masking a portion of the data in the second sample environment perception data. For example, masking the health value of the second sample virtual object and setting it to -1.0 to obtain the second sample data. The aforementioned preset index values ​​can be understood as anthropomorphic indicators of the first and second sample virtual objects, such as sidestepping rate: being able to attack enemies after properly sidestepping; crouching rate: being able to conceal oneself by properly crouching and standing. Position; Evasion Rate: The percentage of time spent undetected in each game; Shooting Rate: The frequency of shooting effectively in each game; Silent Movement Rate: The frequency of using silent movement to hide one's voiceprint and avoid detection; Percentage of attacking after hiding for 3 seconds; Dodge after attacking for 2.5 seconds; Shooting Rate from Hidden Areas: Shooting from places where the enemy cannot see; Cover Utilization Rate: Effectively using pillars, boxes, and other cover for evasion; Movement Dispersion Rate: Flexible and varied movement strategies, avoiding staying in the same area for extended periods; Exposure Rate When Unable to Shoot: Virtual props are unavailable or the terrain obstructs the view, causing the virtual character to expose its head but be unable to shoot; Fire and Retreat Rate: The virtual character actively retreats to find cover while firing; Danger Zone Stay Rate: The virtual character is in a dangerous, open area without cover; Large Turning Rate: The virtual character's orientation cannot change abruptly to avoid poor anthropomorphism.

[0151] Understandable, Figure 12 This is a schematic diagram of another optional virtual object control method according to an embodiment of this application. The above-mentioned non-full processing method is as follows: Figure 12 As shown, taking the first sample virtual object as an example, a mask is added to the feature extraction network layer related to the second sample virtual object. When the second sample virtual object is visible, normal training logic is used. When the second sample virtual object is invisible, since some data in the second sample environment perception data corresponding to the second sample virtual object is filled with fake data -1.0, it does not have actual physical meaning. The output of the feature network layer connected to the current network layer is uniformly masked to 0, and the gradient is intercepted and backpropagated to the previous network layer. That is, network A is trained when the second sample virtual object is visible, and network B is trained when the second sample virtual object is invisible. The network parameters of A and B are frozen and cannot be trained or updated only when the second sample virtual object is invisible. This realizes a dynamic neural network structure throughout the training process. By performing non-full processing on the first sample environment perception data, the first sample data is obtained. By performing non-full processing on the second sample environment perception data, the second sample data is obtained. This can more realistically simulate the game perception input of human players and ensure the anthropomorphism of virtual objects at the input data level of the initial action inference model.

[0152] Furthermore, if the first sample virtual object is visible to the second sample virtual object, then the first sample feature can be determined by performing a feature extraction operation on the first sample data. The first sample feature includes the data that is not masked in the second sample environmental perception data. Similarly, if the second sample virtual object is visible to the first sample virtual object, then the second sample feature can be determined by performing a feature extraction operation on the second sample data. The second sample feature includes the data that is not masked in the first sample environmental perception data.

[0153] For example, the first sample features can be used as input to the initial action inference model to obtain the preset index value of the first sample virtual object, and the second sample features can be used as input to the initial action inference model to obtain the preset index value of the second sample virtual object. The preset index value is used to adjust and update the initial action inference model until the initial action inference model training is completed. For example, the preset index value is a cover utilization rate of 80%. When the cover utilization rate reaches 80%, the initial action inference model is considered to have been successfully trained and can obtain the expected preset index values ​​of the first sample virtual object and the second sample virtual object. The initial action inference model is then determined as the target action inference model, so as to achieve the purpose of accurately predicting action package data using the target action inference model.

[0154] As an optional solution, the acquisition of environmental awareness data corresponding to the first virtual object in the target virtual scene includes: acquiring first environmental awareness data corresponding to frame i when the service in the target virtual scene reaches frame i; acquiring second environmental awareness data corresponding to frame i+n when the service in the target virtual scene reaches frame i+n, wherein the service continues to execute preset business logic from frame i to frame i+n, and i and n are both positive integers; determining action package data based on the environmental awareness data and controlling the first virtual object to execute a target action according to the action package data includes: determining first action package data based on the first environmental awareness data and controlling the first virtual object to execute a first target action according to the first action package data at frame i+j; determining second action package data based on the second environmental awareness data and controlling the first virtual object to execute a second target action according to the second action package data at frame i+n+j, wherein j is a non-negative integer.

[0155] In an exemplary embodiment, the first environmental perception data corresponds to the data corresponding to the business logic executed by the first virtual object in the i-th frame, and the second environmental perception data corresponds to the data corresponding to the business logic executed by the first virtual object in the i+n-th frame. In other words, the generation time of the second environmental perception data is later than the generation time of the first environmental perception data.

[0156] Specifically, during the process of generating the corresponding first action package data using the first environmental perception data, the above-mentioned business will continuously execute the preset business logic. Similarly, during the process of generating the corresponding second action package data using the second environmental perception data, the above-mentioned business will also continuously execute the preset business logic. In other words, the above-mentioned preset business logic will not change with the first environmental perception data and the second environmental perception data. For example, the preset business logic is that the virtual props in the target virtual scene will increase by one every 2 frames. In the i-th frame, the number of virtual props is 10 and n is 2. In the i+2-th frame, the number of virtual props is 12.

[0157] Furthermore, the first action package data can be determined based on the first environmental perception data, so that the first virtual object executes the target action corresponding to the first action package data in the (i+j)th frame. The time interval between obtaining the first environmental perception data and the first virtual object executing the target action indicated by the first action package data can be represented as the (i+j)th frame. Here, j can be 0 or other positive integers. If j is 0, it means that the first virtual object executes the target action corresponding to the first action package data in the (i)th frame.

[0158] Similarly, the second action package data can be determined based on the second environmental perception data, so that the first virtual object executes the target action corresponding to the second action package data in the i+n+j frame. The time interval between obtaining the second environmental perception data and the first virtual object executing the target action indicated by the second action package data can be represented as the i+j frame. Here, j can be 0 or other positive integers. If j is 0, it means that the first virtual object executes the target action corresponding to the second action package data in the i+n frame.

[0159] also, Figure 13 This is a schematic diagram illustrating another optional virtual object control method according to an embodiment of this application, such as... Figure 13 As shown, during the initial action inference model training phase, when generating the corresponding first sample action package data using the first sample environment perception data, the aforementioned business logic will be terminated. In other words, during the time when the first sample virtual object executes the sample target action indicated by the first sample action package data after generating the first sample action package data based on the first sample environment perception data, the aforementioned business will be suspended. For example, the business's preset logic is that a virtual prop in the target virtual scene is added every 2 frames. In the i-th frame, the number of virtual props is 10, and n is 2. In the i+2-th frame, the number of virtual props is still 10. At this time, when the first sample virtual object executes the sample target action indicated by the first sample action package data, the preset business logic will be automatically restored. That is, in the next i+3-th frame, the number of virtual props will become 11.

[0160] As an optional approach, the above-mentioned determination of action package data based on the environmental perception data and control of the first virtual object to perform the target action according to the action package data includes: when the environmental perception data acquired in the i-th frame indicates that the first virtual object needs to be controlled to launch a virtual prop for the first time within a preset time period, obtaining the maximum aiming deviation angle preset for the first virtual object, where i is a positive integer; determining a first confidence region based on the distance between the first virtual object and the target virtual object and the maximum aiming deviation angle, where the radius of the first confidence region is r, and r is greater than 0; sampling a first confidence point from the first confidence region, and determining a second confidence region with the first confidence point as the center and r / x as the radius; sampling a second confidence point from the second confidence region, where the second confidence point represents the aiming crosshair of the first virtual object corresponding to the environmental perception data acquired in the i-th frame.

[0161] Optionally, in this embodiment of the application, taking a shooting game as an example, the aforementioned virtual items may include, but are not limited to, virtual shooting items, such as... Figure 11As shown, the first virtual object launches a virtual item for the first time within a preset time period. The maximum aiming deviation angle mentioned above can be expressed as... Figure 11 The θ shown represents the maximum angle at which the muzzle will deviate when a normal human player shoots. Shooting beyond the angle θ will become meaningless. θ can be, but is not limited to, the arctangent function defined as 40 / 1500. In this case, the maximum pathfinding radius is tan(θ)*d. This application does not impose any restrictions on θ.

[0162] For example, a first confidence region is obtained based on the distance between the first virtual object and the target virtual object and the maximum aiming deviation angle. This first confidence region can be understood as a coarse-grained target shooting confidence region in which the first virtual object can shoot at the second virtual object. Then, the fireability ballistics of each part of the first virtual object are tested, and the parts are sorted according to the hit priority of each part in the fireability test results. The body part with the highest hit priority is used as the aiming reference point.

[0163] Furthermore, the aiming reference point can be used as the first confidence point mentioned above. Assuming x is 3, then r / 3 represents the radius of the second confidence region. By sampling each point in the second confidence region, the second confidence point is obtained. This second confidence point can be understood as the aiming crosshair of the first virtual object aiming at the second virtual object in the i-th frame. This further narrows the shooting confidence domain in which the first virtual object can shoot the second virtual object. If the second confidence point is not within the first confidence region, the firing action of the first virtual object is terminated, and the system waits for the next acquisition of the first environmental perception data to ensure that the shooting action of the first virtual object is an effective shot within the second confidence region, thereby improving the shooting hit rate of the first virtual object.

[0164] As an optional approach, after obtaining the second confidence point from the second confidence region, the method further includes: if the environmental perception data obtained in the (i+m)th frame indicates that the first virtual object needs to launch a virtual prop for the kth time, a third confidence region is determined with the third confidence point determined in the (k-1)th frame as the center and r / x as the radius, where m is a positive integer and m is less than the preset duration, k is a positive integer greater than or equal to 2, and when k = 2, the third confidence point is the second confidence point; a target confidence point is obtained by sampling from the third confidence region, where the target confidence point represents the aiming crosshair of the first virtual object corresponding to the environmental perception data obtained in the (i+m)th frame.

[0165] Optionally, in this embodiment, the first virtual object can fire virtual items multiple times. For example, if k is 3, when the first virtual object fires a virtual item for the third time, the third confidence point determined in the second instance can be used as the center, and a third confidence region can be determined with r / x as the radius. This ensures that when the first virtual object fires a virtual item for the third time, the corresponding shot landing point is located in the third confidence region. In other words, the shooting confidence region determined in the previous instance is determined by the previous confidence point. For example, if K is 3, then the third confidence point is the second confidence point mentioned above, and when the first virtual object fires a virtual item for the second time, the corresponding shot landing point is located in the second confidence region.

[0166] Furthermore, by sampling each point in the third confidence region to obtain the target confidence point, which can be understood as the aiming reticle of the first virtual object aiming at the second virtual object in the i+m frame, when the first virtual object continuously fires virtual props, the horizontal and vertical offset angles of the virtual prop muzzle can be calculated based on each determined aiming reticle to simulate the recoil control shooting effect of a human player. At the same time, since the target confidence point is obtained by random sampling in the third confidence region, the muzzle produces a scattering effect, thereby eliminating the shooting performance of the first virtual object in the first-person perspective of the replay footage can restore the shooting effect of a human player. The human-like shooting is greatly improved compared with the original behavior tree AI shooting method, and the human-like shooting of the AI ​​is enhanced from the bottom execution side of the game client.

[0167] For example, the virtual object control method proposed in this application can be applied to the general training scenarios of non-player characters in games. Specifically, based on a large-scale reinforcement learning training framework and combined with a proximal policy optimization algorithm, a general training scheme for producing highly human-like AI for shooting games can be designed in complex 3D perception scenes and high-dimensional coupled structured action spaces. Human-like auxiliary enhancement processing is performed at three dimensions: the perception input end, the network decision end, and the underlying execution end. This achieves a training scheme for constructing high-intensity and highly human-like AI in complex 3D scenes and high-dimensional discrete action spaces. In other words, based on multimodal environmental perception input, a non-full perception method aligned with human players is adopted. To handle environmental information across different modalities, a dynamic neural network structure is used during model training to synchronously adapt to incomplete input information, ensuring the human-like nature of the perception layer. Furthermore, considering the gameplay habits of high-level players in shooting games, different action decision frequencies are configured for different action heads without sacrificing action completeness and flexibility. This prunes the action exploration space at the same moment, reducing the difficulty of strategy exploration and training. Additionally, it includes large and small circle-assisted shooting modes, combined with realistic atomic firing logic featuring recoil and scattering, truly achieving the muzzle shake and scattering performance of human players. Thus, at the game client's underlying execution level, it meets the requirements for highly human-like AI. These human-like enhancements make the AI's behavior closer to human operating habits, better assisting business teams in content production.

[0168] Specifically, assuming the game application scenario is a farm map, which is the map with the highest player activity, and the game has already deployed virtual objects with behavior tree versions for accompanying players and filling up match counts, the existing behavior tree version AI logic is easily recognized by players, causing them to lose the fun of the battle. Therefore, through the embodiments of this application, a high-intensity, highly human-like AI can be produced using reinforcement learning technology for deployment in high-end matches. The virtual object can move across the entire map, make reasonable use of cover in combat and kiting, move flexibly and naturally, fire with human-like behavior, and the game elimination replay can be aligned with the human player's perspective. The gameplay shows a high degree of human-likeness and diversity, truly achieving the goal of highly human-like accompanying players, thereby driving game activity and player participation.

[0169] Furthermore, since virtual objects can be deployed in any scene on the map, and any map area may be in combat with the player, high demands are placed on the model's generalization ability. The embodiments of this application have made detailed optimizations to the model training combat logic and scene development. This is reflected in the fact that the training scene covers the entire map, several key combat zones are defined, and one of the combat zones is selected with equal probability when each game starts to increase the probability of the main event occurring. The two sides in each combat zone are completely symmetrical, and the spawn points of the two sides are selected according to specific logic. In order to avoid one side being defeated at the start of the game, there must be obstacles between their spawn points to ensure sample diversity. During training, a combat zone is first randomly selected. For the selected combat zone, the combination of spawn points inside is further randomly selected. If one side is defeated or the game times out, the game ends, thereby improving training efficiency.

[0170] For example, the overall structure of the model training stage in this application embodiment may include, but is not limited to: a game client, a game environment that incorporates AIServiceSDK (Artificial Intelligence Service Software Development Kit) for interface calls, and a Dedicated Server (DS) that provides single-game services. After the game state is input into the network, it can output instruction action packets to the game client for execution to ensure the model training effect. The game client sends a start request to the AI ​​server, which returns the corresponding match configuration data based on training requirements. After obtaining the data, the game client constructs a match and generates a reinforcement learning agent (an entity that can perceive the environment and take actions). Each game frame sample in the request response is processed into state data and sent to the reinforcement learning framework for interaction. The AI ​​training service performs action prediction on the request frame of the current game, and the network actions are concatenated into executable action instructions and sent back to the client's game engine for execution. The AI ​​training service continues to wait for the request response of the next game frame sample. During the model training phase, business-side statistical indicators are periodically reported to detect the training effect. At the end of each game, the game client calls the relevant logic interface to destroy the agent of this game and automatically starts the next round of game with the same configuration. This process is repeated to collect data and provide it to the reinforcement learning framework for training.

[0171] In an exemplary embodiment, the model training phase may include, but is not limited to, using a synchronous inference mode. Each game (DS) during the training phase contains no human players, various business behavior tree AIs, or irrelevant game business logic. Only two training AIs are running in DS. The game business logic is very clean and pure, requiring minimal server performance for collision and ray detection. During the time interval between the AI ​​sending a request response for the current game frame and receiving the action response, DS is in a synchronous waiting state. Only after receiving the action response will DS logic continue to run. Therefore, the DS during the training phase is a continuous loop of state requests to action responses, with an action request made every n game frames. The synchronous inference method during model training allows the DS's game business logic and reinforcement learning training to run on the same timeline. The model deployment phase may include, but is not limited to, using an asynchronous inference mode. Each game (DS) during the deployment phase contains a large number of online players simultaneously. Numerous game interaction logics arise between players and between players and various AIs. If synchronous inference is still used during the deployment phase, each time DS sends a game frame request response, a large amount of data will be generated within DS to construct state data. Collision and ray detection consume a significant portion of server performance in related technologies, and this resource consumption increases proportionally with the number of AI deployed in each game after deployment. Furthermore, in synchronous inference, the game logic (DS) is in a synchronous waiting state during the process of receiving the action response after each state request. During this time, the game logic itself does not run; it only resumes after receiving the action response. This causes players to experience frame stuttering during this synchronous waiting window (which may only be tens of milliseconds). This stuttering is further amplified by the model inference time and the response delay caused by network fluctuations. This application's embodiment isolates the DS game logic from the inference service logic, ensuring that the inference service does not block the DS's own game logic. Simultaneously, it removes a large amount of collision and ray detection work that would normally only occur on the DS side, migrating it directly to the inference service. The inference service pre-stores a large amount of static collision data and performs ray detection queries on each basic state request response from the DS, thus avoiding the significant performance consumption caused by dynamically performing ray detection on the DS side each time. Developers only need to focus on the inference service logic, improving development efficiency.

[0172] For example, an asynchronous inference mode can be adopted during the online phase. Asynchronous inference has two timelines. The DS side sends basic environmental awareness data to the inference service at regular intervals (e.g., every n frames, n=4). The inference service side performs static collision and ray detection queries on the sent environmental awareness data according to the first-in-first-out principle, and then processes it into network state data for inference calculation. Then, it continues to return action instruction packets to the DS side according to the first-in-first-out principle. After the DS stably sends the state request response every n frames, it does not need to wait for the action response packet and continues to run its own game business logic. At the same time, it checks for action response packets every frame. If there is one, it immediately executes the latest action. Furthermore, the speed of action response packets from the inference service side is affected by factors such as inference computation time and network fluctuation latency. This can cause the response speed to fail to keep up with the packet sending speed. As the number of response packets accumulates, the timeliness of the actions predicted by the model deteriorates due to untimely response, rendering the inference service meaningless. Through the embodiments of this application, it is ensured that the response speed of the inference service is greater than or equal to the packet sending speed, at least achieving a dynamic balance between sending and receiving packets. Specifically, the response speed of action packet data can be accelerated by increasing the request interval frame number n to reduce the packet sending speed of environmental perception data, or by optimizing the model structure to reduce inference time. If the number of response packets is too large, it can be rolled back to behavior tree AI as a fallback. Asynchronous inference can prevent data from blocking the main thread, ensuring that factors such as neural network inference time and action response delay caused by network fluctuations do not affect other game business logic. It can also reduce the server load on the business side and ensure the security of model inference. If inference fails due to factors such as network fluctuations, it can be switched to ordinary behavior tree AI through business logic to ensure the stable and normal operation of the game logic. Using this hybrid reinforcement learning access logic has the characteristics of ease of use, flexibility, and security.

[0173] For example, in a three-dimensional scene similar to the real world, there are a large number of static, dynamic and environmental perception-related objects, and the details are very realistic. The game control panel also includes many joystick actions that can be executed simultaneously. In such a three-dimensional, complex and ever-changing game environment, if it is necessary to train a virtual object with high strength and strong anthropomorphism, it is necessary to perform anthropomorphism enhancement processing on three levels: environmental perception, network decision-making, and underlying execution. To this end, during the training process, the embodiments of this application can use the following three methods to assist in enhancing the anthropomorphism of the strategy.

[0174] (1) Environmental perception side: Incomplete transformation:

[0175] Specifically, extending from a 2D environment to a 3D environment adds state dimensions such as height and z-axis orientation to environmental entities, significantly increasing the state space. Large maps and realistic environmental scenes in 3D environments, such as doors, windows, stairs, buildings, and passageways, further increase the difficulty of 3D environment perception. Considering the completeness and computational complexity of 3D environment perception, a general, multimodal environment perception scheme is designed. This scheme comprises multiple scalar information items of progressively increasing complexity, three layers of surrounding rays, and a 2D depth map for complete 3D environment information. Simultaneously, to align with and replicate the game perception of human players, the strategy is made more human-like at the environmental perception input level. A non-cheating, non-full-scale perception method is adopted to simulate the perception input of real players to the game environment, as detailed below:

[0176] a) First virtual object information: full perception method, because humans also perceive their own information in the real world in full.

[0177] b) Target Virtual Object Information: This is a non-full-scale perception method, which needs to be discussed in two cases: when the target virtual object is visible and when it is invisible. This is because even in the real world, humans cannot fully obtain information about visible targets, let alone information about invisible targets. The specific approach is as follows:

[0178] 1) Visibility of the target virtual object: Obtain basic information about the target virtual object, including orientation and position, but it is necessary to remove overly precise information such as health, number of bullets and range of virtual items. This is because even human players in the real world do not know the real-time number of bullets and precise health of visible target enemies. Removing this almost cheating information not only further simulates and aligns with the perceptual input of human players, but also prevents the model from learning opportunistic behaviors by using this precise information, thereby polluting the model's strategy training and convergence.

[0179] 2) Target virtual object is invisible: Basic information about the target virtual object, such as whether the target virtual object can see the first virtual object, must be masked. Only the estimated position of the target enemy is used to replace the precise position of the target enemy. Based on this, the relative position and relative angle between them are calculated. Because when the target enemy is invisible, human players also do not know the various basic information of the target enemy or whether the target enemy can see them. All masked information is filled with a special value of -1.0 as a placeholder. This non-full perception method deeply restores the game perception of human players, ensuring human-likeness from the source of environmental perception, and providing data support and potential for subsequent model training of human-like strategies.

[0180] 3) Multiple Ray Information: A comprehensive perception approach, including depth maps and various body rays. As mentioned in the previous environment interaction module, the DS and inference services are isolated. Ray detection and querying are separated from the DS and moved to the inference service. This avoids excessive performance overhead on the business side during model training. Furthermore, ray detection and querying in the inference service can be divided into simple collision queries and complex collision queries. Simple collision query data is mainly used for path perception, where accuracy requirements are not high, and rays are relatively sparse, such as three-layer surround rays and horizontal rays from various parts. Complex collision query data is used for scenarios with higher detail requirements, such as depth maps, overhead rays, rays for determining visibility, and rays for determining shootability. Replacing rays with simple collision data for those with low accuracy requirements further reduces the performance consumption of the inference service.

[0181] Multimodal environmental perception-based non-full processing more realistically simulates the game perception input of human players, ensuring human-likeness at the neural network input data level. Furthermore, non-full processing is also performed at the training level within the neural network to synchronously adapt to the non-full perception input information. Specifically, this non-full processing method involves masking the feature extraction network layer related to target enemy information. When the target enemy is visible, normal training logic applies. When the target enemy is invisible, since the target enemy information is filled with fake data (-1.0) and lacks actual physical meaning, the output of the connected feature network layer is uniformly masked to 0, and gradients are intercepted and backpropagated to the previous network layer. In other words, network A is trained when the target enemy is visible, and network B is trained when the target enemy is invisible. The parameters of network A and B are frozen and cannot be trained or updated when the target enemy is invisible, achieving a dynamic neural network structure throughout the training process.

[0182] (2) Network Decision-Making Side - Multi-Frequency Prediction:

[0183] 3D game scenes contain multiple controllable joystick actions, and 3D environmental entities have richer actions than 2D environmental entities, such as controlling the viewpoint, crouching, jumping, etc. Different types of actions can be executed simultaneously, leading to an explosion of action space combinations, increasing the difficulty of AI exploration and training. To address the challenge of action space complexity, a multi-scale cornering design based on player data priors is constructed to simulate and approximate the player's viewing and aiming habits. Additionally, action scale designs under different virtual item states are constructed to simulate the differences in player vision and perception under different virtual items and scoped states, ultimately aligning with player behavior in the 3D environment. To reduce action complexity, a three-finger manipulation modeling scheme is adopted. The three-finger operation rules, representing the upper limit of human operation, are pruned to remove a large number of illegal and invalid parallel operations, thereby achieving a dual improvement in training efficiency and human-likeness. Based on the composability of human left and right hands and sub-actions, several orthogonal sub-actions are defined, which can occur simultaneously.

[0184] By combining the actual performance of the model with the gameplay habits of high-level human players, appropriate action decision frequencies are tuned for different action heads. This enhances the human-likeness of the strategy at the network decision-making level. Action sampling does not need to be performed simultaneously between different action heads, i.e., multi-frequency prediction. This allows for the selection of the optimal action combination from a vast action space in a shorter time. The advantages of multi-frequency prediction are as follows:

[0185] a) It takes into account both high-frequency and low-frequency motion sampling heads, ensuring steering flexibility while effectively alleviating model twitching and shaking, and enhancing the human-like performance of the model.

[0186] b) Prune the action exploration space at the same time to effectively reduce the difficulty of strategy exploration and training without losing the completeness and flexibility of the model's actions;

[0187] c) The sampling frequency of the action head can be manually adjusted according to the actual effect of the model, which is convenient for training differentiated style models;

[0188] Because shooting is a highly time-sensitive action, the decision-making frequency is crucial. A two-frame decision interval is used to capture potentially fast-moving enemy targets, allowing for a rapid reaction and engagement. However, some action heads don't require such a high decision-making frequency. For example, in movement actions like sprinting, immediately switching to still walking after sprinting would cause the model to appear jerky and unnatural, reducing its human-likeness. Therefore, in actual combat, sprinting should be sustained for at least a period of time. Movement action heads only need a lower prediction frequency during training. This allows different actions to have their own different sampling frequencies, improving human-likeness and significantly reducing the difficulty of model training. Furthermore, besides multi-frequency prediction between different action heads, the prediction frequency of the same action head also dynamically changes during training, reflecting differences between combat and non-combat scenarios. The prediction frequencies of actions strongly correlated with combat will change during model combat. In non-combat scenarios, AI focuses more on pathfinding and does not require fast, high-frequency action logic. However, in combat scenarios, victory or defeat can often be decided within seconds, so AI is more concerned with micro-management. At this time, actions related to movement require a higher decision frequency so that it can perform a human-like playstyle of coming out to fire two shots and then quickly retreating to hide, coming out to fire two more shots and then retreating to hide again. This is to better simulate the playstyle of advanced players. The dynamic multi-frequency prediction mechanism can not only improve the model's human-likeness, but also greatly reduce the difficulty of model training.

[0189] (3) Low-level execution side - shooting anthropomorphism:

[0190] For first-person shooter (FPS) games, shooting feel is a direct and crucial element for players, forming the basis for a successful game. In FPS games, shooting involves controlling a virtual crosshair to move the crosshair, ignoring bullet velocity and bullet drop. Players aim and fire when the crosshair aligns with an enemy model, dealing damage. The aiming actions of a typical human player can be categorized into four types:

[0191] 1) Positioning: Keep the crosshair on the moving target enemy by continuously measuring and inputting small amounts of data; 2) Jump-flash: Move the crosshair to the target point accurately in one go by instantaneous reaction and rapid measurement and input, and decide whether to use the scope (ADS or hip fire) according to the situation; 3) Stabilization: Ensure that the crosshair stays on the relatively fixed target by continuously fine-tuning the measurement input (recoil control); 4) Prediction: Move the crosshair to the point where the predicted bullet trajectory destination and the target enemy coincide by observing the target movement characteristics.

[0192] Aiming maneuvers vary depending on factors such as virtual items, distance, player actions, and aiming method. The emphasis on these four maneuvers also differs across games. Choosing the appropriate aiming method is a key aspect of a normal human player's gaming skills. Even when the crosshair is perfectly aimed at the target enemy, bullets sometimes fly in other directions. This bullet deviation is called bullet scattering. Generally, during sustained fire, bullets are not distributed at a single point in the crosshair but are randomly arranged within a certain range. This range is called the virtual item scattering, and the degree of scattering is related to the type of weapon, especially noticeable in hip-fire. In FPS games, after a normal human player fires a shot, the crosshair of a weapon does not remain in its original position for a short period but shifts slightly. This is due to the recoil of the weapon. Weapon recoil can be broadly divided into vertical recoil and horizontal recoil. The former causes the crosshair to shift upwards, while the latter causes it to shift left and right. The deviation of the crosshair will cause the bullet to deviate. The recoil of shooting items is one of the key factors affecting whether normal human players can hit the target. In order to reduce the impact of recoil, normal players need to control the recoil (moving the muzzle and bullet trajectory in the opposite direction) to make the fired bullets as concentrated as possible.

[0193] In conclusion, normal human players, when playing FPS games, first choose the appropriate aiming method based on the type of virtual item they have. During firing, due to the effects of bullet spread and virtual item recoil, they need to control recoil to improve their accuracy. This makes recoil control an essential skill for advanced FPS players and an important factor in evaluating the anthropomorphism of their behavior. Current behavior tree AI shooting methods do not support realistic virtual item movement logic and do not require ballistic detection. They only require manually configured hit rate parameters. If the hit rate parameter is configured to 100%, it can reliably hit enemies. If it is less than 100%, the behavior tree AI will hit enemies probabilistically according to the configured hit rate. In other words, the behavior tree AI's shooting logic is outcome-based; it does not cause damage through real bullets but directly deducts the target enemy's health based on probability. This outcome-based shooting leads to the following problems:

[0194] a) Result-oriented shooting does not have the effects of bullet scattering and recoil from shooting tools, and cannot simulate the bullet deviation distribution when the player shoots;

[0195] b) The inability to perform the recoil control shooting operation mentioned above by human players results in a serious lack of human-like shooting anthropomorphism in the first-person perspective of the behavior tree in the elimination replay;

[0196] c) The strength of the behavior tree is highly dependent on the adjustment of the hit rate parameter. If the hit rate is too high, the AI ​​will exhibit cheating behavior of getting headshots with every shot. If the hit rate is too low, it will affect the strength of the behavior tree AI, causing the behavior tree AI to lose its balance between strength and human-likeness, further affecting the player's gaming experience.

[0197] Therefore, in this embodiment, in order to realistically reproduce the recoil control and shooting operation of human players, a shooting logic supporting realistic ballistic detection with scattering and recoil effects of shooting props was developed. This truly aligns with the shooting style of human players at the game client execution level, ensuring that the anthropomorphic shooting is established at the underlying execution logic. Simultaneously, to achieve the recoil control and shooting operation of human players and simulate the scattering effect of shooting props, a shooting method with randomness within a small circle and truncation outside a large circle is used. Figure 11 In this context, θ is defined as the arctangent function of 40 / 1500. This value of 40 / 1500 is a reference value given by the business team based on game experience. d represents the distance between the AI's gun muzzle position and the target enemy's position, and θ represents the maximum angle at which the gun muzzle will deviate when a normal human player fires. Shooting beyond this angle becomes meaningless. Figure 11 The circle with a large radius represents the shooting confidence region M at different distances d. Shooting within this region is considered normal shooting, while shooting outside this region is considered invalid shooting. The AI ​​model determines whether to issue a firing command. If not, the neural network freely chooses the horizontal and numerical muzzle angles. If the model continuously predicts firing commands, manual assistance is needed to calculate the horizontal and vertical muzzle angles to continuously pull the crosshair back to the target enemy. Each pullback and shift should simulate the recoil effect of a human player's shooting, and each shift and pullback should simulate the recoil control effect of a human player's shooting. Specific steps include:

[0198] Using the target enemy's location as the first confidence point, such as Figure 11As shown, using r = (40 / 1500) × d as the first pathfinding radius, the coarse-grained target firing confidence region of the target enemy at the corresponding distance d is determined. The fireability trajectory of each body part of the target enemy is detected. The fireability detection results of each part are then sorted according to the hit priority of each part. The body part with the highest hit priority is used as the aiming base point p0, with p0 as the center of a small circle and a radius of r / x. The coefficient x can be adjusted according to the actual effect. This small circle radius is used to determine the fine-grained part firing confidence region N1. A point p1 is randomly selected in N1 as the aiming reticle for this shot. If point p1 exceeds the range of the coarse-grained target firing confidence region M, it is truncated to ensure that this shot is an effective shot within M. For the next shot, p1 is used as the center of the small circle for the next shot, and r is used as the aiming reticle. The fine-grained shooting confidence region N2 is determined with a radius of / 3. A point p2 is randomly selected in N2 as the shooting crosshair. If point p2 exceeds the range of the coarse-grained target shooting confidence region M, it is truncated to ensure that the next shot is an effective shot within M. This process is repeated to obtain the coordinates of the crosshair position for each shot. Based on these crosshair coordinates, the horizontal and vertical offset angles of the AI's muzzle are calculated for each shot. When continuously predicting firing, the recoil control shooting effect of a human player can be simulated. At the same time, due to the random point selection mechanism within the small circle, the muzzle produces a scattering effect, thus eliminating the shooting performance of the AI ​​in the first-person perspective of the replay footage can restore the shooting effect of a human player. The human-like shooting is greatly improved compared to the original behavior tree AI shooting method, truly enhancing the human-like shooting of the AI ​​from the bottom execution side of the game client.

[0199] For example, Figure 14 This is a schematic diagram illustrating another optional virtual object control method according to an embodiment of this application, such as... Figure 14 As shown, comparing the bullet hole distribution maps of human and AI target practice, it can be seen that the bullet holes fired by the AI ​​are more dispersed than those fired by human players. While ensuring effective shooting, the AI ​​can simulate the bullet spread effect and recoil control of human players. To more realistically reproduce human shooting effects and demonstrate the bullet spread and recoil control capabilities of real players, continuous shooting comparison experiments can be conducted. Figure 15 This includes AI shooting effects, where the bullet holes are close together between each two shots, within the range of normal player operation deviations, and also include some random disturbances. The shooting method of randomness within small circles and truncation outside large circles realizes the AI's muzzle pullback, deviation, and pullback in a continuous shooting process. It is precisely because of the continuity of this disturbance and deviation that it ultimately simulates the effect of shooting props scattering and recoil control when human players shoot, indirectly increasing the human-like behavior of AI shooting.

[0200] Furthermore, after training begins at the start of the game, DS acquires runtime state data and divides it into different modal perception data according to Table 1. The environmental information of each modality is independently extracted by different types of neural networks, which can more accurately and precisely represent the current game situation. Figure 15 This is a schematic diagram illustrating another optional virtual object control method according to an embodiment of this application, such as... Figure 15 As shown, the input to the inference network includes not only the output hidden state encoding information of the LSTM (Long Short-Term Memory Neural Network), but also global information. The global information mainly consists of absolute or relative information related to the first virtual object and the target virtual object. The target virtual object information is re-acquired using a full-scale perception approach, meaning that even after the target enemy is out of sight, the real information of the target enemy is still obtained. This global input feature, which includes "cheating" information, can reduce the value estimation bias and variance of the value network, which is beneficial to the policy convergence and stability. This approach is only used in the model training phase because in the actual online deployment phase, only the policy network is needed, not the value network. Autoregressive embeddings are introduced into the model design. Each action head samples actions and performs one-hot encoding (converting categorical variables into binary vectors) to create an embedding vector. This vector is then superimposed on the input vector of the current layer as the input for the next action head, creating a causal dependency layer by layer. This makes subsequent action head decisions more reasonable and natural. The actions sampled by the first action head are one-hot encoded, passed through a fully connected layer to form an embedding vector, and then concatenated with the LSTM output to form a new embedding vector, which is output to the second action head. The second action head samples actions and performs one-hot encoding to form an embedding vector, which is then superimposed on the second layer's input vector to form the third layer. The input is passed on in this way until the last action head. The advantages of the training structure design of this target action reasoning model are: 1) The highly complex neural network integrates multimodal input information such as lists, images, scalars, and booleans; 2) Through the design of autoregressive network action heads, the dependencies between the structured action spaces of the game are decoupled; 3) Interactive hierarchical action masking accelerates training efficiency and also improves the anthropomorphism of the strategy; 4) Different types of inputs are activated separately, which improves the accuracy of the network's input perception; 5) When the target enemy is not visible, the enemy's incomplete perception input is closer to the real world. The dynamic network structure adapts to this incomplete perception input, thereby ensuring that the model can learn human-like combat operations.

[0201] In summary, this application's embodiments present human-like enhancement techniques for training FPS game AI from three levels: environmental perception, network decision-making, and underlying execution. First, a non-full-scale perception method is used to align the training data source with human players' game perception methods, and combined with the non-full-scale processing logic within the neural network, a dynamic neural network structure is achieved during training. Second, the dynamic multi-frequency decision-making mechanism not only alleviates AI twitching and jittering, enhancing its human-like performance, but also significantly prunes the action space dimension, accelerating training efficiency while reducing training difficulty. Abandoning the outcome-based shooting approach used in behavior tree AI, an atomic shooting method with scattering and recoil similar to that of real players is developed, and a shooting method with large and small circles is used to truly achieve the scattering and recoil control effects of shooting props when human players shoot, further enhancing the human-like performance from the underlying execution level. Based on large-scale reinforcement learning training and the three anthropomorphic enhancement processing methods mentioned above, the recognition of complex obstacles is more flexible. Compared with the behavior tree version, the reinforcement learning AI can identify the location of cover and go to take cover without pre-constructing cover points, and can use various unconventional cover to avoid combat; 2) The behavior is flexible and varied, and there are different combat performances in the same fighting scenario, which can improve replayability; 3) The behavior space completely covers the player's behavior space, completes all actions that the player can perform, and thus goes to places that the behavior tree cannot reach and completes operations that the behavior tree cannot perform; 4) The firing anthropomorphism and behavior anthropomorphism have reached a level that the behavior tree cannot achieve. In the game elimination replay, it can also be aligned with the human player's perspective, truly achieving the purpose of virtual objects.

[0202] Furthermore, given the active user base and large-scale real-world battle data on the product line, it's possible to use human data to guide early behavioral patterns. In the initial training phase, the model may engage in many repetitive and ineffective explorations, easily falling into unpredictable special states that can have irreversible long-term effects on strategy learning, causing the strategy to become trapped in local optima and difficult to escape. Supervised learning allows the AI ​​to start from the level of a human player, and reinforcement learning can then be used to further explore strategies, improving exploration efficiency and effectively preventing early strategy learning from getting stuck in local optima. Subsequent human-like training will also be much more efficient. In the actual deployment of game interaction, asynchronous online inference will have communication latency issues, including but not limited to communication latency between the mobile device and the DS, and communication latency between the DS and the online inference service. By isolating the DS and the online inference service, within the time t from when the DS sends a state request s to the online inference service to when the DS receives the action response a, the DS has already run for time t. At this point, the received action a... The action may no longer be applicable to the latest DS state. In short, the timeliness of action 'a' has decreased. If this is compounded by communication latency between the mobile device and the DS, as well as network fluctuations, some time-sensitive sub-action responses will perform poorly when executed on the latest mobile device, leading to a decline in model inference performance. This is a drawback of asynchronous online inference. To address this, we can consider introducing state and action latency during model training, allowing the neural network to adapt to the impact of this latency and learn how to make decisions in the presence of latency. When actually deployed, the model can automatically adapt to the latency introduced by asynchronous inference, ensuring that the model's inference performance does not decline too much. Alternatively, in asynchronous inference, after the DS receives the action response, it can further reassemble the action on the DS side, adjusting and adapting the returned action instructions to the latest DS state. This manually offsets the impact of the decreased timeliness of actions caused by network communication latency between the DS and the online inference service. These two methods can alleviate the impact of asynchronous inference on the decline in model inference performance to some extent.

[0203] Furthermore, through the embodiments of this application, it is possible to accelerate the collection of training data for policy learning through large-scale distributed reinforcement learning; and to meet the anthropomorphic requirements of the model input end by performing non-full-scale perception processing on multimodal environment input; to make flexible and natural movements by performing frequency division decision-making on different action heads, thus meeting the anthropomorphic requirements of the model decision end; to achieve the muzzle shake effect when the player shoots by using large and small circles to assist shooting, thus meeting the anthropomorphic requirements of the game execution end; and to further accelerate the model adaptation and deployment by evaluating the anthropomorphicity of the model through a comprehensive reward and punishment mechanism and rich anthropomorphic indicators.

[0204] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0205] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0206] According to another aspect of the embodiments of this application, a control device for a virtual object for implementing the above-described control method for a virtual object is also provided. For example... Figure 16 As shown, the device includes:

[0207] Display module 1602 is used to display the target virtual scene, wherein the target virtual scene represents the game scene in which the target virtual object in the target application is located, and the target virtual object represents the object controlled by the target account logged into the target application;

[0208] The acquisition module 1604 is used to acquire environmental perception data corresponding to the first virtual object in the target virtual scene. The first virtual object represents an object controlled by the system. The part of the environmental perception data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object. If the target virtual object is not visible to the first virtual object, part of the data is set to be occluded by preset parameters.

[0209] The control module 1606 is used to determine action package data based on environmental perception data and control the first virtual object to perform the target action according to the action package data. The action package data includes different types of sub-actions predicted based on environmental perception data, and the target action is determined by combining different types of sub-actions.

[0210] As an optional solution, the above-mentioned device is used to determine action package data based on the environmental perception data in the following manner, and control the first virtual object to perform a target action according to the action package data: predicting a first sub-action based on the environmental perception data to obtain a first set of sub-action confidence scores, wherein one sub-action confidence score in the first set of sub-action confidence scores corresponds to one action parameter of the first sub-action; predicting a second sub-action based on the environmental perception data to obtain a second set of sub-action confidence scores, wherein one sub-action confidence score in the second set of sub-action confidence scores corresponds to one action parameter of the second sub-action. A set of action parameters corresponds to the first sub-action and the second sub-action, which represent sub-actions that do not affect each other. The action parameters corresponding to the confidence values ​​of the sub-actions in the first set of sub-action confidence values ​​that meet preset conditions are determined as the first sub-action data corresponding to the first sub-action. The action parameters corresponding to the confidence values ​​of the sub-actions in the second set of sub-action confidence values ​​that meet the preset conditions are determined as the second sub-action data corresponding to the second sub-action. The action package data is generated based on the first sub-action data and the second sub-action data, and the first virtual object is controlled to perform the target action according to the action package data.

[0211] As an optional solution, the above-mentioned device is used to determine action package data based on the environmental perception data in the following manner, and control the first virtual object to perform the target action according to the action package data: predicting the first sub-action according to the environmental perception data at a first frequency to obtain the confidence level of the first set of sub-actions, wherein the first frequency represents the preset execution frequency of the first sub-action; predicting the second sub-action according to the environmental perception data at a second frequency to obtain the confidence level of the second set of sub-actions, wherein the second frequency represents the preset execution frequency of the second sub-action, and the first frequency and the second frequency are different.

[0212] As an optional embodiment, the above-mentioned apparatus is further configured to: when the first sub-action represents an attack operation, predict the first sub-action according to the environmental perception data at a first frequency to obtain the confidence level of the first set of sub-actions; when the second sub-action represents a sub-action other than the attack operation, predict the second sub-action according to the environmental perception data at a second frequency to obtain the confidence level of the second set of sub-actions, wherein the first frequency is greater than the second frequency; when the first sub-action represents a pathfinding operation, predict the first sub-action according to the environmental perception data at a first frequency to obtain the confidence level of the first set of sub-actions; when the second sub-action represents a sub-action other than the pathfinding operation, predict the second sub-action according to the environmental perception data at a second frequency to obtain the confidence level of the second set of sub-actions, wherein the first frequency is less than the second frequency.

[0213] As an optional solution, the above-mentioned device is used to generate the action package data based on the first sub-action data and the second sub-action data in the following manner, and control the first virtual object to perform the target action according to the action package data: obtaining a predetermined invalid action combination; if the action combination composed of the first sub-action data and the second sub-action data does not belong to the invalid action combination, generating the action package data based on the first sub-action data and the second sub-action data, and controlling the first virtual object to perform the target action according to the action package data; if the action combination composed of the first sub-action data and the second sub-action data belongs to the invalid action combination, deleting the first sub-action data and the second sub-action data from the action package data.

[0214] As an optional solution, the above-mentioned device is used to acquire environmental perception data corresponding to a first virtual object in the target virtual scene in the following manner: when the target virtual object is visible to the first virtual object, acquire the full attribute information of the first virtual object, the aforementioned partial data, the relative position information between the target virtual object and the first virtual object, and the voiceprint perception information of the first virtual object; and determine the full attribute information, the aforementioned partial data, the aforementioned relative position information, and the aforementioned voiceprint perception information as the environmental perception data; when the target virtual object is not visible to the first virtual object, acquire the full attribute information of the first virtual object, the relative position information between the target virtual object and the first virtual object, and the aforementioned voiceprint perception information of the first virtual object; and determine the full attribute information, the aforementioned relative position information, and the aforementioned voiceprint perception information as the environmental perception data, wherein the aforementioned partial data is set to be occluded according to the aforementioned preset parameters.

[0215] As an optional solution, the above-mentioned device is used to determine action package data based on the environmental perception data in the following manner, and control the first virtual object to perform a target action according to the action package data: sending the environmental perception data to a pre-trained target action inference model through the target application, wherein the target action inference model is deployed on an inference server and the target application is deployed on an application server; receiving the action package data returned by the target action inference model in response to the environmental perception data, wherein the packet sending speed of the environmental perception data to the target action inference model is less than or equal to the packet return speed of the target action inference model; and controlling the first virtual object to perform the target action according to the action package data.

[0216] As an optional solution, the above-mentioned device is further used to: adjust the transmission frequency of the environmental perception data to reduce the transmission speed when the packet transmission speed exceeds the packet return speed; and control the first virtual object to execute behavior tree actions through behavior tree logic when the number of accumulated action packet data of the target action inference model exceeds a preset number threshold, wherein the accumulated action packet data represents action packet data that has not been implemented to control the first virtual object to execute according to the accumulated action packet data.

[0217] As an optional solution, the above-mentioned device is further used to: before determining the action package data based on the above-mentioned environmental perception data and controlling the first virtual object to execute the target action according to the above-mentioned action package data, perform distributed training on the initial action inference model to obtain the target action inference model, wherein the game construction process of each distributed training is as follows: dividing the target virtual scene into multiple key regions, wherein each of the multiple key regions has a pre-set division probability; sampling the target key region from the multiple key regions according to the above-mentioned division probability; sampling the first sample position of the first sample virtual object with the target key region as the center and a first pathfinding distance as the radius, wherein the first pathfinding radius represents the maximum pathfinding radius pre-set for the first sample virtual object; and sampling the first sample position of the first sample virtual object with the first sample position as the center and a second pathfinding distance as the radius. The outer radius is defined as the distance, and the inner radius is defined as the third pathfinding distance. The second sample position of the second sample virtual object is obtained by sampling within the annulus. The second pathfinding radius represents the maximum pathfinding radius preset for the second sample virtual object, and the third pathfinding radius represents the minimum pathfinding radius preset for the second sample virtual object. If there are virtual obstacles between the first sample position and the second sample position, a first sample virtual object is generated at the first sample position, and a second sample virtual object is generated at the second sample position. Based on the first sample environment perception data corresponding to the first sample virtual object and the second sample environment perception data corresponding to the second sample virtual object, the first sample virtual object and the second sample virtual object are controlled to engage in a game until the game ends.

[0218] As an optional approach, the aforementioned apparatus is used to perform distributed training on the initial action inference model to obtain the target action inference model in the following manner: performing non-full processing on the first sample environment perception data and the second sample environment perception data to obtain first sample data and second sample data; when the first sample virtual object is visible to the second sample virtual object, performing feature extraction on the first sample data to obtain first sample features, wherein the first sample features include partial data of the second sample virtual object; when the second sample virtual object is not visible to the first sample virtual object, performing feature extraction on the second sample data to obtain second sample features, wherein the second sample features include partial data of the first sample virtual object masked by the aforementioned preset parameters; inputting the first sample features and the second sample features into the initial action inference model respectively, determining preset index values ​​corresponding to the first sample virtual object and the second sample virtual object respectively, and updating the initial action inference model according to the preset index values ​​to determine the target action inference model, wherein during the updating process of the initial action inference model, the output mask of the feature network layer associated with the aforementioned partial data is uniformly masked as the target value.

[0219] As an optional solution, the above-mentioned device is used to acquire environmental awareness data corresponding to the first virtual object in the target virtual scene in the following manner: when the service in the target virtual scene reaches the i-th frame, acquire the first environmental awareness data corresponding to the i-th frame; when the service in the target virtual scene reaches the i+n-th frame, acquire the second environmental awareness data corresponding to the i+n-th frame, wherein the service continues to execute preset service logic from the i-th frame to the i+n-th frame, and i and n are both positive integers; the above-mentioned determination of action package data based on the environmental awareness data and control of the first virtual object to execute the target action according to the action package data includes: determining the first action package data based on the first environmental awareness data and controlling the first virtual object to execute the first target action according to the first action package data in the i+j-th frame; determining the second action package data based on the second environmental awareness data and controlling the first virtual object to execute the second target action according to the second action package data in the i+n+j-th frame, wherein j is a non-negative integer.

[0220] As an optional solution, the above-mentioned device is used to determine action package data based on the environmental perception data in the following manner, and control the first virtual object to perform the target action according to the action package data: when the environmental perception data acquired in the i-th frame indicates that the first virtual object needs to be controlled to launch a virtual prop for the first time within a preset time period, the maximum aiming deviation angle preset for the first virtual object is obtained, where i is a positive integer; a first confidence region is determined based on the distance between the first virtual object and the target virtual object and the maximum aiming deviation angle, where the radius of the first confidence region is r, and r is greater than 0; a first confidence point is sampled from the first confidence region, and a second confidence region is determined with the first confidence point as the center and r / x as the radius; a second confidence point is sampled from the second confidence region, where the second confidence point represents the aiming crosshair of the first virtual object corresponding to the environmental perception data acquired in the i-th frame.

[0221] As an optional solution, the above-mentioned device is further configured to: after obtaining the second confidence point from the second confidence region, if the environmental perception data obtained in the (i+m)th frame indicates that the first virtual object needs to launch a virtual prop for the kth time, determine the third confidence region with the third confidence point determined in the (k-1)th frame as the center and r / x as the radius, where m is a positive integer and m is less than the preset duration, k is a positive integer greater than or equal to 2, and when k=2, the third confidence point is the second confidence point; obtain the target confidence point from the third confidence region, where the target confidence point represents the aiming crosshair of the first virtual object corresponding to the environmental perception data obtained in the (i+m)th frame.

[0222] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0223] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0224] According to one aspect of this application, a computer program product is provided, the computer program product comprising a computer program.

[0225] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0226] Figure 17 A schematic block diagram of a computer system architecture for implementing an electronic device according to embodiments of the present application is shown.

[0227] It should be noted that, Figure 17 The computer system 1700 of the electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.

[0228] like Figure 17 As shown, the computer system 1700 includes a central processing unit (CPU) 1701, which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) 1702 or programs loaded from storage section 1708 into random access memory (RAM) 1703. The RAM 1703 also stores various programs and data required for system operation. The CPU 1701, ROM 1702, and RAM 1703 are interconnected via a bus 1704. An input / output interface 1705 (I / O interface) is also connected to the bus 1704.

[0229] The following components are connected to the input / output interface 1705: an input section 1706 including a keyboard, mouse, etc.; an output section 1707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1708 including a hard disk, etc.; and a communication section 1709 including a network interface card such as a local area network card, modem, etc. The communication section 1709 performs communication processing via a network such as the Internet. A drive 1710 is also connected to the input / output interface 1705 as needed. Removable media 1711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on the drive 1710 as needed so that computer programs read from them can be installed into the storage section 1708 as needed.

[0230] Specifically, according to embodiments of this application, the processes described in the various method flowcharts can be implemented as computer software programs. For example, embodiments of this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1709, and / or installed from removable medium 1711. When the computer program is executed by central processing unit 1701, it performs various functions defined in the system of this application.

[0231] In such an embodiment, the computer program can be downloaded and installed from a network via communication section 1709, and / or installed from removable media 1711. When the computer program is executed by central processing unit 1701, it performs various functions provided in the embodiments of this application.

[0232] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described control method for virtual objects is also provided. This electronic device may be... Figure 1 The terminal device or server shown. This embodiment uses this electronic device as an example for illustration. Figure 18 As shown, the electronic device includes a memory 1802 and a processor 1804. The memory 1802 stores a computer program, and the processor 1804 is configured to execute the steps of any of the above method embodiments via the computer program.

[0233] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.

[0234] Optionally, in this embodiment, the processor may be configured to execute the methods in the embodiments of this application via a computer program.

[0235] Alternatively, as those skilled in the art will understand, Figure 18 The structure shown is for illustrative purposes only. Figure 18 This does not limit the structure of the aforementioned electronic devices. For example, the electronic device may also include components that are more... Figure 18 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 18 The different configurations shown.

[0236] The memory 1802 can be used to store software programs and modules, such as the program instructions / modules corresponding to the virtual object control method and device in this embodiment. The processor 1804 executes various functional applications and data processing by running the software programs and modules stored in the memory 1802, thereby realizing the aforementioned virtual object control method. The memory 1802 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1802 may further include memory remotely located relative to the processor 1804, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1802 may be used, but is not limited to, to store environmental awareness data, action packet data, and other information. As an example, such as... Figure 18 As shown, the memory 1802 may include, but is not limited to, the display module 1602, acquisition module 1604, and control module 1606 of the control device for the virtual object. Furthermore, it may include, but is not limited to, other module units of the control device for the virtual object, which will not be described further in this example.

[0237] Optionally, the transmission device 1806 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1806 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1806 is a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0238] In addition, the aforementioned electronic device also includes: a display 1808 for displaying the aforementioned environmental perception data and action package data; and a connection bus 1810 for connecting the various module components in the aforementioned electronic device.

[0239] In other embodiments, the aforementioned terminal device or server can be a node in a distributed system, wherein the distributed system can be a blockchain system, which is a distributed system formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer network, and any form of computing device, such as a server, terminal, or other electronic device, can become a node in the blockchain system by joining this peer-to-peer network.

[0240] According to one aspect of this application, a computer-readable storage medium is provided, wherein a processor of an electronic device reads computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the electronic device to perform the virtual object control method provided in various alternative implementations of the control aspect of the virtual object described above.

[0241] Optionally, in this embodiment, the computer-readable storage medium described above may be configured to store methods for performing the embodiments of this application.

[0242] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0243] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0244] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more electronic devices to execute all or part of the steps of the methods described in the various embodiments of this application.

[0245] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0246] In the several embodiments provided in this application, it should be understood that the disclosed application can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between units or modules may be electrical or other forms.

[0247] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0248] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0249] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for controlling a virtual object, characterized in that, include: Displaying a target virtual scene, wherein the target virtual scene represents the game scene in which the target virtual object in the target application is located, and the target virtual object represents an object controlled by the target account logged into the target application; Acquire environmental perception data corresponding to a first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by the system, and the portion of the environmental perception data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object. If the target virtual object is not visible to the first virtual object, the portion of the data is set to be occluded by a preset parameter. Action package data is determined based on the environmental perception data, and the first virtual object is controlled to perform a target action according to the action package data. The action package data includes different types of sub-actions predicted based on the environmental perception data, and the target action is determined by combining the different types of sub-actions.

2. The method according to claim 1, characterized in that, The step of determining action package data based on the environmental perception data and controlling the first virtual object to perform the target action according to the action package data includes: Based on the environmental perception data, the first sub-action is predicted to obtain the confidence scores of the first set of sub-actions, wherein the confidence score of one sub-action in the first set of sub-actions corresponds to an action parameter of the first sub-action. The second sub-action is predicted based on the environmental perception data to obtain the confidence scores of the second set of sub-actions. In this set, the confidence score of one sub-action corresponds to a certain action parameter of the second sub-action. The first sub-action and the second sub-action represent sub-actions that do not affect each other. The action parameters corresponding to the confidence scores of the sub-actions in the first group of sub-actions that meet the preset conditions are determined as the first sub-action data corresponding to the first sub-action. The action parameters corresponding to the confidence scores of the sub-actions in the second group of sub-actions that satisfy the preset conditions are determined as the second sub-action data corresponding to the second sub-action. The action package data is generated based on the first sub-action data and the second sub-action data, and the first virtual object is controlled to perform the target action according to the action package data.

3. The method according to claim 2, characterized in that, The step of determining action package data based on the environmental perception data and controlling the first virtual object to perform the target action according to the action package data includes: Based on the environmental perception data, the first sub-action is predicted according to a first frequency to obtain the confidence level of the first group of sub-actions, wherein the first frequency represents the preset execution frequency of the first sub-action; Based on the environmental perception data, the second sub-action is predicted according to the second frequency to obtain the confidence level of the second set of sub-actions, wherein the second frequency represents the preset execution frequency of the second sub-action, and the first frequency and the second frequency are different.

4. The method according to claim 3, characterized in that, The method further includes: When the first sub-action represents an attack operation, the first sub-action is predicted according to the environmental perception data at a first frequency to obtain the confidence level of the first group of sub-actions; when the second sub-action represents other sub-actions besides the attack operation, the second sub-action is predicted according to the environmental perception data at a second frequency to obtain the confidence level of the second group of sub-actions, wherein the first frequency is greater than the second frequency. When the first sub-action represents a pathfinding operation, the first sub-action is predicted according to the environmental perception data at a first frequency to obtain the confidence level of the first set of sub-actions; when the second sub-action represents a sub-action other than the pathfinding operation, the second sub-action is predicted according to the environmental perception data at a second frequency to obtain the confidence level of the second set of sub-actions, wherein the first frequency is less than the second frequency.

5. The method according to claim 2, characterized in that, The step of generating the action package data based on the first sub-action data and the second sub-action data, and controlling the first virtual object to execute the target action according to the action package data, includes: Obtain a predetermined combination of invalid actions; If the action combination formed by the first sub-action data and the second sub-action data does not belong to the invalid action combination, the action package data is generated according to the first sub-action data and the second sub-action data, and the first virtual object is controlled to perform the target action according to the action package data; If the action combination consisting of the first sub-action data and the second sub-action data belongs to the invalid action combination, the first sub-action data and the second sub-action data shall be deleted from the action package data.

6. The method according to claim 1, characterized in that, The step of obtaining environmental perception data corresponding to the first virtual object in the target virtual scene includes: When the target virtual object is visible to the first virtual object, acquire the full attribute information of the first virtual object, the partial data, the relative position information between the target virtual object and the first virtual object, and the voiceprint perception information of the first virtual object; and determine the full attribute information, the partial data, the relative position information, and the voiceprint perception information together as the environmental perception data. When the target virtual object is not visible to the first virtual object, the full attribute information corresponding to the first virtual object, the relative position information between the target virtual object and the first virtual object, and the voiceprint perception information of the first virtual object are obtained; the full attribute information, the relative position information, and the voiceprint perception information are jointly determined as the environmental perception data, wherein the partial data is set to be occluded according to the preset parameters.

7. The method according to claim 1, characterized in that, The step of determining action package data based on the environmental perception data and controlling the first virtual object to perform the target action according to the action package data includes: The target application sends the environmental perception data to a pre-trained target action inference model, wherein the target action inference model is deployed on an inference server and the target application is deployed on an application server. The system receives the action packet data returned by the target action inference model in response to the environmental perception data, wherein the packet sending speed of the environmental perception data to the target action inference model is less than or equal to the packet return speed of the target action inference model returning the action packet data. Control the first virtual object to perform the target action according to the action package data.

8. The method according to claim 7, characterized in that, The method further includes: If the packet sending speed exceeds the packet return speed, the transmission frequency of the environmental perception data is adjusted to reduce the packet sending speed. If the number of accumulated action package data in the target action inference model exceeds a preset threshold, the first virtual object is controlled to perform a behavior tree action through behavior tree logic, wherein the accumulated action package data represents the action package data sent to the application server through the inference server.

9. The method according to claim 7, characterized in that, Before determining the action package data based on the environmental perception data and controlling the first virtual object to perform the target action according to the action package data, the method further includes: The initial action reasoning model is trained in a distributed manner to obtain the target action reasoning model. The game construction process for each distributed training session is as follows: The target virtual scene is divided into multiple key regions, wherein each key region in the multiple key regions has a pre-set division probability; The target key region is obtained by sampling from the multiple key regions according to the defined probability. Centered on the target key area and with a first pathfinding distance as the radius, the first sample position of the first sample virtual object is sampled, wherein the first pathfinding radius represents the maximum pathfinding radius preset for the first sample virtual object; Using the first sample position as the center, the second pathfinding distance as the outer radius, and the third pathfinding distance as the inner radius, the second sample position of the second sample virtual object is obtained by sampling within the ring. The second pathfinding radius represents the maximum pathfinding radius preset for the second sample virtual object, and the third pathfinding radius represents the minimum pathfinding radius preset for the second sample virtual object. When there are virtual obstacles between the first sample position and the second sample position, a first sample virtual object is generated at the first sample position and a second sample virtual object is generated at the second sample position. Based on the first sample environment perception data corresponding to the first sample virtual object and the second sample environment perception data corresponding to the second sample virtual object, the first sample virtual object and the second sample virtual object are controlled to play against each other until the game between the first sample virtual object and the second sample virtual object ends.

10. The method according to claim 9, characterized in that, The process of distributively training the initial action reasoning model to obtain the target action reasoning model includes: The first sample environmental perception data and the second sample environmental perception data are processed in a non-full manner to obtain the first sample data and the second sample data. When the first sample virtual object is visible to the second sample virtual object, feature extraction is performed on the first sample data to obtain the first sample feature, wherein the first sample feature includes part of the data of the second sample virtual object; When the second sample virtual object is not visible to the first sample virtual object, feature extraction is performed on the second sample data to obtain the second sample features, wherein the second sample features include part of the data of the first sample virtual object after being masked by the preset parameters; The first sample feature and the second sample feature are respectively input into the initial action inference model to determine the preset index values ​​corresponding to the first sample virtual object and the second sample virtual object, and the initial action inference model is updated according to the preset index values ​​to determine the target action inference model. During the update of the initial action inference model, the output mask of the feature network layer associated with the partial data is uniformly masked as the target value.

11. The method according to claim 1, characterized in that, The step of obtaining environmental awareness data corresponding to the first virtual object in the target virtual scene includes: obtaining first environmental awareness data corresponding to the i-th frame when the service in the target virtual scene reaches the i-th frame; obtaining second environmental awareness data corresponding to the i+n frame when the service in the target virtual scene reaches the i+n frame, wherein the service continues to execute preset business logic from the i-th frame to the i+n frame, and i and n are both positive integers; The step of determining action package data based on the environmental perception data and controlling the first virtual object to perform a target action according to the action package data includes: determining first action package data based on the first environmental perception data and controlling the first virtual object to perform a first target action according to the first action package data at frame i+j; determining second action package data based on the second environmental perception data and controlling the first virtual object to perform a second target action according to the second action package data at frame i+n+j, where j is a non-negative integer.

12. The method according to claim 1, characterized in that, The step of determining action package data based on the environmental perception data and controlling the first virtual object to perform the target action according to the action package data includes: In the case where the environmental perception data obtained in the i-th frame indicates that the first virtual object needs to be controlled to launch a virtual prop for the first time within a preset time period, the maximum aiming deviation angle preset for the first virtual object is obtained, where i is a positive integer; A first confidence region is determined based on the distance between the first virtual object and the target virtual object and the maximum aiming deviation angle, wherein the radius of the first confidence region is r, and r is greater than 0; A first confidence point is obtained by sampling from the first confidence region, and a second confidence region is determined with the first confidence point as the center and r / x as the radius; A second confidence point is obtained by sampling from the second confidence region, wherein the second confidence point represents the aiming crosshair of the first virtual object corresponding to the environmental perception data acquired in the i-th frame.

13. The method according to claim 12, characterized in that, After obtaining the second confidence point from the second confidence region, the method further includes: When the environmental perception data obtained in the (i+m)th frame indicates that the first virtual object needs to be controlled to launch a virtual prop for the kth time, the third confidence region is determined with the third confidence point determined in the (k-1)th frame as the center and r / x as the radius, where m is a positive integer and m is less than the preset duration, k is a positive integer greater than or equal to 2, and when k = 2, the third confidence point is the second confidence point; The target confidence point is obtained by sampling from the third confidence region, wherein the target confidence point represents the aiming crosshair of the first virtual object corresponding to the environmental perception data acquired in the (i+m)th frame.

14. A control device for a virtual object, characterized in that, include: The display module is used to display a target virtual scene, wherein the target virtual scene represents the game scene in which the target virtual object in the target application is located, and the target virtual object represents an object controlled by the target account logged into the target application; The acquisition module is used to acquire environmental perception data corresponding to a first virtual object in the target virtual scene, wherein the first virtual object represents an object controlled by the system, and the part of the environmental perception data associated with the target virtual object is determined by whether the target virtual object is visible to the first virtual object. If the target virtual object is not visible to the first virtual object, the part of the data is set to be occluded by a preset parameter. The control module is used to determine action package data based on the environmental perception data and control the first virtual object to perform a target action according to the action package data. The action package data includes different types of sub-actions predicted based on the environmental perception data, and the target action is determined by combining the different types of sub-actions.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program, wherein the computer program can be executed by an electronic device to perform the method described in any one of claims 1 to 13.

16. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program performs the steps of the method described in any one of claims 1 to 13.

17. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method described in any one of claims 1 to 13 through the computer program.