Control method, device, and apparatus for virtual object, and storage medium

CN116983639BActive Publication Date: 2026-09-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211111940.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-13
Publication Date
2026-09-29
Estimated Expiration
2042-09-13

AI Technical Summary

Technical Problem

[0004]但是,AI通常是根据预设的逻辑对虚拟对象进行控制,智能化程度较低,导致与其他玩家之间的沟通协作配合度低

Benefits of technology

[0027]本申请实施例中,为了使虚拟对局中的主控虚拟对象与AI虚拟对象能够实现对局互动交流,设计了一套自然指令与元指令之间的转换机制,在该机制下,AI能够将主控虚拟对象发送的自然指令转换为元指令,从而基于元指令控制AI虚拟对象执行相应操作,实现AI虚拟对象对主控虚拟对象的指令响应,或者,AI能够基于玩家视角生成元指令,并将元指令转换为自然指令并发送给主控终端,使真实玩家能够基于自然指令对主控对象进行控制,实现AI与真实玩家之间的双向互动交流,提高了AI虚拟对象的智能程度以及真实性,优化了己方存在AI虚拟对象时的对局体验。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116983639B_ABST
    Figure CN116983639B_ABST
Patent Text Reader

Abstract

Embodiments of the application disclose a virtual object control method and device, equipment and a storage medium, and belong to the technical field of artificial intelligence. The method comprises the following steps: receiving a natural instruction issued by a main control virtual object in a virtual environment, the natural instruction being triggered by a control terminal of the main control virtual object, the natural instruction being used to instruct other virtual objects belonging to the same camp as the main control virtual object to perform an operation, the other virtual objects including an AI virtual object controlled by AI; converting the natural instruction into a first meta-instruction, the first meta-instruction adopting an instruction format supported by AI identification; determining a first operation performed by the AI virtual object in the virtual environment based on the first meta-instruction and an AI observation value of the AI virtual object on the virtual environment; and controlling the AI virtual object to perform the first operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for controlling virtual objects. Background Technology

[0002] In multiplayer online games, players often need to team up to control virtual objects in a virtual environment and work together to complete the game.

[0003] In related technologies, artificial intelligence (AI) is used to replace some players in the team, and the game is completed through the cooperation between AI and the other players.

[0004] However, AI typically controls virtual objects based on preset logic, resulting in a low level of intelligence and poor communication, collaboration, and cooperation with other players. Summary of the Invention

[0005] This application provides a method, apparatus, device, and storage medium for controlling virtual objects, enabling two-way interactive communication between AI and real players. The technical solution is as follows:

[0006] On one hand, embodiments of this application provide a method for controlling a virtual object, the method comprising:

[0007] Receive natural instructions issued by the master virtual object in the virtual environment. The natural instructions are triggered by the control terminal of the master virtual object. The natural instructions are used to instruct other virtual objects belonging to the same camp as the master virtual object to perform operations. The other virtual objects include AI virtual objects controlled by AI.

[0008] The natural instructions are converted into first meta-instructions, which adopt an instruction format that supports AI recognition.

[0009] Based on the first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment, the first operation performed by the AI ​​virtual object in the virtual environment is determined;

[0010] Control the AI ​​virtual object to perform the first operation.

[0011] On the other hand, embodiments of this application provide a method for controlling a virtual object, the method comprising:

[0012] Based on the master control virtual object's master control observation value of the virtual environment, a third-dimensional instruction is generated. The third-dimensional instruction is generated by AI and adopts an instruction format that AI supports and can recognize. The AI ​​is used to control AI virtual objects in the virtual environment that belong to the same camp as the master control virtual object.

[0013] The third-dimensional instruction is converted into a natural instruction, which is used to instruct the master virtual object to perform an operation;

[0014] The natural command is sent to the control terminal of the master virtual object.

[0015] On the other hand, embodiments of this application provide a control device for a virtual object, the device comprising:

[0016] The instruction receiving module is used to receive natural instructions issued by the master virtual object in the virtual environment. The natural instructions are triggered by the control terminal of the master virtual object. The natural instructions are used to instruct other virtual objects belonging to the same camp as the master virtual object to perform operations. The other virtual objects include AI virtual objects controlled by AI.

[0017] The instruction conversion module is used to convert the natural instruction into a first meta-instruction, wherein the first meta-instruction adopts an instruction format that is supported and recognized by AI;

[0018] An operation determination module is used to determine, based on the first meta-instruction and the AI ​​observation values ​​of the AI ​​virtual object on the virtual environment, the first operation performed by the AI ​​virtual object in the virtual environment;

[0019] The control module is used to control the AI ​​virtual object to perform the first operation.

[0020] On the other hand, embodiments of this application provide a control device for a virtual object, the device comprising:

[0021] The instruction generation module is used to generate third-dimensional instructions based on the master observation values ​​of the master virtual object on the virtual environment. The third-dimensional instructions are generated by AI and adopt an instruction format that AI supports and can recognize. The AI ​​is used to control AI virtual objects in the virtual environment that belong to the same camp as the master virtual object.

[0022] The instruction conversion module is used to convert the third-dimensional instruction into a natural instruction, which is used to instruct the master virtual object to perform an operation.

[0023] The instruction sending module is used to send the natural instructions to the control terminal of the main control virtual object.

[0024] On the other hand, embodiments of this application provide a computer device including a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the virtual object control method as described above.

[0025] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the virtual object control method as described above.

[0026] On the other hand, embodiments of this application provide a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the virtual object control method provided in the above aspects.

[0027] In this embodiment, to enable interactive communication between the master virtual object and the AI ​​virtual object in a virtual game, a conversion mechanism between natural commands and meta-commands is designed. Under this mechanism, the AI ​​can convert the natural commands sent by the master virtual object into meta-commands, thereby controlling the AI ​​virtual object to perform corresponding operations based on the meta-commands, realizing the AI ​​virtual object's command response to the master virtual object. Alternatively, the AI ​​can generate meta-commands based on the player's perspective, convert the meta-commands into natural commands, and send them to the master terminal, enabling the real player to control the master object based on natural commands, realizing two-way interactive communication between the AI ​​and the real player, improving the intelligence and realism of the AI ​​virtual object, and optimizing the game experience when the player's side has an AI virtual object. Attached Figure Description

[0028] Figure 1 A schematic diagram of an implementation environment provided by one embodiment of this application is shown;

[0029] Figure 2 This is a flowchart of a virtual object control method provided in an exemplary embodiment of this application;

[0030] Figure 3 This is a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application;

[0031] Figure 4 This is a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application;

[0032] Figure 5 This is a schematic diagram of instruction conversion provided in an exemplary embodiment of this application;

[0033] Figure 6 This is a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application;

[0034] Figure 7 This is a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application;

[0035] Figure 8 This is an exemplary embodiment of the present application illustrating how a real player instructs an AI virtual object to perform an operation by inputting operation commands into a control terminal;

[0036] Figure 9 This is a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application;

[0037] Figure 10 This is a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application;

[0038] Figure 11 This is a schematic diagram illustrating an exemplary embodiment of this application, showing how AI actively instructs a real player to control a virtual object to perform operations;

[0039] Figure 12 This is a structural block diagram of a virtual object control device provided in an exemplary embodiment of this application;

[0040] Figure 13 This is a structural block diagram of a control device for a virtual object provided in another exemplary embodiment of this application;

[0041] Figure 14 A schematic diagram of the structure of a computer device provided in an exemplary embodiment of this application is shown. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0043] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0044] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0045] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learn-by-doing.

[0046] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0047] The solutions provided in this application involve technologies such as machine learning in artificial intelligence, and are specifically illustrated through the following embodiments.

[0048] Please refer to Figure 1 The diagram illustrates an implementation environment provided in one embodiment of this application. This implementation environment may include: a first terminal 110, a server 120, and a second terminal 130.

[0049] The first terminal 110 runs an application 111 that supports a virtual environment. This application 111 can be a multiplayer online battle arena (MOBA) program. When the first terminal runs the application 111, the user interface of the application 111 is displayed on the screen of the first terminal 110. The application 111 can be any type of game, such as a first-person shooter (FPS) game or a simulation game (SLG). In this embodiment, the application 111 is used as an example of a multiplayer online battle arena (MOBA) game. The first terminal 110 is the terminal used by the first user 112. The first user 112 uses the first terminal 110 to control a first virtual object located in the virtual environment to perform activities. The first virtual object can be referred to as the master virtual object controlled by the first user 112. The activities of the first virtual object include, but are not limited to, at least one of the following: adjusting body posture, crawling, walking, running, riding, flying, jumping, driving, picking up, shooting, attacking, throwing, and releasing skills. Schematic, the first virtual object is a first virtual character, such as a realistic character or an anime character.

[0050] The second terminal 130 runs an application 131 that supports a virtual environment. This application 131 can be a multiplayer online battle arena (MOBA) program. When the second terminal 130 runs the application 131, the user interface of the application 131 is displayed on the screen of the second terminal 130. This client can be any type of FPS game or SLG game; in this embodiment, the application 131 is an MOBA game as an example. The second terminal 130 is the terminal used by the second user 132. The second user 132 uses the second terminal 130 to control a second virtual object located in the virtual environment. The second virtual object can be referred to as the main virtual character controlled by the second user 132. Illustratively, the second virtual object is a second virtual character, such as a lifelike character or an anime character.

[0051] Optionally, the first virtual object and the second virtual object reside in the same virtual world. Optionally, the first virtual object and the second virtual object may belong to the same faction, the same team, the same organization, have a friend relationship, or have temporary communication permissions. Optionally, the first virtual object and the second virtual object may belong to different factions, different teams, different organizations, or have an adversarial relationship.

[0052] Optionally, in addition to the first virtual object and the second virtual object, the virtual objects in the same camp may also include AI virtual objects, wherein the AI ​​virtual objects are controlled by AI. Optionally, the AI ​​program controlling the AI ​​virtual objects may be pre-stored in the first terminal 110 and the second terminal 130 through application 111 and application 131. Optionally, when the first terminal 110 is running application 111, the AI ​​virtual objects may join the camp of the first virtual object through random matching or through user-selected settings.

[0053] Optionally, the applications installed on the first terminal 110 and the second terminal 130 are the same, or the applications installed on the two terminals are the same type of application on different operating system platforms (Android or iOS). The first terminal 110 can refer to one of a plurality of terminals, and the second terminal 130 can refer to another of a plurality of terminals. This embodiment only uses the first terminal 110 and the second terminal 130 as examples. The device types of the first terminal 110 and the second terminal 130 may be the same or different. The device types include at least one of the following: smartphones, tablets, e-book readers, Moving Picture Experts Group Audio Layer III (MP3) players, Moving Picture Experts Group Audio Layer IV (MP4) players, laptops, and desktop computers.

[0054] Figure 1 Only two terminals are shown in the diagram, but in different embodiments, multiple other terminals can access the server 120. Optionally, one or more terminals may also be terminals corresponding to developers, on which a development and editing platform for applications supporting virtual environments is installed. Developers can edit and update applications on these terminals and transmit the updated application installation packages to the server 120 via wired or wireless networks. The first terminal 110 and the second terminal 130 can download the application installation packages from the server 120 to update the applications.

[0055] The first terminal 110, the second terminal 130, and other terminals are connected to the server 120 via a wireless network or a wired network.

[0056] Server 120 includes at least one of the following: a single server, a server cluster consisting of multiple servers, a cloud computing platform, and a virtualization center. Server 120 is used to provide background services for applications that support a 3D virtual environment. Optionally, server 120 undertakes the primary computing task, and the terminal undertakes the secondary computing task; or, server 120 undertakes the secondary computing task, and the terminal undertakes the primary computing task; or, server 120 and the terminal use a distributed computing architecture for collaborative computing.

[0057] Optionally, the AI ​​program can also run on server 120. When creating a virtual game, server 120 can add AI virtual objects to the game and control the AI ​​virtual objects through the AI ​​program.

[0058] In an illustrative example, server 120 includes memory 121, processor 122, user account database 123, battle service module 124, and user-facing input / output interface (I / O interface) 125. The processor 122 loads instructions stored in server 120 and processes data in user account database 123 and battle service module 124. User account database 123 stores data about user accounts used by the first terminal 110, second terminal 130, and other terminals, such as user account avatars, nicknames, combat power indices, and the service area where the user account is located. Battle service module 124 provides multiple battle rooms for users to play, such as 1v1, 3v3, and 5v5 battles. User-facing I / O interface 125 establishes communication and exchanges data with the first terminal 110 and / or second terminal 130 via wireless or wired network.

[0059] For ease of description, the following embodiments will be described in detail using a computer device as the execution subject. This computer device can be a terminal or a server.

[0060] In one possible implementation, some virtual objects within the same faction are controlled by real players via a control terminal, while others are controlled by AI as AI virtual objects. Optionally, real players can instruct AI virtual objects to perform operations by inputting operation commands into the control terminal; similarly, AI can instruct real players to control virtual objects by sending commands to the control terminal. The specific process is described in the following embodiments.

[0061] Please refer to Figure 2 This document illustrates a flowchart of a virtual object control method provided in an exemplary embodiment of this application. This embodiment uses the method applied to a computer device as an example for illustration. The method includes the following steps:

[0062] Step 201: Receive natural instructions issued by the master virtual object in the virtual environment. The natural instructions are triggered by the control terminal of the master virtual object. The natural instructions are used to instruct other virtual objects belonging to the same camp as the master virtual object to perform operations. Among the other virtual objects are AI virtual objects controlled by AI.

[0063] A virtual object refers to an object that is controlled in a three-dimensional virtual environment. For example, a virtual object can be at least one of a virtual character, a virtual animal, an anime character, a virtual vehicle, and a virtual animal.

[0064] A virtual environment is a three-dimensional environment in which virtual objects reside during the operation of an application on a terminal. Optionally, in this embodiment, the virtual environment is observed using a camera model.

[0065] Optionally, the camera model automatically follows the virtual character in the virtual world. That is, when the position of the virtual object changes in the virtual world, the camera model changes its position accordingly, and the camera model always remains within a preset distance range of the virtual object. Optionally, during automatic following, the relative positions of the camera model and the virtual object do not change.

[0066] A camera model refers to a 3D model located around a virtual object in the virtual world. When using a first-person perspective, the camera model is located near or at the head of the virtual object. When using a third-person perspective, the camera model can be located behind the virtual object and bound to it, or it can be located at any position at a preset distance from the virtual object. The camera model allows observation of the virtual object from different angles. Optionally, when the third-person perspective is a first-person over-the-shoulder view, the camera model is located behind the virtual object (e.g., the head and shoulders of the virtual object). Optionally, in addition to first-person and third-person perspectives, other perspectives are also possible, such as a top-down perspective. When using a top-down perspective, the camera model can be located above the head of the virtual object; the top-down perspective is an aerial view of the virtual world. Optionally, the camera model is not actually displayed in the virtual world; that is, it is not displayed in the virtual world shown in the user interface.

[0067] A virtual match is a game in a virtual environment in which at least two virtual objects from different factions battle against each other. Optionally, a virtual match can be a two-player match or a multiplayer match. Optionally, each faction in a virtual match must contain at least two virtual objects, and at least one of these two virtual objects must be a master virtual object, and may also contain at least one AI virtual object. Optionally, the master virtual object is controlled by a real player through a control terminal.

[0068] Optionally, AI virtual objects are controlled by AI, and real players can send operation commands to other virtual objects (including AI virtual objects) of the same faction through a control terminal.

[0069] In one possible implementation, when a real player controlling the master virtual object instructs other virtual objects of the same faction as the master virtual object to perform operations via a control terminal, the control terminal of the master virtual object triggers natural instructions based on the real player to instruct other virtual objects of the same faction as the master virtual object to perform operations. In cases where the other virtual objects include AI virtual objects controlled by AI, the computer device receives natural instructions issued by the master virtual object in the virtual environment.

[0070] Optionally, natural commands can be voice commands, marker commands, text commands, or other commands obtained from real players; this application embodiment does not limit this.

[0071] For illustrative purposes, voice commands can be spoken by real players, such as "Gather in the top lane to clear the minion wave," text commands can be entered by real players on the control terminal, and marking commands can be marked by real players on the control terminal's display screen.

[0072] In an illustrative example, the control terminal of the master virtual object receives a voice command triggered by a real player. This voice command instructs other virtual objects belonging to the same faction as the master virtual object to perform an operation. Thus, the control terminal of the master virtual object triggers the voice command, and the computer device receives the voice command.

[0073] Step 202: Convert natural instructions into first-ary instructions, which adopt an instruction format that is supported by AI.

[0074] In one possible implementation, since natural commands are generated based on real players, direct recognition by AI may result in command recognition errors. Therefore, in order to facilitate AI to control AI virtual objects to perform corresponding operations based on natural commands, the computer device converts natural commands into first-level commands, wherein the first-level commands adopt a command format that AI supports for recognition.

[0075] Optionally, the instruction format that the AI ​​can recognize can be a computer language, and the first meta-instruction includes at least the instructions from natural instructions that instruct the AI ​​virtual object to perform an operation.

[0076] Step 203: Based on the first meta-instruction and the AI ​​observation values ​​of the AI ​​virtual object on the virtual environment, determine the first operation performed by the AI ​​virtual object in the virtual environment.

[0077] In one possible implementation, since different virtual objects are located in different positions in the virtual environment, the operations they perform in response to the same instruction may also be different. Therefore, in order to control the AI ​​virtual objects to perform operations, the computer device needs to first determine the AI ​​observation values ​​of the AI ​​virtual objects to the virtual environment.

[0078] In one possible implementation, the computer device determines the AI ​​observation value of the AI ​​virtual object to the virtual environment based on the AI ​​virtual object's observation perspective in the virtual environment.

[0079] Optionally, AI observations may include the state information of AI virtual objects in the virtual environment, the faction situation in the virtual environment, the state information of other virtual objects, the function information of virtual props, etc., which are not limited in this embodiment.

[0080] Optionally, AI observations are obtained based on the observation range of AI virtual objects in the virtual environment. Optionally, the observation range of AI virtual objects in the virtual environment can be determined by simulating the observation range of a real player in the virtual environment.

[0081] In one possible implementation, based on data on the observation range of a certain number of real players in the virtual environment, the computer device determines the observation range of the AI ​​virtual object in the virtual environment.

[0082] In one possible implementation, based on the first meta-instruction and the AI ​​observations of the AI ​​virtual object on the virtual environment, the computer device determines the first operation performed by the AI ​​virtual object in the virtual environment. Optionally, the first operation may include movement operations, skill release operations, item use operations, etc., that the AI ​​virtual object needs to perform.

[0083] Step 204: Control the AI ​​virtual object to perform the first operation.

[0084] In one possible implementation, the computer device controls the AI ​​virtual object to perform the first operation based on specific instruction information of the first operation.

[0085] Optionally, the server can determine the corresponding operation instruction based on the first operation and send it to each terminal, so that each terminal can control the AI ​​virtual object to perform the first operation according to the operation instruction; alternatively, the terminal can directly determine the corresponding operation instruction based on the first operation and control the AI ​​virtual object to perform the first operation. Since the AI ​​is consistent in each terminal, the first operation that controls the AI ​​virtual object to perform is also consistent.

[0086] Indicatively, the natural command is the voice command "Go to the top lane to clear the minion wave" input by a real player. In this case, the computer device controls the AI ​​virtual object to move to the top lane and perform the operation of clearing the minion wave.

[0087] In summary, in this embodiment of the application, in order to enable the master virtual object and the AI ​​virtual object to achieve interactive communication in the virtual game, a conversion mechanism between natural commands and meta commands is designed. Under this mechanism, the AI ​​can convert the natural commands sent by the master virtual object into meta commands, thereby controlling the AI ​​virtual object to perform corresponding operations based on the meta commands. This enables the AI ​​virtual object to respond to the commands of the master virtual object, improves the intelligence and realism of the AI ​​virtual object, and optimizes the game experience when the AI ​​virtual object is present on the player's side.

[0088] In one possible implementation, the AI ​​can instruct other virtual objects to perform corresponding operations by sending operation commands to the control terminals of other virtual objects belonging to the same faction as the AI ​​virtual object. In other words, the AI ​​actively instructs the real player to control the virtual objects to perform operations.

[0089] Please refer to Figure 3 This document illustrates a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application. This embodiment uses the method applied to a computer device as an example for illustration. The method includes the following steps:

[0090] Step 301: Based on the master control virtual object's master control observation value of the virtual environment, generate a third-dimensional instruction. The third-dimensional instruction is generated by AI and adopts an instruction format that AI supports and can recognize. The AI ​​is used to control AI virtual objects in the virtual environment that belong to the same camp as the master control virtual object.

[0091] Optionally, in a virtual game within a virtual environment, the same faction may include both master virtual objects controlled by real players via control terminals and AI virtual objects controlled by AI.

[0092] In one possible implementation, in order to generate operation instructions that conform to the master virtual object based on the state of the master virtual object in the virtual environment, the computer device controls the AI ​​to generate third-dimensional instructions based on the master virtual object's master observation value of the virtual environment, wherein the third-dimensional instructions adopt an instruction format that the AI ​​supports and can recognize.

[0093] In one possible implementation, the computer device determines the master observation value of the virtual object on the virtual environment by simulating the observation value of a real player on the virtual environment.

[0094] Optionally, the master observation value may include the state information of the master virtual object in the virtual environment, the faction situation in the virtual environment, the state information of other virtual objects, the function information of virtual props, etc., and this application embodiment does not limit this.

[0095] Optionally, since different virtual objects in the same faction play different roles and have different functions, the observation values ​​corresponding to different virtual objects are also different.

[0096] In one possible implementation, the computer device determines the observation value of the master virtual object on the virtual environment based on the function corresponding to the master virtual object in the camp, thereby controlling the AI ​​to generate third-dimensional instructions.

[0097] Step 302: Convert the third-dimensional instruction into a natural instruction, which is used to instruct the master virtual object to perform an operation.

[0098] In one possible implementation, since the third-dimensional instructions are generated by AI and use an AI-supported instruction format, real players cannot directly read the operation instruction information contained in the third-dimensional instructions. Therefore, in order to improve the accuracy of the expression of instruction intent and facilitate real players to obtain accurate information in the third-dimensional instructions, the computer device converts the third-dimensional instructions into natural instructions, which are used to instruct the master virtual object.

[0099] In one possible implementation, the computer device extracts the operation instruction information contained in the third-order instruction, which may include displacement instruction information, time instruction information, etc., and then combines the extracted operation instruction information into natural instructions.

[0100] Optionally, natural commands can be voice commands, marker commands, text commands, or other commands that can be obtained by real players; this application embodiment does not limit this.

[0101] In one possible implementation, the same tertiary instruction can be converted into different forms of natural instructions, and the computer device can randomly select one of these forms to convert the tertiary instruction.

[0102] In one possible implementation, since the effects of different natural instructions expressed by different events may vary, the computer device can select the form that best expresses the effect based on the event type. For example, for events such as clearing minion waves or contesting map resources, the computer device can use marking instructions to mark them in the virtual environment.

[0103] In one possible implementation, when the natural instruction is a text instruction, the computer device can fill the location information and event information in the third-order instruction into the text instruction template for sending to the control terminal; when the natural instruction is a voice instruction, the computer device can further convert the text information into voice to obtain a voice instruction.

[0104] Step 303: Send natural commands to the control terminal of the master virtual object.

[0105] In one possible implementation, the computer device sends natural instructions to the control terminal of the master virtual object.

[0106] Optionally, the control terminal may display natural instructions on the screen in the form of text, play natural instructions in the form of voice, or give specific instructions in the form of tags in the virtual environment. This application embodiment does not limit this.

[0107] In summary, in this embodiment of the application, in order to enable the main virtual object and the AI ​​virtual object to achieve interactive communication in the virtual game, a conversion mechanism between natural commands and meta commands is designed. Under this mechanism, the AI ​​can generate meta commands based on the player's perspective, convert the meta commands into natural commands, and send them to the main control terminal. This allows the real player to control the main object based on natural commands, realizing two-way interactive communication between the AI ​​and the real player. This improves the intelligence and realism of the AI ​​virtual object and optimizes the game experience when the player's side has an AI virtual object.

[0108] In one possible implementation, when a real player instructs an AI virtual object to perform an operation by inputting operation commands into a control terminal, since the necessary information for controlling the virtual object to perform the operation includes location, event, and operation duration, in order to more accurately convert the natural instructions into first-level instructions and control the AI ​​virtual object to accurately perform the first operation, the computer device needs to determine a first-level instruction based on the natural instructions, which includes at least location information, event information, and time information.

[0109] Please refer to Figure 4 This document illustrates a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application. This embodiment uses the method applied to a computer device as an example for illustration. The method includes the following steps:

[0110] Step 401: Receive natural instructions issued by the master virtual object in the virtual environment. The natural instructions are triggered by the control terminal of the master virtual object. The natural instructions are used to instruct other virtual objects belonging to the same camp as the master virtual object to perform operations. Among the other virtual objects are AI virtual objects controlled by AI.

[0111] For a detailed implementation of this step, please refer to step 201. This embodiment will not elaborate further here.

[0112] Step 402: Determine location information, event information, and time information based on natural instructions. Location information is used to indicate the location of the operation, event information is used to indicate the event to which the operation belongs, and time information is used to indicate the duration of the operation.

[0113] In one possible implementation, in order to obtain the specific operation information contained in the natural instructions that instruct other virtual objects to perform in the virtual environment, the computer device determines the location information, event information, and time information contained in the natural instructions.

[0114] Optionally, location information is used to indicate the location where the operation is performed. The virtual object can perform the corresponding displacement operation in the virtual environment based on the location information. For example, the location information could be "in the middle lane".

[0115] Optionally, event information is used to indicate the event to which the operation belongs. Virtual objects can execute corresponding events in the virtual environment based on event information. For example, event information could be "clear minion waves".

[0116] Optionally, time information is used to indicate the duration of the operation. The virtual object can perform the operation for the corresponding duration in the virtual environment based on the time information. For example, the time information could be "10 seconds".

[0117] In one possible implementation, since location information and event information can be specific and explicit, while time information may vary depending on the skills possessed by different virtual objects or the specific operation instructions received by different virtual objects, in order to more accurately instruct other virtual objects to perform corresponding operations, the computer device first determines configuration information, which includes the correspondence between location information, event information and time information. Schematic, the configuration information can be as shown in Table 1.

[0118] Table 1

[0119] On the road Clear the minion wave 20S jungle Competition for resources 30S

[0120] Optionally, different combinations of location information and event information can correspond to different time information. In one possible implementation, the time information can be changed based on the distance between the virtual object's current location in the virtual environment and the location information contained in the natural command, or it can be changed based on the difference between the virtual object's current force value and the force value required to execute the event contained in the natural command.

[0121] Optionally, the correspondence between location information, event information, and time information included in the configuration information can be preset. In one possible implementation, the computer device obtains time information based on the historical command response records of different master virtual objects for location information and event information.

[0122] In one possible implementation, the computer device first extracts location information and event information from natural instructions, and further determines the time information corresponding to the location information and event information based on configuration information.

[0123] In one possible implementation, in order to obtain location information and event information more accurately, the computer device divides the map of the virtual environment into grids and configures the event types of events in the virtual environment. Thus, when extracting location information and event information from natural instructions, the location information includes the grid identifier of the grid in which the operation is performed, and the event information includes the event type identifier of the event to which the operation is performed.

[0124] In one possible implementation, the computer device represents location information as "L" by dividing the virtual environment map into M = m × m grids, with each grid corresponding to a grid identifier. Event information is represented as "E," and N channels are set for each grid to represent the event type, corresponding to N event type identifiers. Further, the computer device represents location information using M-dimensional one-hot encoding and event information using N-dimensional one-hot encoding. When an instruction exists, the encoding for location L and event E is set to 1, and the rest are set to 0. For time information, the computer device represents time information as "T." By obtaining the historical instruction response times of different master virtual objects to location and event information, the octet of the historical instruction response time can be taken as the time information.

[0125] Step 403: The triple consisting of location information, event information, and time information is determined as the first instruction.

[0126] In one possible implementation, the computer device represents the location information, event information, and time information as a triple, thereby determining the triple as the first instruction.

[0127] In one possible implementation, the computer device represents location information as "L", event information as "E", and time information as "T", thus forming a triplet.<L,E,T> .

[0128] Indicative, such as Figure 5 As shown, the computer device can extract the location information and event information contained in the first meta-instruction 503 from the natural instruction 502 through the instruction converter 501, and perform strategy conversion on the first meta-instruction 503 according to the specific scenario in the virtual environment to obtain a macro strategy 504 applicable to the current game. The macro strategy 504 contains regional intent and resource intent, so that the regional intent "top lane" is represented by location information, and the resource intent "clearing the minion wave" in the virtual environment is represented by event information. The first meta-instruction 503 can be represented in the form of a triple as <top lane, clear the minion wave, 20s>.

[0129] Step 404: Input the first meta-instruction and the AI ​​observation values ​​of the AI ​​virtual object on the virtual environment into the policy network to obtain the first operation output by the policy network. The policy network is trained based on the sample meta-instruction, sample observation values ​​and sample operation.

[0130] In one possible implementation, since the operations performed by the virtual object under the same instruction may be different, the completion effect of the instruction may not be achieved. Therefore, in order to control the AI ​​virtual object more accurately and reasonably according to the first instruction, the computer device inputs the first instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment into the policy network, and outputs the first operation through the policy network.

[0131] The policy network is trained based on sample meta-instructions, sample observations, and sample operations. The computer device inputs the sample meta-instructions and sample observations into the policy network to obtain the predicted operations and determines the actual sample operations corresponding to the sample meta-instructions. Thus, the actual sample operations serve as supervision for the predicted operations, thereby training the policy network.

[0132] Optionally, sample meta-instructions, sample observations, and sample operations can be obtained from historical data of computer devices performing operations based on operation instructions and observations through different virtual objects.

[0133] Optionally, the policy network can also be trained using the Montgomery Algorithm (MGG), but this embodiment of the application does not limit this.

[0134] Step 405: Control the AI ​​virtual object to perform the first operation.

[0135] For a detailed implementation of this step, please refer to step 204. This embodiment does not limit the specific implementation of this step.

[0136] In the above embodiments, the computer device extracts information from natural instructions to obtain location information and event information, and determines the time information based on the correspondence between location information, event information and time information, thereby constituting a first meta-instruction. The first operation is obtained through the policy network based on the first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment, which improves the accuracy of converting natural instructions into first meta-instructions, improves the intelligence and realism of the AI ​​virtual object, and optimizes the game experience when the player has an AI virtual object.

[0137] In one possible implementation, multiple real players in the same faction may issue operation instructions to the AI ​​virtual object through their respective control terminals, resulting in multiple first-level instructions. However, the AI ​​virtual object can only execute one operation at a time. Therefore, the computer device needs to determine one of the at least two first-level instructions as the target instruction, and then control the AI ​​virtual object to execute the corresponding operation based on the target instruction.

[0138] Please refer to Figure 6 This document illustrates a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application. This embodiment uses the method applied to a computer device as an example for illustration. The method includes the following steps:

[0139] Step 601: Receive natural instructions issued by the master virtual object in the virtual environment. The natural instructions are triggered by the control terminal of the master virtual object. The natural instructions are used to instruct other virtual objects belonging to the same camp as the master virtual object to perform operations. Among the other virtual objects are AI virtual objects controlled by AI.

[0140] Step 602: Convert natural instructions into first-ary instructions, which adopt an instruction format that is supported by AI.

[0141] The specific implementation of steps 601 and 602 can be referred to steps 201 and 202, and this embodiment does not limit them.

[0142] Step 603: Determine the instruction value corresponding to at least two first-order instructions, where the instruction value is the virtual environment reward generated by the AI ​​virtual object executing the first-order instruction.

[0143] In one possible implementation, in the presence of at least two first-ary instructions, in order to obtain the optimal first-ary instruction from the at least two first-ary instructions, the computer device determines the instruction value corresponding to each of the at least two first-ary instructions based on the instruction value corresponding to the first-ary instruction, thereby selecting the at least two first-ary instructions, wherein the instruction value is the virtual environment reward generated by the AI ​​virtual object executing the first-ary instruction.

[0144] Optionally, virtual environment rewards can increase the faction's win rate in the game, and may include virtual economy, virtual combat power, virtual items, etc., which are not limited in this embodiment of the application.

[0145] To illustrate, taking MOBA games as an example, virtual environment rewards could include the growth of the faction's economy after executing commands, the probability of acquiring map resources, the state of virtual objects, and so on.

[0146] In one possible implementation, the computer device selects at least two first-order instructions via an instruction selector. First, the computer device inputs at least two first-order instructions and the AI ​​observation values ​​of the AI ​​virtual object on the virtual environment into the instruction selector, thereby obtaining the instruction values ​​corresponding to the at least two first-order instructions output by the instruction selector.

[0147] In one possible implementation, to improve the accuracy of the instruction value output by the instruction selector, the computer device trains the instruction selector. The computer device inputs sample meta-instructions and sample observations of the sample virtual environment into the instruction selector to obtain the instruction prediction value. Based on the actual virtual environment reward generated by executing the sample meta-instruction, the computer device determines the actual instruction value of the sample meta-instruction. The computer device then feeds back the actual instruction value to the instruction selector, using the actual instruction value as supervision for the instruction prediction value, thus training the instruction selector.

[0148] Optionally, the sample meta-instructions, sample observations of the sample virtual environment, and actual virtual environment returns can be obtained in advance by the computer device from historical data of virtual environment returns obtained by executing sample meta-instructions according to the control of the virtual object.

[0149] Optionally, since executing sample instructions requires a certain amount of operation time, the computer device uses the virtual environment rewards accumulated over a period of time during the execution of sample instructions as the actual value of the instructions to supervise the training of the instruction selector.

[0150] Optionally, when the computer device is a server, the instruction selector can be trained directly by the server, or the instruction selector can be trained by other devices and then deployed on the server. This embodiment only illustrates the case where the instruction selector is trained by the server.

[0151] Step 604: Determine the first meta-instruction with the highest instruction value as the target meta-instruction.

[0152] In one possible implementation, in order to improve the virtual environment reward generated by the AI ​​virtual object performing operations, the computer device determines the first-level instruction with the highest instruction value as the target first-level instruction based on the instruction values ​​corresponding to at least two first-level instructions.

[0153] To illustrate, taking a MOBA game as an example, if the virtual object executes two first-order instructions respectively, the faction's economy grows by 500 and 750 respectively, then the computer device will determine the first-order instruction corresponding to the economic growth of 750 as the target instruction.

[0154] Step 605: Based on the target meta-instruction and the AI ​​observation values ​​of the AI ​​virtual object on the virtual environment, determine the first operation performed by the AI ​​virtual object in the virtual environment.

[0155] In one possible implementation, the computer device determines the first operation to be performed by the AI ​​virtual object in the virtual environment based on the target meta-instruction and the AI ​​virtual object's AI observations of the virtual environment.

[0156] Similarly, the computer device can output the first operation through the policy network. In a specific implementation, step 404 can be performed, which will not be described in detail here.

[0157] Step 606: Control the AI ​​virtual object to perform the first operation.

[0158] For a detailed implementation of this step, please refer to step 204. This embodiment will not elaborate further here.

[0159] In the above embodiments, when there are at least two first-level instructions, the computer device determines the first-level instruction with the highest corresponding instruction value by judging the instruction value of each of the at least two first-level instructions, and then controls the AI ​​virtual object to perform the first operation based on the target first-level instruction. This increases the virtual value reward obtained by the AI ​​virtual object in performing the first operation, improves the accuracy of AI decision-making and the realism of the AI ​​virtual object, and optimizes the game experience when the player has an AI virtual object.

[0160] In one possible implementation, in addition to receiving natural instructions triggered by the control terminal of the master virtual object and converting them into first-level instructions, the AI ​​itself can also determine its own strategy based on the real-time game situation. When there are instructions sent by other objects and instructions obtained based on its own determined strategy, the computer device needs to make a comprehensive judgment to determine whether to respond to the instructions sent by other objects.

[0161] Please refer to Figure 7 This document illustrates a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application. This embodiment uses the method applied to a computer device as an example for illustration. The method includes the following steps:

[0162] Step 701: Receive natural instructions issued by the master virtual object in the virtual environment. The natural instructions are triggered by the control terminal of the master virtual object. The natural instructions are used to instruct other virtual objects belonging to the same camp as the master virtual object to perform operations. Among the other virtual objects are AI virtual objects controlled by AI.

[0163] Step 702: Convert natural instructions into first-ary instructions, which adopt an instruction format that AI supports for recognition.

[0164] The specific implementation of steps 701 and 702 can be referred to steps 201 and 202, and this embodiment does not limit them.

[0165] Step 703: Generate a second meta-instruction based on the AI ​​observation values ​​of the AI ​​virtual object on the virtual environment.

[0166] In one possible implementation, the computer device generates a second-dimensional instruction to instruct the AI ​​virtual object to perform an operation based on the AI ​​virtual object's AI observations of the virtual environment.

[0167] In one possible implementation, in order to make the second-level instructions more consistent with the state of the AI ​​virtual object in the virtual environment and to improve the win rate of the operation corresponding to the second-level instructions in the game, the computer device obtains the second-level instructions through the meta controller.

[0168] In one possible implementation, the computer device inputs the AI ​​observations of the AI ​​virtual object on the virtual environment into the meta controller, thereby obtaining the second meta instruction output by the meta controller.

[0169] In one possible implementation, to make the second meta-instruction output by the meta-controller more consistent with the AI ​​virtual object's AI observations of the virtual environment, the computer device trains the meta-controller. First, the computer device inputs the sample observations of the AI ​​virtual object on the sample virtual environment into the meta-controller to obtain the predicted meta-instructions, and uses the sample meta-instructions corresponding to the sample virtual environment as supervision for the predicted meta-instructions to train the meta-controller.

[0170] Optionally, the sample observations of the AI ​​virtual object on the sample virtual environment and the sample meta-instructions corresponding to the sample virtual environment can be obtained in advance by the computer device based on the historical operation instructions determined by the virtual object on the virtual environment.

[0171] Optionally, when the computer device is a server, the meta-controller can be trained directly by the server, or it can be trained by other devices and then deployed on the server. This embodiment only illustrates the case where the meta-controller is trained by the server.

[0172] Step 704: Determine the instruction value corresponding to the first-level instruction and the second-level instruction respectively. The instruction value is the virtual environment reward generated by the AI ​​virtual object executing the first-level instruction or the second-level instruction.

[0173] In one possible implementation, in order to determine the target meta-instruction from the first meta-instruction and the second meta-instruction, the computer device needs to first determine the instruction value corresponding to the first meta-instruction and the second meta-instruction, wherein the instruction value is the virtual environment reward generated by the AI ​​virtual object executing the first meta-instruction or the second meta-instruction.

[0174] Optionally, virtual environment rewards can increase the faction's win rate in the game, and may include virtual economy, virtual combat power, virtual items, etc., which are not limited in this embodiment of the application.

[0175] Similarly, the computer device can determine the instruction value corresponding to the first instruction and the second instruction through the instruction selector. For specific implementation, please refer to step 603. This embodiment will not be described in detail here.

[0176] Step 705: Determine the meta-instruction with the highest instruction value as the target meta-instruction.

[0177] In one possible implementation, in order to improve the virtual environment reward generated by the AI ​​virtual object performing operations, the computer device compares the instruction value corresponding to the first meta-instruction and the instruction value corresponding to the second meta-instruction, and determines the meta-instruction with the highest instruction value as the target meta-instruction.

[0178] Step 706: If the target meta-instruction is the first meta-instruction, determine the first operation performed by the AI ​​virtual object in the virtual environment based on the first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment.

[0179] Step 707: Control the AI ​​virtual object to perform the first operation.

[0180] The specific implementation of steps 706 and 707 can be found in steps 605 and 606, and will not be repeated here.

[0181] Step 708: If the target meta-instruction is a second meta-instruction, determine the second operation performed by the AI ​​virtual object in the virtual environment based on the second meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment.

[0182] In one possible implementation, when the target meta-instruction is a second meta-instruction, the computer device determines the second operation to be performed by the AI ​​virtual object in the virtual environment based on the second meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment.

[0183] Similarly, the computer device can output the second operation through the policy network. In a specific implementation, step 404 can be performed, which will not be described in detail here.

[0184] Step 709: Control the AI ​​virtual object to perform the second operation.

[0185] Furthermore, the computer device controls the AI ​​virtual object to perform the second operation based on the specific instructions of the second operation.

[0186] In the above embodiments, by generating second-level instructions from the AI ​​observation values ​​of AI virtual objects on the virtual environment, the feasibility of AI virtual objects actually executing instructions in the virtual environment is further considered. Based on the corresponding instruction value, the target meta-instruction is determined from the first and second meta-instructions, which further ensures the virtual value reward obtained by AI virtual objects in executing the corresponding operations, improves the intelligence and realism of AI virtual objects, and optimizes the game experience when one's own AI virtual objects are present.

[0187] Please refer to Figure 8 It illustrates a schematic diagram of an exemplary embodiment of this application, showing how a real player instructs an AI virtual object to perform operations by inputting operation commands to a control terminal.

[0188] In a schematic manner, the process by which a real player instructs an AI virtual object to perform operations by inputting operation commands into a control terminal can specifically include three stages: the meta-command conversion stage, the meta-command selection stage, and the meta-command execution stage.

[0189] First, during the meta-instruction translation phase, the real player's observations of the virtual environment are based on the master virtual object. The computer device sends text commands, signal commands, and other natural commands to the control terminal of the master virtual object, thereby receiving the natural commands issued by the master virtual object in the virtual environment, and converting the natural commands into first-level commands through the instruction converter 801. Meanwhile, computer devices generate AI observation values ​​O of the virtual environment based on AI virtual objects. t The second meta-instruction m is generated by the meta-controller 802. t Thus, the computer device will have at least one first-order instruction. Secondary instruction m t This forms a candidate meta-instruction set 803.

[0190] Secondly, during the meta-instruction selection phase, the computer device uses AI observations O of the virtual environment based on AI virtual objects. t The instruction selector 804 selects at least one first meta-instruction and one second meta-instruction from the candidate meta-instruction set 803, and determines the meta-instruction with the highest instruction value as the target meta-instruction C. t .

[0191] Furthermore, during the meta-instruction execution phase, the computer device will process the AI ​​virtual object's observations of the virtual environment. t and target meta-instruction C t The input is to the policy network 805, and the first operation a is output through the policy network 805. t Thus, the computer device controls the AI ​​virtual object to perform the first operation a. t .

[0192] Among them, the cumulative virtual environment reward obtained by the computer equipment during the process of controlling the AI ​​virtual object to perform the first operation. Supervised training of instruction selector 804 is performed, with feedback from a virtual environment. t Supervised training was performed on the policy network 805.

[0193] In one possible implementation, when the AI ​​instructs the real player to control the virtual object by sending commands to the control terminal, in order to more reasonably determine the third-dimensional command, the computer device further determines the master observation value of the master virtual object based on the visible range of the master virtual object in the virtual environment, and further specifies the process of converting the third-dimensional command into a natural command.

[0194] Please refer to Figure 9 This document illustrates a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application. This embodiment uses the method applied to a computer device as an example for illustration. The method includes the following steps:

[0195] Step 901: Based on the visible range of the master virtual object in the virtual environment, determine the master observation value of the master virtual object to the virtual environment.

[0196] In one possible implementation, since different virtual objects have different visible ranges in the virtual environment, the computer device needs to determine the master observation value of the master virtual object in the virtual environment based on the visible range of the master virtual object in the virtual environment.

[0197] Optionally, the visible range may vary based on the virtual object's position in the virtual environment, may be affected by the virtual object's harassed state, or may vary based on the virtual object's corresponding role in the faction. This application embodiment does not limit this.

[0198] To illustrate, the visible range of a virtual object located in the mid lane is the virtual environment of the mid lane area, and the visible range of a virtual object located in the jungle is the virtual environment of the jungle area.

[0199] In one possible implementation, the computer device determines the visible range of the main virtual object in the virtual environment based on the visible range of the virtual environment as observed by real players through the perspective of the main virtual object, using a large amount of real player data.

[0200] In one possible implementation, the computer device first determines the visible range of the master virtual object in the virtual environment, thereby obtaining status information, faction situation, virtual item information, etc. of other virtual objects from the visible range of the virtual environment, and then determines the master observation value of the master virtual object on the virtual environment.

[0201] Step 902: Input the master control virtual object's master control observation value of the virtual environment into the meta controller to obtain the third meta instruction output by the meta controller.

[0202] In one possible implementation, to make the third-dimensional instructions more consistent with the state of the master virtual object in the virtual environment and to improve the win rate of the operation corresponding to the third-dimensional instructions in the game, the computer device outputs the third-dimensional instructions through the meta-controller. First, the computer device inputs the master observation value of the master virtual object on the virtual environment into the meta-controller, thereby obtaining the third-dimensional instructions output by the meta-controller.

[0203] In one possible implementation, in order to make the third-level instructions output by the meta controller more consistent with the master virtual object's master observation of the virtual environment, the computer device trains the meta controller.

[0204] In one possible implementation, the computer device inputs the sample observations of the AI ​​virtual object on the sample virtual environment into the meta controller to obtain the prediction meta instructions, and uses the sample meta instructions corresponding to the sample virtual environment as supervision for the prediction meta instructions to train the meta controller.

[0205] Optionally, the sample observations of the AI ​​virtual object on the sample virtual environment and the sample meta-instructions corresponding to the sample virtual environment can be obtained in advance by the computer device based on the historical operation instructions determined by the virtual object on the virtual environment.

[0206] In one possible implementation, the third instruction is determined by a triple consisting of location information, event information, and time information. The location information is used to indicate the location of the operation, the event information is used to indicate the event to which the operation belongs, and the time information is used to indicate the duration of the operation.

[0207] Similarly, location information can be represented by "L", event information by "E", and time information by "T". That is, a third-ary instruction can be represented as a triple.<L,E,T> The specific representation of location information, event information, and time information can be found in step 402, and will not be elaborated upon in this embodiment.

[0208] Step 903: Extract location information and event information from the third-order instruction.

[0209] In one possible implementation, since the timing of a real player controlling a virtual object to perform an operation according to instructions is an uncertain variable, the computer device only extracts location and event information from the third-party instructions.

[0210] In one possible implementation, the computer device uses a grid to represent location information through M-dimensional one-hot encoding and event information through N-dimensional one-hot encoding. Thus, during the extraction of location and event information, the computer device obtains the actual location information and event type in the virtual environment based on the encoding information corresponding to the third-order instruction.

[0211] Step 904: Based on location information and event information, determine the target instruction type from the candidate instruction types. The candidate instruction types include at least one of voice instructions, text instructions, marker instructions, and signal instructions.

[0212] In one possible implementation, when sending natural commands to the control terminal of the master virtual object, different operation instruction information may be suitable for different instruction types. For example, when the location information of the master virtual object is displayed on the screen of the current master virtual object's control terminal, the computer device can directly instruct the real player by marking it in the virtual environment; when the location information of the master virtual object is not displayed on the screen of the current master virtual object's control terminal, the computer device can instruct the real player by voice or text.

[0213] In one possible implementation, in order to improve the effectiveness of instruction transmission, the computer device first determines the target instruction type from candidate instruction types based on location information and event information, wherein the candidate instruction types may include at least one of voice instructions, text instructions, marker instructions, and signal instructions.

[0214] Optionally, in order to improve the efficiency of real players in obtaining natural commands, the computer device can determine the target command type by combining voice commands with marked commands, or by combining text commands with marked commands. This application embodiment does not limit this.

[0215] In an illustrative example, the location information is "gather in the mid lane" and the event information is "clear the minion wave". The computer device can instruct the real player with the location information as a marker command and with the event information as a voice command.

[0216] Optionally, the target instruction type can also be preset by the actual player according to their own needs.

[0217] Step 905: Generate a natural instruction of the target instruction type based on the location information and event information.

[0218] In one possible implementation, the computer device generates natural instructions based on location information and event information, according to a determined target instruction type.

[0219] In an illustrative example, the location information is "gathering in the mid lane" and the event information is "clearing the minion wave". The computer device generates a marker signal for the mid lane position in the virtual environment map and converts "clearing the minion wave" into a voice signal.

[0220] Step 906: Send natural commands to the control terminal of the master virtual object.

[0221] For a detailed implementation of this step, please refer to step 303. This embodiment will not elaborate further here.

[0222] In the above embodiments, the computer device determines the master observation value of the master virtual object based on the visible range of the master virtual object in the virtual environment, thereby generating a third-dimensional instruction. By extracting the position information and event information in the third-dimensional instruction, the type of natural instruction is determined, thereby generating a natural instruction. This improves the efficiency of the master virtual object in executing the operation corresponding to the third-dimensional instruction, enhances the intelligence and realism of the AI ​​virtual object, and optimizes the game experience when the player has an AI virtual object.

[0223] In one possible implementation, in addition to generating third-order instructions based on the master observations of the virtual environment by the master virtual object, in order to consider the overall situation of the same faction, the AI ​​can also generate fourth-order instructions based on the AI ​​observations of the virtual environment by the AI ​​virtual object according to the real-time game situation, to provide strategy instructions to the master virtual object, so that the computer device can then comprehensively consider the instruction values ​​corresponding to the third-order instructions and the fourth-order instructions to determine the target instruction.

[0224] Please refer to Figure 10 This document illustrates a flowchart of a method for controlling a virtual object provided in another exemplary embodiment of this application. This embodiment uses the method applied to a computer device as an example for illustration. The method includes the following steps:

[0225] Step 1001: Based on the master control virtual object's master control observation value of the virtual environment, generate a third-dimensional instruction. The third-dimensional instruction is generated by AI and adopts an instruction format that AI supports and can recognize. The AI ​​is used to control AI virtual objects in the virtual environment that belong to the same camp as the master control virtual object.

[0226] For a detailed implementation of this step, please refer to step 301. This embodiment will not elaborate further here.

[0227] Step 1002: Generate the fourth meta-instruction based on the AI ​​observation values ​​of the AI ​​virtual object on the virtual environment.

[0228] In one possible implementation, the computer device generates a fourth-order instruction to instruct the master virtual object to perform an operation based on the AI ​​observations of the virtual environment by the AI ​​virtual object.

[0229] In one possible implementation, in order to make the fourth meta-instruction more consistent with the state of the AI ​​virtual object in the virtual environment and improve the win rate of the operation corresponding to the fourth meta-instruction in the game, the computer device obtains the fourth meta-instruction through the meta-controller. For specific implementation details, please refer to the specific description of generating the second meta-instruction in step 703. This embodiment will not be repeated here.

[0230] Step 1003: Determine the instruction value corresponding to the third-order instruction and the fourth-order instruction respectively. The instruction value is the virtual environment reward generated by the master virtual object executing the third-order instruction or the fourth-order instruction.

[0231] In one possible implementation, in order to determine the target meta-instruction from the tertiary and quaternary instructions, the computer device needs to first determine the instruction value corresponding to each of the tertiary and quaternary instructions, wherein the instruction value is the virtual environment reward generated by the master virtual object executing the tertiary or quaternary instruction.

[0232] Optionally, virtual environment rewards can increase the faction's win rate in the game, and may include virtual economy, virtual combat power, virtual items, etc., which are not limited in this embodiment of the application.

[0233] In one possible implementation, the computer device selects between tertiary and quaternary instructions via an instruction selector. Optionally, the computer device inputs the tertiary instruction and the master observation value of the master virtual object on the virtual environment into the instruction selector to obtain the instruction value of the tertiary instruction output by the instruction selector; and inputs the quaternary instruction and the master observation value of the master virtual object on the virtual environment into the instruction selector to obtain the instruction value of the quaternary instruction output by the instruction selector.

[0234] In one possible implementation, to improve the accuracy of the instruction value output by the instruction selector, the computer device trains the instruction selector. First, the computer device inputs sample meta-instructions and sample observations of the sample virtual environment into the instruction selector to obtain the instruction prediction value. Based on the actual virtual environment reward generated by executing the sample meta-instruction, the actual instruction value of the sample meta-instruction is determined. Thus, the instruction selector is trained using the actual instruction value as supervision for the instruction prediction value.

[0235] Optionally, since executing sample metadata instructions requires a certain amount of operation time, the computer device uses the virtual environment rewards accumulated over a period of time during the execution of sample metadata instructions as the actual value of the instructions.

[0236] Step 1004: Determine the meta-instruction with the highest instruction value as the target meta-instruction.

[0237] In one possible implementation, in order to improve the virtual environment reward generated by the execution of operations by the master virtual object, the computer device determines the meta-instruction with the highest instruction value as the target meta-instruction based on the instruction values ​​corresponding to the third and fourth meta-instructions.

[0238] Step 1005: If the target meta-instruction is a tertiary instruction, convert the tertiary instruction into a natural instruction.

[0239] In one possible implementation, when the target meta-instruction is a tertiary instruction, the computer device converts the tertiary instruction into a natural instruction. The specific conversion method can be referred to in steps 903 to 905, which will not be elaborated here in this embodiment.

[0240] Step 1006: Send natural commands to the control terminal of the master virtual object.

[0241] For a detailed implementation of this step, please refer to step 303. This embodiment will not elaborate further here.

[0242] Step 1007: If the target meta-instruction is a fourth-level meta-instruction, convert the fourth-level meta-instruction into a natural instruction.

[0243] In one possible implementation, when the target meta-instruction is a fourth-level instruction, the computer device converts the fourth-level instruction into a natural instruction. The specific conversion method can be referred to in steps 903 to 905, which will not be elaborated here in this embodiment.

[0244] Step 1008: If there are at least two AI virtual objects, determine the target AI virtual object corresponding to the fourth meta-instruction.

[0245] In one possible implementation, there may be at least two AI virtual objects in the same faction. In order to ensure that real players can know the AI ​​virtual object corresponding to the fourth-order instruction, the computer device determines the target AI virtual object corresponding to the fourth-order instruction.

[0246] In one possible implementation, the computer device determines the corresponding target AI virtual object through the AI ​​observation value corresponding to the fourth-order instruction.

[0247] Step 1009: Using the target AI virtual object as the instruction sender, send natural instructions to the control terminal of the master virtual object.

[0248] In one possible implementation, in order to enhance the realism of the interaction between virtual objects during the game, the computer device sends natural commands to the control terminal of the master virtual object, with the target AI virtual object as the command sender.

[0249] In an illustrative example, the computer device displays the icon of the virtual object corresponding to the target AI virtual object on the screen of the control terminal of the master virtual object, and presents natural commands to the real player through a dialog box.

[0250] In the above embodiments, in addition to generating third-order instructions based on the master observation values ​​of the virtual environment by the master virtual object, the computer device also generates fourth-order instructions based on the AI ​​observation values ​​of the virtual environment by the AI ​​virtual object, and determines the target instruction by comparing the instruction values ​​corresponding to the third-order instructions and the fourth-order instructions, thereby improving the intelligence and realism of the AI ​​virtual object and optimizing the game experience when the player has an AI virtual object.

[0251] Please refer to Figure 11 It illustrates a schematic diagram of an exemplary embodiment of this application, showing an AI actively instructing a real player to control a virtual object to perform an operation.

[0252] Indicatively, the process by which AI instructs other virtual objects to perform corresponding operations by sending operation instructions to the control terminals of other virtual objects belonging to the same camp as the AI ​​virtual object can specifically include three stages: the meta-instruction generation stage, the meta-instruction selection stage, and the meta-instruction sending stage.

[0253] First, in the meta-instruction generation phase, the computer device uses the master control virtual object to observe the master control values ​​of the virtual environment. Third-level instructions are generated through the meta controller 1101. Meanwhile, computer devices generate AI observation values ​​O of the virtual environment based on AI virtual objects. t The fourth instruction m is generated by the meta controller 1101. t Thus, the computer device will execute third-order instructions. and at least one fourth-order instruction m t To form a candidate meta-instruction set 1102.

[0254] Secondly, during the meta-instruction selection phase, the computer device uses the master control virtual object to observe the master control values ​​of the virtual environment. The instruction selector 1103 selects from the candidate meta-instruction set 1102 a third meta-instruction and at least one fourth meta-instruction, and determines the meta-instruction with the highest instruction value as the target meta-instruction C. t .

[0255] Furthermore, during the meta-instruction sending phase, the computer device uses the instruction converter 1104 to send the target meta-instruction C. t Convert it into natural commands and send the natural commands to the control terminal of the master virtual object.

[0256] Please refer to Figure 12 The diagram illustrates a structural block diagram of a control device for a virtual object provided in an exemplary embodiment of this application. The device includes:

[0257] The instruction receiving module 1201 is used to receive natural instructions issued by the master virtual object in the virtual environment. The natural instructions are triggered by the control terminal of the master virtual object. The natural instructions are used to instruct other virtual objects belonging to the same camp as the master virtual object to perform operations. The other virtual objects include AI virtual objects controlled by AI.

[0258] The instruction conversion module 1202 is used to convert the natural instruction into a first meta-instruction, wherein the first meta-instruction adopts an instruction format that is supported by AI.

[0259] The operation determination module 1203 is used to determine the first operation performed by the AI ​​virtual object in the virtual environment based on the first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment;

[0260] The control module 1204 is used to control the AI ​​virtual object to perform the first operation.

[0261] Optionally, the instruction conversion module 1202 includes:

[0262] An information determination unit is used to determine location information, event information, and time information based on the natural instruction. The location information is used to indicate the location of the operation, the event information is used to indicate the event to which the operation belongs, and the time information is used to indicate the duration of the operation.

[0263] The instruction determination unit is used to determine the triple consisting of the location information, the event information, and the time information as the first instruction.

[0264] Optionally, the information determining unit is used for:

[0265] Extract the location information and the event information from the natural instructions;

[0266] Based on the configuration information, the time information corresponding to the location information and the event information is determined. The configuration information includes the correspondence between the location information, the event information and the time information, and the time information is obtained based on the statistical analysis of historical instruction response records of different master virtual objects.

[0267] Optionally, the map of the virtual environment is divided into grids, and the event types of events in the virtual environment are configured;

[0268] The location information includes the grid identifier of the grid where the operation is performed, and the event information includes the event type identifier of the event to which the operation belongs.

[0269] Optionally, if at least two first meta-instructions exist, before determining the first operation performed by the AI ​​virtual object in the virtual environment based on the first meta-instructions and the AI ​​observations of the AI ​​virtual object on the virtual environment, the apparatus further includes:

[0270] The value determination module is used to determine the instruction value corresponding to at least two first meta instructions, wherein the instruction value is the virtual environment reward generated by the AI ​​virtual object executing the first meta instruction;

[0271] The instruction determination module is used to determine the first meta-instruction with the highest instruction value as the target meta-instruction.

[0272] The operation determination module 1203 is used for:

[0273] Based on the target meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment, the first operation performed by the AI ​​virtual object in the virtual environment is determined.

[0274] Optionally, the value determination module is used for:

[0275] The first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment are input into the instruction selector to obtain the instruction value output by the instruction selector.

[0276] Optionally, the device further includes:

[0277] The value prediction module is used to input sample meta-instructions and sample observations of the sample virtual environment into the instruction selector to obtain the instruction prediction value.

[0278] The value determination module is used to determine the actual value of the sample element instruction based on the actual virtual environment reward generated by executing the sample element instruction.

[0279] A training module is used to supervise the training of the instruction selector by using the actual value of the instruction as the predicted value of the instruction.

[0280] Optionally, before determining the first operation performed by the AI ​​virtual object in the virtual environment based on the first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment, the device further includes:

[0281] The instruction generation module is used to generate a second-level instruction based on the AI ​​observation values ​​of the AI ​​virtual object on the virtual environment;

[0282] The value determination module is used to determine the instruction value corresponding to the first meta-instruction and the second meta-instruction respectively. The instruction value is the virtual environment reward generated by the AI ​​virtual object executing the first meta-instruction or the second meta-instruction.

[0283] The instruction determination module is used to determine the meta-instruction with the highest instruction value as the target meta-instruction.

[0284] The operation determination module 1203 is used for:

[0285] When the target meta-instruction is the first meta-instruction, the first operation performed by the AI ​​virtual object in the virtual environment is determined based on the first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment.

[0286] Optionally, the operation determination module 1203 is further configured to:

[0287] When the target meta-instruction is the second meta-instruction, based on the second meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment, the second operation performed by the AI ​​virtual object in the virtual environment is determined;

[0288] The control module 1204 is used to control the AI ​​virtual object to perform the second operation.

[0289] Optionally, the instruction generation module is used for:

[0290] The AI ​​observation values ​​of the AI ​​virtual object on the virtual environment are input into the meta controller, and the second meta instruction is output by the meta controller.

[0291] Optionally, the device further includes:

[0292] The instruction prediction module is used to input the sample observation values ​​of the AI ​​virtual object to the sample virtual environment into the meta controller to obtain the predicted meta instructions;

[0293] The training module is further configured to train the meta-controller by using the sample meta-instructions corresponding to the sample virtual environment as supervision for the prediction meta-instructions.

[0294] Optionally, the operation determination module 1203 is used for:

[0295] The first meta-instruction and the AI ​​observations of the AI ​​virtual object on the virtual environment are input into the policy network to obtain the first operation output by the policy network. The policy network is trained based on sample meta-instructions, sample observations, and sample operations.

[0296] In summary, in this embodiment of the application, in order to enable the master virtual object and the AI ​​virtual object to achieve interactive communication in the virtual game, a conversion mechanism between natural commands and meta commands is designed. Under this mechanism, the AI ​​can convert the natural commands sent by the master virtual object into meta commands, thereby controlling the AI ​​virtual object to perform corresponding operations based on the meta commands. This enables the AI ​​virtual object to respond to the commands of the master virtual object, improves the intelligence and realism of the AI ​​virtual object, and optimizes the game experience when the AI ​​virtual object is present on the player's side.

[0297] Please refer to Figure 13 This illustrates a structural block diagram of a control device for a virtual object provided in another exemplary embodiment of this application, the device comprising:

[0298] The instruction generation module 1301 is used to generate third-dimensional instructions based on the master control virtual object's master control observation values ​​of the virtual environment. The third-dimensional instructions are generated by AI and adopt an instruction format that AI supports and can recognize. The AI ​​is used to control AI virtual objects in the virtual environment that belong to the same camp as the master control virtual object.

[0299] Instruction conversion module 1302 is used to convert the third-dimensional instruction into a natural instruction, wherein the natural instruction is used to instruct the master virtual object to perform an operation;

[0300] The instruction sending module 1303 is used to send the natural instruction to the control terminal of the master virtual object.

[0301] Optionally, the instruction generation module 1301 is used for:

[0302] Based on the visible range of the master virtual object in the virtual environment, determine the master observation value of the master virtual object to the virtual environment;

[0303] The master control virtual object's master control observation value of the virtual environment is input into the meta controller to obtain the third meta instruction output by the meta controller.

[0304] Optionally, the device further includes:

[0305] The instruction prediction module is used to input the sample observation values ​​of the AI ​​virtual object to the sample virtual environment into the meta controller to obtain the predicted meta instructions;

[0306] The training module is used to train the meta-controller by using the sample meta-instructions corresponding to the sample virtual environment as supervision for the prediction meta-instructions.

[0307] Optionally, before converting the third-order instruction into a natural instruction, the apparatus further includes:

[0308] The instruction generation module 1301 is used to generate a fourth-order instruction based on the AI ​​observation values ​​of the AI ​​virtual object on the virtual environment.

[0309] The value determination module is used to determine the instruction value corresponding to the third-order instruction and the fourth-order instruction respectively. The instruction value is the virtual environment reward generated by the master virtual object executing the third-order instruction or the fourth-order instruction.

[0310] The instruction determination module is used to determine the meta-instruction with the highest instruction value as the target meta-instruction.

[0311] The instruction conversion module 1302 is used for:

[0312] If the target meta-instruction is the third meta-instruction, the third meta-instruction is converted into the natural instruction.

[0313] Optionally, the instruction conversion module 1302 is further configured to:

[0314] If the target meta-instruction is the fourth meta-instruction, the fourth meta-instruction is converted into the natural instruction.

[0315] Optionally, the instruction sending module 1303 is used for:

[0316] In the presence of at least two AI virtual objects, the target AI virtual object corresponding to the fourth meta-instruction is determined;

[0317] Using the target AI virtual object as the instruction sender, the natural instruction is sent to the control terminal of the master virtual object.

[0318] Optionally, the value determination module is used for:

[0319] The third-order instruction and the master control virtual object's master control observation value of the virtual environment are input into the instruction selector to obtain the instruction value of the third-order instruction output by the instruction selector;

[0320] The fourth-element instruction and the master control virtual object's master control observation value of the virtual environment are input into the instruction selector to obtain the instruction value of the fourth-element instruction output by the instruction selector.

[0321] Optionally, the device further includes:

[0322] The value prediction module is used to input sample meta-instructions and sample observations of the sample virtual environment into the instruction selector to obtain the instruction prediction value.

[0323] The value determination module is used to determine the actual value of the sample element instruction based on the actual virtual environment reward generated by executing the sample element instruction.

[0324] The training module is used to train the instruction selector under supervision, using the actual value of the instruction as the predicted value of the instruction.

[0325] Optionally, the third instruction is determined by a triple consisting of location information, event information, and time information, wherein the location information is used to indicate the location of the operation, the event information is used to indicate the event to which the operation belongs, and the time information is used to indicate the duration of the operation.

[0326] The instruction conversion module 1302 is used for:

[0327] Extract the location information and the event information from the third-level instruction;

[0328] Based on the location information and the event information, a target instruction type is determined from the candidate instruction types, wherein the candidate instruction types include at least one of voice instructions, text instructions, marker instructions, and signal instructions;

[0329] Based on the location information and the event information, the natural instruction of the target instruction type is generated.

[0330] In summary, in this embodiment of the application, in order to enable the main virtual object and the AI ​​virtual object to achieve interactive communication in the virtual game, a conversion mechanism between natural commands and meta commands is designed. Under this mechanism, the AI ​​can generate meta commands based on the player's perspective, convert the meta commands into natural commands, and send them to the main control terminal. This allows the real player to control the main object based on natural commands, realizing two-way interactive communication between the AI ​​and the real player. This improves the intelligence and realism of the AI ​​virtual object and optimizes the game experience when the player's side has an AI virtual object.

[0331] It should be noted that the apparatus provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their implementation process can be found in the method embodiments, which will not be repeated here.

[0332] Please refer to Figure 14 This illustration shows a schematic diagram of a computer device provided in an exemplary embodiment of this application. The computer device 1400 can be implemented as a terminal or server as described in the above embodiments. Specifically, the computer device 1400 includes a Central Processing Unit (CPU) 1401, a system memory 1404 including a random access memory 1402 and a read-only memory 1403, and a system bus 1405 connecting the system memory 1404 and the CPU 1401. The computer device 1400 also includes a basic input / output system (I / O system) 1406 to facilitate information transfer between various devices within the computer, and a mass storage device 1407 for storing an operating system 1413, application programs 1414, and other program modules 1415.

[0333] The basic input / output system 1406 includes a display 1408 for displaying information and an input device 1409 for user input, such as a mouse or keyboard. Both the display 1408 and the input device 1409 are connected to the central processing unit 1401 via an input / output controller 1410 connected to the system bus 1405. The basic input / output system 1406 may also include the input / output controller 1410 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1410 also provides output to a display screen, printer, or other types of output devices.

[0334] The mass storage device 1407 is connected to the central processing unit 1401 via a mass storage controller (not shown) connected to the system bus 1405. The mass storage device 1407 and its associated computer-readable media provide non-volatile storage for the computer device 1400. That is, the mass storage device 1407 may include computer-readable media (not shown) such as a hard disk or drive.

[0335] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage media are not limited to the above-mentioned types. The system memory 1404 and mass storage device 1407 described above can be collectively referred to as memory.

[0336] The memory stores one or more programs, which are configured to be executed by one or more central processing units 1401. The one or more programs contain instructions for implementing the methods described above, and the central processing unit 1401 executes the one or more programs to implement the methods provided in the various method embodiments described above.

[0337] According to various embodiments of this application, the computer device 1400 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 1400 can be connected to a network 1412 via a network interface unit 1411 connected to the system bus 1405, or the network interface unit 1411 can be used to connect to other types of networks or remote computer systems (not shown).

[0338] This application also provides a computer-readable storage medium storing at least one instruction, which is loaded and executed by a processor to implement the virtual object control method provided in the above embodiments.

[0339] Optionally, the computer-readable storage medium may include ROM, RAM, solid-state drives (SSDs), or optical discs, etc. The RAM may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).

[0340] This application provides a computer program product including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the virtual object control method described in the above embodiments.

[0341] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0342] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for controlling a virtual object, characterized in that, The method includes: Receive natural instructions issued by the master virtual object in the virtual environment. The natural instructions are triggered by the control terminal of the master virtual object. The natural instructions are used to instruct other virtual objects belonging to the same camp as the master virtual object to perform operations. The other virtual objects include AI virtual objects controlled by AI. Based on the natural instructions, location information, event information, and time information are determined. The location information is used to indicate the location of the operation, the event information is used to indicate the event to which the operation belongs, and the time information is used to indicate the duration of the operation. The triple consisting of the location information, the event information, and the time information is determined as the first instruction, and the first instruction adopts an instruction format that supports AI recognition. Based on the first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment, the first operation performed by the AI ​​virtual object in the virtual environment is determined; Control the AI ​​virtual object to perform the first operation.

2. The method according to claim 1, characterized in that, The determination of location information, event information, and time information based on the natural instructions includes: Extract the location information and the event information from the natural instructions; Based on the configuration information, the time information corresponding to the location information and the event information is determined. The configuration information includes the correspondence between the location information, the event information and the time information, and the time information is obtained based on the statistical analysis of historical instruction response records of different master virtual objects.

3. The method according to claim 1, characterized in that, The map of the virtual environment is divided into grids, and the event types of events in the virtual environment are configured. The location information includes the grid identifier of the grid where the operation is performed, and the event information includes the event type identifier of the event to which the operation belongs.

4. The method according to claim 1, characterized in that, In the presence of at least two first meta-instructions, before determining the first operation performed by the AI ​​virtual object in the virtual environment based on the first meta-instructions and the AI ​​observations of the AI ​​virtual object on the virtual environment, the method further includes: Determine the instruction value corresponding to at least two first meta instructions, wherein the instruction value is the virtual environment reward generated by the AI ​​virtual object executing the first meta instruction; The first meta-instruction corresponding to the highest instruction value is determined as the target meta-instruction; The step of determining the first operation performed by the AI ​​virtual object in the virtual environment based on the first meta-instruction and the AI ​​virtual object's AI observations of the virtual environment includes: Based on the target meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment, the first operation performed by the AI ​​virtual object in the virtual environment is determined.

5. The method according to claim 4, characterized in that, Determining the instruction value corresponding to each of the at least two first meta instructions includes: The first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment are input into the instruction selector to obtain the instruction value output by the instruction selector.

6. The method according to claim 5, characterized in that, The method further includes: The sample meta-instruction and the sample observations of the sample virtual environment are input into the instruction selector to obtain the instruction prediction value; The actual value of the sample meta-instruction is determined based on the actual virtual environment feedback generated by executing the sample meta-instruction. The instruction selector is trained under supervision where the actual value of the instruction is used as the predicted value of the instruction.

7. The method according to claim 1, characterized in that, Before determining the first operation performed by the AI ​​virtual object in the virtual environment based on the first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment, the method further includes: Based on the AI ​​observations of the virtual environment by the AI ​​virtual object, a second meta-instruction is generated; Determine the instruction value corresponding to the first meta-instruction and the second meta-instruction respectively, wherein the instruction value is the virtual environment reward generated by the AI ​​virtual object executing the first meta-instruction or the second meta-instruction; The meta-instruction with the highest instruction value is identified as the target meta-instruction; The step of determining the first operation performed by the AI ​​virtual object in the virtual environment based on the first meta-instruction and the AI ​​virtual object's AI observations of the virtual environment includes: When the target meta-instruction is the first meta-instruction, the first operation performed by the AI ​​virtual object in the virtual environment is determined based on the first meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment.

8. The method according to claim 7, characterized in that, The method further includes: When the target meta-instruction is the second meta-instruction, based on the second meta-instruction and the AI ​​observation value of the AI ​​virtual object on the virtual environment, the second operation performed by the AI ​​virtual object in the virtual environment is determined; Control the AI ​​virtual object to perform the second operation.

9. The method according to claim 7, characterized in that, The generation of second-dimensional instructions based on the AI ​​observations of the virtual environment by the AI ​​virtual object includes: The AI ​​observation values ​​of the AI ​​virtual object on the virtual environment are input into the meta controller to obtain the second meta instruction output by the meta controller.

10. The method according to claim 9, characterized in that, The method further includes: The sample observations of the AI ​​virtual object on the sample virtual environment are input into the meta controller to obtain the prediction meta instruction; The meta-controller is trained using the sample meta-instructions corresponding to the sample virtual environment as supervision for the prediction meta-instructions.

11. The method according to claim 1, characterized in that, The step of determining the first operation performed by the AI ​​virtual object in the virtual environment based on the first meta-instruction and the AI ​​virtual object's AI observations of the virtual environment includes: The first meta-instruction and the AI ​​observations of the AI ​​virtual object on the virtual environment are input into the policy network to obtain the first operation output by the policy network. The policy network is trained based on sample meta-instructions, sample observations, and sample operations.

12. A method for controlling a virtual object, characterized in that, The method includes: Based on the master virtual object's master observation values ​​of the virtual environment, a third-dimensional instruction is generated. The third-dimensional instruction is generated by AI and adopts an instruction format that AI supports and can recognize. The AI ​​is used to control AI virtual objects in the virtual environment that belong to the same camp as the master virtual object. The third-dimensional instruction is determined by a triple consisting of location information, event information, and time information. The location information is used to indicate the location of the operation, the event information is used to indicate the event to which the operation belongs, and the time information is used to indicate the duration of the operation. The third-dimensional instruction is converted into a natural instruction, which is used to instruct the master virtual object to perform an operation; The natural command is sent to the control terminal of the master virtual object.

13. The method according to claim 12, characterized in that, The generation of third-dimensional instructions based on the master control observation values ​​of the virtual environment from the master control virtual object includes: Based on the visible range of the master virtual object in the virtual environment, determine the master observation value of the master virtual object to the virtual environment; The master control virtual object's master control observation value of the virtual environment is input into the meta controller to obtain the third meta instruction output by the meta controller.

14. The method according to claim 13, characterized in that, The method further includes: The sample observations of the AI ​​virtual object on the sample virtual environment are input into the meta controller to obtain the prediction meta instruction; The meta-controller is trained using the sample meta-instructions corresponding to the sample virtual environment as supervision for the prediction meta-instructions.

15. The method according to claim 12, characterized in that, Before converting the tertiary instruction into a natural instruction, the method further includes: Based on the AI ​​observations of the virtual environment by the AI ​​virtual object, a fourth-order instruction is generated; Determine the instruction value corresponding to each of the third-order instruction and the fourth-order instruction, wherein the instruction value is the virtual environment reward generated by the master virtual object executing the third-order instruction or the fourth-order instruction; The meta-instruction with the highest instruction value is identified as the target meta-instruction; The process of converting the third-order instruction into a natural instruction includes: If the target meta-instruction is the third meta-instruction, the third meta-instruction is converted into the natural instruction.

16. The method according to claim 15, characterized in that, The method further includes: If the target meta-instruction is the fourth meta-instruction, the fourth meta-instruction is converted into the natural instruction.

17. The method according to claim 16, characterized in that, Sending the natural command to the control terminal of the master virtual object includes: In the presence of at least two AI virtual objects, the target AI virtual object corresponding to the fourth meta-instruction is determined; Using the target AI virtual object as the instruction sender, the natural instruction is sent to the control terminal of the master virtual object.

18. The method according to claim 15, characterized in that, Determining the instruction value corresponding to the third-order instruction and the fourth-order instruction respectively includes: The third-order instruction and the master control virtual object's master control observation value of the virtual environment are input into the instruction selector to obtain the instruction value of the third-order instruction output by the instruction selector; The fourth-element instruction and the master control virtual object's master control observation value of the virtual environment are input into the instruction selector to obtain the instruction value of the fourth-element instruction output by the instruction selector.

19. The method according to claim 18, characterized in that, The method further includes: The sample meta-instruction and the sample observations of the sample virtual environment are input into the instruction selector to obtain the instruction prediction value; The actual value of the sample meta-instruction is determined based on the actual virtual environment feedback generated by executing the sample meta-instruction. The instruction selector is trained under supervision where the actual value of the instruction is used as the predicted value of the instruction.

20. The method according to claim 12, characterized in that, The process of converting the third-order instruction into a natural instruction includes: Extract the location information and the event information from the third-level instruction; Based on the location information and the event information, a target instruction type is determined from the candidate instruction types, wherein the candidate instruction types include at least one of voice instructions, text instructions, marker instructions, and signal instructions; Based on the location information and the event information, the natural instruction of the target instruction type is generated.

21. A control device for a virtual object, characterized in that, The device includes: The instruction receiving module is used to receive natural instructions issued by the master virtual object in the virtual environment. The natural instructions are triggered by the control terminal of the master virtual object. The natural instructions are used to instruct other virtual objects belonging to the same camp as the master virtual object to perform operations. The other virtual objects include AI virtual objects controlled by AI. The instruction conversion module is used to determine location information, event information, and time information based on the natural instruction. The location information is used to indicate the location of the operation, the event information is used to indicate the event to which the operation belongs, and the time information is used to indicate the duration of the operation. The triple consisting of the location information, the event information, and the time information is determined as the first instruction. The first instruction adopts an instruction format that supports AI recognition. An operation determination module is used to determine, based on the first meta-instruction and the AI ​​observation values ​​of the AI ​​virtual object on the virtual environment, the first operation performed by the AI ​​virtual object in the virtual environment; The control module is used to control the AI ​​virtual object to perform the first operation.

22. A control device for a virtual object, characterized in that, The device includes: The instruction generation module is used to generate tertiary instructions based on the master control virtual object's master control observation values ​​of the virtual environment. The tertiary instructions are generated by AI and adopt an instruction format that AI supports and can recognize. The AI ​​is used to control AI virtual objects in the virtual environment that belong to the same camp as the master control virtual object. The tertiary instructions are determined by a triple consisting of location information, event information, and time information. The location information is used to indicate the location of the operation, the event information is used to indicate the event to which the operation belongs, and the time information is used to indicate the duration of the operation. The instruction conversion module is used to convert the third-dimensional instruction into a natural instruction, which is used to instruct the master virtual object to perform an operation. The instruction sending module is used to send the natural instructions to the control terminal of the main control virtual object.

23. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one instruction, the at least one instruction being loaded and executed by the processor to implement the control method of the virtual object as described in any one of claims 1 to 11, or to implement the control method of the virtual object as described in any one of claims 12 to 20.

24. A computer-readable storage medium, characterized in that, The readable storage medium stores at least one instruction, which is loaded and executed by a processor to implement the control method of the virtual object as described in any one of claims 1 to 11, or to implement the control method of the virtual object as described in any one of claims 12 to 20.

Citation Information

Patent Citations

  • Smart-remote protocol

    CN102681840A

  • Metadata searching method, device and equipment and computer readable storage medium

    CN113010476A