Methods and apparatuses for commanding non-player character, and device and medium

By combining natural language commands with environmental awareness information, the problem of inaccurate non-player character control in existing technologies has been solved, enabling flexible and precise non-player character control and improving human-computer interaction efficiency and gaming experience.

WO2026031797A1PCT designated stage Publication Date: 2026-02-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102841
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-06-23
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

In existing technologies, it is difficult to achieve precise control of non-player characters through simple basic commands, especially when complex strategies or precise task execution are required. Players find it difficult to accurately control non-player characters to reach virtual props or specific locations.

Method used

By combining natural language commands with environmental awareness information, and through collaborative processing between the terminal and the server, the system identifies target entities and controls non-player characters to perform virtual activities related to the target. It utilizes a large language model and a hot word system to improve the accuracy of speech recognition and enhances the user experience by incorporating environmental sound effects.

Benefits of technology

It enables flexible and precise control of non-player characters, improves human-computer interaction efficiency, and enhances the immersion and interactive level of the game experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102841_12022026_PF_FP_ABST
    Figure CN2025102841_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of human-computer interaction. Disclosed are methods and apparatuses for commanding a non-player character, and a device and a medium. A method for commanding a non-player character comprises: displaying at least one of a master virtual character and a non-player character (210); receiving a natural language command, wherein the natural language command is used for commanding the non-player character (220); determining a target entity, wherein the target entity is an entity in a virtual environment that matches description in the natural language command and is perceived by the non-player character or the master virtual character (230); and in response to a behavioral intent of the natural language command, controlling the non-player character to execute a virtual activity related to the target entity (240). The method can realize intelligent control over non-player characters.
Need to check novelty before this filing date? Find Prior Art

Description

Method, device, equipment and medium for commanding non-player character

[0001] The present application claims priority to the Chinese patent application No. 202411095474.9, filed on August 9, 2024, and entitled "Method, device, equipment and medium for commanding non-player character", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] Embodiments of the present application relate to the field of human-computer interaction, in particular to a method, device, equipment and medium for commanding non-player character. BACKGROUND

[0003] AI(Artificial Intelligence, artificial intelligence) control of NPC(Non-Player Character, non-player character) has become a core technology for significantly enhancing the immersion and interaction level of game experience. Under this technical framework, players can issue basic instructions such as "attack" or "follow" to directly affect the behavior patterns of NPCs. The AI system is responsible for interpreting these simple instructions and adjusting the corresponding activities of NPCs accordingly.

[0004] However, when faced with tasks that require complex strategies or precise execution, simply relying on such basic instructions is difficult to achieve. For example, when a player wants to control an NPC to pick up a virtual prop, the player can only issue simple instructions such as "forward", "backward", "left", and "right" multiple times to command the NPC to move towards the location of the virtual prop to approach the virtual prop, but cannot accurately control the non-player character to reach the location of the virtual prop. SUMMARY

[0005] Embodiments of the present application provide a method, device, equipment and medium for commanding non-player character, which can achieve flexible control of non-player character using natural language commands. The technical solution is as follows:

[0006] In one aspect, a method for commanding non-player character is provided, the method is executed by a terminal, and the method comprises:

[0007] displaying at least one of a master virtual character and a non-player character;

[0008] receiving a natural language command, the natural language command being used to command the non-player character;

[0009] determining a target entity, the target entity being an entity in a virtual environment that matches a description in the natural language command and is perceived by the non-player character or the master virtual character;

[0010] in response to a behavior intention of the natural language command, control the non-player character to perform a virtual activity related to the target entity.

[0011] In another aspect, a method for commanding a non-player character is provided, the method being performed by a server, the method comprising:

[0012] receiving a natural language command for commanding a non-player character;

[0013] determining a target entity, the target entity being an entity in a virtual environment that matches a description in the natural language command and is perceived by the non-player character or a master virtual character;

[0014] in response to a behavior intention of the natural language command, controlling the non-player character to perform a virtual activity related to the target entity.

[0015] In another aspect, a device for commanding a non-player character is provided, the device comprising:

[0016] a display module configured to display at least one of a master virtual character and a non-player character;

[0017] an interaction module configured to receive a natural language command for commanding the non-player character;

[0018] a first determination module configured to determine a target entity, the target entity being an entity in a virtual environment that matches a description in the natural language command and is perceived by the non-player character or the master virtual character;

[0019] a first control module configured to, in response to a behavior intention of the natural language command, control the non-player character to perform a virtual activity related to the target entity.

[0020] In another aspect, a device for commanding a non-player character is provided, the device comprising:

[0021] a receiving module configured to receive a natural language command for commanding a non-player character;

[0022] a second determination module configured to determine a target entity, the target entity being an entity in a virtual environment that matches a description in the natural language command and is perceived by the non-player character or a master virtual character;

[0023] a second control module configured to, in response to a behavior intention of the natural language command, control the non-player character to perform a virtual activity related to the target entity.

[0024] In another aspect, a computer device is provided, the computer device comprising a processor and a memory having stored therein at least one instruction, at least one program, a code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the non-player character commanding method as in the above aspect.

[0025] In another aspect, a computer readable storage medium is provided, the computer readable storage medium having stored therein at least one instruction, at least one program, a code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by a processor to implement the non-player character commanding method as in the above aspect.

[0026] In another aspect, the embodiments of the present application provide a computer program product or computer program, the computer program product or computer program comprising computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the non-player character commanding method provided in the above optional implementation manners.

[0027] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:

[0028] A way of performing flexible control on a non-player character through a natural language command is provided. A terminal can receive a natural language command, identify a target entity indicated in the natural language command based on information perceived by a master virtual character and / or a non-player character from a virtual environment, and control the non-player character to perform a virtual activity related to the target entity. The target entity is an entity in the virtual environment that matches a description in the natural language command and is perceived by the master virtual character or the non-player character. For example, the natural language command is "move behind the truck", and a truck that the master virtual character can see in a field of view of the master virtual character at the time of issuing the natural language command is determined as the "truck" indicated in the natural language command, and the non-player character is controlled to accurately move behind the truck. Or, the natural language command is "is there a stream nearby", and the non-player character is controlled to explore in a direction of a source of a stream sound effect heard by the non-player character from the virtual environment, and whether the non-player character discovers a stream is identified based on a visual picture of the non-player character, and a feedback of the natural language command is generated based on the exploration result. With this method, the target entity indicated in the natural language command can be accurately determined from many entities in the virtual environment according to environmental perception information, for a natural language command with a relatively vague description, the target entity indicated in the natural language command can be inferred based on information perceived by the master virtual character or the non-player character from the virtual environment at the time of issuing the natural language command, and the non-player virtual character is controlled to perform a behavior activity related to the target entity indicated in the natural language command, to achieve flexible control on the non-player character using natural language. Compared with a way of repeatedly using simple basic instructions to control the non-player character in the related art, the natural language command in this method contains more information content and has better information transmission, and the flexibility and accuracy of control on the non-player character are higher, to improve human-computer interaction efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0029] FIG. 1 is a structural block diagram of a computer system provided by an example embodiment of the present application;

[0030] FIG. 2 is a method flowchart of a non-player character directing method provided by another example embodiment of the present application;

[0031] FIG. 3 is a schematic diagram of a non-player character directing method provided by another example embodiment of the present application;

[0032] FIG. 4 is a method flowchart of a non-player character directing method provided by another example embodiment of the present application;

[0033] FIG. 5 is a method flowchart of a non-player character directing method provided by another example embodiment of the present application;

[0034] FIG. 6 is a flow chart of a method for commanding a non-player character according to another exemplary embodiment of the present application;

[0035] FIG. 7 is a flow chart of a method for commanding a non-player character according to another exemplary embodiment of the present application;

[0036] FIG. 8 is a diagram for commanding a non-player character according to another exemplary embodiment of the present application;

[0037] FIG. 9 is a flow chart of a method for commanding a non-player character according to another exemplary embodiment of the present application;

[0038] FIG. 10 is a flow chart of a method for commanding a non-player character according to another exemplary embodiment of the present application;

[0039] FIG. 11 is a flow chart of a method for commanding a non-player character according to another exemplary embodiment of the present application;

[0040] FIG. 12 is a flow chart of a method for commanding a non-player character according to another exemplary embodiment of the present application;

[0041] FIG. 13 is a flow chart of a method for commanding a non-player character according to another exemplary embodiment of the present application;

[0042] FIG. 14 is a diagram for commanding a non-player character according to another exemplary embodiment of the present application;

[0043] FIG. 15 is a flow chart of a method for commanding a non-player character according to another exemplary embodiment of the present application;

[0044] FIG. 16 is a flow chart of a method for commanding a non-player character according to another exemplary embodiment of the present application;

[0045] FIG. 17 is a diagram for commanding a non-player character according to another exemplary embodiment of the present application;

[0046] FIG. 18 is a flow chart of a method for commanding a non-player character according to another exemplary embodiment of the present application;

[0047] FIG. 19 is a diagram for a space entity according to another exemplary embodiment of the present application;

[0048] FIG. 20 is a block diagram of a device for commanding a non-player character according to another exemplary embodiment of the present application;

[0049] FIG. 21 is a block diagram of a device for commanding a non-player character according to another exemplary embodiment of the present application;

[0050] FIG. 22 is a block diagram of a terminal according to another exemplary embodiment of the present application. DETAILED DESCRIPTION

[0051] Firstly, the terms involved in the embodiments of the present application are briefly introduced:

[0052] Virtual environment: a virtual environment displayed (or provided) by an application when running on a terminal. The virtual environment can be a simulation environment of the real world, or a semi-simulation and semi-fictional environment, or a purely fictional environment. The virtual environment can be any one of a two-dimensional virtual environment, a 2.5-dimensional virtual environment, and a virtual environment, which is not limited in the embodiments of the present application.

[0053] Virtual character: a virtual character refers to a movable object in a virtual environment. The movable object can be a virtual person, a virtual animal, an animation character, etc., such as a person displayed in a virtual environment. Alternatively, the virtual character is a three-dimensional model created based on animation skeleton technology. Each virtual character has its own shape and volume in a three-dimensional virtual world, and occupies a part of the space in the three-dimensional virtual world.

[0054] User interface (UI) control: any visible control or element on the user interface of an application. For example, pictures, input boxes, text boxes, buttons, labels, etc. Some of the UI controls respond to user operations.

[0055] 3D (three-dimensional) model: in computer graphics and video game design, a 3D model is a three-dimensional object defined by three-dimensional geometric data. These data can be created manually or generated by a 3D scanner. A complete 3D model allows users to view the object from any angle, and may contain additional information such as textures, colors, and surface material properties, etc. In electronic games, this type of model is usually used to create the image of a player or NPC (non-player character).

[0056] First-person shooting game (FPS): a shooting game in which the user can view the virtual environment in the first-person perspective. The picture of the virtual environment in the game is the picture of the virtual environment observed from the perspective of the virtual character controlled by the user. In the game, virtual characters of at least two teams compete in a single round of battle mode in the virtual environment. The virtual characters survive in the virtual environment by avoiding attacks launched by other virtual characters and dangers existing in the virtual environment (such as a toxic gas ring, a swamp, etc.). When the virtual character's life value in the virtual environment is zero, the life of the virtual character in the virtual environment ends, and the virtual character who survives in the virtual environment is the winning team. Alternatively, the competition mode of the battle (match) can include a single-player battle mode, a two-player team battle mode, or a multi-player team battle mode, and the competition mode is not limited in the embodiments of the present application.

[0057] Figure 1 shows a structural block diagram of a computer system according to an example embodiment of the present application. The computer system includes a terminal 110 and a server 120.

[0058] The terminal 110 is installed and runs a client 111 supporting a virtual environment, which is a client of an application program. When the terminal runs the client 111, a user interface of the client 111 is displayed on a screen of the terminal 110. The application program can be any one of a battle royale shooting game, a virtual reality (VR) application program, an augmented reality (AR) program, a three-dimensional map program, a virtual reality game, an augmented reality game, a first-person shooting game (FPS), a third-person shooting game (TPS), a multiplayer online battle arena game (MOBA), and a simulation game (SLG).

[0059] In this embodiment, the client is taken as an example of a MOBA game. The terminal 110 is a terminal used by a first user 112, and in a game session, the first user 112 uses the terminal 110 to control a first virtual character in a virtual environment to perform activities, which can be referred to as a master virtual character of the first user 112 in the game session. The activities of the first virtual character include, but are not limited to, at least one of adjusting a body posture, crawling, walking, running, riding, flying, jumping, driving, picking up, shooting, attacking, and throwing.

[0060] Only one terminal is shown in Figure 1, but there are a plurality of other terminals that can access the server 120 in different embodiments. Optionally, there is also one or more terminals that are corresponding terminals of developers, and a development and editing platform supporting the virtual environment is installed on the terminals. The developers can edit and update the client on the terminals, and transmit the updated client installation package to the server 120 through a wired or wireless network. The terminal 110 can download the client installation package from the server 120 to update the client.

[0061] The terminal 110 and the other terminals are connected to the server 120 through a wireless network or a wired network.

[0062] The server 120 comprises at least one of a server, a plurality of servers, a cloud computing platform, and a virtualization center. The server 120 is configured to provide background services for a client supporting a virtual environment. Optionally, the server 120 undertakes a major computing task, and the terminal undertakes a minor computing task; or the server 120 undertakes a minor computing task, and the terminal undertakes a major computing task; or the server 120 and the terminal adopt a distributed computing architecture to perform collaborative computing.

[0063] In an illustrative example, the server 120 comprises a processor, a user account database, a battle service module, and a user-oriented input / output interface (I / O interface). The processor is configured to load instructions stored in the server 120 and process data in the user account database and the battle service module. The user account database is configured to store data of user accounts used by the terminal 110 and other terminals, such as an avatar of a user account, a nickname of a user account, a battle power index of a user account, and a service area in which a user account is located. The battle service module is configured to provide a plurality of battle rooms for users to perform battles, such as 1V1 battles, 3V3 battles, and 5V5 battles. The user-oriented I / O interface is configured to exchange data with the terminal 110 through a wireless network or a wired network.

[0064] The embodiments of the present application provide a scheme for a user to control a non-player character in a virtual environment by using a voice command in a natural language form. In the scheme, the user usually has one or more NPCs as teammates in addition to a main virtual character controlled by the user. The user can use a human conversation manner that is relatively random and non-mechanical to instruct the NPCs to make a feedback on a first entity in the virtual environment as expected by the user, so as to instruct the NPCs to complete a task in cooperation with the main virtual character. Referring to FIG. 1, a user controls a main virtual character 10 to play a game, and the main virtual character 10 has an NPC teammate 20. There is a car 30 in a visual range of the main virtual character 10, and the user says a voice instruction in a natural language form: "No. 2, go to hide behind the car 30", and then the NPC teammate 20 will move to hide behind the car 30. It should be noted that there can be multiple cars in the scene shown in FIG. 1, and the NPC teammate 20 will accurately understand the car said by the user as the car 30 in the visual range of the user. That is, the NPC teammate 20 has a relatively intelligent natural language understanding capability.

[0065] In combination with FIG. 1, the scheme comprises at least one of the following five stages:

[0066] Stage one: preprocessing of spatial data;

[0067] The server pre-processes a first entity in the virtual environment, and constructs spatial data of the virtual environment.

[0068] The spatial data of the virtual environment includes a matching label of the first entity and a spatial position of the first entity. In some examples, the spatial data further includes a perspective image of the first entity. The first entity is any object appearing in the virtual environment, such as the car, the wall, the box, etc. in FIG. 1.

[0069] For the matching label of the first entity: first, obtain attribute information of the first entity in the virtual environment, such as at least one of the name, size, etc. of the first entity; and obtain a perspective image of the first entity, the perspective image including images obtained by observing the first entity from at least two perspectives. Observing the first entity from multiple perspectives can ensure that the perspective image carries comprehensive appearance information of the first entity.

[0070] The server invokes a multi-modal model to perform prediction on the attribute information and the perspective image of the first entity, to obtain an object label of the first entity, the object label being used to describe the first entity in at least one dimension; such as describing the first entity in the dimensions of material, transparency, color, shape, etc. The server invokes a large language model to perform label optimization on the object label, to obtain a matching label of the scene item; the large language model has a text generation capability, and performs rewriting on the object label input to the large language model in a manner conforming to spoken language expressions, to realize label optimization of the object label, and obtain the matching label of the scene item.

[0071] For the spatial position of the first entity: the spatial position includes at least one of the following of the first entity: a coordinate position (Location), such as coordinate information of a center point or a preset point of the first entity in the virtual environment; a rotation (Rotation), such as a direction in which a front surface of the first entity faces in the virtual environment; a bounding box (Bounding Box) used to indicate a size of the first entity in the virtual environment; cover points (Cover points) used to indicate a recommended position point of a virtual character when the virtual character approaches the first entity, to realize that the first entity can mask the virtual character.

[0072] For the perspective image of the first entity, the perspective image is obtained in the process of predicting the matching label of the first entity, and the perspective image includes images obtained by observing the first entity from at least two perspectives.

[0073] Stage two: speech recognition and intent recognition;

[0074] In the process in which the user controls the master virtual character, the terminal receives a natural language command input by the user in speech; the server performs speech recognition and intent recognition on the natural language command, and analyzes the input instruction of the user;

[0075] The voice recognition is to perform text conversion on a natural language command input by a user in a voice to obtain instruction text corresponding to the natural language command, and convert voice information into text information. A coding network is called to perform feature coding on the instruction text to obtain feature representation of the natural language command in a hidden layer space; a text segmentation network is called to perform segmentation on the feature representation to obtain a plurality of clauses in the instruction text, and then perform intent recognition on each of the plurality of clauses. In a case where the natural language command input by the user in the voice is a long sentence, the segmentation of the instruction text corresponding to the natural language command can be implemented, and the demand for analyzing the input instruction of the user in a long sentence scenario can be met.

[0076] The intent recognition of each clause is introduced as follows; the input instruction of the user is obtained through the intent recognition, and the input instruction of the user includes an entity, a semantic type, a subject type and an intent type in a clause.

[0077] Stage three: query of spatial data;

[0078] Taking the input instruction of the user as input and taking spatial data of a virtual environment and runtime information corresponding to a master virtual role as reference, an AI model is used to determine a target entity in the virtual space and a control instruction for an NPC;

[0079] For example, first, the input instruction in a text form is parsed to parse an entity in the text instruction. Then, the spatial data of the virtual environment is taken as a query range, and a real-time position and orientation of the master virtual role are taken as query reference conditions to query a target entity position corresponding to the input instruction in the virtual environment. So that the NPC is subsequently controlled to move to a vicinity of the target entity position, or the input instruction indicated behavior is performed on the target entity position.

[0080] Stage four: voice feedback of the NPC;

[0081] The NPC will make voice feedback to the user. The types of voice feedback include at least one of immediate feedback, instruction execution feedback and dynamic feedback. The immediate feedback refers to that the NPC makes voice feedback to a natural language command immediately after the user issues the natural language command to the NPC, such as “received”, “start execution”, “OK” and the like; the instruction execution feedback refers to that the NPC makes voice feedback to an execution result of a control instruction indicated by an intent of the natural language command after the NPC executes the control instruction; and the dynamic feedback refers to that the NPC makes voice feedback according to real-time environmental perception when the user does not issue a natural language command to the NPC.

[0082] The voice feedback is obtained by performing text-to-speech on the text content. The text content is inferred by a large language model in combination with the spatial data and dynamic runtime information. Using a large language model to dynamically generate the text content of the voice feedback can break the monotony, mechanicalness, and repeatability of fixed template voice feedback. On the other hand, it does not need to pre-save too many pre-produced template voices, and can avoid the problem of too large data volume on the client side.

[0083] Stage five: behavior control of the NPC;

[0084] The natural language command usually further includes a control instruction, which is an instruction executable by a behavior tree of the NPC; the control instruction indicates a subsequent action to be performed by the NPC, which can be a single action or a sequence of actions. In the case of a sequence instruction of multiple actions, the sequence instruction of multiple actions is used to instruct the NPC to perform a sequence of actions, and the data structure for storing the sequence instruction of multiple actions is in the form of a list. The cache is executed when the NPC's behavior tree executes the current instruction. The cache is executed when the NPC's behavior tree executes the current instruction.

[0085] Hotword updating function:

[0086] For stage two, the scheme also introduces a hotword system to improve the accuracy of voice recognition.

[0087] The hotword system is used to provide scene hotwords related to the scene involved in the natural language command in the voice recognition process, so that when a natural language command of a voice input is received, the hotword library can be searched for words as an instruction analysis result based on the currently running virtual environment, so that the matching degree between the instruction analysis result and the virtual environment is higher.

[0088] The virtual environment includes at least one of a virtual battle scene, a virtual transaction scene, a virtual office scene, a virtual kitchen scene, and the like. The frequently used words in multiple virtual environments, specific entities existing in virtual environments (such as virtual buildings and virtual characters existing only in specific scene types), and the like are pre-analyzed as scene hotwords, and a scene hotword library composed of multiple scene hotwords is obtained, and the multiple scene hotwords correspond to virtual environments. In addition, different scene language models are pre-trained for different virtual environments, such as a scene language model in a battle scene that pays more attention to words related to battles, and a scene language model in a transaction scene that pays more attention to words related to transactions; the scene language model can more specifically analyze the natural language command received in the virtual environment.

[0089] For example, in the process of user controlling a non-player character, a natural language command is received, based on the game state, object position and current task completion, it is determined that a first virtual environment in which a target virtual character (at least one of a master virtual character and a non-player character) is currently located is a virtual battle scene, a battle scene language model corresponding to the virtual battle scene is obtained, and based on a field of view range of the target virtual character, first scene hot words in the virtual battle scene are obtained, including "truck", "virtual grass", "virtual house", "enter", "open truck", "defend", "desert", "attack", "oak tree", "stable", and the like, decoding processing is performed on the natural language command by a pre-trained natural language analysis model (including an acoustic network, a language network, a preset dictionary, and the like) under the constraint of the first scene hot words, and an instruction analysis result is output to accurately control the non-player character. The instruction analysis result is as follows.

[0090] (1) The instruction analysis result is "No. 1, go to the front and defend", which can be used to control No. 1 to move to the front and make a defensive posture, avoiding the enemy virtual character attacking the master virtual character first. Without the constraint of the first scene hot words, the natural language command is easily identified as "No. 1, go to the front and make a sound", which affects the virtual combat plan of the player.

[0091] (2) The instruction analysis result is "No. 2, go to the desert and explore", which can be used to control No. 2 to move to the desert position to explore whether there is an enemy virtual character or a virtual treasure chest and the like. Without the constraint of the first scene hot words, the natural language command is easily identified as "No. 2, go to the mountain and explore", which makes No. 2 move to the wrong position, and the human-computer interaction efficiency is low.

[0092] (3) The instruction analysis result is "No. 1 and No. 2, attack me", which can be used to control No. 1 and No. 2 to perform virtual attacks on the virtual characters that can be attacked nearby. Without the constraint of the first scene hot words, the natural language command is easily identified as "No. 1 and No. 2, supply me", which makes No. 1 and No. 2 make an error behavior that does not meet the player's expectation, which cannot protect the master virtual character, and easily makes the enemy virtual character win the virtual game.

[0093] (4) The instruction analysis result is "No. 1, go to find the oak tree nearby", which can be used to control No. 1 to find the oak tree in the nearby area, so as to complete the game task or find the oak tree that can be used to avoid attacks. Without the constraint of the first scene hot words, the natural language command is easily identified as "No. 1, go to find the project book nearby", which makes No. 1 search for an entity that does not meet the player's expectation, and cannot meet the virtual combat demand of the player.

[0094] (5) The instruction analysis result is "1 and 2, go to the stable over there", which can be used to control No. 1 and No. 2 to search for a stable nearby and move to the location of the stable. Without the constraint of the first scene hotword, the natural language command is easily identified as "1 and 2, go over there, it will be done soon", so that No. 1 and No. 2 mistakenly think that the host virtual character currently wants to complete the virtual game by himself, and thus cannot provide better game assistance for the host virtual character.

[0095] (6) The instruction analysis result is "No. 2, pick up the prop in front", which can be used to control No. 2 to search for a prop in front and perform a picking action on the prop. Without the constraint of the first scene hotword, the natural language command is easily identified as "No. 2, pick up the crab in front", so that No. 2 cannot accurately find the entity that the host virtual character wants to pick up, which may prompt the player that "the crab cannot be found", or pick up the wrong entity "crab" for the player, thereby interfering with the virtual combat process of the player.

[0096] Ambient sound effect function:

[0097] In addition, for the ambient sound effect of the entire virtual environment, the scheme also provides a spatial audio enhancement scheme.

[0098] The audio played by the terminal includes ambient audio and NPC audio. The ambient audio is audio generated based on the scene characteristics of the virtual environment in which the host virtual object is currently located, so as to give the user a sense of being there. The NPC audio is audio generated based on the characteristics of the NPC, so that the user can intuitively feel the mood and physical state of the character through the sound heard.

[0099] Ambient audio: identify the scene elements in the virtual environment where the host virtual object is located, generate / choose appropriate element sound effects from the sound effect library in real time according to the scene elements, and synthesize the element sound effects to obtain the ambient audio. For example, the virtual environment is a forest at night, and the scene elements include trees, owls, insects, etc. The ambient audio includes rustling of leaves, owl calls, insect calls, etc.

[0100] NPC audio: determine the NPC contained in the virtual environment, generate the corresponding character voice based on the character type, current mood and current behavior of the NPC. For example, when the NPC is a middle-aged man who is running, the generated character voice is a deep male voice with a breathing sound effect when running.

[0101] When the terminal plays ambient audio and / or NPC audio to the user, the terminal performs audio enhancement processing on the ambient audio and / or NPC audio to improve the realism of the audio.

[0102] In combination with the above description of the implementation environment, the method for commanding the non-player character provided by the embodiments of the present application is described.

[0103] The method is executed by an application program, and when the application program is an application program of a single-player game, the method is executed by a terminal.

[0104] When the application program is a multi-player online application program, the method is executed by a client of the application program installed on the terminal, or by the terminal and a server in cooperation.

[0105] When the method is executed by the terminal and the server in cooperation, the terminal is configured to display a picture, receive a natural language command, and report the natural language command to the server; the server is configured to perform identification of a target entity according to the natural language command, and generate a control instruction of the non-player character, and send the control instruction to the terminal; and the terminal is configured to control the non-player character to move in the virtual environment according to the control instruction, and render and display a corresponding picture.

[0106] When the method is executed by the terminal and the server in cooperation, the terminal is configured to only display a picture and receive a trigger operation and a natural language command; the terminal reports the received trigger operation and the natural language command to the server, and the server performs identification of a target entity and generates a control instruction; the server renders a picture in which the non-player character moves in the virtual environment according to the control instruction, and sends the picture to the terminal; and the terminal receives and displays the picture. The terminal and the server described above can be referred to as computer devices.

[0107] FIG. 2 is a flowchart of a method for commanding a non-player character provided by an example embodiment of the present application. The method can be executed by the terminal in FIG. 1 described above. The method includes the following steps.

[0108] In step 210, at least one of a master virtual character and a non-player character is displayed.

[0109] The method can be executed by a client installed on the terminal, and the client is a client of an application program, and the application program is an application program supporting a virtual environment. For example, the application program can be an application program of a shooting game.

[0110] It should be noted that the calculation in the embodiments of the present application can be completely executed by the client, for example, when the game application program is a single-player game application program, all calculations are completed by the client; or can be completely executed by the server, for example, in a cloud game scenario, all calculations of the application program are completed by the server, and the client is only responsible for display and collection of operation behaviors of a user; or can be completed by the client and the server in cooperation, for example, the client is responsible for part of the calculation, and the server is responsible for another part of the calculation.

[0111] Optionally, the terminal displays a game session interface, the game session interface including a virtual environment picture, and the virtual environment picture displays a host virtual character and a non-player character. The virtual environment picture is a picture obtained by observing the virtual environment from the perspective of the host virtual character.

[0112] The host virtual character is a virtual character directly controlled by the terminal. The terminal can directly control the behavior of the host virtual character according to a preset control logic and according to a received control instruction. For example, the host virtual character is controlled to move forward when a forward instruction is received, and the host virtual character is controlled to jump up when a jump instruction is received.

[0113] For example, the terminal displays a control control (for example, a movement control, a shooting control, a throwing control, a jumping control, etc.) corresponding to the host virtual character, and the user controls the host virtual character to move in the virtual environment by triggering the control control. Alternatively, the terminal is connected to an input device (for example, a mouse, a keyboard, a camera, a sensor, etc.), and the user issues a control instruction through the input device, and controls the host virtual character to move in the virtual environment according to the control instruction.

[0114] The non-player character is a virtual character controlled by a game application program, which is controlled by a server or an artificial intelligence program in a client. The non-player character can be commanded by a user through a natural language command. In the embodiments of the present application, the game application program can control the behavior of the non-player character according to the received natural language command. The game application program can understand the intention of the natural language command, generate a control instruction according to the intention, and control the non-player character to execute the intention indicated by the natural language command.

[0115] It should be noted that "command" and "control" are different control methods. The player commanding the non-player character means that the player issues an intention (for example, a control intention or a behavior expectation) of the non-player character through a natural language command, and the AI understands the intention in the natural language command and generates a control instruction of the intention according to the autonomous behavior ability of the non-player character, and controls the non-player character to execute the activity indicated by the intention through the control instruction.

[0116] And the player controls the host virtual character, that is, the player directly gives a control operation, each control operation corresponds to a control instruction, and the host virtual character directly performs activities according to the control instruction.

[0117] In an alternative embodiment, the master virtual character is a virtual character controlled by the control instruction generated by the trigger UI control; the non-player character is a virtual character controlled by the AI according to the natural language command. When no natural language command is received, the AI can control the non-player character to move in the virtual environment; when a natural language command is received, the AI can control the non-player character to move in the virtual environment according to the intention of the natural language command.

[0118] For example, the non-player character is a teammate of the master virtual character, or the non-player character is a pet of the master virtual character, or the non-player character is a valet of the master virtual character, or the non-player character is an intelligent robot controlled by the master virtual character, or the non-player character is a clone of the master virtual character.

[0119] In an alternative embodiment, the game application program can provide multiple types of non-player characters, and different types of non-player virtual characters can have different personalities, different behavior habits or different autonomous behavior capabilities.

[0120] For example, the game application program can provide a sniper non-player character, and the shooting accuracy of the sniper non-player character can reach 99%. Or the game application program can provide a collection enthusiast non-player character, and the collection enthusiast non-player character will prefer to perform resource collection without considering the natural language command of the player once finding a virtual resource to be collected in the virtual environment. Or the game application program can provide a road-losing non-player character, and the road-losing non-player character needs to always follow the master virtual character, and once the master virtual character leaves the visual range of the road-losing non-player character, the road-losing non-player character will lose the following target and stand still and cannot move.

[0121] For example, in a team mode game match, the player needs to team up with at least one other player to play the game. When the player lacks a team-up teammate, the game application program can provide a non-player character as the teammate of the player to team up with the player to play the game match. In the game match, the player can command the non-player virtual character to move by issuing a voice command (natural language command) to make the non-player virtual character cooperate with the master virtual character to complete the game match.

[0122] For another example, in a game match, the player can carry a pet to explore the virtual environment through the pet. Then the player can command the pet to move in the virtual environment through a voice command or a text command to explore unknown areas in the virtual environment, so as to facilitate the master virtual character to make strategic decisions and deployment.

[0123] Alternatively, in a game match of a MOBA game, the host virtual character is in a team with four other virtual characters controlled by players. When one of the teammates is offline and cannot control his virtual character, for example, the player of the first virtual character is offline and cannot control the first virtual character, the host virtual character can give voice commands to direct the first virtual character during the offline period of the teammate, so that the first virtual character assists other teammates to participate in team battles, or the first virtual character takes accurate actions to respond to changes in the game.

[0124] At step 220, a natural language command is received, the natural language command being used to direct the non-player character.

[0125] The natural language command is a command instruction conveyed using a natural language of human beings. For example, the natural language command is a command instruction conveyed using a Chinese language, or a command instruction conveyed using an English language. The terminal receives the natural language command issued by the player, and controls the non-player character according to an intention corresponding to the natural language command. In some examples, the natural language command is issued through voice spoken by the player or through voice synthesized by the player. For example, the terminal receives voice audio of the player, converts the voice audio into text, and obtains the natural language command. Alternatively, the terminal receives text of the natural language command input by the player.

[0126] The player can use oral language or written language and the like to issue the natural language command. The game application program invokes a large language model to perform understanding on the natural language command, extracts an intention expressed in the natural language command, and generates a control instruction according to the intention to control activities of the non-player character.

[0127] For example, the natural language command can be “come to me”, and the game application program can parse an intention of the natural language command as “controlling the non-player virtual character to move to a position of the host virtual character”. The game application program obtains the position of the host virtual character, generates a navigation route of the non-player virtual character moving to the position of the host virtual character, and controls the non-player virtual character to move to the position of the host virtual character according to the navigation route.

[0128] For another example, the natural language command can be “go to pick up the treasure chest”, and the game application program can parse an intention of the natural language command as “controlling the non-player virtual character to move to a position of the treasure chest and pick up the treasure chest”. The game application program can obtain the position of the treasure chest, control the non-player virtual character to move to the position of the treasure chest, and perform an operation of picking up the treasure chest after the non-player virtual character arrives at the position of the treasure chest.

[0129] Optionally, the natural language command is used to control the non-player character to perform an activity related to the target entity. For example, the target entity can be a moving destination of the non-player character, the target entity can be an object that the non-player virtual character needs to observe, the target entity can also be a target that the non-player virtual character needs to attack, or the target entity is an object that the non-player virtual character needs to interact with.

[0130] The target entity is a virtual object existing in the virtual environment. The target entity is an entity described in the natural language command. The entity can refer to a virtual object fixed in the virtual environment, for example, the target entity can be a virtual building, a virtual terrain, a virtual vehicle, a virtual plant, a virtual prop, a virtual item, etc. The entity can also refer to a virtual object that can perform interaction with the virtual character in the virtual environment, for example, the target entity can be an interaction point (for example, a door, a window, a cabinet, a cellar, etc.) in the virtual environment, a virtual prop, a virtual character, a virtual light source, etc.

[0131] For example, the natural language command can be “please help me open the door”, and the target entity can be “the door”, and the non-player character needs to perform an activity related to the door “open the door”. Or, the natural language command can be “is the kitchen safe”, and the target entity can be “the kitchen”, and the non-player character needs to perform an activity related to the kitchen “check whether there is an enemy virtual character or other dangerous situation in the kitchen”.

[0132] Optionally, the natural language command includes target entity information of the target entity. The target entity information can be a description text of the target entity. For example, the target entity information can include at least one of the following contents: the name, type, location, and characteristics of the target entity. The game application can find the target entity from a plurality of entities in the virtual environment based on the target entity information in the natural language.

[0133] Step 230, determining the target entity, the target entity is an entity in the virtual environment that matches the description in the natural language command and is perceived by the non-player character or the master virtual character.

[0134] The game application performs analysis on the natural language command, extracts the target entity information therein, determines the target entity from the entity set of the virtual environment based on the target entity information, and then controls the non-player character to perform an activity related to the target entity according to the intention of the natural language command.

[0135] The target entity is determined from the virtual environment in combination with the environment perception information, and the environment perception information includes information perceived by at least one of the master virtual character and the non-player character from the virtual environment.

[0136] Since the target entity information in the natural language command is usually ambiguous, for example, the natural language command can be "move to the red truck", and the target entity information therein is "red truck", and there can be many red trucks in the virtual environment, and the target entity cannot be accurately determined from the virtual environment according to the target entity information in the natural language command.

[0137] Therefore, the embodiments of the present application provide a method of determining a target entity in combination with target entity information and environment perception information. For the above example, the "red truck" expressed by the player in the natural language command should be a red truck that the player can see, and therefore, in combination with the visual field range of the virtual character controlled by the player, the red truck located in the visual field range of the virtual character controlled by the player can be screened from the multiple red trucks in the virtual environment, and the red truck is the target entity indicated by the player in the natural language command.

[0138] Since the natural language command is issued by the player based on the perception of the virtual environment, in order to accurately identify the target entity indicated in the natural language command, the game application program will combine the environment perception information of the player when issuing the natural language command to identify the target entity in the natural language command. Based on the perception of the virtual environment by the player, the entity closest to the target entity information that the player can perceive in the virtual environment is inferred as the target entity.

[0139] For example, the virtual environment includes multiple candidate entities matching the natural language command, and the target entity is an entity screened from the multiple candidate entities based on the environment perception information of the virtual character controlled by the player or the non-player character. For example, the game application program first selects multiple candidate entities matching the target entity information in the natural language command from the entity set, and then selects an entity that can be perceived by the virtual character controlled by the player or the non-player character from the multiple candidate entities as the target entity according to the environment perception information.

[0140] For example, the environment perception information includes at least one of the following:

[0141] · visual information perceived within the visual field range of the virtual character controlled by the player;

[0142] · auditory information perceived within the auditory range of the virtual character controlled by the player;

[0143] · information perceived by the virtual character controlled by the player with a perception skill or a perception virtual prop;

[0144] · visual information perceived within the visual field range of the non-player character;

[0145] · auditory information perceived within the auditory range of the non-player character;

[0146] • information perceived by the non-player character using a perception skill or a perception virtual prop.

[0147] For example, the environment perception information of the host virtual character can include: a picture or entity that the host virtual character can see when observing the virtual environment, an environmental sound that the host virtual character can hear, a direction of the source of the environmental sound, a type and a sound size of the environmental sound, perception information obtained by the host virtual character using a perception skill (e.g., visual information, auditory information, sound wave information, light reflection information, etc.), and perception information obtained by the host virtual character using a perception virtual prop (e.g., a sensing signal of a sensor, a positioning signal of a positioning prop, etc.).

[0148] The environment perception information of the non-player character can include: a picture or entity that the non-player character can see when observing the virtual environment, an environmental sound that the non-player character can hear, a direction of the source of the environmental sound, a type and a sound size of the environmental sound, perception information obtained by the non-player character using a perception skill (e.g., visual information, auditory information, sound wave information, light reflection information, etc.), and perception information obtained by the non-player character using a perception virtual prop (e.g., a sensing signal of a sensor, a positioning signal of a positioning prop, etc.).

[0149] For example, as shown in FIG. 3, when the natural language command is “come to the stairs in front of me here”, the game application program can determine that the target entity is the stairs 301 in front of the host virtual character in combination with the field of view range of the host virtual character.

[0150] It should be noted that the environment perception information can include information perceived by the host virtual character and the non-player character in real time when the natural language command is received, or information perceived by the host virtual character and the non-player character historically before the natural language command is received. That is, the target entity can be an entity currently seen by the player, or an entity historically seen by the player. Therefore, the game application program needs to combine the real-time environment perception information and the historical environment perception information of the host virtual character and / or the non-player character to perform the identification of the target entity.

[0151] For example, when the natural language command is “let's go back to the hotel we just passed by”, the game application program needs to query the hotel that the host virtual character has passed by according to the historical environment perception information of the host virtual character.

[0152] Step 240, in response to the behavior intention of the natural language command, controlling the non-player character to perform a virtual activity related to the target entity.

[0153] Exemplarily, the large language model is invoked to identify an intention of the natural language command, a control instruction or a control instruction sequence is generated according to the intention, and the non-player character is controlled to complete an activity based on the target entity according to the control instruction or the control instruction sequence.

[0154] To sum up, the method provided in the embodiment identifies the target entity indicated in the natural language command based on information perceived by the host virtual character and / or the non-player character from the virtual environment. For example, the natural language command is “move behind the truck”, and according to the visual field range of the host virtual character when the natural language command is issued, the truck that can be seen by the host virtual character in the visual field range is determined as the “truck” indicated in the natural language command, and then the non-player character is controlled to accurately move behind the truck. Or, the natural language command is “is there a stream nearby”, and according to the stream sound effect heard by the non-player character from the virtual environment, the non-player character is controlled to explore in the direction of the source of the stream sound effect, and based on the visual picture of the non-player character, it is identified whether the non-player character finds a stream, and a feedback of the natural language command is generated based on the exploration result. By using this method, the target entity indicated in the natural language command can be accurately determined from many entities in the virtual environment according to the environmental perception information. For the received natural language command with relatively vague description, the target entity indicated in the natural language command can be inferred in combination with the information perceived by the host virtual character from the virtual environment when the command is issued, and the non-player virtual character is controlled to perform the behavior activity related to the target entity indicated in the natural language command, so as to realize the flexible control of the non-player character using natural language.

[0155] FIG. 4 shows a method flowchart of a method for commanding a non-player character provided in an example embodiment of the application. The method can be performed by the server in FIG. 1 described above. The method comprises:

[0156] Step 220, receiving a natural language command, the natural language command being used to command a non-player character.

[0157] Exemplarily, the identification of the target entity can be performed by the client or by the server.

[0158] For example, the client reports the received natural language command to the server. The server performs analysis on the natural language command to obtain target entity information. Subsequently, the method shown in FIG. 4 is performed to determine the target entity indicated by the target entity information from the entity set of the virtual environment according to the environmental perception information.

[0159] Step 250, determining a target entity, the target entity being an entity in the virtual environment that matches the description in the natural language command and is perceived by the non-player character or the host virtual character.

[0160] Step 260, in response to the behavior intention of the natural language command, controlling the non-player character to perform a virtual activity related to the target entity.

[0161] Exemplarily, the method for the server to determine the target entity from the entity set is the same as the embodiment shown in FIG. 2, the difference is that the embodiment shown in FIG. 2 is executed by the client, and the embodiment shown in FIG. 4 is executed by the server, and the specific method will not be repeated here.

[0162] Exemplarily, the server can determine a candidate entity list matched with the target entity information of the target entity from the entity set according to the target entity information of the target entity; determine the perception range of the master virtual character and / or the non-player character as the search range according to the environment perception information; and determine the entity in the search range in the candidate entity list as the target entity.

[0163] Alternatively, the server determines the perception range of the master virtual character and / or the non-player character as the search range according to the environment perception information, determines the entity in the search range as the candidate entity, and determines the candidate entity with the highest similarity to the target entity information of the target entity as the target entity.

[0164] In summary, the method provided in the embodiment identifies the target entity indicated in the natural language command based on the information perceived by the master virtual character and / or the non-player character from the virtual environment. For example, the natural language command is "move behind the truck", and according to the field of view of the master virtual character when the natural language command is issued, the truck that the master virtual character can see in the field of view is determined as the "truck" indicated in the natural language command, and then the non-player character is controlled to move accurately behind the truck. Alternatively, the natural language command is "is there a stream nearby", and according to the stream sound effect heard by the non-player character from the virtual environment, the non-player character is controlled to explore the direction of the source of the stream sound effect, and based on the visual picture of the non-player character, whether the non-player character discovers a stream is identified, and a feedback of the natural language command is generated based on the exploration result. By using this method, the target entity indicated in the natural language command can be accurately determined from the many entities in the virtual environment according to the environment perception information. For the received natural language command with relatively vague description, the target entity indicated in the natural language command can be inferred in combination with the information perceived by the master virtual character from the virtual environment when the command is issued, and the non-player virtual character is controlled to perform the behavior activity related to the target entity indicated in the natural language command, so as to realize the flexible control of the non-player character using natural language.

[0165] Exemplarily, an example embodiment for determining a target entity from many entities in a virtual environment is given.

[0166] FIG. 5 shows a flowchart of a method for commanding a non-player character according to an example embodiment of the present application. The method can be performed by the terminal in FIG. 1. Based on the embodiment shown in FIG. 2, step 230 can include steps 231-233.

[0167] In step 231, a set of entities of the virtual environment is obtained, the set of entities including entities in the virtual environment and entity information.

[0168] Optionally, the game application infers the target entity based on the target entity information, the environment perception information, and preprocessed virtual environment data. The preprocessed virtual environment data includes the set of entities of the virtual environment.

[0169] The set of entities includes entity information of each entity in the virtual environment. The entity information includes at least one of the following: name, type, location, feature, text label, text embedding vector of the text label, image, visual label, and image embedding vector of the visual label.

[0170] The text label can be at least one word obtained by tokenizing a feature (e.g., a description text of an appearance of an entity), and the text embedding vector corresponding to each text label is obtained by calling an embedding model.

[0171] The text label is used to introduce inherent properties of the entity in the virtual environment. On the one hand, the text label realizes the description of the entity in the text mode, and provides semantic information for the prediction of the visual label of the entity while describing the entity. On the other hand, the text label is a description of the entity in the virtual environment, and in the case of reuse of an entity model, the text label can more accurately describe the inherent properties of the entity in the virtual environment. For example, in the virtual environment, to avoid redesigning the entity model of the entity, at least one of scaling, stretching, and rotating is performed on the object model of the virtual table, and the transformed entity model is deployed in the virtual environment as a virtual bench. In the case of reuse of the model, the text label can describe the inherent properties of the entity in the virtual environment.

[0172] In an optional implementation, the text label includes at least one of a name of the entity in the virtual environment and a size of the entity in the virtual environment. The size of the entity in the virtual environment is used to indicate the virtual space occupied by the entity in the virtual environment, and in the case of reuse of the entity model, the size of the entity model is changed to affect the description of the entity. The name in the virtual environment can reflect the purpose of the entity in the virtual environment, and in the case of reuse of the entity model, two different entities are deployed based on the same entity model, which affects the description of the entity.

[0173] The image can include an image obtained by observing a three-dimensional model of the entity from at least one direction, for example, the image can include a three-view image of the entity. The embedding vector of the image is an embedding vector of a visual feature label. The visual feature label is obtained by calling a multi-modal model to perform feature recognition based on the image and the text label of the entity.

[0174] For example, the image is used to describe the style of the entity; based on the richness of natural language, the player can describe the entity in the virtual environment from multiple semantic angles when describing the entity; when constructing the visual label of the entity, it is necessary to obtain comprehensive description information of the entity, and the image carries the appearance style of the entity such as color, texture, shape and the mutual positional relationship between each part on the picture modal, and can comprehensively describe the entity from the picture modal (or visual modal).

[0175] In an optional implementation, the image of the entity includes images obtained by observing the entity from at least two perspectives. Observing the entity from at least two perspectives in the image can avoid the problem that the spatial structure of the entity is blocked and cannot present all the appearance styles in the entity from a single perspective.

[0176] It should be noted that the image can be obtained in the virtual environment, such as taking the image of the entity in the entity; the image can also be obtained outside the virtual environment, such as taking the image of the entity in the development interface during the development and design of the entity.

[0177] For example, as shown in FIG. 6, the virtual environment is a three-dimensional virtual environment, and to control the non-player character to move to any position in the virtual environment, it is necessary to search all entities in the virtual environment according to the natural language command, so it is necessary to first preprocess the virtual environment data to generate an entity set that can be searched. Optionally, the entities in the entity set can include two categories: a region / building 403 and other entity objects 404. The entity information of the region / building can be manually labeled by artificial, because the region / building 403 usually has a large range, and there can be subspaces inside, for example, there can be rooms in a building, and the positions, orientations, etc. of the region, building and subspaces need to be manually labeled. The other entity objects 404 can automatically export the entity information such as position, orientation, bounding box and occlusion point by the game engine. These entity information exists in the game engine assets and can be directly recorded.

[0178] For example, the entity set includes first-class entities and second-class entities, the first-class entities include at least one of a region and a building, and the second-class entities include entity objects. The text label of the first-class entity is obtained by manual labeling, for example, the text label of a certain room in a certain building is manually labeled as: a certain region, a certain building, a certain floor, a room. The text label of the second-class entity is automatically generated by calling a large language model.

[0179] The text label generation method of the second type of entity can include: obtaining at least one image of the entity, inputting the at least one image into a large language model to obtain a description text of the entity; performing word segmentation on the description text to obtain at least one text label of the entity, each text label including one word obtained after word segmentation processing. Performing a vectorization operation on the text label to obtain an embedding vector corresponding to the text label.

[0180] For example, as shown in FIG. 7, the text label generation process for an oil drum includes: taking 9 images of the oil drum 501 from different perspectives as shown in FIG. 8. Then, a large language model is called to output a description text 405 of the oil drum “a metal oil drum” according to the 9 images. At least one text label is obtained by performing word segmentation processing 406 on the description text, for example: “a metal oil drum” is segmented into “metal” and “oil drum”, which represents all the features of the oil drum. Word segmentation can be performed using a large language model, and an embedding model 407 is called to perform a vectorization operation on each word in the segmentation result to generate an Embedding vector, and record the Embedding vector index corresponding to each text label of the oil drum.

[0181] The embedding vector is a high-dimensional vector data of the text label, and the embedding vector of the text label is generated in advance, so that it is not necessary to repeatedly perform vector generation on the text label of each entity when matching the target entity based on the target entity information from the entity set, which can improve the matching efficiency of the target entity information and the entity information.

[0182] It should be noted that this method does not directly use the description text as the text label, but uses the word segmentation result of the description text as the text label, because the description text is usually of varying lengths, and for longer description texts, the embedding vector has poor effect when actually searching and comparing similarity. For example, searching “a truck” for “a blue small truck with rusted body”, the “a blue small truck with rusted body” in the recall result is probably behind “a car”, but in fact, for the search “truck”, no matter how complex the additional description is, it should prioritize searching “truck”, so when performing similarity search, the comparison should be one by one feature word, not the description text.

[0183] Further, in order to further extract the visual features of the entity, the method provided by the embodiment of the present application can further extract hidden visual features in the image of the entity based on the text label and the image of the entity. Taking the first entity as an example, as shown in FIG. 9, the method comprises: obtaining image data of at least one perspective image of the first entity 408; obtaining at least one text label of the first entity 409; calling a multi-modal model 410 to extract visual features of the first entity based on the at least one perspective image and the at least one text label of the first entity, to obtain a visual label of the first entity, the visual label being used to describe the visual features of the first entity in at least one dimension; and converting the visual label to obtain an image embedding vector of the first entity 411.

[0184] Since the text label is manually annotated or generated by a large language model based on human language characteristics, the text label extracted based on human language habits may ignore part of the features of the entity. For example, when manually annotating the text label of an oil drum, only more noticeable text labels such as “metal” and “rust” may be annotated, and detailed features such as “rust”, “blue”, “yellow”, and “right lower corner paint drop” may be ignored. Therefore, the method provided by the embodiment of the present application also uses a multi-modal model to extract more visual features from the text label and the image of the entity based on the multi-modal model, and performs image similarity matching with the target entity based on the visual features to improve the recognition accuracy of the target entity.

[0185] For example, the multi-modal model has the ability to perform model prediction on different modal text information and picture information. In the embodiment, the input parameters of the multi-modal model are the text label and the image of the scene object; the multi-modal model predicts the visual label of the scene object from the two modalities of the text modality and the picture modality, and the visual label is used to describe the visual features of the scene object in at least one dimension.

[0186] Exemplarily, the multi-modal model comprises an artificial neural network (ANN), and the network of the artificial neural network comprises at least one of the following: a convolutional neural network (CNN), a recurrent neural network (RNN), a temporal convolutional network (TCN), a long short-term memory (LSTM), a multilayer perceptron (MLP), and a support vector machine (SVM). The network structure of the multi-modal network is not limited in the application. Exemplarily, the multi-modal model has the ability to extract visual features of a scene object from input information of a text mode and a picture mode, and obtain visual labels of the scene object. Exemplarily, the multi-modal model can be a classification model that assigns a scene object to a preset type label, or a prediction model that predicts a label capable of expressing visual features of the scene object for the scene object.

[0187] The training method of the multi-modal model can be: inputting the sample entity image and the sample label into the pre-trained multi-modal model, performing fine-tuning on the pre-trained multi-modal model according to the loss of the predicted label output by the pre-trained multi-modal model and the sample label, to obtain a multi-modal model capable of outputting visual feature labels based on input text labels and images. The visual feature labels comprise more detailed description texts extracted from the images. For example, the text description of “a metal oil drum” can only be associated with two text labels “metal” and “oil drum”, but the Clip (multi-modal) model can also identify visual information hidden in the image of the oil drum, such as “rust”, “blue”, and “yellow” visual feature labels. These visual features do not appear in the text labels, and therefore, performing visual feature search in combination with the Clip model can further improve the accuracy of entity search.

[0188] Exemplarily, the multi-modal model comprises at least one of a visual question answering model and a picture description model.

[0189] 1. The multi-modal model comprises a visual question answering model.

[0190] The game application constructs a question sentence for the perspective image, and the question sentence carries a text label of the first entity. The perspective image of the first entity and the question sentence are input into the visual question answering model to obtain an answer sentence, and the answer sentence is taken as a visual label of the first entity.

[0191] Exemplarily, the question sentence carries a text label of the entity; the question sentence is used to guide the visual question answering model to convert the image into a visual label of the entity; and the question sentence provides supplementary information of the entity in a text manner while guiding the conversion of the visual label of the entity.

[0192] Exemplarily, the question sentence includes at least two sub-sentences; the perspective image of the first entity and a first sentence in the at least two sub-sentences are input into the visual question answering model to obtain a first answer sub-sentence; the above steps are repeated until at least two answer sub-sentences corresponding to the at least two sub-sentences are obtained, and the at least two sub-sentences are used to inquire the visual features of the perspective image from multiple dimensions; and sentence aggregation is performed on the at least two answer sub-sentences to obtain an answer sentence of the first entity.

[0193] Exemplarily, expected information of the first entity is obtained, the expected information is used to indicate a description dimension of the first entity expected in the answer sentence, and / or an expected format of the answer sentence; a question sentence of the perspective image is constructed according to the expected information and the text label; wherein a first sub-part in the question sentence is supplementary introduction information of the first entity, and carries a text label of the first entity; and a second sub-part in the question sentence is an answer guiding sentence for the visual question answering model, and carries the expected information.

[0194] Exemplarily, the expected information is used to indicate a description dimension of a scene object expected in the visual label, and / or an expected format of the visual label; in an example, the description dimension of the entity expected in the visual label indicated by the expected information includes, but is not limited to, at least one of a type description, a material, a transparency, a color, a surface feature, and a shape of the entity. In another example, the expected format of the visual label indicated by the expected information is at least one of a Comma Separated Values (CSV), a JavaScript Object Notation (JSON), an eXtensible Markup Language (XML), and an Excel format.

[0195] Exemplarily, a first sub-part in the question sentence is supplementary introduction information of the entity, and carries a text label of the entity; and a second sub-part in the question sentence is an answer guiding sentence for the visual question answering model, and carries the expected information.

[0196] In one example, the question sentence is: "Below is a rendering image of a three-dimensional model named A, with a size of B; please understand what the three-dimensional model is through the model name, image, and size information, and summarize the overall characteristics, and finally output the visual label in Json format." Among them, the semicolon ";" in the question sentence is divided, the first half is the first subpart, which is the supplementary introduction information of the entity. The supplementary introduction information of the entity includes the text label of the entity obtained from the entity set, such as the name of the entity in the virtual environment A and the size of the entity in the virtual environment B.

[0197] The second half is the second subpart, which is the answer guiding sentence for the visual question and answer model. Further, the question sentence includes a preset sentence template, and the text label of the entity is filled into the preset sentence template to obtain the question sentence. Among them, the filling position of the text label of the entity is the position of the words "A" and "B" in the question sentence; wherein the name of the entity in the virtual environment in the text label is marked as asset name (asset name), and the size of the entity in the virtual environment in the text label is marked as asset size (asset size); fill into the preset sentence template to obtain the question sentence.

[0198] For example, the image of the entity and the question sentence are input into the visual question and answer model, and the question sentence prompts the guiding information of the visual label generated by the visual question and answer model in a text manner, while providing supplementary information of the entity and further describing the entity. For example, the image of the entity includes images obtained by observing the entity from nine perspectives. The question sentence and the image of the entity are input into the visual question and answer model to obtain the visual label of the entity. For example, the visual label of the entity includes: "Type description: wooden double bed; Material: [wood]; Transparency: opaque; Color: [brown]; Surface feature: [carving]; Shape: [cuboid]".

[0199] Further, the visual label of the entity is used to supplement the entity set; after the visual label is supplemented, the entity set includes the image of the entity, the visual label, and the spatial position of the entity in the virtual environment, and the spatial position includes the coordinate position, the orientation information, the bounding box information, and the mask point information. In various embodiments of the present application, the entity set is also referred to as the spatial information library of the virtual environment.

[0200] In an optional implementation, the visual question answering model takes the image and the first sentence as input parameters, and the first answer sub-sentence obtained is an answer sentence under the guidance of the first sentence. For example, at least two sub-sentences are used to inquire the visual features of the appearance image from multiple dimensions; the first answer sub-sentence is used to answer the visual features of the appearance image inquired in one dimension corresponding to the first sentence. Each of the at least two sub-sentences is input into the visual question answering model respectively with the image, and the corresponding answer sub-sentence is obtained, that is, at least two answer sub-sentences corresponding to the at least two sub-sentences are obtained. The sentence aggregation is performed on the at least two answer sub-sentences, and part or all contents in the at least two answer sub-sentences are extracted to obtain the visual label of the entity. For example, the sentence aggregation performed on the at least two answer sub-sentences can be based on artificial neural network to perform the sentence aggregation on each answer sub-sentence input; or words in a preset position of the answer sub-sentence can be intercepted to realize the sentence aggregation.

[0201] In the embodiment, the visual question answering model (VQA) can be implemented as at least one of the following: a large language model with vision ability (LLAVA), a mini generative pre-trained transformer (MiniGPT4), a cognitive visual language model (CogVLM), a generative pre-trained transformer with vision (GPT4V), and a Gemini model.

[0202] 2. The multi-modal model includes a picture description model.

[0203] The game application inputs the perspective image of the first entity into the picture description model to predict a description text of the first entity; and performs sentence aggregation on the description text and the text label to extract a visual label of the first entity.

[0204] Exemplarily, the image captioning model has the capability of describing the input image in text. By invoking the image captioning model, the image in the graphical modality is converted into the description text in the textual modality. Exemplarily, the image captioning model includes, but is not limited to, at least one of the following: a show and tell model, a show, attend and tell (SAAT) model, a bootstrap latent integer picture (BLIP) model, and a transformer-based image captioning (TICM) model.

[0205] Exemplarily, part or all of the content in the description text and the text label is determined as the visual label of the entity in the perspective image; exemplarily, the sentence aggregation is performed on the description text and the text label, which can be based on an artificial neural network to perform the sentence aggregation on the input description text and the text label, or can be to intercept words at a preset position of the description text and the text label to achieve the sentence aggregation.

[0206] Exemplarily, the text label of the first entity includes at least one of a name of the first entity in the virtual environment and a size of the first entity in the virtual environment; and / or, the perspective image of the first entity includes images obtained by observing the first entity from at least two perspectives.

[0207] Optionally, the game application performs a quasi-spontaneous rewriting on the visual label of the first entity to obtain a matching label conforming to the spontaneous expression of the natural language.

[0208] Exemplarily, as introduced above, the visual label is used to describe the visual features of the entity in at least one dimension, and the image of the entity presents rich visual features of the entity. However, in the spontaneous expression of the natural language, the description of the entity cannot cover the visual features of the entity in various dimensions. For example, when describing a virtual bed in the virtual environment in the spontaneous expression, usually, the color, material, and placement of the virtual bed are concerned, while usually, the process, such as carving and embossing, used for the head of the bed and the side plate of the bed is neglected. The purpose of performing the quasi-spontaneous rewriting on the visual label of the entity is to obtain a matching label closer to the spontaneous expression; it can be understood that performing the quasi-spontaneous rewriting can be to delete part of the content in the visual label, or to change the visual label into a label having the same semantics but different text expression.

[0209] For example, the visual label of the first entity is input into a large language model to predict a matching label conforming to the spontaneous expression of the natural language, and the large language model carries prior knowledge of the spontaneous expression of the natural language.

[0210] The game application obtains a first sample label pair, the first sample label pair including a first label before being rewritten in a quasi-spoken manner and a second label obtained by rewriting in a quasi-spoken manner; constructs a rewriting guide sentence according to the first sample label pair and the visual label, the rewriting guide sentence having a natural semantic meaning of rewriting the visual label with reference to the first sample label pair; inputs the rewriting guide sentence into a large language model to predict a matching label conforming to a spoken expression of a natural language.

[0211] For example, the quasi-spoken rewriting performed on the visual label is based on calling a natural language model. The natural language model carries prior knowledge of a spoken expression of a natural language. For example, the natural language model is implemented as a large language model (LLM).

[0212] Further, the calling of the natural language model to perform the quasi-spoken rewriting includes:

[0213] obtaining a first sample label pair;

[0214] constructing a rewriting guide sentence according to the first sample label pair and the visual label;

[0215] inputting the rewriting guide sentence into the natural language model to predict a matching label conforming to a spoken expression of a natural language;

[0216] In this example, the first sample label pair includes a first label before being rewritten in a quasi-spoken manner and a second label obtained by rewriting in a quasi-spoken manner. Taking a virtual bed in a virtual environment as an example, the first label in the first sample label pair is: “type description: wooden double bed; material: [wood]; transparency: opaque; color: [brown]; surface feature: [carving]; shape: [cuboid]”. The second label is: “bed, brown, wooden, cuboid”.

[0217] The rewritten guiding sentence has a natural semantic rewriting of the visual label with reference to the first sample label pair. In one example, the visual label that needs to be paralinguistically rewritten is: "Type description: green plant potted; material: [plant]; transparency: opaque; color: [brown, green]; surface feature: [smooth]; shape: [irregular shape]". Correspondingly, according to the first sample label pair and the visual label, the guiding sentence obtained is: Example: rewrite "Type description: wooden double bed; material: [wood]; transparency: opaque; color: [brown]; surface feature: [carving]; shape: [cuboid]" as "bed, brown, wooden, cuboid". According to the above example, "Type description: green plant potted; material: [plant]; transparency: opaque; color: [brown, green]; surface feature: [smooth]; shape: [irregular shape]" is rewritten.

[0218] For example, the spatial position of the first entity in the virtual environment is obtained, and the spatial position of the first entity is determined as auxiliary information of the visual label of the first entity. For example, at least one of the coordinate position, orientation information, bounding box information, and cover point information of the first entity in the virtual environment is obtained; wherein the coordinate position is used to indicate the position of the first entity in the virtual environment, the orientation information is used to indicate the direction faced by the first entity in the virtual environment, the bounding box information is used to indicate the size of the first entity in the virtual environment, and the cover point information indicates the recommended virtual character position when the virtual character approaches the first entity.

[0219] For example, the spatial position of the first entity in the virtual environment is used to indicate the deployment situation of the first entity in the virtual environment, and the spatial information is used to indicate the position, size, and the like of the first entity in the virtual environment after the first entity is deployed in the virtual environment.

[0220] In an optional implementation, the coordinate position (Location) is used to indicate the position of the first entity in the virtual environment, such as the coordinate information of the center point or the preset point of the first entity in the virtual environment; the orientation (Rotation) is used to indicate the direction faced by the first entity in the virtual environment, such as the direction faced by the front of the first entity in the virtual environment; the bounding box (Bounding Box) is used to indicate the size of the first entity in the virtual environment; and the cover point (Cover points) is used to indicate the recommended virtual character position when the virtual character approaches the first entity, so as to realize that the first entity can mask the virtual character.

[0221] Exemplarily, the plurality of type information of the first entity is stored in a table manner, for example, in a table, a first column of the table records an asset name of the first entity, and indicates a name of a mapping resource and a name of a skeleton model resource of the first entity. A second column records an entity name of the first entity, and is used to indicate a name of the first entity in a virtual environment. A third column stores a picture storage path of an appearance image of the first entity. A fourth column is a spatial position of the appearance image, also referred to as tactical information, including coordinates and an orientation of the appearance image. A fifth column is an asset scaling ratio of the first entity, and is used to indicate a size of the asset. A sixth column is a picture name of the appearance image of the first entity. A seventh column is the appearance image of the first entity.

[0222] In step 232, semantic understanding is performed on the natural language command to obtain behavior intention and target entity information of the natural language command.

[0223] The game application parses the natural language command to obtain the target entity information, and the target entity information includes at least one of the following: an entity type, an entity name, an entity position, and an entity feature.

[0224] Optionally, the game application calls a large language model to perform parsing on the natural language command to obtain the target entity information of the target entity. An embedding model is called to perform vectorization processing on the target entity information to obtain a target embedding vector of the target entity information.

[0225] For example, the natural language command is “come here in front of the blue truck”, and the large language model can parse the target entity information from the natural language command, including the position information “in front of” and the entity name “blue truck”.

[0226] Exemplarily, a named entity recognition technology in the field of natural language processing can also be called to extract the target entity information from the natural language command. The named entity recognition technology is used to recognize entities with specific meanings in a text, for example, to recognize names, place names, position words, and adjectives.

[0227] For example, as shown in FIG. 12, the natural language command 412 is “find a paper box behind the red sofa on the first floor of the motel”, and the named entity recognition technology can be used to recognize the following from the natural language command: the building noun “motel on the first floor”; the object nouns “sofa” and “paper box”; the position word “behind”; and the adjective “red”. After logical construction, a hierarchical scene query calling form 413 with a search type, content, position word, and constraint information is formed. The search type is used to narrow the data retrieval range based on similarity matching, the search content is a specific entity description, the adjective is included in the description text for query, and the position word is used to form a constraint of the search range. The floor constraint is determined by the player's position, and in an indoor scene, the up and down range is narrowed to avoid finding entities that are not visible across floors.

[0228] At step 233, the target entity is queried from the entity set according to the target entity information and the environment perception information.

[0229] In an optional embodiment, the client queries the target entity from the entity set of the virtual environment according to the target entity information indicated by the natural language command and the environment perception information. In another optional embodiment, the client reports the natural language command to the server, and the server queries the target entity from the entity set of the virtual environment according to the target entity information indicated by the natural language command and the environment perception information.

[0230] In an optional embodiment, after obtaining the entity set, the target entity can be matched from the entity set according to the target entity information in the natural language command. As shown in FIG. 10, the game application performs command analysis 401 on the input natural language command to obtain target entity information, performs target entity query 402 according to the preprocessed virtual environment data and the real-time position and orientation of the host virtual character, finds the target entity in the preprocessed virtual environment data that is located within the visual range of the host virtual character and matches the target entity information, and determines the position of the target entity to determine the target moving position of the non-player character.

[0231] In an optional embodiment, as shown in FIG. 11, step 233 can include steps 2331 to 2332. The method can be executed by a terminal or by a server.

[0232] At step 2331, the similarity between the target entity information and the entity information of each entity in the entity set is calculated.

[0233] For example, the entity set includes entity information of at least one entity, and the entity information of each entity can include at least one of a text embedding vector of a text label and an embedding vector of a visual feature. The target entity information can be calculated for similarity with the text embedding vector and the image embedding vector, respectively, and the entity with a similarity higher than a threshold value is determined as the target entity. The threshold value can be set according to actual technical needs.

[0234] The method of performing similarity matching with the text embedding vector and the method of performing similarity matching with the image embedding vector are described below, respectively.

[0235] 1. The similarity includes text similarity between the target entity information and the text embedding vector.

[0236] Taking a first entity in the entity set as an example, the game application (client or server) tokenizes the target entity information to obtain at least one target entity label; converts the at least one target entity label into at least one target embedding vector; obtains entity information of the first entity, the entity information including a text embedding vector, the text embedding vector being an embedding vector converted based on a text label of the first entity; respectively calculates text parent similarities of the at least one target embedding vector and the text embedding vector to obtain at least one text parent similarity corresponding to the at least one target embedding vector respectively; and determines a sum of the at least one text parent similarity as a text similarity of the target entity information and the entity information of the first entity.

[0237] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponding to text embedding vector 1 and entity 2 corresponding to text embedding vector 2. Then, similarity 1 of target entity label 1 and text embedding vector 1 is calculated, similarity 2 of target entity label 2 and text embedding vector 1 is calculated, and a sum of similarity 1 and similarity 2 is determined as a text similarity of the target entity information and entity 1. Similarity 3 of target entity label 1 and text embedding vector 2 is calculated, similarity 4 of target entity label 2 and text embedding vector 2 is calculated, and a sum of similarity 3 and similarity 4 is determined as a text similarity of the target entity information and entity 2.

[0238] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponding to text embedding vector 1 and entity 2 corresponding to text embedding vector 2. Then, similarity 1 of target entity label 1 and text embedding vector 1 is calculated, similarity 2 of target entity label 2 and text embedding vector 1 is calculated, and a sum of similarity 1 and similarity 2 is determined as a text similarity of the target entity information and entity 1. Similarity 3 of target entity label 1 and text embedding vector 2 is calculated, similarity 4 of target entity label 2 and text embedding vector 2 is calculated, and a sum of similarity 3 and similarity 4 is determined as a text similarity of the target entity information and entity 2.

[0239] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponding to text embedding vector 1 and entity 2 corresponding to text embedding vector 2. Then, similarity 1 of target entity label 1 and text embedding vector 1 is calculated, similarity 2 of target entity label 2 and text embedding vector 1 is calculated, and a sum of similarity 1 and similarity 2 is determined as a text similarity of the target entity information and entity 1. Similarity 3 of target entity label 1 and text embedding vector 2 is calculated, similarity 4 of target entity label 2 and text embedding vector 2 is calculated, and a sum of similarity 3 and similarity 4 is determined as a text similarity of the target entity information and entity 2.

[0240] For example, for the input target entity information "blue car", first perform word segmentation processing on the target entity information to decompose into target entity labels "blue" and "car", representing that the query target entity has two features. Then calculate the similarity of each feature with each entity in the entity set, take the maximum value of each feature, and finally take the sum of the similarity scores corresponding to all query features, which represents the text similarity between the target entity information and the queried entity.

[0241] As shown in FIG. 13, the target entity information contains two target entity labels "blue" and "car", which are converted into two target embedding vectors. Then the entity information in the entity set is obtained, for example, the entity set includes entity 1 and entity 2, the text labels of entity 1 include "metal", "old", "car", "truck", "blue", "broken", and the text labels of entity 2 include "metal", "scratch", "car", "broken", "money truck", and "black".

[0242] The text similarity between the target entity information and entity 1 is calculated: first, calculate the similarity of the target entity label "blue" with each text label of entity 1, and take the maximum value, for example, the similarity value between the target entity label "blue" and the text label "blue" of entity 1 is 1 (1 is the maximum value); then calculate the similarity of the target entity label "car" with each text label of entity 1, and take the maximum value, for example, the similarity value between the target entity label "car" and the text label "car" of entity 1 is 1, then sum the two similarities of the target entity labels "car" and "blue" to get the final similarity between the target entity information and entity 1 as 2.

[0243] Similarly, the text similarity between the target entity information and entity 2 is calculated: first, calculate the similarity of the target entity label "blue" with each text label of entity 2, and take the maximum value, for example, the similarity value between the target entity label "blue" and the text label "black" of entity 1 is 0.91; then calculate the similarity of the target entity label "car" with each text label of entity 1, and take the maximum value, for example, the similarity value between the target entity label "car" and the text label "car" of entity 2 is 1 (1 is the maximum value), then sum the two similarities of the target entity labels "car" and "blue" to get the final similarity between the target entity information and entity 2 as 1.91.

[0244] It can be seen that the similarity between the target entity information and entity 1 is 2, which is higher than the similarity between the target entity information and entity 2, which is 1.91.

[0245] 2. The similarity includes image similarity between the target entity information and the image embedding vector.

[0246] Taking a first entity in the entity set as an example, the game application (client or server) tokenizes the target entity information to obtain at least one target entity label; converts the at least one target entity label into at least one target embedding vector; obtains entity information of the first entity, the entity information including an image embedding vector, the image embedding vector being an embedding vector extracted based on an image of the first entity; respectively calculates image parent similarities of the at least one target embedding vector and the text embedding vector to obtain at least one image parent similarity corresponding to the at least one target embedding vector respectively; and determines a sum of the at least one image parent similarity as an image similarity of the target entity information and the entity information of the first entity.

[0247] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponding to image embedding vector 1 and entity 2 corresponding to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and a sum of similarity 1 and similarity 2 is determined as the image similarity of the target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and a sum of similarity 3 and similarity 4 is determined as the image similarity of the target entity information and entity 2.

[0248] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponding to image embedding vector 1 and entity 2 corresponding to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and a sum of similarity 1 and similarity 2 is determined as the image similarity of the target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and a sum of similarity 3 and similarity 4 is determined as the image similarity of the target entity information and entity 2.

[0249] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponding to image embedding vector 1 and entity 2 corresponding to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and a sum of similarity 1 and similarity 2 is determined as the image similarity of the target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and a sum of similarity 3 and similarity 4 is determined as the image similarity of the target entity information and entity 2.

[0250] In an optional embodiment, the calculation of the image similarity can use a target image embedding vector, that is, the target image embedding vector is used instead of the target embedding vector in the above method. The target image embedding vector can be obtained by tokenizing the target entity information to obtain at least one target entity label; and inputting the at least one target entity label into a multi-modal model to obtain the target image embedding vector.

[0251] As shown in FIG. 14, based on the target entity information "red door", through the image similarity search, features that do not exist in many text labels can be found, for example, the "red" feature does not exist in the text labels of many "red door" entities, but through the image similarity search, entities with the red door feature can still be searched and the images of these entities are returned. For example, as shown in (1) of FIG. 14, it is a wooden door, and the image similarity is 0.4898; as shown in (2) of FIG. 14, it is a door made of wood board, and the image similarity is 0.4895; as shown in (3) of FIG. 14, it is a metal door, and the image similarity is 0.4539; as shown in (4) of FIG. 14, it is a rusty metal door, and the image similarity is 0.4494; as shown in (5) of FIG. 14, it is a rusty metal door, and the image similarity is 0.4605; as shown in (6) of FIG. 14, it is a door frame made of wood, the top of the door has a horizontal beam, and the left and right sides of the door each has a vertical column, and the top of the vertical column has a horizontal beam, and the image similarity is 0.4581; as shown in (7) of FIG. 14, it is a rusty metal door, and the middle of the door has a rectangular window, and the image similarity is 0.4439; as shown in (8) of FIG. 14, it is a rusty metal door, and the image similarity is 0.4439.

[0252] 3. The similarity includes a text similarity and an image similarity.

[0253] In a case where the first entity and the target entity correspond to a text similarity and an image similarity, the game application determines an average of the text similarity and the image similarity as the similarity of the first entity and the target entity.

[0254] Alternatively, in a case where the first entity and the target entity correspond to a text similarity and an image similarity, a larger value of the text similarity and the image similarity is determined as the similarity of the first entity and the target entity.

[0255] Further alternatively, in a case where the first entity and the target entity correspond to a text similarity and an image similarity, a sum of the text similarity and the image similarity is determined as the similarity of the first entity and the target entity.

[0256] Alternatively, weighting coefficients of the text similarity and the image similarity are preset, and in a case where the first entity and the target entity correspond to the text similarity and the image similarity, a weighted average of the text similarity and the image similarity is determined as the similarity between the first entity and the target entity. In addition, different weighting coefficients can be configured for different types of entities. For example, in a case where entities (such as a microwave oven and a safe) that are easily confused in appearance but have a large difference in text expression exist in the virtual environment, the player can confuse the two and make a mistake. When performing similarity calculation of such entities, the weight of the image similarity is configured to be higher, and it is easier to accurately determine the user's real intention.

[0257] At step 2332, the target entity is determined from the entity set according to the similarity and the environment perception information.

[0258] In an optional embodiment, as shown in FIG. 15, the game application first performs similarity searching 414 according to the target entity information to obtain at least one candidate entity with a similarity higher than a threshold value, and performs sorting according to the similarity; then, based on the position, orientation, direction, and the like of the master virtual character and / or the non-player character, filtering and sorting of the perception are performed 415, and finally the target entity is screened.

[0259] In an optional embodiment, the game application determines a search range according to the environment perception information; and screens the target entity from the entities perceived in the search range according to the similarity.

[0260] The determining of the search range according to the environment perception information can include at least one of the following:

[0261] The visual field range of the master virtual character is determined as the search range according to the position of the master virtual character and the direction of the master virtual character;

[0262] The hearing range of the master virtual character is determined as the search range according to the position of the master virtual character;

[0263] The detection range of the perception virtual prop used by the master virtual character is determined as the search range;

[0264] The visual field range of the non-player character is determined as the search range according to the position of the non-player character and the direction of the non-player character;

[0265] The hearing range of the non-player character is determined as the search range according to the position of the non-player character;

[0266] The detection range of the perception virtual prop used by the non-player character is determined as the search range;

[0267] According to the position of the master virtual role, the orientation of the master virtual role, the position of the non-player role and the orientation of the non-player role, the overlapping part of the visual range of the master virtual role and the non-player role is determined as the search range;

[0268] According to the position of the master virtual role and the position of the non-player role, the overlapping part of the auditory range of the master virtual role and the non-player role is determined as the search range.

[0269] In another optional embodiment, the game application determines the search range according to the target entity information and the environment perception information; and filters the target entity from the entities perceived in the search range according to the similarity.

[0270] For example, in the case that the target entity information includes the position information of the target entity; the search range is determined according to the position information of the target entity and at least one of the visual range, the auditory range of the master virtual role and / or the non-player role, and the search range of the virtual prop indicated by the environment perception information. The search range is the estimated range in which the target virtual role and / or the non-player role can perceive the target entity. For example, the game application determines the perception range according to the environment perception information. The perception range includes at least one of the visual range of the master virtual role, the auditory range of the master virtual role, the visual range of the non-player role, the auditory range of the non-player role, and the detection range of the virtual prop; and the area range in the perception range that matches the position information is determined as the search range.

[0271] Alternatively, the game application determines the x entities with the highest similarity from the entity set according to the similarity to form a candidate entity list; and determines the entities in the search range as the target entity according to the target entity information and the environment perception information. x is a positive integer.

[0272] For example, the game application determines the entity with the highest similarity in the search range as the target entity.

[0273] It should be noted that when there are multiple entities with the same similarity in the search range, the multiple entities can be sorted according to the distance between the entities and the master virtual role. For example, in the case that the number of entities with the highest similarity in the search range is at least two, the entity with the highest similarity and closest to the master virtual role in the search range is determined as the target entity.

[0274] Exemplarily, when the natural language command includes a fuzzy position description of the target entity, the game application program can perform text matching according to the position of the entity recorded in the entity set and the position described in the natural language command, and query the target entity whose position is more consistent with the natural language command. When the target entity information includes the position information of the target entity, the game application program can call the large language model to generate a real-time relative position description of each entity according to the position of each entity recorded in the entity set, the current position of the host virtual character and / or the non-player character, the real-time relative position description is used to describe the relative position relationship between the entity and the host virtual character / non-player character, and the game application program calculates the similarity of the text including the position information and the real-time relative position description.

[0275] For example, as shown in FIG. 16, the search range is determined according to the position and orientation 416 of the host virtual character. For example, the target entity information indicates that the target entity is located in front of the host virtual character, and the search range as shown in FIG. 16 can be constructed, which is a fan-shaped area in front of the host virtual character with a certain field of view size; filter out entities not in the fan-shaped range, that is, filter out entity 1 and entity 4, then the entities in the search range are entity 2, entity 3 and entity 5. Then, sorting is performed according to the similarity of the entities in the search range, and when there are two entities in the search range with the highest and equal similarity scores, that is, the similarity scores of entity 2 and entity 3 are both 2 points, then the entity 2 closer to the host virtual character is selected as the target entity.

[0276] In an optional embodiment, the game application program determines the entity with the highest similarity in the search range and higher than the threshold value as the target entity. Exemplarily, the number of search ranges is at least one. For example, the search range includes at least one of the following: the field of view range of the host virtual character, the hearing range of the host virtual character, the detection range of the host virtual character for perceiving virtual props, the field of view range of the non-player character, the hearing range of the non-player character, the detection range of the non-player character for perceiving virtual props.

[0277] Exemplarily, the game application program can sequentially traverse at least one search range described above, or the game application program can sequentially traverse at least one search range determined according to the natural language command. For example, the natural language command is “Can you see the truck in front of you? Move to the truck”, then the game application program can preferentially traverse the field of view range of the host virtual character, and then traverse the field of view range of the non-player character, and match the entity with the highest similarity and higher than the threshold value from the field of view range of the non-player character.

[0278] Exemplarily, when there are at least two search ranges, the traversal order of the at least two search ranges can be preset, or the traversal order of the at least two search ranges can be determined according to the intention indicated in the natural language command. For example, when the intention is to search for an item, the visual field range is traversed preferentially. When the intention is to search for a combat scene, the auditory range is traversed preferentially.

[0279] If there are multiple levels of nested instructions in the natural language command, the first target entity is searched first, and then the next target entity is searched according to the position of the previous target entity. In the case where the natural language command includes a second target entity determined based on the position of the first target entity, the first target entity is queried from the entity set of the virtual environment according to the first target entity information of the first target entity indicated by the natural language command and the environment perception information; and then the second target entity is queried from the entity set of the virtual environment according to the second target entity information of the second target entity, the first target entity and the environment perception information.

[0280] In the multiple levels of nested instructions, at least two entities are included in the natural language command, and at least one entity of the at least two entities is used as a reference entity to determine a target entity. That is, the reference entity is used to limit the target entity.

[0281] For example, the natural language command describes the positional relationship between the reference entity and the target entity, such as the natural language command is “move to the blue car behind the red car”, the red car is the reference entity, and the blue car is the target entity, and the target entity is located behind the reference entity.

[0282] Alternatively, the natural language command describes the similarity between the reference entity and the target entity, such as the natural language command is “find out if there is another red chest nearby”, the reference entity is the red chest currently visible to the host virtual character, and the target entity is the red chest not discovered by the host virtual character nearby, and the target entity and the reference entity have the same appearance.

[0283] Exemplarily, the target entity information includes reference entity information of the reference entity, and the reference entity is used to determine the target entity; the game application queries the reference entity from the entity set according to the reference entity information and the environment perception information; and queries the target entity from the entity set according to the reference entity, the target entity information and the environment perception information.

[0284] For example, the target entity information includes: reference entity information, position relationship between the reference entity and the target entity, and feature text of the target entity; the game application calculates the similarity between the feature text and each entity information in the entity set; determines the search range according to the environmental perception information; and screens the target entity from the entities in the entity set within the search range according to the similarity and the position relationship.

[0285] Optionally, the target entity information can include reference entity information of n reference entities, the i th reference entity in the n reference entities is used to determine / limit the i+1 th reference entity, n is an integer greater than 1, and i is a positive integer less than n. Then the game application queries the i th reference entity from the entity set according to the reference entity information of the i th reference entity and the environmental perception information; queries the i+1 th reference entity from the entity set according to the reference entity information of the i th reference entity and the i+1 th reference entity and the environmental perception information, until the n th reference entity is queried; and queries the target entity from the entity set according to the n th reference entity, the target entity information and the environmental perception information.

[0286] For example, as shown in FIG. 17, when the natural language command is “go to the broken car on the left side of the red truck in front”, the game application will find the nearest red truck 503 based on the position and orientation of the master virtual role, and then find the nearest car wreckage 504 on the left side of the red truck 503.

[0287] In summary, the method provided in the embodiment proposes a game scene perception and query method based on deep learning, which can initiate a query positioning within a spatial range to a game scene in the form of natural language. Combined with pre-built static game scene data and real-time game data, the method can control non-player characters, for example, the player inputs: “surround the house in front” and “hide behind the stone”. The method can analyze the entity positions of the “house” and “stone” based on the current position of the player, and make the non-player character move to a reasonable position next to the entity. In addition, the method can also adapt to complex and ambiguous instructions, such as “go to the red truck inside the house” and “go upstairs”.

[0288] The method provided in the embodiment can be applied to a product of voice command of non-player characters in shooting games, and can complete the scheduling function of thousands of target positions in the game scene. It can adapt to any text instructions input by the player, analyze the action mode of the non-player character, and the target position to be reached, wherein the action mode corresponds to the fixed instructions of the traditional scheme, and the position to be reached is the fine control of the non-player character. Finally, accurate and complex control of the non-player character is realized.

[0289] The method provided by the embodiment can greatly improve the game experience of a player controlling a non-player character, the player can issue any instruction to direct teammates to complete the scheduling of thousands of target positions in a game scene, and the degree of freedom is high. Compared with the related art, the method saves a large amount of game design and development costs.

[0290] The method provided by the embodiment realizes the description information of the entity in the text dimension and the image dimension, and predicts the visual label of the entity by calling the multi-modal model to perform prediction on the text label and the image. In the process of predicting the visual label of the entity, the description information of the entity in multiple dimensions is used, so that the characteristics of the entity can be fully described from the natural language semantics in the text dimension and the visual information in the image dimension, and the accuracy of the visual label is ensured. The multi-modal model is called to perform prediction on the text label and the image, so that the labeling of the visual label of the entity can be realized in batches, and the efficiency of obtaining the visual label of the entity is improved.

[0291] The method provided by the embodiment realizes the description information of the entity in the text dimension and the image dimension, and predicts the visual label of the entity by calling the multi-modal model to perform prediction on the text label and the image. In the process of predicting the visual label of the entity, the description information of the entity in multiple dimensions is used, so that the characteristics of the entity can be fully described from the natural language semantics in the text dimension and the visual information in the image dimension, and the accuracy of the visual label is ensured. The multi-modal model is called to perform prediction on the text label and the image, so that the labeling of the visual label of the entity can be realized in batches, and the efficiency of obtaining the visual label of the entity is improved.

[0292] The method provided by the embodiment realizes the description information of the entity in the text dimension and the image dimension, and predicts the visual label of the entity by calling the multi-modal model to perform prediction on the text label and the image. In the process of predicting the visual label of the entity, the description information of the entity in multiple dimensions is used, so that the characteristics of the entity can be fully described from the natural language semantics in the text dimension and the visual information in the image dimension, and the accuracy of the visual label is ensured. Further, the visual label is subjected to pseudo-spontaneous rewriting to obtain a matching label that is more similar to a spontaneous expression, so that the accuracy of the matching label in describing the entity in a spontaneous expression scene is improved. Further, the spatial position of the entity in a virtual environment is obtained, and the position, size and the like of the scene in the virtual environment are described, so that more basis is provided for screening the entity in the virtual environment.

[0293] An example embodiment of the present application provides a method for commanding a non-player character. The method comprises:

[0294] Step 1, display at least one of a main control virtual character and an NPC located in a virtual environment;

[0295] For example, the master virtual role is a virtual role directly controlled by the user in the virtual environment. There is one or more non-player characters (NPCs) in the virtual environment; the NPCs and the master virtual role belong to the same virtual camp and are teammates of the master virtual role, and the NPCs follow the master virtual role to move in the virtual environment. In addition, the NPCs can also be follower roles, pet roles, etc. controlled by the master virtual role. In addition, the NPCs can also be neutral roles and perform cooperative actions with the master virtual role only when certain conditions are met.

[0296] Step 2, obtaining a natural language command;

[0297] For example, the natural language command includes a natural semantic instruction for the NPC. The natural language command for instructing the NPC in the virtual scene can be directly input by the user or extracted from voice information input by the user, and the application does not limit the obtaining method of the natural language command.

[0298] The natural language command includes at least one of a behavior intention and a scene entity; the behavior intention is used to indicate a type of virtual activity performed by the NPC, such as indicating a type of virtual activity performed in the virtual environment; the scene entity includes a first scene object, such as indicating a type of virtual activity performed on a scene object in the virtual environment. The behavior intention and the scene entity included in the natural language command are part of the words in the natural language command.

[0299] In one example, the natural language command controls the NPC from two aspects of performing a type of virtual activity and performing the virtual activity on a scene object in the virtual environment; it is a complex instruction for controlling the NPC.

[0300] Step 3, controlling the NPC to perform a virtual activity in response to the natural language command;

[0301] For example, the virtual activity performed by the NPC is determined based on environmental perception information of the master virtual role and / or the NPC. In some examples, the master virtual role and / or the NPC obtain environmental perception information within a perception range of visual, auditory, action trajectory, etc. For example, the virtual activity is close to the current perception situation of the virtual environment of the master virtual role and / or the NPC, and the NPC is controlled with reference to the current perception situation of the virtual environment of the virtual role and / or the NPC. For example, the NPC has autonomous behavior capability, and the user only instructs the NPC. The user gives a command, and the NPC understands and executes the command based on its own autonomous behavior capability.

[0302] To sum up, the method provided in the embodiment controls the complex instructions of the NPC from two aspects of executing what type of virtual activity and executing the virtual activity on what scene object in the virtual environment, determines the virtual activity to be executed by taking the environment perception information of the host virtual character and / or the NPC as reference, and ensures the NPC to execute the virtual activity in cooperation with the host virtual character.

[0303] FIG. 18 shows a method flowchart of a method for commanding a non-player character according to an example embodiment of the present application. The method can be performed by the terminal in FIG. 1.

[0304] Step 201: obtaining a spatial data set of a virtual environment;

[0305] The spatial data set of the virtual environment includes a visual label of a scene object, and the visual label is used to describe the visual features of the scene object in at least one dimension, such as material, transparency, color, shape, etc.

[0306] Step 202: displaying at least one of a host virtual character and an NPC in the virtual environment;

[0307] For example, the host virtual character is a virtual character directly controlled by a user in the virtual environment. There is one or more NPC in the virtual environment; the NPC and the host virtual character belong to the same virtual camp and are teammates of the host virtual character.

[0308] Step 203: obtaining a natural language command in the form of speech;

[0309] The natural language command includes a natural semantic for commanding the NPC, and the command information includes a behavior intention and a scene entity; the behavior intention is used to indicate the type of virtual activity executed by the NPC, and the scene entity is used to indicate the target entity of the virtual activity;

[0310] Step 204: converting the natural language command in the form of speech into a natural language command in the form of text;

[0311] For example, the natural language command in the form of speech is subjected to automatic speech recognition (ASR) processing to determine the natural language command in the form of text. The automatic speech recognition processing usually includes calling acoustic models, language models, etc. to realize the recognition of the pronunciation, vocabulary and syntax structure, etc. of the natural language command in the form of speech, and convert it into the natural language command in the form of text.

[0312] Step 205: performing intention recognition on the natural language command to obtain a first classification label;

[0313] For example, the first classification label corresponds to a first intention indicating a command intention for the non-player role in one intention dimension; such as at least one of the following: indicating which non-player role is commanded by the command text among a plurality of non-player roles, indicating whether the virtual activity performed by the non-player role is related to initiating a virtual attack, and indicating what virtual activity is performed by the non-player role.

[0314] Step 206: performing entity recognition on the natural language command to obtain a target entity;

[0315] The target entity is determined from the virtual environment in combination with environment perception information, and the environment perception information includes information perceived by at least one of the master virtual role and the non-player role from the virtual environment. The target entity can be an entity currently seen by the player, or an entity historically seen by the player.

[0316] Step 207: in response to the natural language command, controlling the NPC to perform a virtual activity according to the first intention corresponding to the first classification label, or controlling the NPC to perform a virtual activity associated with the target entity, or controlling the NPC to perform a virtual activity associated with the target entity according to the first intention corresponding to the first classification label;

[0317] For example, the NPC is controlled to perform a virtual activity according to the indication of the first classification label and / or the target entity.

[0318] Step 208: broadcasting feedback information of the NPC;

[0319] For example, the game application program can generate corresponding feedback information in real time according to the environment perception information of the non-player role, and perform broadcasting. For example, the game application program can generate corresponding feedback information in real time according to the environment perception information triggering the broadcasting condition when the environment perception information of the non-player role triggers the broadcasting condition; or the game application program can generate corresponding feedback information in real time according to the environment perception information of the non-player role when the non-player role or the master virtual role triggers the broadcasting condition.

[0320] The preprocessing stage of step 201 can be implemented as sub-step 1 to sub-step 3:

[0321] Sub-step 1: obtaining attribute text of scene objects in the virtual scene;

[0322] Exemplarily, the attribute text is used to introduce the inherent attribute of the scene object in the virtual scene. On the one hand, the attribute text realizes the description of the scene object in the text mode, and provides semantic information for the prediction of the visual label of the scene object while describing the scene object. On the other hand, the attribute text is the description of the scene object in the virtual scene, and in the case of a large number of scene objects in the virtual scene and the reuse of object models, the inherent attribute of the scene object in the virtual scene can be more accurately described.

[0323] In an optional implementation, the attribute text includes at least one of a name of the scene object in the virtual scene and a size of the scene object in the virtual scene.

[0324] Substep 2: Obtain an appearance image of the scene object in the virtual scene;

[0325] Exemplarily, the appearance image is used to describe the style of the scene object. The appearance image carries the appearance style such as color, texture, shape and the mutual positional relationship between the subparts of the scene object in the picture mode, and can comprehensively describe the scene object from the picture mode (or visual mode). In an optional implementation, the appearance image of the scene object includes images obtained by observing the scene object from at least two perspectives.

[0326] Substep 3: calling a multi-modal model to perform prediction on the attribute text and the appearance image of the scene object to obtain a visual label of the scene object;

[0327] Exemplarily, the multi-modal model has the ability to perform model prediction on text information and picture information of different modes. In this embodiment, the input parameters of the multi-modal model are the attribute text and the appearance image of the scene object. The multi-modal model predicts the visual label of the scene object from the two modes of the text mode and the picture mode, and the visual label is used to describe the visual features of the scene object in at least one dimension.

[0328] In some embodiments, substep 3 can be implemented as substep 31 and substep 32:

[0329] Substep 31: constructing a question sentence for the appearance image;

[0330] Substep 32: inputting the appearance image of the scene object and the question sentence into a visual question and answer model to obtain an answer sentence, and taking the answer sentence as the visual label of the scene object;

[0331] Optionally, the multi-modal model includes a visual question and answer model; the attribute text is carried in the question sentence; the question sentence is used to guide the visual question and answer model to convert the appearance image into the visual label of the scene object on the one hand, and the question sentence provides supplementary information of the scene object in the form of text while guiding the conversion of the visual label of the scene object on the other hand.

[0332] In an optional implementation, the sub-step 31 can be implemented as the following sub-step 311 and sub-step 312.

[0333] The sub-step 311 obtains expected information of the scene object.

[0334] For example, the expected information is used to indicate a description dimension of the scene object expected in the visual label, and / or an expected format of the visual label; in an example, the expected information used to indicate the description dimension of the scene object expected in the visual label includes but is not limited to at least one of a type description, a material, a transparency, a color, a surface feature, and a shape of the scene object. In another example, the expected information used to indicate the expected format of the visual label is at least one of a Comma Separated Values (CSV), a JavaScript Object Notation (JSON), and an eXtensible Markup Language (XML).

[0335] The sub-step 312 constructs a question sentence of the appearance image according to the expected information and the attribute text.

[0336] For example, a first sub-part in the question sentence is supplementary introduction information of the scene object, carrying the attribute text of the scene object; a second sub-part in the question sentence is a response guiding sentence for the visual question and answer model, carrying the expected information.

[0337] Optionally, the spatial data set includes a matching label of the scene object in the virtual scene; correspondingly, the visual label of the scene object is executed with the quasi-spontaneous rewriting to obtain the matching label conforming to the spontaneous language expression.

[0338] For example, as introduced above, the visual label is used to describe the visual feature of the scene object in at least one dimension, and the appearance image of the scene object presents rich visual features of the scene object. However, the description of the scene object in the spontaneous language expression cannot cover the visual features of the scene object in each dimension. The purpose of executing the quasi-spontaneous rewriting on the visual label of the scene object is to obtain the matching label more similar to the spontaneous language expression; it can be understood that the quasi-spontaneous rewriting can be deleting part of the content in the visual label, or changing the visual label to a label with the same semantics but different text expression.

[0339] In an optional implementation, the quasi-spoken rewriting is performed by invoking a natural language model; the visual label of the scene object is input into the natural language model, and a matching label conforming to the spoken expression of the natural language is predicted. The quasi-spoken rewriting on the visual label is implemented based on the invocation of the natural language model. For example, the natural language model carries prior knowledge of the spoken expression of the natural language. For example, the natural language model is implemented as a large language model (LLM).

[0340] Optionally, the spatial data set further comprises spatial information of the scene object, and correspondingly further comprises:

[0341] The spatial position of the scene object in the virtual scene is obtained, and the spatial position of the scene object is determined as auxiliary information of the visual label of the scene object.

[0342] For example, the spatial position of the scene object in the virtual scene is used to indicate the deployment of the scene object in the virtual scene, and the spatial information is used to indicate the position, size, etc. of the scene object in the virtual scene after the deployment of the scene object in the virtual scene.

[0343] Further, at least one of the coordinate position, the orientation information, the bounding box information, and the cover point information of the scene object in the virtual scene is obtained.

[0344] The coordinate position (Location) is used to indicate the position of the scene object in the virtual scene, such as the coordinate information of the center point or the preset point of the scene object in the virtual scene; the orientation (Rotation) is used to indicate the direction faced by the scene object in the virtual scene, such as the direction faced by the front of the scene object in the virtual scene; the bounding box (Bounding Box) is used to indicate the size of the scene object in the virtual scene; and the cover point (Cover points) is used to indicate the recommended position point of the virtual character when the virtual character approaches the scene object, so that the scene object can mask the virtual character.

[0345] For the intent recognition stage of step 205, a sub-step 4 can be implemented:

[0346] Sub-step 4: invoking at least one hierarchical prediction network to perform intent recognition on the natural language command to obtain a first classification label of the natural language command commanding the NPC;

[0347] The hierarchical prediction network can be configured to predict the classification label corresponding to the command text. The hierarchical prediction network can include at least two sub-networks, an upper network and a lower network, which are cascaded. The lower network can further perform the prediction of the classification label based on the prediction result output by the upper network. The at least two sub-networks can be constructed based on a tree structure of multiple classification labels. The hierarchical prediction network can be configured to construct at least two sub-networks corresponding to the hierarchical structure of the tree structure of the classification labels, and split the prediction task of the classification labels into prediction sub-tasks performed by the at least two sub-networks, so as to reduce the complexity of the classification prediction of each sub-network.

[0348] In an optional implementation, each of the at least one hierarchical prediction network is configured to predict a command intention of the command text in one intention dimension for the non-player character. The intention dimension can include at least one of a subject dimension, a semantic dimension, and a behavior dimension. In an example, the i-th sub-network in each of the at least one hierarchical prediction network is configured to predict a first-level behavior label of the command text in one intention dimension, and the i+1-th sub-network in the hierarchical prediction network is configured to predict a second-level behavior label corresponding to the first-level behavior label. i is a positive integer. The high-level behavior label (e.g., the first-level behavior label) predicted by the high-level sub-network (e.g., the i-th sub-network) in the hierarchical prediction network can include multiple sub-labels or subordinate low-level labels. The corresponding low-level sub-network (e.g., the i+1-th sub-network) needs to be called to further perform the classification prediction (e.g., to predict the second-level behavior label). The prediction task of the classification labels can be split into prediction sub-tasks performed by the at least two sub-networks, so as to reduce the complexity of the classification prediction of each sub-network.

[0349] Further, the intention dimension can include the subject dimension, and the hierarchical prediction network can include a hierarchical structure of a subject prediction network. The subject prediction network can be configured to predict the subject type in the natural language command. The subject type can be used to indicate the identity of the NPC commanded by the natural language command. In an example, the subject prediction network can be configured to predict which of the non-player characters is commanded by the command text. The number of the non-player characters commanded by the command text can be one or more.

[0350] That is, the game application inputs the natural language command into the subject prediction network to obtain a subject type output by the network, and the subject type is used to determine an action execution subject indicated in the natural language command. Optionally, the game application can also determine the NPC commanded by the natural language command, i.e., the action execution subject, according to the subject type output by the subject prediction network and the environment perception information. For example, according to the environment perception information, the NPC of the subject type closest to the master virtual character is determined as the action execution subject; or, according to the environment perception information, the NPC of the subject type within the visual range of the master virtual character is determined as the action execution subject; or, according to the environment perception information, the NPC of the subject type to which the visual focus of the master virtual character is directed is determined as the action execution subject.

[0351] And / or, the intent dimension includes a semantic dimension, and the hierarchical prediction network includes a hierarchical semantic prediction network having the capability of predicting a semantic type in the natural language command; the semantic type is used to indicate a control manner of the NPC for initiating a virtual attack; and, for example, the subject prediction network is used to predict whether a virtual activity performed by the non-player character commanded by the command text is related to initiating a virtual attack.

[0352] And / or, the intent dimension includes a behavior dimension, and the hierarchical prediction network includes a hierarchical behavior prediction network having the capability of predicting a behavior intent type in the natural language command; the behavior intent type is used to indicate a behavior manner of the NPC for performing a virtual activity. For example, the behavior prediction network is used to predict what virtual activity is performed by the non-player character.

[0353] Correspondingly, in step 207, in response to the natural language command, the NPC is controlled to perform a virtual activity according to a first intent corresponding to the first classification label;

[0354] For example, the first intent corresponding to the first classification label indicates a command intent for the non-player character in one intent dimension; such as at least one of the following: indicating which character of the non-player characters is commanded by the command text, indicating whether a virtual activity performed by the non-player character is related to initiating a virtual attack, and indicating what virtual activity is performed by the non-player character. The NPC is controlled to perform a virtual activity according to the indication of the first classification label.

[0355] The target entity recognition process for step 206 can be implemented as sub-step 5:

[0356] Sub-step 5: querying a target entity from an entity set of the virtual environment according to target entity information of the target entity indicated by the natural language command and the environment perception information;

[0357] The target entity is determined from the virtual environment in combination with environmental perception information, and the environmental perception information includes information perceived by the host virtual character and / or non-player character from the virtual environment.

[0358] Since the target entity information in the natural language command is usually ambiguous, for example, the natural language command can be "move to the red truck", and the target entity information is "red truck", and there can be many red trucks in the virtual environment, and the target entity cannot be accurately determined from the virtual environment according to the target entity information in the natural language command.

[0359] Therefore, the embodiment of the present application provides a method for determining a target entity in combination with target entity information and environmental perception information. For the above example, the "red truck" expressed by the player in the natural language command should be a red truck that the player can see, and therefore, in combination with the visual range of the host virtual character, the red truck located in the visual range of the host virtual character can be filtered from the multiple red trucks in the virtual environment, and the red truck is the target entity indicated by the player in the natural language command.

[0360] Since the natural language command is issued by the player based on the perception of the virtual environment, in order to accurately identify the target entity indicated in the natural language command, the game application program will combine the environmental perception information of the player when issuing the natural language command to identify the target entity in the natural language command. Based on the player's perception of the virtual environment, the entity closest to the target entity information that the player can perceive in the virtual environment is inferred as the target entity.

[0361] For example, the environmental perception information of the host virtual character can include: the picture or entity that the host virtual character can see when observing the virtual environment, the environmental sound that the host virtual character can hear, the source direction of the environmental sound, the type and sound size of the environmental sound, the perception information obtained by the host virtual character by using the perception skill (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and the perception information obtained by the host virtual character by using the perception virtual prop (for example, the sensing signal of the sensor, the positioning signal of the positioning prop, etc.).

[0362] For example, the environmental perception information of the host virtual character can include: the picture or entity that the host virtual character can see when observing the virtual environment, the environmental sound that the host virtual character can hear, the source direction of the environmental sound, the type and sound size of the environmental sound, the perception information obtained by the host virtual character by using the perception skill (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and the perception information obtained by the host virtual character by using the perception virtual prop (for example, the sensing signal of the sensor, the positioning signal of the positioning prop, etc.).

[0363] The environment perception information of the non-player character can include: a picture or entity that the non-player character can see when observing the virtual environment, an environmental sound that the non-player character can hear, a source direction of the environmental sound, a type and sound size of the environmental sound, perception information obtained by the non-player character by using a perception skill (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and perception information obtained by the non-player character by using a perception virtual prop (for example, a sensing signal of a sensor, a positioning signal of a positioning prop, etc.).

[0364] It should be noted that the environment perception information can include information that the host virtual character and the non-player character perceive in real time when the natural language command is received, or information that the host virtual character and the non-player character perceive historically before the natural language command is received. That is, the target entity can be an entity that the player currently sees, or an entity that the player has historically seen. Therefore, the game application needs to combine the real-time environment perception information and the historical environment perception information of the host virtual character and / or the non-player character to perform identification of the target entity.

[0365] For example, when the natural language command is “Let’s go back to the hotel we just passed by”, the game application needs to query the hotel that the host virtual character has passed by according to the historical environment perception information of the host virtual character.

[0366] In an optional embodiment, the client queries the target entity from the entity set of the virtual environment according to the target entity information indicated by the natural language command and the environment perception information. In another optional embodiment, the client reports the natural language command to the server, and the server queries the target entity from the entity set of the virtual environment according to the target entity information indicated by the natural language command and the environment perception information.

[0367] Optionally, the game application infers the target entity according to the target entity information, the environment perception information, and preprocessed virtual environment data. The preprocessed virtual environment data includes the entity set of the virtual environment.

[0368] The entity set includes entity information of each entity in the virtual environment. The entity information includes at least one of the following: a name, a type, a location, a feature, a text label, an embedding vector of the text label, an image, and an embedding vector of the image.

[0369] The text label can be at least one word obtained by tokenizing a feature (for example, a description text of an appearance of an entity), and an embedding model is called to obtain an embedding vector corresponding to each text label.

[0370] The image can include an image obtained by observing a three-dimensional model of the entity from at least one direction, for example, the image can include a three-view image of the entity. The embedding vector of the image is an embedding vector of a visual feature label. The visual feature label is obtained by calling a multi-modal model to perform feature recognition based on the image and the text label of the entity.

[0371] For example, the entity set includes a first type of entity and a second type of entity, the first type of entity includes at least one of a region and a building, and the second type of entity includes an entity object. The text label of the first type of entity is manually labeled, for example, the text label of a certain room of a certain building is manually labeled as: a certain region, a certain building, a certain floor, and a room. The text label of the second type of entity is automatically generated by calling a large language model.

[0372] The method for generating the text label of the second type of entity can include: obtaining at least one image of the entity, inputting the at least one image into the large language model to obtain a description text of the entity; performing word segmentation on the description text to obtain at least one text label of the entity, each text label including one word obtained after word segmentation; and performing a vectorization operation on the text label to obtain an embedding vector corresponding to the text label.

[0373] The embedding vector is a high-dimensional vector data of the text label, and the embedding vector of the text label is generated in advance, so that it is not necessary to repeatedly perform vector generation on the text label of each entity when matching the target entity based on the target entity information from the entity set, and the matching efficiency of the target entity information and the entity information can be improved.

[0374] It should be noted that the method does not directly use the description text as the text label, but uses the word segmentation result of the description text as the text label, because the description text is usually of different lengths, and for a longer description text, its embedding vector has poor effect in actual search and comparison of similarity. For example, searching for “a truck” in “a rusty blue small truck with a rusty body” will probably be ranked after “a car” in the recall result, but in fact, for the search of “truck”, no matter how complex the additional description is, “truck” should be searched first, so when performing similarity search, the comparison should be based on individual feature words, not the description text.

[0375] In addition, in order to further extract the visual features of the entity, the method provided in the embodiments of the present application can further extract hidden visual features in the image of the entity based on the text label and the image of the entity. Taking the first entity as an example, the method includes: obtaining at least one perspective image of the first entity; obtaining at least one text label of the first entity; calling a multi-modal model to extract visual features of the first entity based on the at least one perspective image and the at least one text label of the first entity, and obtaining an image embedding vector of the first entity.

[0376] Since the text label is manually annotated or generated by a large language model based on human language characteristics, the text label extracted based on human language habits may ignore some features of the entity. For example, when manually annotating the text label for an oil drum, the more noticeable text labels such as "metal" and "rust" may be annotated, and the detailed features such as "rust", "blue", "yellow", and "right lower corner paint drop" may be ignored. Therefore, the method provided in the embodiments of the present application also uses a multi-modal model to extract more visual features from the text label and the image of the entity based on the multi-modal model, and performs image similarity matching with the target entity based on the visual features to improve the recognition accuracy of the target entity.

[0377] The training method of the multi-modal model can be: inputting the sample entity image and the sample label into the pre-trained multi-modal model, performing fine-tuning on the pre-trained multi-modal model according to the loss of the predicted label output by the pre-trained multi-modal model and the sample label, and obtaining a multi-modal model capable of outputting visual feature labels based on input text labels and images. The visual feature label includes more detailed description text extracted from the image. For example, the text description of "a metal oil drum" can only be associated with the two text labels "metal" and "oil drum", but the Clip (multi-modal) model can also identify hidden visual information from the image of the oil drum, such as the visual feature labels "rust", "blue", and "yellow". These visual features are not present in the text label, so combining the Clip model to perform visual feature search can further improve the accuracy of entity search.

[0378] In an optional embodiment, the game application program can query the target entity by the following method.

[0379] (1) Analyzing the natural language command to obtain target entity information; the target entity information includes at least one of the following: entity type, entity name, entity position, and entity feature.

[0380] Optionally, the game application program calls a large language model to perform analysis on the natural language command to obtain the target entity information of the target entity. The embedding model is called to perform vectorization processing on the target entity information to obtain a target embedding vector of the target entity information.

[0381] For example, the natural language command is "come here in front of the blue truck", and the target entity information analyzed by the large language model includes the position information "in front of" and the entity name "blue truck".

[0382] Exemplarily, a named entity recognition technique in the field of natural language processing can also be called to extract target entity information from the natural language command. The named entity recognition technique is used to identify entities with specific meanings in text, for example, to identify names, place names, adverbs, adjectives, etc.

[0383] For example, the natural language command is "find a paper box behind the red sofa on the first floor of the motel". Using the named entity recognition technique, the building name "motel first floor", the article name "sofa" and "paper box", the adverb "behind", and the adjective "red" can be identified. After logical construction, a hierarchical scene query call form with search type, content, adverb, and constraint information is formed. Among them, the search type is used to narrow the data retrieval range based on similarity matching, the search content is the specific entity description, and the adjectives are combined with the description text for query. The floor constraint is determined by the player's location, which needs to narrow the range up and down in the indoor scene, thereby avoiding finding entities that are not visible across floors.

[0384] (2) Calculate the similarity of the target entity information and the entity information of each entity in the entity set.

[0385] Exemplarily, the entity set includes entity information of at least one entity, and the entity information of each entity can include at least one of a text embedding vector of a text label and an embedding vector of a visual feature. Then the target entity information can be calculated with the text embedding vector and the image embedding vector respectively. The entity with a similarity higher than a threshold is determined as the target entity.

[0386] The methods of performing similarity matching with the text embedding vector and the image embedding vector are given below respectively.

[0387] 1) The similarity includes a text similarity of the target entity information and the text embedding vector.

[0388] Taking a first entity in the entity set as an example, the game application program (client or server) tokenizes the target entity information to obtain at least one target entity label; converts the at least one target entity label into at least one target embedding vector; obtains the entity information of the first entity, which includes a text embedding vector converted based on the text label of the first entity; respectively calculates the text parent similarity of the at least one target embedding vector and the text embedding vector to obtain at least one text parent similarity corresponding to the at least one target embedding vector respectively; and determines the sum of the at least one text parent similarity as the text similarity of the target entity information and the entity information of the first entity.

[0389] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to text embedding vector 1, and entity 2 corresponds to text embedding vector 2. Then, similarity 1 of target entity label 1 and text embedding vector 1 is calculated, similarity 2 of target entity label 2 and text embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the text similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and text embedding vector 2 is calculated, similarity 4 of target entity label 2 and text embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the text similarity of target entity information and entity 2.

[0390] For example, the entity information of the first entity includes at least one text embedding vector, and the at least one target embedding vector includes a first target embedding vector. The game application program calculates text sub-similarities of the first target embedding vector and the at least one text embedding vector respectively, and obtains at least one text sub-similarity. The highest value in the at least one text sub-similarity is determined as a text parent-similarity corresponding to the first target embedding vector.

[0391] For example, the number of target entity labels is one: target entity label 1. The number of entities in the entity set is also one: entity 1. Entity 1 corresponds to text embedding vector 1 and text embedding vector 3. Then, similarity 1 of target entity label 1 and text embedding vector 1 is calculated, similarity 5 of target entity label 1 and text embedding vector 3 is calculated, and the greater value of similarity 1 and similarity 5 is determined as the similarity of target entity information (target entity label 1) and entity 1.

[0392] For example, for the input target entity information "blue car", first perform word segmentation processing on the target entity information to decompose into target entity labels "blue" and "car", representing two features of the query target entity. Then, the similarity of each feature with each entity in the entity set is calculated, the maximum value of the similarity of each feature is taken, and finally the sum of the similarity scores corresponding to all query features is taken, which represents the text similarity of the target entity information and the queried entity.

[0393] For example, the target entity information includes two target entity labels "blue" and "car", which are converted into two target embedding vectors. Then, the entity information in the entity set is obtained, for example, the entity set includes entity 1 and entity 2, the text labels of entity 1 include "metal", "old", "car", "truck", "blue", and "damaged", and the text labels of entity 2 include "metal", "scratch", "car", "damaged", "money truck", and "black".

[0394] The similarity between the target entity information and the text of entity 1 is calculated: first, the similarity between the target entity label "blue" and each text label of entity 1 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "blue" and the text label "blue" of entity 1 is 1 (1 is the maximum value); then, the similarity between the target entity label "car" and each text label of entity 1 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "car" and the text label "car" of entity 1 is 1, and then the sum of the two similarities of the target entity labels "car" and "blue" is taken to obtain the final similarity between the target entity information and the entity information of entity 1, which is 2.

[0395] Similarly, the similarity between the target entity information and the text of entity 2 is calculated: first, the similarity between the target entity label "blue" and each text label of entity 2 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "blue" and the text label "black" of entity 1 is 0.91; then, the similarity between the target entity label "car" and each text label of entity 1 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "car" and the text label "car" of entity 2 is 1 (1 is the maximum value), and then the sum of the two similarities of the target entity labels "car" and "blue" is taken to obtain the final similarity between the target entity information and the entity information of entity 2, which is 1.91.

[0396] It can be seen that the similarity between the target entity information and entity 1 is 2, which is higher than the similarity between the target entity information and entity 2, which is 1.91.

[0397] 2) The similarity includes an image similarity between the target entity information and the image embedding vector.

[0398] Taking the first entity in the entity set as an example, the game application (client or server) tokenizes the target entity information to obtain at least one target entity label; converts the at least one target entity label into at least one target embedding vector; obtains the entity information of the first entity, and the entity information includes an image embedding vector, which is an embedding vector extracted based on the image of the first entity; calculates the image parent similarity between the at least one target embedding vector and the text embedding vector respectively to obtain at least one image parent similarity corresponding to the at least one target embedding vector respectively; and determines the sum of the at least one image parent similarity as the image similarity between the target entity information and the entity information of the first entity.

[0399] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to image embedding vector 1, and entity 2 corresponds to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the image similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the image similarity of target entity information and entity 2.

[0400] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to image embedding vector 1, and entity 2 corresponds to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the image similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the image similarity of target entity information and entity 2.

[0401] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to image embedding vector 1, and entity 2 corresponds to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the image similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the image similarity of target entity information and entity 2.

[0402] In an optional embodiment, the calculation of image similarity can use target image embedding vectors, that is, using target image embedding vectors instead of target embedding vectors in the above method. The target image embedding vector can be obtained by the following method: tokenizing the target entity information to obtain at least one target entity label; inputting the at least one target entity label into a multi-modal model to obtain a target image embedding vector.

[0403] 3) The similarity includes text similarity and image similarity.

[0404] In the case where the first entity and the target entity correspond to text similarity and image similarity, the game application determines the average of the text similarity and the image similarity as the similarity between the first entity and the target entity.

[0405] Alternatively, in the case where the first entity and the target entity correspond to text similarity and image similarity, the greater value of the text similarity and the image similarity is determined as the similarity between the first entity and the target entity.

[0406] Alternatively, in a case where the first entity and the target entity correspond to a text similarity and an image similarity, a sum of the text similarity and the image similarity is determined as the similarity of the first entity and the target entity.

[0407] (3) determining the target entity from the entity set according to the similarity and the environmental perception information.

[0408] In an optional embodiment, the game application first performs similarity searching according to the target entity information to obtain at least one candidate entity with a similarity higher than a threshold value, and performs sorting according to the similarity; then performs filtering and sorting of the perception based on the position, orientation, direction, etc. of the master virtual character and / or the non-player character, and finally screens the target entity.

[0409] For example, the game application determines a searching range according to the target entity information and the environmental perception information; and screens the target entity from the entities perceived in the searching range according to the similarity.

[0410] Alternatively, the game application determines x entities with the highest similarity from the entity set according to the similarity to form a candidate entity list; and determines the entities in the searching range from the x entities as the target entity according to the target entity information and the environmental perception information. x is a positive integer.

[0411] For example, the game application determines the entity with the highest similarity in the searching range as the target entity.

[0412] It should be noted that when there are multiple entities with equal similarity in the searching range, the multiple entities can be sorted according to the distance between the entities and the master virtual character. For example, in a case where the number of entities with the highest similarity in the searching range is at least two, the entity with the highest similarity in the searching range and closest to the master virtual character is determined as the target entity.

[0413] In an optional embodiment, the game application determines the entity with the highest similarity in the searching range and higher than a threshold value as the target entity. For example, the number of searching ranges is at least one. For example, the searching range includes at least one of the following: the field of view range of the master virtual character, the hearing range of the master virtual character, the detection range of the virtual prop perceived by the master virtual character, the field of view range of the non-player character, the hearing range of the non-player character, and the detection range of the virtual prop perceived by the non-player character.

[0414] For example, the game application can sequentially traverse the at least one search range described above, or the game application can sequentially traverse the at least one search range determined according to the natural language command. For example, the natural language command is "Can you see the truck in front of you? Move to the truck", the game application can traverse the visual range of the master virtual role first, and then traverse the visual range of the non-player character, and match the entity with the highest similarity and higher than the threshold.

[0415] For example, when there are at least two search ranges, the traversal order of the at least two search ranges can be preset, or the traversal order of the at least two search ranges can be determined according to the intention indicated in the natural language command. For example, when the intention is to search for an item, the visual range is traversed first. When the intention is to search for a combat scene, the auditory range is traversed first.

[0416] If there are multiple levels of nested instructions in the natural language command, the first target entity is searched first, and then the next target entity is searched according to the position of the previous target entity. In the case where the natural language command includes a second target entity determined based on the position of the first target entity, the first target entity is queried from the entity set of the virtual environment according to the first target entity information of the first target entity indicated by the natural language command and the environmental perception information; and then the second target entity is queried from the entity set of the virtual environment according to the second target entity information of the second target entity, the first target entity and the environmental perception information.

[0417] In response to the natural language command, the NPC is controlled to perform a virtual activity associated with the target entity.

[0418] For example, the behavior intention of the natural language command is recognized by calling a large language model, a control instruction or a control instruction sequence is generated according to the behavior intention, and the non-player character is controlled to complete an activity based on the target entity according to the control instruction or the control instruction sequence.

[0419] The NPC feedback information broadcast of step 208 can be implemented as sub-step 6:

[0420] Sub-step 6: Broadcast the feedback information of the non-player character, and the feedback text corresponding to the feedback information is non-fixed text generated based on the environmental perception information of the non-player character;

[0421] Exemplarily, the game application program can generate corresponding feedback information in real time according to the environmental perception information of the non-player character, and perform broadcasting. For example, the game application program can generate corresponding feedback information in real time according to the environmental perception information triggering the broadcasting condition when the environmental perception information of the non-player character triggers the broadcasting condition; or the game application program can generate corresponding feedback information in real time according to the environmental perception information of the non-player character when the non-player character or the master virtual character triggers the broadcasting condition.

[0422] Exemplarily, the feedback text is obtained by calling a large language model to perform reasoning based on static entity data of a three-dimensional virtual environment and dynamic environmental perception information of a non-player character. The large language model generates different feedback texts for different perception situations of the non-player character.

[0423] Exemplarily, the feedback text can also be obtained by calling a large language model to perform reasoning based on static entity data of a three-dimensional virtual environment and dynamic environmental perception information of a non-player character, and the feedback text conforms to the personality characteristics of the non-player character. For example, when the non-player character is a robust uncle image, the feedback text can adopt a bold tone; when the non-player character is a reporter, the feedback text can adopt a news report tone.

[0424] It should be noted that due to the randomness of the generation result of the large language model, in the same scene, when the environmental perception information of the non-player character is the same, the generated feedback text can also be different.

[0425] The environmental perception information can include information perceived by the non-player character in real time, or information perceived by the non-player character historically.

[0426] Exemplarily, the feedback information is obtained by the game application program according to the environmental perception information of the non-player character and the static entity data of the three-dimensional virtual environment. The feedback information can be a voice form broadcast or a text form broadcast. The static entity data includes data of relatively unchanged entities in the three-dimensional virtual environment, for example, includes model information of various buildings, terrains, and vehicles in the three-dimensional virtual environment.

[0427] Exemplarily, the feedback information can include at least one of the following: immediate feedback, execution feedback, and dynamic feedback. The immediate feedback and the execution feedback are feedback information generated for natural language commands, and the dynamic feedback is feedback information generated by the non-player character spontaneously.

[0428] 1. The immediate feedback is feedback information generated immediately when a natural language command is received. The immediate feedback can include a reply broadcast and an immediate broadcast. The reply broadcast is used to recover the query in the natural language command; and the immediate broadcast is used to feed back the receiving situation of the natural language command.

[0429] For example, the client can instruct the non-player character to act in the three-dimensional virtual environment according to the intention of the natural language command. Then, the client can generate instant feedback corresponding to the natural language command when receiving the natural language command, and the instant feedback is used to indicate that the natural language command has been received, or the instant feedback is used to reply to the natural language command immediately.

[0430] In an optional embodiment, the feedback information includes a reply broadcast. The reply broadcast includes a reply to the inquiry natural language command generated for the inquiry natural language command proposed by the player. For example, the inquiry natural language command includes an inquiry about the location of the target entity, and the corresponding reply broadcast should include the query result of the location of the target entity.

[0431] The terminal broadcasts the reply broadcast of the non-player character in the case that the behavior intention of the received natural language command is inquiry; the reply broadcast includes the reply content to the natural language command generated according to the environmental perception information of the non-player character to the three-dimensional virtual environment.

[0432] The natural language command is used to inquire information from the non-player character. This natural language command can be referred to as an inquiry natural language command. The inquiry natural language command is a natural language command with an inquiry intention. The inquiry natural language command contains at least one question.

[0433] The natural language command is a command indication conveyed using human natural language. The terminal receives the natural language command issued by the player, and instructs the non-player character according to the intention corresponding to the natural language command. For example, the terminal receives the voice audio of the player, converts the voice audio into text, and obtains the natural language command. Or, the terminal receives the text of the natural language command input by the player.

[0434] The player can use oral or written language to issue natural language commands. The game application program calls a large language model to perform understanding on the natural language command, extracts the intention expressed in the natural language command, generates control instructions according to the intention, and instructs the activity of the non-player character.

[0435] For example, the natural language command can be “come to me”, and the game application program can parse the intention of the natural language command as “control the non-player character to move to the location of the master virtual character”. Then, the game application program obtains the location of the master virtual character, generates a navigation route for the non-player character to move to the location of the master virtual character, and controls the non-player character to move to the location of the master virtual character according to the navigation route.

[0436] For example, the natural language command can be "go pick up the treasure chest", and the game application can parse the intention of the natural language command as "control the non-player character to move to the location of the treasure chest and pick up the treasure chest". The game application can obtain the location of the treasure chest, control the non-player character to move to the location of the treasure chest, and perform the operation of picking up the treasure chest after reaching the location of the treasure chest.

[0437] For example, in the case that the inquiry natural language command includes an intelligence inquiry to the target entity, based on the perception of the target entity by the non-player character, a first reply broadcast is broadcasted; the first reply broadcast includes the intelligence perception result of the target entity by the non-player character.

[0438] Alternatively, in the case that the inquiry natural language command includes a location inquiry to the target entity, based on the perception of the target entity by the non-player character, a second reply broadcast is broadcasted, and the second reply broadcast includes the location information of the target entity.

[0439] For example, the large language model is called to perform parsing on the received natural language command, and the intention of the natural language command is obtained. When the intention is an inquiry type, it can be determined that the natural language command is an inquiry natural language command. When the natural language command is an inquiry natural language command, the large language model can perform reasoning according to the static entity data and the environmental perception information according to the intention of the natural language command, and obtain the reply text. That is, the large language model is called to perform parsing on the received natural language command, and the intention of the natural language command is output as an inquiry, and the reply text of the natural language command. The game application calls the text-to-speech service according to the processing logic of the inquiry intention, converts the reply text into a reply broadcast, and the reply broadcast is an audio form of broadcast. The reply audio is sent to the client to perform broadcast.

[0440] For example, when the large language model identifies the intention of the natural language command as an inquiry, since there can be thousands of questions from the user, it is impossible to store the reply to each question locally on the client side. Therefore, the game application calls the text-to-speech service to generate a reply broadcast in real time according to the reply text returned by the large language model, so that the game application can generate a corresponding reply broadcast to reply to the player's question in real time.

[0441] For example, the game application (client or server) calls the large language model to parse the inquiry natural language command, and based on the static environment data and the environmental perception information of the three-dimensional virtual environment, the reply text is obtained by reasoning. Based on the reply text, a human voice audio is generated to obtain a reply broadcast.

[0442] For example, the client reports the received inquiry natural language command to the server, the server invokes the large language model to analyze the inquiry natural language command, infers the reply text based on the static environment data and the environment perception information of the three-dimensional virtual environment, generates the human voice audio based on the reply text to obtain the reply broadcast, and sends the reply broadcast to the client.

[0443] 2. The execution feedback is generated when the natural language command is executed according to the intention indicated by the natural language command. The execution feedback can also be referred to as execution broadcast.

[0444] For example, when the natural language command is used to instruct the non-player character to execute the target task, the client can generate the execution feedback during the execution of the target task by the non-player character, and the execution feedback is used to indicate the execution state of the target task; or the client can generate the execution feedback after the non-player character completes the target task, and the execution feedback is used to indicate the execution result of the target task.

[0445] In an optional embodiment, the feedback information includes feedback broadcast (which can also be referred to as execution feedback). The feedback broadcast includes the broadcast of the task execution state and the task execution result generated for the task natural language command proposed by the player. For example, the task natural language command includes moving to the target entity, and the corresponding feedback broadcast can include moving to the target entity, or having arrived at the target entity.

[0446] The terminal broadcasts the feedback broadcast in the case that the behavior intention of the received natural language command is to execute the task; the feedback broadcast includes the execution state after the non-player character executes the natural language command according to the environment perception information of the three-dimensional virtual environment.

[0447] The natural language command is used to instruct the non-player character to execute the task. This natural language command can be referred to as a task natural language command. The task natural language command includes the description text of the task, for example, the description text of the task can include at least one of the following: task name, action required to be executed by the non-player character, method for executing the task, task target, and location of the task target.

[0448] For example, in the case that the intention of the task natural language command includes searching for the target entity, the first feedback broadcast is broadcasted, and the first feedback broadcast includes the search result of the target entity by the non-player character.

[0449] Or, in the case that the intention of the task natural language command includes using the virtual prop, the second feedback broadcast is broadcasted, and the second feedback broadcast is used to indicate that the virtual prop is ready.

[0450] Or, in the case where the intention of the task natural language command includes using a virtual prop, a third feedback broadcast is broadcasted, the third feedback broadcast being used to indicate the use result of the virtual prop.

[0451] Or, in the case where the intention of the task natural language command includes controlling the movement of the non-player character, a fourth feedback broadcast is broadcasted, the fourth feedback broadcast being used to indicate the movement result of the movement.

[0452] Or, in the case where the intention of the task natural language command includes performing an interaction with a target entity, a fifth feedback broadcast is broadcasted, the fifth feedback broadcast being used to indicate at least one of the search result of the non-player character on the target entity, the interaction result of the non-player character and the target entity.

[0453] Or, in the case where the intention of the task natural language command includes an item interaction, a sixth feedback broadcast is broadcasted, the sixth feedback broadcast being used to indicate the item interaction result of the non-player character performing the item interaction.

[0454] For example, the large language model is called to perform parsing on the received natural language command to obtain the intention of the natural language command. When the intention is a task, it can be determined that the natural language command is a task natural language command. When the natural language command is a task natural language command, the large language model can perform reasoning according to the static entity data and the environmental perception information according to the intention of the natural language command to obtain a task instruction sequence, the task instruction sequence being used to control the non-player character to perform the task. The game application program receives the task instruction sequence, and controls the non-player character to perform the task according to the task instructions in the task instruction sequence. During the control of the non-player character according to the task instructions, the corresponding execution condition execution feedback broadcast can be obtained according to the feedback broadcast logic corresponding to different task instructions.

[0455] For example, since the instructions that can be performed by the non-player character in the three-dimensional virtual environment are limited and traversable, the client can locally store the feedback broadcast corresponding to each character instruction. When the non-player character performs the corresponding instruction, the client can read the local feedback broadcast to perform voice broadcast.

[0456] Or, when there is variable content in the feedback broadcast corresponding to a certain task instruction, for example, the target entity in the feedback broadcast is variable, or the position information in the feedback broadcast is variable; the client can also generate a broadcast voice of the variable content according to the target entity or the position information returned by the large language model, and splice the broadcast voice of the variable content with the broadcast voice of the non-variable content stored locally to obtain the final feedback broadcast.

[0457] For example, when the task natural language command is "find the nearby treasure chest", calling the large language model to perform parsing on the task natural language command can obtain that its intention is a task, the task target is to find a target entity, and the target entity is a treasure chest; based on the static entity data and the environment perception information, inference can be performed to obtain that there is a treasure chest in the left front of the host virtual role. In the case where the target entity can be found, the feedback report corresponding to the task target is "find aaa at xxx", where xxx is the location of the target entity, and aaa is the name of the target entity. The game application program can call the online text-to-speech technology to generate the location voice according to the location information "left front" in the inference result; call the online text-to-speech technology to generate the name voice according to the name "treasure chest" of the target entity, and then splice the location voice and the name voice into the feedback report to obtain the final feedback report.

[0458] For example, the game application program (client or server) calls the large language model to parse the task natural language command, infers the task execution instruction based on the static environment data and the environment perception information of the three-dimensional virtual environment, controls the non-player character to execute the task according to the task execution instruction, generates the feedback text according to the execution state of the non-player character executing the task, and generates the human voice audio based on the feedback text to obtain the feedback report.

[0459] For example, the client reports the task natural language command to the server; the server parses the task natural language command, infers the task execution instruction based on the static environment data and the environment perception information of the three-dimensional virtual environment, controls the non-player character to execute the task according to the task execution instruction, generates the feedback text according to the execution state of the non-player character executing the task, generates the human voice audio based on the feedback text to obtain the feedback report, and sends the feedback report to the client; the client receives and reports the feedback report.

[0460] 3. The dynamic feedback is the feedback information generated spontaneously according to the environment perception information of the non-player character. The dynamic feedback can also be referred to as dynamic report.

[0461] For example, the client can spontaneously generate the dynamic feedback based on the perception of the three-dimensional virtual environment without receiving the natural language command, and the dynamic feedback is used to indicate the abnormal situation found by the non-player character in the three-dimensional virtual environment. For example, the abnormal situation can include finding an enemy virtual character, finding that the state of a friendly virtual character has changed, finding a dangerous situation, finding a battle trace or a looting trace, etc.

[0462] In an alternative embodiment, the feedback information comprises dynamic broadcasting (may also be referred to as "dynamic feedback"). The dynamic broadcasting comprises broadcasting of a perceived abnormal situation of the non-player character. For example, when the non-player character perceives that it is under attack, the broadcasting is that it is under attack; when the non-player character perceives that there is a dangerous situation ahead, the broadcasting is that there is danger ahead.

[0463] The terminal broadcasts the dynamic broadcasting when the environmental perception information of the non-player character in the three-dimensional virtual environment satisfies a dynamic broadcasting condition.

[0464] The dynamic broadcasting is a broadcasting generated according to an abnormal situation when the game application determines that there is an abnormal situation that needs to be broadcasted based on the environmental perception information of the non-player character. The dynamic broadcasting comprises reminder information of the abnormal situation. For example, the abnormal situation can comprise at least one of the following: discovery of an enemy virtual character, discovery of a change in the state of a friendly virtual character, discovery of a new trace.

[0465] For example, in a case where the non-player character perceives an enemy virtual character, a first dynamic broadcasting is broadcasted; the first dynamic broadcasting comprises a position of the enemy virtual character perceived by the non-player character.

[0466] Alternatively, in a case where the non-player character perceives a dangerous situation, a second dynamic broadcasting is broadcasted; the second dynamic broadcasting is used to prompt the dangerous situation.

[0467] Alternatively, in a case where the non-player character perceives a change in the state of a friendly virtual character, a third dynamic broadcasting is broadcasted; the third dynamic broadcasting is used to prompt the change in the state of the friendly virtual character.

[0468] Alternatively, in a case where the non-player character discovers a new trace in the three-dimensional virtual environment, a fourth dynamic broadcasting is broadcasted; the fourth dynamic broadcasting is used to prompt the new trace. The new trace can be a battle trace and / or a looting trace.

[0469] For example, since the number of dynamic broadcastings that can be triggered by the non-player character in the three-dimensional virtual environment is limited and traversable, the client can locally store dynamic broadcasting speech, and the client can read the dynamic broadcasting speech from the local to execute the broadcasting when the corresponding dynamic broadcasting speech is triggered. Similarly, some dynamic broadcasting speech can have variable content, and the client can generate the variable content in real time through a text-to-speech technology and splice the dynamic broadcasting template to obtain the final dynamic broadcasting.

[0470] In an alternative embodiment, when generating the feedback information, it can be necessary to query the position of a target entity in the three-dimensional virtual environment, or to determine the target entity indicated in the natural language command from the three-dimensional virtual environment. At this time, an entity query method needs to be used to perform the query of the target entity.

[0471] Before performing the target entity query, static entity data of the three-dimensional virtual environment needs to be constructed in advance. The construction of static entity data can serve both the inference of the large language model and the generation of real-time feedback text on the game application side after the inference result is returned.

[0472] In an alternative embodiment, entities in the three-dimensional virtual environment are traversed, and entity information of each entity is exported from the game engine to obtain the static entity data.

[0473] For example, for in-game scenes, an editor tool is developed to perform full-scene StaticMesh traversal, special Actor traversal (such as interactive doors, which do not belong to the StaticMesh category), vegetation traversal, and export the positions, orientations, and bounding box sizes of entities as static entity data for subsequent tagging by the large language model and inference by the large language model.

[0474] Among them, StaticMesh is a type of static geometry resource in Unreal Engine 4 (UE4) game engine, used to represent immutable three-dimensional models such as buildings, props, etc. Actor is a basic object class in Unreal Engine 4 (UE4) game engine, representing an entity in the game world, such as characters, objects, light sources, etc. An Actor can contain multiple components to achieve different functions.

[0475] For example, spatial data can also be included in the static entity data, which is obtained by manual labeling. Spatial data is more complex than static entities, as there are multiple floors and multiple layers of nesting within a floor in the space (such as a motel area containing a second floor, and the second floor area containing guest rooms). For such data, manual calibration is used to construct. Under the editor, add Volume to the scene, divide the required area of the whole graph and label it accordingly. Volume is a special Actor class in Unreal Engine 4 (UE4) game engine, representing a three-dimensional area with specific functions, such as trigger areas, audio areas, etc.

[0476] For example, as shown in FIG. 19, if there is a second floor area of the recycling tower in the virtual environment, a space entity 801 is created in the second floor area, the space entity 801 is used to surround the second floor area, the space position of the second floor area is determined, and the description text of the space relationship "Farm Map-Recycling Tower-Second Floor Area" is set for the space entity 801.

[0477] For example, as shown in FIG. 19, if there is a second floor area of the recycling tower in the virtual environment, a space entity 801 is created in the second floor area, the space entity 801 is used to surround the second floor area, the space position of the second floor area is determined, and the description text of the space relationship "Farm Map-Recycling Tower-Second Floor Area" is set for the space entity 801.

[0478] After obtaining the static entity data, the entity or space query can be performed based on the static entity data during the game running.

[0479] For example, when the player issues a natural language command, the natural language command will be accompanied by the constructed static entity data and the real-time captured runtime data (including the environmental perception information of the non-player character and / or the host virtual character) to perform an inference request to the large language model. In addition, when the NPC needs to perform dynamic feedback based on the runtime data, an entity query is also performed to generate feedback text based on the current state.

[0480] When the target entity is included in the natural language command, the game application program performs a query of the target entity according to the current position and orientation of the host virtual character, the position and orientation of the teammates or enemies, and according to the target entity information of the target entity described in the natural language command and the environmental perception information of the host virtual character and / or the non-player character, the target entity that matches the target entity information is filtered from at least one candidate entity within the perception range of the host virtual character and / or the non-player character.

[0481] Optionally, the natural language command is used to instruct the non-player character to perform an activity related to the target entity. For example, the target entity can be the movement destination of the non-player character, the target entity can be the object to be observed by the non-player character, the target entity can be the target to be attacked by the non-player character, or the target entity can be the object to be interacted with by the non-player character.

[0482] The target entity is a virtual object existing in the three-dimensional virtual environment. The target entity is an entity described in the natural language command. The entity can refer to a virtual object fixed in the three-dimensional virtual environment, for example, the target entity can be a virtual building, a virtual terrain, a virtual vehicle, a virtual plant, a virtual prop, a virtual item, etc. The entity can also refer to a virtual object that can perform interaction with the virtual character in the three-dimensional virtual environment, for example, the target entity can be an interaction point (for example, a door, a window, a cabinet, a cellar, etc.), a virtual prop, a virtual character, a virtual light source, etc. in the three-dimensional virtual environment.

[0483] For example, the natural language command can be "please help me open the door", and the target entity can be "the door", and the non-player character needs to perform the activity "opening the door" related to the door. Or, the natural language command can be "is the kitchen safe", and the target entity can be "the kitchen", and the non-player character needs to perform the activity "checking whether the kitchen has an enemy virtual character or other dangerous situation" related to the kitchen.

[0484] Optionally, the natural language command includes target entity information of the target entity. The target entity information can be a description text of the target entity. For example, the target entity information can include at least one of the following contents: the name, the type, the location, the feature of the target entity. The game application program can find the target entity from the entities in the three-dimensional virtual environment based on the target entity information in the natural language.

[0485] The game application program performs parsing on the natural language command, extracts the target entity information therein, determines the target entity from the entity set of the three-dimensional virtual environment based on the target entity information, and then controls the non-player character to perform the activity related to the target entity according to the intention of the natural language command.

[0486] The target entity is determined from the three-dimensional virtual environment in combination with the environment perception information, and the environment perception information includes information perceived by at least one of the master virtual character and the non-player character from the three-dimensional virtual environment.

[0487] Since the target entity information in the natural language command is usually ambiguous, for example, the natural language command can be "move to the red truck", and the target entity information therein is "the red truck", and there can be many red trucks in the three-dimensional virtual environment, and the target entity cannot be accurately determined from the three-dimensional virtual environment according to the target entity information in the natural language command.

[0488] Therefore, the embodiment of the present application provides a method for determining a target entity by combining target entity information and environment perception information. In the above example, the "red truck" expressed by the player in the natural language command should be a red truck that the player can see. Therefore, by combining the visual field range of the virtual character controlled by the player, the red truck located in the visual field range of the virtual character controlled by the player can be screened from the plurality of red trucks in the three-dimensional virtual environment, and the red truck is the target entity indicated by the player in the natural language command.

[0489] Since the natural language command is issued by the player based on the perception of the three-dimensional virtual environment, in order to accurately identify the target entity indicated in the natural language command, the game application program identifies the target entity in the natural language command by combining the environment perception information of the player when issuing the natural language command. Based on the perception of the three-dimensional virtual environment by the player, the entity closest to the target entity information that the player can perceive in the three-dimensional virtual environment is inferred as the target entity.

[0490] For example, the three-dimensional virtual environment includes a plurality of candidate entities matching the natural language command, and the target entity is an entity screened from the plurality of candidate entities based on the environment perception information of the virtual character controlled by the player or the non-player character. For example, the game application program first selects a plurality of candidate entities matching the target entity information in the natural language command from the entity set, and then selects an entity that can be perceived by the virtual character controlled by the player or the non-player character from the plurality of candidate entities as the target entity according to the environment perception information.

[0491] For example, the perception range of the virtual character controlled by the player is determined according to the position and orientation of the virtual character controlled by the player, and the perception range is determined as the search range of the target entity. For example, the perception range is a fan-shaped region of a certain visual field size in front of the virtual character controlled by the player; entities not in the fan-shaped range are filtered out. Then, sorting is performed according to the similarity of the target entity information and the entity information of the candidate entity, and entities with a similarity higher than a threshold value are screened as the target entity. When there are two entities with the highest similarity scores in the perception range, for example, the similarity scores of entity 2 and entity 3 are both 2 points, and then the entity 2 closer to the virtual character controlled by the player is selected as the target entity.

[0492] When it is necessary to generate feedback information according to the position information of the target entity, the game application program can perform spatial query to generate position description according to the positional relationship between the virtual character controlled by the player and the target entity.

[0493] The main application scenario of spatial query is to obtain the positions of the virtual character controlled by the player, the non-player character, the enemy virtual character, and the gun sound. For example, a report containing spatial information such as "enemy found on the second floor of the motel" can be made.

[0494] The spatial query logic defines the concept of parent area and child area. The parent area refers to a large area in the three-dimensional virtual environment that covers multiple buildings or multiple isolated spaces, such as a motel with a front yard and a back yard, and containing multiple rooms and a basement. The child area refers to an independent space within the parent area, which can be a closed space or a special functional space, such as a motel room, a kitchen, etc.

[0495] In order to make the location description returned by the spatial query closer to human expression, when the host virtual character is in a different parent area from the queried location, the query result returns complete text information (for example, the host virtual character is outside the motel, and the queried location is on the second floor of the motel, then the return is "the enemy is on the second floor of the motel"). When the host virtual character is in the same parent area as the queried location, but in different child areas, the query result only returns the child area information (for example, the host virtual character is on the first floor of the motel, and the queried location is on the second floor, then the return is "the enemy is on the second floor"). When the host virtual character is in the same parent area as the queried location, and in the same child area, the query result only returns the relative position of the player from the host virtual character's perspective (for example, the host virtual character and the queried location are both on the second floor of the motel, then the return is "the enemy is in the right front position").

[0496] For example, in the case where the host virtual character and the target entity are located in different parent spaces, the location description includes the parent space and the child space where the target entity is located. In the case where the host virtual character and the target entity are located in the same parent space but different child spaces, the location description includes the child space where the target entity is located. In the case where the host virtual character and the target entity are located in the same parent space and the same child space, the location description includes the relative position of the target entity and the host virtual character; wherein one parent space includes at least one child space.

[0497] For the broadcast of the NPC feedback information in step 208, the feedback information is implemented as spatial human voice audio after spatial sound effect processing, and the process of spatial sound effect processing can be implemented as sub-step 7 to sub-step 9:

[0498] Sub-step 7: Obtain human voice audio corresponding to the feedback text;

[0499] In the embodiments of the present application, the generation process of the feedback text is as shown in the foregoing sub-step 6, which will not be repeated here.

[0500] In some embodiments, the process of obtaining corresponding human voice audio based on the feedback text can be implemented as: obtaining the feedback text and the character feature information corresponding to the non-player character; generating human voice audio matching the character feature information according to the feedback text.

[0501] Illustratively, the role characteristic information is used to describe the attribute of the non-player character, that is, the role characteristic information is information used to describe the current state and characteristics of the non-player character. In some embodiments, the attribute of the non-player character includes at least one of a basic attribute and a situational performance attribute of the non-player character. The basic attribute is an attribute pre-configured for the non-player character, that is, an attribute that does not change for the non-player character due to the three-dimensional virtual environment or the current situation, such as age information of the non-player character, male / female information of the non-player character, faction information of the non-player character, and the like. The situational performance attribute is an attribute of the non-player character determined in real time in the situation of the three-dimensional virtual environment, that is, an attribute associated with the current situation, such as emotion information of the non-player character in the current situation, behavior information, and the like.

[0502] Optionally, the role characteristic information includes at least one of role basic information, role emotion information, and role behavior information. The role basic information is used to indicate the basic situation of the non-player character, and the role basic information is information of the basic attribute of the non-player character, such as age information of the non-player character, male / female information of the non-player character, faction information of the non-player character, and the like. The role emotion information is used to indicate the emotional state of the non-player character in the dialogue situation corresponding to the feedback text, and the role emotion information is the situational performance attribute of the non-player character, such as indicating that the non-player character is in an excited state, an angry state, a sad state, and the like. The role behavior information is used to indicate the action performed by the non-player character in the dialogue situation, and the role behavior information is the situational performance attribute of the non-player character, such as being in a running state, being in an attack state, being in a wounded treatment state, and the like.

[0503] In some embodiments, the generation of the human voice audio is implemented by a pre-trained voice generation model, and the voice generation model is used to generate a voice matching the role characteristic information of the non-player character. Illustratively, the feedback text and the role characteristic information are input into the pre-trained voice generation model to obtain the human voice audio as the human voice audio.

[0504] Optionally, the voice generation model can be implemented by a neural network model such as a convolutional neural network, a feedforward neural network, a residual network, a Transformer, a Multimodal Large Language Model (MLLM), and the like, which is not specifically limited here.

[0505] In some embodiments, the voice generation model includes a text encoder, a role encoder, and a decoder. The text encoder is used to perform text encoding on the input feedback text, that is, the feedback text is input into the text encoder to obtain a text encoding representation. The role encoder is used to perform feature encoding on the input role characteristic information, that is, the role characteristic information is input into the role encoder to obtain a role encoding representation.

[0506] Illustratively, after obtaining the text encoded representation and the role encoded representation, fusion is performed on the text encoded representation and the role encoded representation to obtain the encoded representation input to the decoder, i.e., fusing the text encoded representation and the role encoded representation to obtain a joint encoded representation, inputting the joint encoded representation to the decoder to generate the human voice audio.

[0507] In some embodiments, the role encoder comprises at least one of a first sub-encoder, a second sub-encoder, and a third sub-encoder. The first sub-encoder is configured to perform feature encoding based on the base attribute of the non-player character, i.e., in the case that the role feature information comprises role base information, inputting the role base information to the first sub-encoder to obtain a first role encoded representation; the second sub-encoder is configured to perform feature encoding based on the emotional state of the non-player character, i.e., in the case that the role feature information comprises role emotional information, inputting the role emotional information to the second sub-encoder to obtain a second role encoded representation; and the third sub-encoder is configured to perform feature encoding based on the behavior state of the non-player character, i.e., in the case that the role feature information comprises role behavior information, inputting the role behavior information to the third sub-encoder to obtain a third role encoded representation.

[0508] Sub-step 8: based on the relative positional relationship between the non-player character and the master virtual character in the three-dimensional virtual environment, obtaining the sound effect parameter corresponding to the non-player character;

[0509] Illustratively, the relative positional relationship between the non-player character and the master virtual character in the three-dimensional virtual environment is used to indicate the relative relationship between a first position corresponding to the non-player character and a second position corresponding to the master virtual character in the three-dimensional virtual environment.

[0510] Optionally, the relative positional relationship comprises at least one of a positional distance, a positional direction, and a space surrounding condition between the non-player character and the master virtual character, wherein the positional distance is used to indicate the proximity relationship between the non-player character and the master virtual character, the positional direction is used to indicate the angle size relationship between the first position where the non-player character is located and the orientation direction of the master virtual character, and the space surrounding condition is used to indicate the surrounding condition of the space element forming a virtual space in the three-dimensional virtual environment to the master virtual character, and the propagation influence of the space on the audio when the non-player character emits the audio in the virtual space.

[0511] In some embodiments, the determination of the relative positional relationship between the non-player character and the master virtual character can be implemented as follows: obtaining first position information of the non-player character in the three-dimensional virtual environment, obtaining second position information of the master virtual character in the three-dimensional virtual environment, and determining the relative positional relationship between the non-player character and the master virtual character based on the first position information and the second position information.

[0512] In some embodiments, the first position information of the non-player character and the second position information of the host virtual character are position information determined based on a same preset coordinate system. Optionally, the preset coordinate system can be implemented as a world coordinate system corresponding to the three-dimensional virtual environment, or the preset coordinate system can be implemented as a coordinate system established with the host virtual character as the origin.

[0513] In some embodiments, when the sound effect parameter corresponding to the non-player character is pre-generated and stored in a database, the sound effect parameter corresponding to the non-player character is obtained from the database. In other embodiments, the sound effect parameter corresponding to the non-player character can also be generated in real time, i.e., based on the relative orientation relationship, the sound effect parameter corresponding to the non-player character is generated.

[0514] Optionally, the sound effect parameter corresponding to the non-player character can include at least one of the following types:

[0515] The first type is the distance type

[0516] Illustratively, the sound effect parameter of the distance type is used to indicate that the sound effect of the human voice audio is adjusted based on the distance between the non-player character and the host virtual character. In some embodiments, the sound effect parameter of the distance type can adjust the audio volume of the human voice audio to simulate the effect of distance on the size of the sound; and / or, the sound effect parameter of the distance type can adjust the audio time delay of the human voice audio to simulate the effect of distance on the sound propagation time.

[0517] The second type is the direction type

[0518] Illustratively, the sound effect parameter of the direction type is used to indicate that the sound effect of the human voice audio is adjusted based on the direction of the non-player character relative to the host virtual character. In some embodiments, the sound effect parameter of the direction type can adjust the audio volume of the human voice audio to simulate whether the host virtual character is facing the scene audio element; and / or, the sound effect parameter of the direction type can adjust the volume parameters of the human voice audio in different channels to simulate the direction of the scene sound source element emitting audio.

[0519] The third type is the spatial effect type

[0520] Illustratively, the sound effect parameter of the spatial effect type is used to indicate that the sound effect in the virtual space formed by the spatial element is adjusted, i.e., the sound effect performance of the virtual space on the audio emitted by the non-player character is indicated by the sound effect parameter of the spatial effect type. In some embodiments, the sound effect parameter of the spatial effect type can increase the echo effect and the reverberation effect to simulate the effect of sound in space.

[0521] Optionally, the sound effect parameter of the space effect type includes at least one of a reverberation time parameter, a pre-delay parameter, a wet / dry mix parameter, a room width parameter, and a distance effect parameter. The reverberation time parameter is used to indicate the speed of audio decay in the virtual space; the pre-delay parameter is used to indicate the time difference between the direct arrival of audio to the main virtual character and the first reflection arrival to the main virtual character; the wet / dry mix parameter is used to indicate the proportion between directly propagated audio and reflected propagated audio; the room width parameter is used to indicate the degree of diffusion of audio on the horizontal plane in the virtual space; and the distance effect parameter is used to indicate the attenuation of audio with the propagation distance in the propagation process.

[0522] Fourth, a custom type

[0523] Illustratively, the sound effect parameter of the custom type is a parameter for sound effect adjustment customized by the user. Optionally, the user can customize the overall volume of the scene audio, whether to add background music, and the volume of different types of audio.

[0524] In some embodiments, the client provides a custom interface for the user to configure the custom sound effect parameter.

[0525] In some embodiments, the server is configured with a parameter prediction model for performing personalized learning on the custom sound effect parameter of the user. Illustratively, the custom sound effect parameters of a plurality of candidate accounts are obtained, the plurality of custom sound effect parameters are input into the parameter prediction model to be trained, the parameter prediction model to be trained is iteratively trained to obtain the parameter prediction model, and the sound effect parameters generated by the system (e.g., the sound effect parameters of the distance type, the direction type, and the space effect type) are optimized through the parameter prediction model.

[0526] Sub-step 9: adjusting the human voice audio based on the sound effect parameter to generate and play the spatial human voice audio corresponding to the non-player character;

[0527] The spatial human voice audio is used to represent the perception effect of the audio generated by the main virtual character on the non-player character in the relative positional relationship.

[0528] In the embodiments of the present application, the spatial human voice audio corresponding to the non-player character is generated by adjusting the human voice audio according to the sound effect parameter of the non-player character. That is, the spatial human voice audio is the audio data corresponding to the non-player character obtained by adjusting the human voice audio through the sound effect parameter. In some embodiments, the spatial human voice audio includes at least one of the audio adjusted by the sound effect parameter, the audio duration information, the audio playback start timestamp, the audio playback end timestamp, the audio playback condition information, and the element identification information corresponding to the audio.

[0529] Optionally, when the sound effect parameter indicates to perform adjustment on the volume size of the human voice audio of the non-player character, the volume size of the human voice audio of the non-player character on at least one sound channel is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates to perform adjustment on the playing time delay of the human voice audio of the non-player character, the audio playing start time corresponding to the human voice audio of the non-player character is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates to perform adjustment on the sound speed size of the human voice audio of the non-player character, the playing speed of the human voice audio of the non-player character is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates the pitch of the human voice audio of the non-player character, the frequency in the frequency spectrum of the human voice audio of the non-player character is adjusted based on the sound effect parameter. When the sound effect parameter indicates the timbre of the human voice audio of the non-player character, a filter corresponding to the non-player character is determined based on the sound effect parameter, and the timbre corresponding to the human voice audio is adjusted through the filter.

[0530] Optionally, the client performs playing on the spatial human voice audio to realize the broadcast of the feedback information of the non-player character.

[0531] The speech-to-text stage for step 204 can be implemented as sub-step 10 and sub-step 11:

[0532] Sub-step 10: Obtain a plurality of first scene hotwords corresponding to a first virtual scene in which a target virtual character is located.

[0533] The target virtual character includes at least one of an NPC and a master virtual character.

[0534] Optionally, the master virtual character is a virtual character commanded by a player, and the NPC is a virtual character assisting the master virtual character in a virtual game. In general, the virtual game includes one master virtual character and at least one NPC. There is a certain difference between the master virtual character and the NPC.

[0535] Optionally, the master virtual character is a virtual character mainly controlled by a player, and the NPC is a virtual character commanded by a player in the form of voice. For example, the master virtual character is a virtual character manually controlled by a player during a game process, and the manual operation of the terminal interface by the player is used to control the master virtual character. The NPC is a virtual character commanded by a player in the form of voice, and the player occasionally issues a natural language command in the form of voice to command the NPC.

[0536] In some embodiments, the first virtual scene is determined based on the target virtual character, and the target virtual character is at least one of an NPC and a master virtual character, i.e., the process of determining the first virtual scene is implemented as at least one of the following.

[0537] (1) If the master virtual character is selected as the target virtual character, a virtual scene in which the master virtual character is located is taken as the first virtual scene;

[0538] (2) If the NPC is selected as the target virtual character, the virtual scene in which the NPC is located is taken as the first virtual scene; if there are multiple NPCs, the virtual scene including the most NPCs can be taken as the first virtual scene, the NPC closest to the host virtual character can also be taken as the target virtual character, and the virtual scene in which the NPC is located is taken as the first virtual scene, or the NPC with the largest character attribute value (such as at least one of the virtual blood volume, the virtual magic power value, the virtual defense value, etc.) can also be taken as the target virtual character, and the virtual scene in which the NPC is located is taken as the first virtual scene.

[0539] (3) If the host virtual character and the NPC are selected as the target virtual character, the virtual scene in which the host virtual character and the NPC are located can be taken as the first virtual scene, etc.

[0540] In some embodiments, the target virtual character is determined based on the selection of the player; or the target virtual character is set by default by the system.

[0541] In some embodiments, a plurality of first scene hot words are determined based on the first virtual scene.

[0542] The plurality of first scene hot words are scene-related words of the first virtual scene. That is, there is an association between the first scene hot word and the first virtual scene, and the first scene hot word is a word describing the first virtual scene.

[0543] Optionally, the first scene hot word includes a scene state of the first virtual scene, such as the first virtual scene being a virtual restaurant, the first scene hot word including an operating state, a resting state, a closed state, etc.

[0544] Optionally, the first scene hot word includes an element name of an entity in the first virtual scene, such as the first virtual scene being a virtual restaurant, the first scene hot word including a virtual table, a virtual chair, a virtual kitchen utensil, a virtual dish, etc.

[0545] Optionally, the first scene hot word includes an interactive word for interacting with an entity in the first virtual scene, such as the first virtual scene being a virtual battlefield, the first scene hot word including attack, attack virtual character, defense, build barrier, etc.

[0546] In some embodiments, after receiving the natural language command, a plurality of first scene hot words are collected based on the first virtual scene; or after receiving the natural language command, a plurality of first scene hot words corresponding to the first virtual scene are filtered from a plurality of scene hot words obtained in advance based on the first virtual scene, etc.

[0547] In an optional embodiment, a perception range of the target virtual character in the first virtual scene is obtained.

[0548] wherein the perception range is a three-dimensional spatial range in which the target virtual character has a perception of other scene elements.

[0549] Optionally, when the target virtual character is a master virtual character, the perception range of the master virtual character in the first virtual scene is obtained; when the target virtual character is an NPC, the perception range of the NPC in the first virtual scene is obtained; or when the target virtual character is a master virtual character and an NPC, the perception range of the master virtual character in the first virtual scene is obtained and the perception range of the NPC in the first virtual scene is obtained.

[0550] In some embodiments, the perception range comprises at least one of the following.

[0551] (1) a visual perception range (visual field range) of the target virtual character.

[0552] Illustratively, the visual perception range refers to a spatial range that an observer can visually perceive or understand, and is usually used to describe the size of an area that a biological body or a technical device can visually cover or perceive. The visual perception range of the target virtual character refers to a three-dimensional spatial range that the target virtual character can perceive through observation as an observer.

[0553] Optionally, the visual perception range of the target virtual character is obtained based on the position of the target virtual character in the first virtual scene and the visual observation capability.

[0554] (2) an auditory perception range (auditory range) of the target virtual character.

[0555] Illustratively, the auditory perception range refers to a spatial range that a listener can audibly perceive or understand, and is usually used to describe the size of a sound range that a biological body or a technical device can audibly cover or perceive. The auditory perception range of the target virtual character refers to a three-dimensional spatial range that the target virtual character can perceive through auditory perception as a listener.

[0556] Optionally, the auditory perception range of the target virtual character is obtained based on the position of the target virtual character in the first virtual scene and the auditory perception capability.

[0557] (3) an olfactory perception range (olfactory range) of the target virtual character.

[0558] Illustratively, the olfactory perception range refers to a spatial range in which a smell is perceived or identified in the sense of smell. The olfactory perception range of the target virtual character refers to a three-dimensional spatial range that the target virtual character can perceive in the sense of smell.

[0559] Optionally, based on the position of the target virtual role in the first virtual scene and the olfactory perception capability, an olfactory perception range of the target virtual role is obtained.

[0560] (4) a detection range of a perception skill possessed by the target virtual role or a perception virtual prop.

[0561] Illustratively, the perception skill is a capability of the target virtual role to perform scene perception on the first virtual scene, and the perception skill includes at least one of a perception acquisition form and a perception enhancement form.

[0562] Optionally, the perception acquisition form is used to represent a process of acquiring a non-possessed perception skill. After the perception skill is acquired through the perception acquisition form, the target virtual role can perform targeted scene perception on the first virtual scene through the perception skill. Optionally, the perception enhancement form is used to represent a process of enhancing a certain perception capability by possessing the perception skill in the case of possessing the certain perception capability.

[0563] Illustratively, the perception virtual prop is a virtual prop for perceiving the first virtual scene, and the perception virtual prop includes at least one of a perception acquisition form and a perception enhancement form.

[0564] Optionally, different perception skills or perception virtual props each correspond to a detection range, which represents a range interval that can be perceived when the perception skill or the perception virtual prop is applied, and the detection range represents the perception range.

[0565] In some embodiments, the perception range is a spherical range, a ring-shaped range, an irregular three-dimensional space range, etc., and the shape of the perception range is not limited here.

[0566] In an optional embodiment, a plurality of first scene hotwords are obtained based on environmental perception information in the perception range.

[0567] Illustratively, the environmental perception information is used to represent environmental information perceived and obtained by the target virtual role in the perception range. Optionally, the environmental perception information includes entities determined by emitting rays, and can also include region structure information determined by analyzing the geometric structure and layout information of the virtual environment, etc.

[0568] Illustratively, after the perception range corresponding to the target virtual role is determined, based on the range of the perception range in the first virtual scene and the environmental perception information determined by the target virtual role in the perception range based on perception, a plurality of first scene hotwords representing the state of the target virtual role are obtained.

[0569] In an optional embodiment, the environmental perception information includes entities; and entities in the first virtual scene within the perception range are determined.

[0570] The perception range is a part of the three-dimensional space range in the first virtual scene. The first virtual scene includes a large number of entities, and the entities are elements constituting the virtual scene, such as virtual ground, virtual buildings, virtual trees, virtual characters, virtual stones, and the like.

[0571] In some embodiments, in the case of determining the perception range, the entities in the perception range are determined from the plurality of entities corresponding to the first virtual scene.

[0572] In some embodiments, the entity name of the entity is taken as the first scene hotword; or the action name of the interactive action corresponding to the entity is taken as the first scene hotword.

[0573] Sub-step 11: converting the natural language command in the form of speech into a text form based on the plurality of first scene hotwords;

[0574] Illustratively, the first scene hotword is based on the vocabulary highly associated with the first virtual scene, and therefore, in order to better coordinate the NPC with the main control virtual character activity through the natural language command, the first scene hotword can be taken as a constraint condition on the basis of the first virtual scene, so that in the process of converting the natural language command into a text form, the text content of the command analysis result is more suitable for the first virtual scene, and the command analysis result that is not suitable for the first virtual scene is avoided.

[0575] In an optional embodiment, in the case of analyzing the natural language command through the natural language analysis model, the plurality of speech units corresponding to the natural language command are obtained through an acoustic network.

[0576] The speech unit is a basic constituting unit of the pronunciation of the vocabulary.

[0577] In some embodiments, the command feature representation corresponding to the natural language command is extracted; and the command feature representation is analyzed through the acoustic network.

[0578] Illustratively, taking the command feature representation representing a plurality of acoustic features as an example, the target of the acoustic network is to map the continuous acoustic feature sequence to a speech unit sequence, such as a phoneme or a phonetic segment. It learns to predict the possible speech unit (phoneme or phonetic segment) sequence under the given acoustic feature.

[0579] Optionally, the acoustic network outputs one or more possible speech unit sequences, and the speech unit sequence includes a plurality of speech units. Different speech unit sequences may have different speech units, or only the order of the speech units may be different. The plurality of speech unit sequences represent the arrangement of the speech units considered by the language network as the most possible when understanding the input natural language command.

[0580] The natural language analysis model includes an acoustic network, a preset dictionary, and language sub-networks corresponding to the at least two virtual scenes respectively.

[0581] Illustratively, the virtual environment includes at least two virtual scenes. For example, the virtual environment is a large virtual world, which includes virtual kitchen 1, virtual kitchen 2, virtual office, etc., and each of virtual kitchen 1, virtual kitchen 2, virtual office, etc. can be regarded as a virtual scene.

[0582] In an optional embodiment, based on the first virtual scene in which the target virtual role is located, a first language sub-network corresponding to the first virtual scene is determined from the at least two language sub-networks.

[0583] Illustratively, the plurality of virtual scenes correspond to a language sub-network respectively, and different language sub-networks are used for analyzing the corresponding virtual scenes. For example, virtual scene A corresponds to language sub-network 1, and virtual scene B corresponds to language sub-network 2; when analysis based on virtual scene A is needed, language sub-network 1 is used to perform language analysis; when analysis based on virtual scene B is needed, language sub-network 2 is used to perform language analysis, etc.

[0584] Illustratively, the first virtual scene is a virtual scene in the at least two virtual scenes, and in addition to determining the first scene hot word based on the first virtual scene in which the target virtual role is located, a first language sub-network corresponding to the first virtual scene is determined from the at least two language sub-networks based on the first virtual scene, and the first language sub-network is a sub-network layer for performing semantic analysis on the natural language command based on the first virtual scene.

[0585] In an optional embodiment, the sequence matching relationship between the plurality of voice units and the selected vocabulary in the preset dictionary is analyzed by the first language sub-network to convert the natural language command in voice form into text form.

[0586] Optionally, the different language sub-networks are language sub-networks pre-trained based on the corresponding virtual scenes, and the first language sub-network is a language sub-network trained based on the first virtual scene.

[0587] In some embodiments, the matching relationship between the plurality of voice units and the selected vocabulary in the preset dictionary is analyzed to obtain a plurality of candidate vocabulary sequences.

[0588] The selected vocabulary includes at least a plurality of first scene hot words and general vocabulary.

[0589] The process of analyzing the matching relationship is illustratively a process of analyzing the matching between the selected vocabulary and the phonetic unit. Illustratively, each selected vocabulary corresponds to at least one phonetic unit, and by analyzing the matching between the plurality of phonetic units and the selected vocabulary, a conversion probability of the phonetic unit converting into the selected vocabulary is determined, the conversion probability representing the matching probability between the selected vocabulary and the phonetic unit, the greater the conversion probability, the closer the phonetic unit to the selected vocabulary, i.e., the greater the possibility of the phonetic unit expressing the selected vocabulary, and the stronger the matching relationship between the selected vocabulary and the phonetic unit.

[0590] Optionally, the plurality of phonetic units form at least one phonetic unit sequence, and based on the process of analyzing the matching between the phonetic unit and the selected vocabulary, at least one candidate vocabulary sequence corresponding to the at least one phonetic unit sequence is determined, and a plurality of candidate vocabulary sequences are obtained.

[0591] In some embodiments, the sequence semantics of the plurality of candidate vocabulary sequences are analyzed by the first language sub-network, and at least one candidate vocabulary sequence is obtained from the plurality of candidate vocabulary sequences as the text form of the natural language command conversion.

[0592] Illustratively, the sequence semantics are used to represent the semantic information expressed by the candidate vocabulary sequence, which can not only reflect the semantic changes of at least one vocabulary in the candidate vocabulary sequence, but also reflect the sentence semantics of the entire candidate vocabulary sequence.

[0593] Illustratively, the natural language command in the text form is referred to as a command analysis result; the plurality of candidate vocabulary sequences are respectively evaluated by the first language sub-network, so that at least one candidate vocabulary sequence that is more consistent with the current semantic situation is obtained as the command analysis result, i.e., the natural language command in the phonetic form is converted into the text form.

[0594] It is worth noting that the above is only an illustrative example, and the embodiments of the present application are not limited in this regard.

[0595] The following is a device embodiment of the present application. For details not described in detail in the device embodiment, reference can be made to the above method embodiments.

[0596] FIG. 20 is a block diagram of a non-player character directing device according to an example embodiment of the present application. The device is used to implement a terminal. The device includes:

[0597] The display module 1004 is configured to display at least one of the master virtual character and the non-player character.

[0598] The interaction module 1001 is configured to receive a natural language command, the natural language command being used to direct the non-player character.

[0599] The first determining module 1003 is configured to determine a target entity, which is an entity in the virtual environment that matches the description in the natural language command and is perceived by the non-player character or the host virtual character.

[0600] The first control module 1002 is configured to control the non-player character to perform a virtual activity related to the target entity in response to the behavior intention of the natural language command.

[0601] In an optional embodiment, the first determining module 1003 is configured to obtain an entity set of the virtual environment, the entity set including entities and entity information in the virtual environment; perform semantic understanding on the natural language command to obtain the behavior intention and target entity information of the natural language command; and query the target entity from the entity set according to the target entity information and environment perception information.

[0602] In an optional embodiment, the first determining module 1003 is configured to calculate a similarity between the target entity information and each entity information in the entity set; and determine the target entity from the entity set according to the similarity and the environment perception information.

[0603] In an optional embodiment, the first determining module 1003 is configured to determine a search range according to the environment perception information; and filter the target entity from entities in the search range according to the similarity.

[0604] In an optional embodiment, the first determining module 1003 is configured to determine a visual field range of the host virtual character as the search range according to a position of the host virtual character and an orientation of the host virtual character.

[0605] The first determining module 1003 is configured to determine an auditory range of the host virtual character as the search range according to a position of the host virtual character.

[0606] The first determining module 1003 is configured to determine a detection range of a perception virtual prop used by the host virtual character as the search range.

[0607] The first determining module 1003 is configured to determine a visual field range of the non-player character as the search range according to a position of the non-player character and an orientation of the non-player character.

[0608] The first determining module 1003 is configured to determine an auditory range of the non-player character as the search range according to a position of the non-player character.

[0609] The first determining module 1003 is configured to determine the detection range of the perception virtual prop used by the non-player character as the search range.

[0610] The first determining module 1003 is configured to determine, according to the position of the host virtual character, the orientation of the host virtual character, the position of the non-player character, and the orientation of the non-player character, the overlapping part of the visual range of the host virtual character and the non-player character as the search range.

[0611] The first determining module 1003 is configured to determine, according to the position of the host virtual character and the position of the non-player character, the overlapping part of the hearing range of the host virtual character and the non-player character as the search range.

[0612] In an optional embodiment, the target entity information includes position information of the target entity.

[0613] The first determining module 1003 is configured to determine a perception range according to the environment perception information, and the perception range includes at least one of the following: the visual range of the host virtual character, the hearing range of the host virtual character, the visual range of the non-player character, the hearing range of the non-player character, and the detection range of the perception virtual prop.

[0614] The first determining module 1003 is configured to determine a range of an area in the perception range that matches the position information as the search range.

[0615] In an optional embodiment, the first determining module 1003 is configured to determine, as the target entity, an entity in the search range that has the highest similarity.

[0616] The first determining module 1003 is configured to, in a case where the number of entities in the search range that have the highest similarity is at least two, determine, as the target entity, an entity in the search range that has the highest similarity and is closest to the host virtual character.

[0617] In an optional embodiment, the similarity includes a text similarity; and the entity set includes entity information of a first entity.

[0618] The first determining module 1003 is configured to perform word segmentation on the target entity information to obtain at least one target entity label; convert the at least one target entity label into at least one target embedding vector; obtain entity information of the first entity, the entity information including a text embedding vector, the text embedding vector being an embedding vector converted based on a text label of the first entity; calculate a text parent similarity between the at least one target embedding vector and the text embedding vector respectively to obtain at least one text parent similarity corresponding to the at least one target embedding vector respectively; and determine a sum of the at least one text parent similarity as a text similarity between the target entity information and the entity information of the first entity.

[0619] In an optional embodiment, the entity information of the first entity includes at least one text embedding vector; and the at least one target embedding vector includes a first target embedding vector.

[0620] The first determining module 1003 is configured to calculate a text child similarity between the first target embedding vector and the at least one text embedding vector respectively to obtain at least one text child similarity; and determine a highest value in the at least one text child similarity as the text parent similarity corresponding to the first target embedding vector.

[0621] In an optional embodiment, the similarity includes an image similarity; and the entity set includes entity information of a first entity.

[0622] The first determining module 1003 is configured to perform word segmentation on the target entity information to obtain at least one target entity label; convert the at least one target entity label into at least one target embedding vector; obtain entity information of the first entity, the entity information including an image embedding vector, the image embedding vector being an embedding vector extracted based on an image of the first entity; calculate an image parent similarity between the at least one target embedding vector and the text embedding vector respectively to obtain at least one image parent similarity corresponding to the at least one target embedding vector respectively; and determine a sum of the at least one image parent similarity as an image similarity between the target entity information and the entity information of the first entity.

[0623] In an optional embodiment, the entity information of the first entity includes at least one image embedding vector; and the at least one target embedding vector includes a first target embedding vector.

[0624] The first determining module 1003 is configured to calculate an image child similarity between the first target embedding vector and the at least one image embedding vector respectively to obtain at least one image child similarity; and determine a highest value in the at least one image child similarity as the image parent similarity corresponding to the first target embedding vector.

[0625] In an optional embodiment, the entity set includes a first entity;

[0626] The first determining module 1003 is configured to, in the case that the first entity and the target entity correspond to a text similarity and an image similarity, determine an average value of the text similarity and the image similarity as the similarity of the entity information of the first entity and the target entity information.

[0627] In an optional embodiment, the entity set includes a first entity; and the device further includes:

[0628] The first preprocessing module 1008 is configured to acquire at least one perspective image of the first entity, and acquire at least one text label of the first entity; the at least one perspective image is used to describe the style of the first entity; the text label is used to introduce the inherent attribute of the first entity in the virtual environment; call a multi-modal model, extract the visual feature of the first entity based on the at least one perspective image and the at least one text label of the first entity, obtain a visual label of the first entity, and the visual label is used to describe the visual feature of the first entity in at least one dimension; and convert the visual label into an image embedding vector of the first entity.

[0629] In an optional embodiment, the multi-modal model includes a visual question and answer model;

[0630] The first preprocessing module 1008 is configured to construct a question sentence for the perspective image, the question sentence carrying the text label of the first entity; input the perspective image and the question sentence of the first entity into the visual question and answer model to obtain an answer sentence, and take the answer sentence as the visual label of the first entity.

[0631] In an optional embodiment, the question sentence includes at least two sub-sentences;

[0632] The first preprocessing module 1008 is configured to input the perspective image of the first entity and a first sentence in the at least two sub-sentences into the visual question and answer model to obtain a first answer sub-sentence.

[0633] Repeat the above steps until at least two answer sub-sentences corresponding to the at least two sub-sentences are obtained, and the at least two sub-sentences are used to inquire the visual feature of the perspective image from multiple dimensions;

[0634] Perform sentence aggregation on the at least two answer sub-sentences to extract the answer sentence of the first entity.

[0635] In an optional embodiment, the first preprocessing module 1008 is configured to obtain expected information of the first entity, the expected information being used to indicate a description dimension of the first entity expected in the answer statement and / or an expected format of the answer statement; and construct the question statement of the perspective image according to the expected information and the text label.

[0636] In the question statement, a first subpart is supplementary introduction information of the first entity, and carries the text label of the first entity; and a second subpart is an answer guide statement for the visual question and answer model, and carries the expected information.

[0637] In an optional embodiment, the multi-modal model comprises a picture description model.

[0638] The first preprocessing module 1008 is configured to input the perspective image of the first entity into the picture description model to predict description text of the first entity; and perform sentence aggregation on the description text and the text label to extract the visual label of the first entity.

[0639] In an optional embodiment, the text label of the first entity comprises at least one of a name of the first entity in the virtual environment and a size of the first entity in the virtual environment.

[0640] In an optional embodiment, the perspective image of the first entity comprises images obtained by observing the first entity from at least two perspectives.

[0641] In an optional embodiment, the first preprocessing module 1008 is configured to perform pseudo-spokenization rewriting on the visual label of the first entity to obtain a matching label conforming to natural language spoken expression.

[0642] In an optional embodiment, the first preprocessing module 1008 is configured to input the visual label of the first entity into a large language model to predict the matching label conforming to natural language spoken expression, the large language model carrying prior knowledge of natural language spoken expression.

[0643] In an optional embodiment, the first preprocessing module 1008 is configured to obtain a first sample label pair, the first sample label pair comprising a first label before pseudo-spokenization rewriting and a second label obtained through pseudo-spokenization rewriting.

[0644] The first preprocessing module 1008 is configured to construct a rewriting guide statement according to the first sample label pair and the visual label, the rewriting guide statement having a natural semantic of rewriting the visual label with reference to the first sample label pair.

[0645] The first preprocessing module 1008 is configured to input the rewriting guide sentence into a large language model to predict the matching label conforming to the spoken language expression.

[0646] In an optional embodiment, the first preprocessing module 1008 is configured to obtain a spatial position of the first entity in the virtual environment, and determine the spatial position of the first entity as auxiliary information of the visual label of the first entity.

[0647] In an optional embodiment, the first preprocessing module 1008 is configured to obtain at least one of a coordinate position, orientation information, bounding box information, and cover point information of the first entity in the virtual environment.

[0648] The coordinate position is used to indicate the position of the first entity in the virtual environment, the orientation information is used to indicate the direction faced by the first entity in the virtual environment, the bounding box information is used to indicate the size of the first entity in the virtual environment, and the cover point information indicates a recommended virtual character standing point when a virtual character approaches the first entity.

[0649] In an optional embodiment, the target entity information includes reference entity information of a reference entity, and the reference entity is used to determine the target entity.

[0650] The first determining module 1003 is configured to query the reference entity from the entity set according to the reference entity information and the environment perception information.

[0651] The first determining module 1003 is configured to query the target entity from the entity set according to the reference entity, the target entity information, and the environment perception information.

[0652] In an optional embodiment, the target entity information includes reference entity information, a position relationship between the reference entity and the target entity, and feature text of the target entity.

[0653] The first determining module 1003 is configured to calculate the similarity between the feature text and each entity information in the entity set.

[0654] The first determining module 1003 is configured to determine a search range according to the environment perception information.

[0655] The first determining module 1003 is configured to filter the target entity from entities in the entity set within the search range according to the similarity and the position relationship.

[0656] In an alternative embodiment, the entity set includes a space entity, the space entity being used to identify a three-dimensional space region in the virtual environment; the apparatus further includes:

[0657] a first preprocessing module 1008, configured to create the space entity in the virtual environment, the space entity being used to enclose the three-dimensional space region to be identified;

[0658] the first preprocessing module 1008, configured to obtain a text label of the space entity, the text label including description text of a space relationship of the three-dimensional space region;

[0659] The description text of the space relationship includes at least one of the following: a scenario to which the three-dimensional space region belongs, a building to which the three-dimensional space region belongs, a floor in the building where the three-dimensional space region is located, a spatial orientation of the three-dimensional space region on the floor, and a spatial name of the three-dimensional space region.

[0660] FIG. 21 is a block diagram of a non-player character directing apparatus according to an example embodiment of the present application. The apparatus is configured to implement a server. The apparatus includes:

[0661] a receiving module 1005, configured to receive a natural language command, the natural language command being used to direct a non-player character;

[0662] a second determining module 1006, configured to determine a target entity, the target entity being an entity in the virtual environment that matches a description in the natural language command and is perceived by the non-player character or a host virtual character;

[0663] a second control module 1007, configured to control the non-player character to perform a virtual activity related to the target entity in response to a behavior intention of the natural language command.

[0664] In an alternative embodiment, the second determining module 1006 is configured to obtain an entity set of the virtual environment, the entity set including entities and entity information in the virtual environment; perform semantic understanding on the natural language command to obtain the behavior intention of the natural language command and target entity information; and query the target entity from the entity set according to the target entity information and environment perception information.

[0665] In an alternative embodiment, the second determining module 1006 is configured to calculate a similarity between the target entity information and each entity information in the entity set; and determine the target entity from the entity set according to the similarity and the environment perception information.

[0666] In an optional embodiment, the second determining module 1006 is configured to determine a search range according to the environment perception information; and filter the target entity from entities located in the search range according to the similarity.

[0667] In an optional embodiment, the second determining module 1006 is configured to determine a visual field range of the host virtual role as the search range according to the position of the host virtual role and the orientation of the host virtual role.

[0668] The second determining module 1006 is configured to determine an auditory range of the host virtual role as the search range according to the position of the host virtual role.

[0669] The second determining module 1006 is configured to determine a detection range of a perception virtual prop used by the host virtual role as the search range.

[0670] The second determining module 1006 is configured to determine a visual field range of the non-player role as the search range according to the position of the non-player role and the orientation of the non-player role.

[0671] The second determining module 1006 is configured to determine an auditory range of the non-player role as the search range according to the position of the non-player role.

[0672] The second determining module 1006 is configured to determine a detection range of a perception virtual prop used by the non-player role as the search range.

[0673] The second determining module 1006 is configured to determine an overlapping part of the visual field range of the host virtual role and the non-player role as the search range according to the position of the host virtual role, the orientation of the host virtual role, the position of the non-player role, and the orientation of the non-player role.

[0674] The second determining module 1006 is configured to determine an overlapping part of the auditory range of the host virtual role and the non-player role as the search range according to the position of the host virtual role and the position of the non-player role.

[0675] In an optional embodiment, the target entity information includes position information of the target entity.

[0676] The second determining module 1006 is configured to determine a perception range according to the environment perception information, the perception range including at least one of the following: the visual field range of the host virtual role, the auditory range of the host virtual role, the visual field range of the non-player role, the auditory range of the non-player role, and the detection range of the perception virtual prop.

[0677] The second determining module 1006 is configured to determine a range of an area matching the position information in the perception range as the search range.

[0678] In an optional embodiment, the second determining module 1006 is configured to determine an entity with the highest similarity in the search range as the target entity.

[0679] The second determining module 1006 is configured to, in a case where the number of entities with the highest similarity in the search range is at least two, determine an entity with the highest similarity in the search range and closest to the master virtual role as the target entity.

[0680] In an optional embodiment, the second determining module 1006 is configured to, in the calculation of the similarity between the target entity information and the entity information of each entity in the entity set, perform the following steps: tokenizing the target entity information to obtain at least one target entity label; converting the at least one target entity label into at least one target embedding vector; obtaining entity information of the first entity, the entity information including a text embedding vector, the text embedding vector being an embedding vector converted based on a text label of the first entity; respectively calculating text parent similarities between the at least one target embedding vector and the text embedding vector to obtain at least one text parent similarity corresponding to the at least one target embedding vector respectively; and determining a sum of the at least one text parent similarity as a text similarity between the target entity information and the entity information of the first entity.

[0681] In an optional embodiment, the entity information of the first entity includes at least one text embedding vector; and the at least one target embedding vector includes a first target embedding vector.

[0682] The second determining module 1006 is configured to respectively calculate text child similarities between the first target embedding vector and the at least one text embedding vector to obtain at least one text child similarity; and determine a highest value in the at least one text child similarity as the text parent similarity corresponding to the first target embedding vector.

[0683] In an optional embodiment, the similarity includes an image similarity; and the entity set includes entity information of a first entity.

[0684] The second determining module 1006 is configured to perform word segmentation on the target entity information to obtain at least one target entity label; convert the at least one target entity label into at least one target embedding vector; obtain entity information of the first entity, wherein the entity information comprises an image embedding vector, and the image embedding vector is an embedding vector extracted based on an image of the first entity; calculate image parent similarities between the at least one target embedding vector and the text embedding vector respectively to obtain at least one image parent similarity corresponding to the at least one target embedding vector respectively; and determine a sum of the at least one image parent similarity as an image similarity between the target entity information and the entity information of the first entity.

[0685] In an optional embodiment, the entity information of the first entity comprises at least one image embedding vector; and the at least one target embedding vector comprises a first target embedding vector.

[0686] The second determining module 1006 is configured to calculate image child similarities between the first target embedding vector and the at least one image embedding vector respectively to obtain at least one image child similarity; and determine a highest value in the at least one image child similarity as the image parent similarity corresponding to the first target embedding vector.

[0687] In an optional embodiment, the entity set comprises a first entity; and the apparatus further comprises:

[0688] The second preprocessing module 1009 is configured to, in a case where the first entity and the target entity correspond to a text similarity and an image similarity, determine an average value of the text similarity and the image similarity as a similarity between the entity information of the first entity and the target entity information.

[0689] In an optional embodiment, the entity set comprises a first entity.

[0690] The second preprocessing module 1009 is configured to obtain at least one perspective image of the first entity, and obtain at least one text label of the first entity; the at least one perspective image is used to describe a style of the first entity; the text label is used to introduce an inherent attribute of the first entity in the virtual environment; call a multi-modal model, extract a visual feature of the first entity based on the at least one perspective image and the at least one text label of the first entity to obtain a visual label of the first entity, wherein the visual label is used to describe the visual feature of the first entity in at least one dimension; and convert the visual label into an image embedding vector of the first entity.

[0691] In an optional embodiment, the second preprocessing module 1009 is configured to construct a question sentence for the perspective image, the question sentence carrying the text label of the first entity; input the perspective image of the first entity and the question sentence into the visual question answering model to obtain an answer sentence, and take the answer sentence as the visual label of the first entity.

[0692] In an optional embodiment, the question sentence includes at least two sub-sentences.

[0693] The second preprocessing module 1009 is configured to input the perspective image of the first entity and a first sentence in the at least two sub-sentences into the visual question answering model to obtain a first answer sub-sentence.

[0694] Repeat the above steps until at least two answer sub-sentences corresponding to the at least two sub-sentences are obtained, and the at least two sub-sentences are used to inquire the visual features of the perspective image from multiple dimensions.

[0695] Perform sentence aggregation on the at least two answer sub-sentences to extract the answer sentence of the first entity.

[0696] In an optional embodiment, the second preprocessing module 1009 is configured to obtain expected information of the first entity, the expected information being used to indicate a description dimension of the first entity expected in the answer sentence and / or an expected format of the answer sentence; and construct the question sentence of the perspective image according to the expected information and the text label.

[0697] In the question sentence, a first subpart is supplementary introduction information of the first entity, and carries the text label of the first entity; and a second subpart is an answer guiding sentence for the visual question answering model, and carries the expected information.

[0698] In an optional embodiment, the multi-modal model includes a picture description model.

[0699] The second preprocessing module 1009 is configured to input the perspective image of the first entity into the picture description model to predict a description text of the first entity; and perform sentence aggregation on the description text and the text label to extract the visual label of the first entity.

[0700] In an optional embodiment, the text label of the first entity includes at least one of a name of the first entity in the virtual environment and a size of the first entity in the virtual environment.

[0701] And / or, the visual angle image of the first entity includes images obtained by observing the first entity from at least two visual angles.

[0702] In an optional embodiment, the second preprocessing module 1009 is configured to perform pseudo-spoken language rewriting on the visual label of the first entity to obtain a matching label conforming to a spoken expression of natural language.

[0703] In an optional embodiment, the second preprocessing module 1009 is configured to input the visual label of the first entity into a large language model to predict the matching label conforming to the spoken expression of natural language, the large language model carrying prior knowledge of the spoken expression of natural language.

[0704] In an optional embodiment, the second preprocessing module 1009 is configured to obtain a first sample label pair, the first sample label pair including a first label before pseudo-spoken language rewriting and a second label obtained through pseudo-spoken language rewriting; construct a rewriting guide sentence according to the first sample label pair and the visual label, the rewriting guide sentence having a natural semantic of rewriting the visual label with reference to the first sample label pair; and input the rewriting guide sentence into a large language model to predict the matching label conforming to the spoken expression of natural language.

[0705] In an optional embodiment, the second preprocessing module 1009 is configured to obtain a spatial position of the first entity in the virtual environment, and determine the spatial position of the first entity as auxiliary information of the visual label of the first entity.

[0706] In an optional embodiment, the second preprocessing module 1009 is configured to obtain at least one of a coordinate position, orientation information, bounding box information, and cover point information of the first entity in the virtual environment.

[0707] The coordinate position is used to indicate the position of the first entity in the virtual environment, the orientation information is used to indicate the direction faced by the first entity in the virtual environment, the bounding box information is used to indicate the size of the first entity in the virtual environment, and the cover point information indicates a recommended virtual character standing point when a virtual character approaches the first entity.

[0708] In an optional embodiment, the target entity information includes reference entity information of a reference entity, and the reference entity is used to reference determination of the target entity.

[0709] The second determination module 1006 is configured to query the reference entity from the entity set according to the reference entity information and the environment perception information.

[0710] The second determining module 1006 is configured to query the target entity from the entity set according to the reference entity, the target entity information and the environment perception information.

[0711] In an optional embodiment, the target entity information includes reference entity information, a position relationship between the reference entity and the target entity, and feature text of the target entity.

[0712] The second determining module 1006 is configured to calculate a similarity between the feature text and each entity information in the entity set.

[0713] The second determining module 1006 is configured to determine a search range according to the environment perception information.

[0714] The second determining module 1006 is configured to filter the target entity from entities in the entity set within the search range according to the similarity and the position relationship.

[0715] In an optional embodiment, the entity set includes a space entity, and the space entity is used to identify a three-dimensional space region in the virtual environment. The apparatus further includes:

[0716] The second preprocessing module 1009 is configured to create the space entity in the virtual environment, and the space entity is used to surround the three-dimensional space region to be identified.

[0717] The second preprocessing module 1009 is configured to obtain a text label of the space entity, and the text label includes description text of a space relationship of the three-dimensional space region.

[0718] The description text of the space relationship includes at least one of the following: a scenario to which the three-dimensional space region belongs, a building to which the three-dimensional space region belongs, a floor on which the three-dimensional space region is located in the building, a space direction of the three-dimensional space region on the floor, and a space name of the three-dimensional space region.

[0719] It should be noted that the non-player character commanding apparatus provided in the above embodiments is only used as an example for the division of the above functional modules. In actual applications, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the non-player character commanding apparatus and the non-player character commanding method provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0720] The application further provides a terminal comprising a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the non-player character commanding method provided by each of the above method embodiments.

[0721] FIG. 22 shows a structural block diagram of a terminal 900 according to an example embodiment of the application. The terminal 900 can be a smartphone, a tablet computer, an MP3 (Moving Picture Experts Group Audio Layer III) player, an MP4 (Moving Picture Experts Group Audio Layer IV) player, a notebook computer or a desktop computer. The terminal 900 can also be referred to as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, or other names.

[0722] Generally, the terminal 900 comprises a processor 901 and a memory 902.

[0723] The processor 901 can comprise one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 901 can be implemented in the form of at least one of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array). The processor 901 can also comprise a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 901 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content to be displayed on a display screen. In some embodiments, the processor 901 can further comprise an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.

[0724] The memory 902 can include one or more computer-readable storage media. The computer-readable storage media can be non-transitory. The memory 902 can also include high-speed random access memory and can include non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 902 stores instructions for at least one application program, which are executed by the processor 901 to implement the information prompting method of the round chess game provided by the method embodiments of the present application.

[0725] In some embodiments, the terminal 900 can also optionally include a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902, and the peripheral device interface 903 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 903 through a bus, a signal line, or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 904, a display screen 905, a camera component 906, an audio circuit 907, a position component 908, and a power supply 909.

[0726] The peripheral device interface 903 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 901 and the memory 902. In some embodiments, the processor 901, the memory 902, and the peripheral device interface 903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 901, the memory 902, and the peripheral device interface 903 can be implemented on a separate chip or circuit board, and the present embodiment does not limit this.

[0727] The radio frequency circuit 904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 904 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 904 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 904 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G, and 5G), a wireless local area network, and / or a WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 904 can also include NFC (Near Field Communication) related circuitry, which is not limited by the present application.

[0728] Display screen 905 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 905 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 901 for processing. In this case, display screen 905 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 905, which serves as the front panel of terminal 900; in other embodiments, there may be at least two display screens 905, respectively disposed on different surfaces of terminal 900 or in a folded design; in still other embodiments, display screen 905 may be a flexible display screen, disposed on a curved or folded surface of terminal 900. Furthermore, display screen 905 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 905 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0729] The camera assembly 906 is used to acquire images or videos. Optionally, the camera assembly 906 includes a front-facing camera and a rear-facing camera. Typically, the ...

Claims

1. A method for commanding a non-player character, the method being performed by a terminal, the method comprising: displaying at least one of a master virtual character and a non-player character; receiving a natural language command for commanding the non-player character; determining a target entity, the target entity being an entity in a virtual environment that matches a description in the natural language command and is perceived by the non-player character or the master virtual character; controlling the non-player character to perform a virtual activity related to the target entity in response to an action intent of the natural language command.

2. The method of claim 1, wherein, The determining the target entity comprises: obtaining an entity set of the virtual environment, the entity set comprising entities and entity information in the virtual environment; performing semantic understanding on the natural language command to obtain the action intent of the natural language command and target entity information; querying the target entity from the entity set according to the target entity information and environment perception information.

3. The method of claim 2, wherein, The querying the target entity from the entity set according to the target entity information and environment perception information comprises: calculating a similarity between the target entity information and each entity information in the entity set; determining the target entity from the entity set according to the similarity and the environment perception information.

4. The method of claim 3, wherein, The determining the target entity from the entity set according to the similarity and the environment perception information comprises: determining a search range according to the environment perception information; screening the target entity from entities in the search range according to the similarity.

5. The method of claim 4, wherein, The determining the search range according to the environment perception information comprises at least one of: determining a visual field range of the master virtual character as the search range according to a position of the master virtual character and an orientation of the master virtual character; determining a hearing range of the master virtual character as the search range according to the position of the master virtual character; determining a detection range of a perception virtual prop used by the master virtual character as the search range; determining a visual field range of the non-player character as the search range according to a position of the non-player character and an orientation of the non-player character; determining a hearing range of the non-player character as the search range according to the position of the non-player character; determining a detection range of a perception virtual prop used by the non-player character as the search range; determining an overlapping part of the visual field ranges of the master virtual character and the non-player character as the search range according to the position of the master virtual character, the orientation of the master virtual character, the position of the non-player character and the orientation of the non-player character; determining an overlapping part of the hearing ranges of the master virtual character and the non-player character as the search range according to the position of the master virtual character and the position of the non-player character.

6. The method of claim 4, wherein, The screening the target entity from entities in the search range according to the similarity comprises one of: determining an entity with the highest similarity in the search range as the target entity; In a case where a number of entities with the highest similarity in the search range is at least two, an entity with the highest similarity in the search range and closest to the master virtual role is determined as the target entity.

7. The method according to any one of claims 3 to 6, wherein, The similarity includes a text similarity; and the entity set includes entity information of the first entity. The calculating the similarity between the target entity information and each entity information in the entity set includes: tokenizing the target entity information to obtain at least one target entity label; converting the at least one target entity label into at least one target embedding vector; obtaining entity information of the first entity, the entity information including a text embedding vector, the text embedding vector being an embedding vector converted based on a text label of the first entity; respectively calculating a text parent similarity between the at least one target embedding vector and the text embedding vector to obtain at least one text parent similarity corresponding to the at least one target embedding vector respectively; determining a sum of the at least one text parent similarity as the text similarity between the target entity information and the entity information of the first entity.

8. The method of any one of claims 3 to 6, wherein, The similarity includes an image similarity; and the entity set includes entity information of the first entity. The calculating the similarity between the target entity information and each entity information in the entity set includes: tokenizing the target entity information to obtain at least one target entity label; converting the at least one target entity label into at least one target embedding vector; obtaining entity information of the first entity, the entity information including an image embedding vector, the image embedding vector being an embedding vector extracted based on an image of the first entity; respectively calculating an image parent similarity between the at least one target embedding vector and the text embedding vector to obtain at least one image parent similarity corresponding to the at least one target embedding vector respectively; determining a sum of the at least one image parent similarity as the image similarity between the target entity information and the entity information of the first entity.

9. The method according to any one of claims 2 to 6, wherein, The entity set includes the first entity; and the method further includes: obtaining at least one perspective image of the first entity and at least one text label of the first entity, the at least one perspective image being used to describe a style of the first entity, and the text label being used to introduce an inherent attribute of the first entity in the virtual environment; calling a multi-modal model to extract a visual feature of the first entity based on the at least one perspective image and the at least one text label of the first entity, to obtain a visual label of the first entity, the visual label being used to describe the visual feature of the first entity in at least one dimension; converting the visual label into an image embedding vector of the first entity.

10. The method of claim 9, wherein, The multi-modal model includes a visual question answering model. The calling the multi-modal model to extract the visual feature of the first entity based on the at least one perspective image and the at least one text label of the first entity to obtain the visual label of the first entity includes: construct a question sentence for the perspective image, the question sentence carrying the text label of the first entity; input the perspective image of the first entity and the question sentence into the visual question answering model to obtain an answer sentence, and take the answer sentence as the visual label of the first entity.

11. The method of claim 10, wherein, The question sentence includes at least two sub-sentences. The inputting the perspective image of the first entity and the question sentence into the visual question answering model to obtain an answer sentence includes: input the perspective image of the first entity and a first sentence in the at least two sub-sentences into the visual question answering model to obtain a first answer sub-sentence; repeat the above steps until at least two answer sub-sentences corresponding to the at least two sub-sentences are obtained, and the at least two sub-sentences are used to inquire the visual features of the perspective image from multiple dimensions; perform sentence aggregation on the at least two answer sub-sentences to extract the answer sentence of the first entity.

12. The method of claim 9, wherein, The multi-modal model includes a picture description model; The calling the multi-modal model to extract the visual features of the first entity based on the at least one perspective image of the first entity and the at least one text label to obtain the visual label of the first entity includes: input the perspective image of the first entity into the picture description model to predict a description text of the first entity; perform sentence aggregation on the description text and the text label to extract the visual label of the first entity.

13. The method of claim 9, wherein, The method further includes: performing pseudo-spoken language rewriting on the visual label of the first entity to obtain a matching label conforming to a natural language spoken expression.

14. The method of claim 2, wherein, The target entity information includes reference entity information of a reference entity, and the reference entity is used to reference determination of the target entity; The querying the target entity from the entity set according to the target entity information and the environment perception information includes: querying the reference entity from the entity set according to the reference entity information and the environment perception information; querying the target entity from the entity set according to the reference entity, the target entity information and the environment perception information.

15. The method of claim 2, wherein, The entity set includes a space entity, and the space entity is used to identify a three-dimensional space region in the virtual environment; the method further includes: creating the space entity in the virtual environment, and the space entity is used to surround a three-dimensional space region to be identified; obtaining a text label of the space entity, and the text label includes description text of a space relationship of the three-dimensional space region; The description text of the space relationship includes at least one of the following: a scene to which the three-dimensional space region belongs, a building to which the three-dimensional space region belongs, a floor on which the three-dimensional space region is located in the building, a space orientation of the three-dimensional space region on the floor, and a space name of the three-dimensional space region.

16. A non-player character directing method, the method is executed by a server, and the method includes: receiving a natural language command, the natural language command being used to direct a non-player character; determining a target entity that matches the description in the natural language command and is perceived by the non-player character or the host virtual character in the virtual environment; controlling the non-player character to perform a virtual activity related to the target entity in response to the behavior intention of the natural language command.

17. The method of claim 16, wherein, The determining a target entity comprises: obtaining an entity set of the virtual environment, the entity set comprising entities and entity information in the virtual environment; performing semantic understanding on the natural language command to obtain the behavior intention of the natural language command and target entity information; querying the target entity from the entity set according to the target entity information and environment perception information.

18. The method of claim 17, wherein, The querying the target entity from the entity set according to the target entity information and environment perception information comprises: calculating a similarity between the target entity information and entity information of each entity in the entity set; determining the target entity from the entity set according to the similarity and the environment perception information.

19. The method of claim 18, wherein, The determining the target entity from the entity set according to the similarity and the environment perception information comprises: determining a search range according to the environment perception information; screening the target entity from entities in the search range according to the similarity.

20. The method of claim 19, wherein, The determining a search range according to the environment perception information comprises at least one of: determining a visual field range of the host virtual character as the search range according to a position of the host virtual character and an orientation of the host virtual character; determining a hearing range of the host virtual character as the search range according to the position of the host virtual character; determining a detection range of a perception virtual prop used by the host virtual character as the search range; determining a visual field range of the non-player character as the search range according to a position of the non-player character and an orientation of the non-player character; determining a hearing range of the non-player character as the search range according to the position of the non-player character; determining a detection range of a perception virtual prop used by the non-player character as the search range; determining an overlapping part of the visual field ranges of the host virtual character and the non-player character as the search range according to the position of the host virtual character, the orientation of the host virtual character, the position of the non-player character and the orientation of the non-player character; determining an overlapping part of the hearing ranges of the host virtual character and the non-player character as the search range according to the position of the host virtual character and the position of the non-player character.

21. The method of claim 19, wherein, The screening the target entity from entities in the search range according to the similarity comprises one of: determining an entity with the highest similarity in the search range as the target entity; in a case where a number of entities with the highest similarity in the search range is at least two, determining an entity with the highest similarity and closest to the host virtual character in the search range as the target entity.

22. The method of any one of claims 18 to 21, wherein, The similarity includes a text similarity; and the entity set includes entity information of the first entity. The calculating the similarity between the target entity information and each entity information in the entity set comprises: performing word segmentation on the target entity information to obtain at least one target entity label; converting the at least one target entity label into at least one target embedding vector; obtaining entity information of the first entity, the entity information including a text embedding vector, the text embedding vector being an embedding vector converted based on a text label of the first entity; respectively calculating text parent similarities between the at least one target embedding vector and the text embedding vector to obtain at least one text parent similarity corresponding to each of the at least one target embedding vector; determining a sum of the at least one text parent similarity as the text similarity between the target entity information and the entity information of the first entity.

23. The method of any one of claims 18 to 21, wherein, The similarity includes an image similarity; and the entity set includes entity information of the first entity. The calculating the similarity between the target entity information and each entity information in the entity set comprises: performing word segmentation on the target entity information to obtain at least one target entity label; converting the at least one target entity label into at least one target embedding vector; obtaining entity information of the first entity, the entity information including an image embedding vector, the image embedding vector being an embedding vector extracted based on an image of the first entity; respectively calculating image parent similarities between the at least one target embedding vector and the text embedding vector to obtain at least one image parent similarity corresponding to each of the at least one target embedding vector; determining a sum of the at least one image parent similarity as the image similarity between the target entity information and the entity information of the first entity.

24. The method of any one of claims 17 to 21, wherein, The entity set includes the first entity; and the method further comprises: obtaining at least one perspective image of the first entity and at least one text label of the first entity, the at least one perspective image being used to describe a style of the first entity, and the text label being used to introduce inherent properties of the first entity in the virtual environment; calling a multi-modal model to extract visual features of the first entity based on the at least one perspective image and the at least one text label of the first entity, to obtain a visual label of the first entity, the visual label being used to describe the visual features of the first entity in at least one dimension; converting the visual label into an image embedding vector of the first entity.

25. The method of claim 24, wherein, The multi-modal model includes a visual question answering model; and the calling the multi-modal model to extract the visual features of the first entity based on the at least one perspective image and the at least one text label of the first entity to obtain the visual label of the first entity comprises: constructing a question sentence for the perspective image, the question sentence carrying the text label of the first entity; inputting the perspective image and the question sentence of the first entity into the visual question answering model to obtain an answer sentence, and taking the answer sentence as the visual label of the first entity.

26. The method of claim 25, wherein, The question sentence includes at least two sub-sentences; The inputting the visual angle image of the first entity and the question sentence into the visual question answering model includes: Inputting the visual angle image of the first entity and a first sentence in the at least two sub-sentences into the visual question answering model to obtain a first answer sub-sentence; Repeating the above steps until at least two answer sub-sentences corresponding to the at least two sub-sentences are obtained, and the at least two sub-sentences are used to inquire the visual features of the visual angle image from multiple dimensions; Performing sentence aggregation on the at least two answer sub-sentences to extract the answer sentence of the first entity.

27. The method of claim 24, wherein, The multi-modal model includes a picture description model; The calling the multi-modal model to extract the visual features of the first entity based on the at least one visual angle image of the first entity and the at least one text label to obtain the visual label of the first entity includes: Inputting the visual angle image of the first entity into the picture description model to predict a description text of the first entity; Performing sentence aggregation on the description text and the text label to extract the visual label of the first entity.

28. The method of claim 24, wherein, The method further includes: Performing pseudo-spoken language rewriting on the visual label of the first entity to obtain a matching label conforming to a natural language spoken expression.

29. The method of claim 17, wherein, The target entity information includes reference entity information of a reference entity, and the reference entity is used to reference determination of the target entity; The querying the target entity from the entity set according to the target entity information and environment perception information includes: Querying the reference entity from the entity set according to the reference entity information and the environment perception information; Querying the target entity from the entity set according to the reference entity, the target entity information and the environment perception information.

30. The method of claim 17, wherein, The entity set includes a space entity, and the space entity is used to identify a three-dimensional space region in the virtual environment; the method further includes: Creating the space entity in the virtual environment, and the space entity is used to surround a three-dimensional space region to be identified; Obtaining a text label of the space entity, and the text label includes description text of a space relationship of the three-dimensional space region; The description text of the space relationship includes at least one of the following: a scene to which the three-dimensional space region belongs, a building to which the three-dimensional space region belongs, a floor on which the three-dimensional space region is located in the building, a space orientation of the three-dimensional space region on the floor, and a space name of the three-dimensional space region.

31. A non-player character directing device, the device comprising: a display module configured to display at least one of a master virtual character and a non-player character; an interaction module configured to receive a natural language command, the natural language command configured to direct the non-player character; a first determination module configured to determine a target entity, the target entity being an entity in a virtual environment that matches a description in the natural language command and is perceived by the non-player character or the master virtual character. The first control module is configured to control the non-player character to perform a virtual activity related to the target entity in response to the behavior intention of the natural language command. 32.A non-player character commanding apparatus, the apparatus comprising: a receiving module configured to receive a natural language command, the natural language command being used to command a non-player character; a second determining module configured to determine a target entity, the target entity being an entity in a virtual environment that matches a description in the natural language command and is perceived by the non-player character or a host virtual character; a second control module configured to control the non-player character to perform a virtual activity related to the target entity in response to the behavior intention of the natural language command. 33.A computer device, comprising a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the non-player character commanding method according to any one of claims 1 to 30. 34.A computer readable storage medium, the computer readable storage medium storing at least one program, the at least one program being loaded and executed by a processor to implement the non-player character commanding method according to any one of claims 1 to 30. 35.A computer program product or a computer program, the computer program product or the computer program comprising computer instructions stored in a computer readable storage medium; a processor of a computer device reads the computer instructions from the computer readable storage medium, and executes the computer instructions, so that the computer device performs to implement the non-player character commanding method according to any one of claims 1 to 30.

Citation Information

Patent Citations

  • Virtual object control method and device, equipment and storage medium

    CN116983639A

  • Non-player character interaction method and system

    CN118079400A

  • Non-player character command method and device, equipment and medium

    CN118987625A

  • Audio-visual games and game computer programs embodying interactive speech recognition and methods related thereto

    US20060040718A1

  • Techniques for creating dynamic game activities for games

    US20160121218A1