Broadcasting method and apparatus for non-player character, device, medium, and product

By receiving natural language commands and combining them with environmental awareness information to generate NPC feedback, the limitations of NPC voice broadcasts under fixed instructions are solved, improving the interaction efficiency of NPCs and the utilization of computing resources in the game.

WO2026031852A1PCT designated stage Publication Date: 2026-02-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/104664
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-06-27
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

In existing technologies, voice broadcasts by non-player characters (NPCs) in games are only performed under fixed commands, which cannot provide effective feedback in complex scenarios, resulting in low efficiency of human-computer interaction between players and NPCs.

Method used

By receiving natural language commands and combining environmental awareness information from the main virtual character and non-player characters, non-fixed text feedback is generated, enabling flexible voice broadcasting for NPCs.

Benefits of technology

It enriches the feedback information of NPCs, improves the efficiency of human-computer interaction, and reduces game time and computing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025104664_12022026_PF_FP_ABST
    Figure CN2025104664_12022026_PF_FP_ABST
Patent Text Reader

Abstract

A broadcasting method and apparatus for a non-player character, a device, a medium, and a product, relating to the field of human-computer interaction. The method comprises: displaying at least one of a master virtual character and a non-player character which are located in a three-dimensional virtual environment (210); receiving a natural language command for the non-player character (220); and broadcasting feedback information of the non-player character in response to the natural language command (230).
Need to check novelty before this filing date? Find Prior Art

Description

Non-player character broadcasting method, device, equipment, medium and product

[0001] The present application claims priority to the Chinese patent application No. 202411095466.4, filed on August 9, 2024, and entitled "Non-player character broadcasting method, device, equipment, medium and product", the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] Embodiments of the present application relate to the field of human-computer interaction, and in particular to a non-player character broadcasting method, device, equipment, medium and product. BACKGROUND

[0003] In the field of game design and development, AI(Artificial Intelligence) controlled NPC(Non-Player Character) has become a key technology to improve game immersion and interactivity.

[0004] In related technologies, in order to facilitate the feedback of the activity state of the non-player character to the player, the game application program provides a voice broadcasting capability. The game application program packages the pre-recorded voice broadcast into the client, and triggers the voice broadcast when receiving a fixed instruction.

[0005] The method in related technologies can only make a simple voice broadcast when receiving a fixed instruction, and in more complex scenarios, the player cannot obtain effective feedback from the NPC, and the human-computer interaction efficiency between the player and the NPC is too low. SUMMARY

[0006] Embodiments of the present application provide a non-player character broadcasting method, device, equipment, medium and product. The technical solution is as follows:

[0007] On the one hand, a non-player character broadcasting method is provided, and the method comprises:

[0008] displaying at least one of a master virtual character and a non-player character located in a three-dimensional virtual environment;

[0009] receiving a natural language command for the non-player character;

[0010] in response to the natural language command, broadcasting feedback information of the non-player character, the feedback text corresponding to the feedback information being a non-fixed text generated based on environment perception information, the environment perception information including information perceived by at least one of the master virtual character and the non-player character to the three-dimensional virtual environment.

[0011] In another aspect, a method for broadcasting a non-player character is provided, the method comprising:

[0012] receiving a natural language command for a non-player character sent by a client, the client having a control authority of a master virtual character;

[0013] obtaining environment perception information of at least one of the master virtual character and the non-player character;

[0014] determining feedback text of the non-player character according to the natural language command and the environment perception information, the feedback text being non-fixed text generated based on the environment perception information of the non-player character;

[0015] sending, to the client, information for playing feedback information based on the feedback text, the feedback information being human voice audio corresponding to the feedback text.

[0016] In another aspect, a device for broadcasting a non-player character is provided, the device comprising:

[0017] a display module configured to display at least one of a master virtual character and a non-player character located in a three-dimensional virtual environment;

[0018] a receiving module configured to receive a natural language command for the non-player character;

[0019] a broadcasting module configured to broadcast feedback information of the non-player character in response to the natural language command, the feedback information corresponding to feedback text being non-fixed text generated based on environment perception information, the environment perception information including information perceived by at least one of the master virtual character and the non-player character to the three-dimensional virtual environment.

[0020] In another aspect, a device for broadcasting a non-player character is provided, the device comprising:

[0021] a receiving module configured to receive a natural language command for a non-player character sent by a client, the client having a control authority of a master virtual character;

[0022] an obtaining module configured to obtain environment perception information of at least one of the master virtual character and the non-player character;

[0023] a determining module configured to determine feedback text of the non-player character according to the natural language command and the environment perception information, the feedback text being non-fixed text generated based on the environment perception information of the non-player character;

[0024] The sending module is configured to send, to the client, relevant information for playing feedback information corresponding to the feedback text based on the feedback text, the feedback information being human voice audio corresponding to the feedback text.

[0025] In another aspect, a computer device is provided, which includes a processor and a memory having stored therein at least one instruction, at least one program, a code set or instruction set, which is loaded and executed by the processor to implement the non-player character broadcasting method as described in the above aspects.

[0026] In another aspect, a computer readable storage medium is provided, which has stored therein at least one instruction, at least one program, a code set or instruction set, which is loaded and executed by a processor to implement the non-player character broadcasting method as described in the above aspects.

[0027] In another aspect, the embodiments of the present application provide a computer program product or computer program, which includes computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the non-player character broadcasting method provided in the above optional implementation manners.

[0028] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:

[0029] When a user feeds back relevant information of a three-dimensional virtual environment to a non-player character in the three-dimensional virtual environment through a natural language command, according to the information perceived by the non-player character and / or a master virtual character from the three-dimensional virtual environment, feedback information corresponding to the non-player character is generated and broadcasted, so that the feedback information can cover multiple scenarios. When the non-player character is controlled through a natural language command, the feedback information is broadcasted according to the environmental perception information of the non-player character and / or the master virtual character, which can enrich the content that can be fed back by the feedback information, so that the broadcast is not limited to a fixed triggering scenario. The game application program can flexibly broadcast the feedback information according to the environmental perception information in a game match, enrich the feedback information broadcast corresponding to the non-player character, and improve the human-computer interaction efficiency between the non-player character and the player.

[0030] And, since the user can instruct the non-player character to feedback the information perceived from the three-dimensional virtual scene through the natural language command, the user can quickly acquire the information in the three-dimensional virtual scene, and thus the user can perform appropriate operations according to the known three-dimensional virtual scene information, accelerate the game session process, reduce the game session time, thereby reducing the cost of computing resources required in implementing the game session and reducing the power consumption of the device. BRIEF DESCRIPTION OF DRAWINGS

[0031] FIG. 1 is a structural block diagram of a computer system according to an example embodiment of the present application;

[0032] FIG. 2 is a method flowchart of a non-player character broadcasting method according to an example embodiment of the present application;

[0033] FIG. 3 is a schematic diagram of feedback information corresponding to a non-player character according to an example embodiment of the present application;

[0034] FIG. 4 is a method flowchart of a non-player character broadcasting method according to an example embodiment of the present application;

[0035] FIG. 5 is a method flowchart of a non-player character broadcasting method according to an example embodiment of the present application;

[0036] FIG. 6 is a method flowchart of a non-player character broadcasting method according to an example embodiment of the present application;

[0037] FIG. 7 is a schematic diagram of a three-dimensional virtual environment according to an example embodiment of the present application;

[0038] FIG. 8 is a schematic diagram of a three-dimensional virtual environment according to an example embodiment of the present application;

[0039] FIG. 9 is a schematic diagram of a position and orientation of a host virtual character according to an example embodiment of the present application;

[0040] FIG. 10 is a schematic diagram of spatial hierarchy of a target entity according to an example embodiment of the present application;

[0041] FIG. 11 is a method flowchart of a non-player character broadcasting method according to an example embodiment of the present application;

[0042] FIG. 12 is a method flowchart of a non-player character broadcasting method according to an example embodiment of the present application;

[0043] FIG. 13 is a method flowchart of a non-player character broadcasting method according to an example embodiment of the present application;

[0044] FIG. 14 is a schematic diagram of a grid structure according to an example embodiment of the present application;

[0045] FIG. 15 is a schematic diagram of reverberation sound waves according to an example embodiment of the present application;

[0046] FIG. 16 is a flowchart of a method for broadcasting a non-player character according to an example embodiment of the present application;

[0047] FIG. 17 is a schematic diagram of a speech generation model according to an example embodiment of the present application;

[0048] FIG. 18 is a flowchart of a method for broadcasting a non-player character according to an example embodiment of the present application;

[0049] FIG. 19 is a block diagram of a device for broadcasting a non-player character according to an example embodiment of the present application;

[0050] FIG. 20 is a block diagram of a device for broadcasting a non-player character according to an example embodiment of the present application;

[0051] FIG. 21 is a structural block diagram of a terminal according to an example embodiment of the present application. DETAILED DESCRIPTION

[0052] First, terms involved in the embodiments of the present application are briefly introduced:

[0053] Three-dimensional virtual environment: is a three-dimensional virtual environment displayed (or provided) by an application when running on a terminal. The three-dimensional virtual environment can be a simulation environment of the real world, or a semi-simulation and semi-fictional environment, or a purely fictional environment. The three-dimensional virtual environment can be any one of a two-dimensional three-dimensional virtual environment, a 2.5-dimensional three-dimensional virtual environment, and a three-dimensional virtual environment, which is not limited by the embodiments of the present application.

[0054] Virtual character: a virtual character refers to a movable object in a three-dimensional virtual environment. The movable object can be a virtual person, a virtual animal, an animation character, etc., such as a person displayed in a three-dimensional virtual environment. Alternatively, the virtual character is a three-dimensional model created based on animation skeleton technology. Each virtual character has its own shape and volume in the three-dimensional virtual world, and occupies a part of the space in the three-dimensional virtual world.

[0055] User interface (UI) control: refers to any visible control or element that can be seen on the user interface of an application. For example, pictures, input boxes, text boxes, buttons, labels, etc. Some of the UI controls respond to user operations.

[0056] 3D (Three-Dimensional) Model: In computer graphics and video game design, a 3D model is a three-dimensional object defined by three-dimensional geometric data. These data can be created manually or generated by 3D scanners. A complete 3D model allows users to view the object from any angle and may contain additional information such as textures, colors, and surface material properties. In electronic games, this type of model is often used to create the appearance of players or NPCs (non-player characters).

[0057] Natural Language Processing (NLP): A subfield of artificial intelligence and computer science that studies how to enable computers to understand, interpret, and generate human natural language, achieving natural language communication between humans and computers.

[0058] Text-to-Speech (TTS): A technology that converts text information into audible speech signals, allowing computers or other devices to read text information with human voice pronunciation.

[0059] Entity Query: A method for querying specific entities (such as characters, objects, etc.) in a game or virtual world, usually used to obtain a list of entities that meet certain conditions for further processing.

[0060] Spatial Query: A method for querying corresponding spaces based on spatial relationships (such as distance, direction, intersection, etc.), used to determine the specific space where the target entity is currently located. Spatial query can return a position description of the target entity's location based on the spatial position relationship between the main virtual character and the target entity. The target entity refers to a specified entity in a three-dimensional virtual scene, which can include at least one of virtual characters, virtual props, virtual buildings, virtual animals, etc. in a three-dimensional virtual scene. The target entity can be selected by the player or automatically detected by the computer device.

[0061] Intent Recognition: A task in the field of natural language processing aimed at recognizing the user's intent from the user's input natural language text, so as to perform corresponding processing and response.

[0062] Large Language Model (LLM): A natural language processing model trained based on a large amount of text data, usually with strong language understanding and generation capabilities, which can be applied to tasks such as machine translation, text summarization, and question-answering systems.

[0063] NPC (Non-Player Character): refers to a character in a game that is controlled by a computer, usually with certain behavior patterns and tasks, can interact with players, provide tasks or act as enemies, etc.

[0064] Figure 1 shows a structural block diagram of a computer system provided by an exemplary embodiment of the present application. The computer system includes a terminal 110 and a server 120.

[0065] The terminal 110 is installed and runs a client 111 supporting a three-dimensional virtual environment, which is a client of an application program. When the terminal runs the client 111, a user interface of the client 111 is displayed on a screen of the terminal 110. The application program can be any one of a battle royale shooting game, a virtual reality (VR) application program, an augmented reality (AR) program, a three-dimensional map program, a virtual reality game, an augmented reality game, a first-person shooting game (FPS), a third-person shooting game (TPS), a multiplayer online battle arena game (MOBA), and a simulation game (SLG).

[0066] In this embodiment, the client is taken as an example of a MOBA game. The terminal 110 is a terminal used by a first user 112, and in a game session, the first user 112 uses the terminal 110 to control a first virtual character located in a three-dimensional virtual environment to perform activities, which can be referred to as a master virtual character of the first user 112 in the game session. The activities of the first virtual character include, but are not limited to, at least one of adjusting a body posture, crawling, walking, running, riding, flying, jumping, driving, picking up, shooting, attacking, and throwing.

[0067] Only one terminal is shown in Figure 1, but there are multiple other terminals that can access the server 120 in different embodiments. Optionally, there is also one or more terminals that are corresponding to developers, on which a development and editing platform supporting the three-dimensional virtual environment is installed, and the developers can edit and update the client on the terminal, and transmit the updated client installation package to the server 120 through a wired or wireless network, and the terminal 110 can download the client installation package from the server 120 to update the client.

[0068] The terminal 110 and other terminals are connected to the server 120 through a wireless network or a wired network.

[0069] The server 120 comprises at least one of a server, multiple servers, a cloud computing platform, and a virtualization center. The server 120 is configured to provide background services for clients supporting a three-dimensional virtual environment. Optionally, the server 120 undertakes major computing work, and the terminal undertakes secondary computing work; or the server 120 undertakes secondary computing work, and the terminal undertakes major computing work; or the server 120 and the terminal adopt a distributed computing architecture for collaborative computing.

[0070] In an illustrative example, the server 120 comprises a processor, a user account database, a battle service module, and a user-oriented input / output interface (I / O interface). The processor is configured to load instructions stored in the server 120 and process data in the user account database and the battle service module. The user account database is configured to store data of user accounts used by the terminal 110 and other terminals, such as avatars of the user accounts, nicknames of the user accounts, battle power indexes of the user accounts, and service areas in which the user accounts are located. The battle service module is configured to provide multiple battle rooms for users to perform battles, such as 1V1 battles, 3V3 battles, and 5V5 battles. The user-oriented I / O interface is configured to establish communication with the terminal 110 through a wireless network or a wired network and exchange data.

[0071] The embodiments of the present application provide a scheme for a user to use a voice form natural language command to instruct a non-player character (NPC) in a virtual environment. In the scheme, the user usually has one or more NPCs as teammates in addition to a master virtual character controlled by the user. The user can use a human conversation manner that is relatively arbitrary and non-mechanical to instruct the NPCs to make feedbacks on scene objects in the virtual environment as desired by the user, so as to instruct the NPCs to complete a task in cooperation with the master virtual character. Referring to FIG. 1, the user controls a master virtual character 10 to play a game, and the master virtual character 10 has an NPC teammate 20. There is a car 30 in a visual range of the master virtual character 10, and the user says a voice form natural language command: “No. 2, go to hide behind the car 30”. Then the NPC teammate 20 will move to hide behind the car 30 by itself. It should be noted that there can be multiple cars in the scene shown in FIG. 1, and the NPC teammate 20 will accurately understand the car said by the user as the car 30 in the visual range of the user. That is, the NPC teammate 20 has relatively intelligent natural language understanding capability.

[0072] Compared with the traditional technology of using mechanical fixed instructions to command the NPC, the embodiments of the present application use relatively arbitrary natural language commands to command the NPC, which can provide the user with more natural, more complex, and more flexible language command capabilities.

[0073] On the other hand, the intelligent natural language understanding capability of the embodiments of the present application is embodied in that the NPC teammates not only consider the literal information of the natural language commands when understanding the natural language commands, but also combine the environmental perception information of the virtual environment of the main control virtual character and / or the NPC itself to assist in understanding the semantics of the natural language commands. That is, the NPC not only considers the information of one modality of the natural language command, but also considers the perception information of other modalities such as visual field, hearing, radar, etc. to assist in understanding and executing the natural language command. Since the natural language command in the form of voice is a command in the form of spoken language, not a command in the form of written language, there may be multiple candidate understanding ways or unclear places when only considering the literal information of the natural language command. The NPC teammates combine the environmental perception information of the virtual environment of the main control virtual character and / or the NPC itself to determine a reasonable understanding way among the multiple candidate understanding ways or eliminate doubts about the unclear places, thereby realizing a relatively intelligent natural language understanding capability. The environmental perception information mentioned above includes but is not limited to at least one of the following: visual information perceived within the visual field of the main control virtual character; hearing information perceived within the hearing range of the main control virtual character; information perceived by a perception skill or a perception device possessed by the main control virtual character; visual information perceived within the visual field of the NPC; hearing information perceived within the hearing range of the NPC; information perceived by a perception skill or a perception device possessed by the NPC.

[0074] In combination with reference to FIG. 1, the scheme includes at least one of the following five stages:

[0075] Stage one: preprocessing of spatial data;

[0076] The server will preprocess the scene objects in the virtual scene and construct the spatial data of the virtual scene.

[0077] The spatial data of the virtual scene includes the visual label of the scene object and the spatial position of the scene object, and in some examples, the spatial data also includes the appearance image of the scene object. The scene object is any object appearing in the virtual scene, such as the car, wall, box, etc. in FIG. 1.

[0078] Firstly, the pre-acquisition process of the visual label of the scene object is introduced. The attribute information of the scene object in the virtual scene is acquired, such as at least one of the name, size and the like of the scene object; and the appearance image of the scene object is acquired, the appearance image including images obtained by observing the scene object from at least two perspectives, and observing the scene object from multiple perspectives can ensure that the appearance image carries comprehensive appearance information of the scene object.

[0079] The server calls the multi-modal model to perform prediction on the attribute information and the appearance image of the scene object, extracts hidden layer features from the attribute information in the text dimension and the appearance image in the image dimension based on the ability of the multi-modal model to process information in the text dimension and the image dimension, and performs decoding on the hidden layer features to predict the visual label of the scene object, the visual label being used to describe the scene object in at least one dimension in a text manner; such as describing the scene object in the dimensions of material, transparency, color, shape, etc. The natural language model is called to perform label optimization on the visual label to obtain the visual label of the scene object; the natural language model has a text generation capability, and performs rewriting on the visual label input into the natural language model in a manner conforming to spoken language expression of natural language, so as to realize label optimization of the visual label and obtain the optimized visual label of the scene object. The optimized visual label conforms to the manner of spoken language expression of natural language, and describes the scene object in at least one dimension in a manner close to spoken language expression.

[0080] Taking a virtual bed in a virtual scene as an example, usually the color, material, and position of the virtual bed are concerned, while the surface features of the head of the bed and the side plate of the bed are usually ignored. The purpose of performing pseudo-spoken language rewriting on the visual label of the scene object is to obtain a matching label closer to spoken language expression; such as a label of the dimensions of color, material, and position of the virtual bed that are concerned in the spoken language expression, and a label of the surface feature dimension of the head of the bed and the side plate of the bed that are ignored in the spoken language expression.

[0081] Next, the spatial position of the scene object is introduced. The spatial position includes at least one of the following of the scene object: a coordinate position (Location), such as coordinate information of a center point or a preset point of the scene object in the virtual scene; a rotation (Rotation), such as a direction in which a front of the scene object faces in the virtual scene; a bounding box (Bounding Box) for indicating a size of the scene object in the virtual scene; cover points for indicating recommended virtual character position points when the virtual character approaches the scene object, so that the scene object can mask the virtual character. The bounding box (UE4 Bounding Box) is a rectangular box used to represent the boundary of a three-dimensional object in the Unreal Engine 4 (UE4) game engine, and is usually used for fast collision detection and spatial query operations.

[0082] Next, the pre-acquisition process of the appearance image of the scene object is introduced. The appearance image is acquired in the process of predicting the visual label of the scene object, and the appearance image includes images obtained by observing the scene object from at least two perspectives.

[0083] Optionally, the spatial data of each scene object in the entire virtual environment is pre-acquired and arranged into a data set for use in the subsequent query process.

[0084] Stage two: speech recognition and intent recognition;

[0085] In the process of controlling the master virtual character by the user, the terminal receives a natural language command input by the user in speech; the server performs speech recognition and intent recognition on the natural language command to analyze the input instruction of the user;

[0086] The speech recognition is text conversion performed on the natural language command input by the user in speech to obtain instruction text corresponding to the natural language command, and converts speech information into text information. A coding network is called to perform feature coding on the instruction text to obtain feature representation of the natural language command in a hidden layer space; a text segmentation network is called to perform segmentation on the feature representation to obtain multiple clauses in the instruction text, and then intent recognition is performed on each of the multiple clauses. In the case that the natural language command input by the user in speech is a long sentence, the segmentation of the instruction text corresponding to the natural language command can be realized, and the demand for analyzing the input instruction of the user in a long sentence scenario can be met.

[0087] The intent recognition of each clause is introduced as follows: the input instruction of the user is obtained through intent recognition, and the input instruction of the user includes at least one of the following: an entity, a semantic type, a subject type, and an intent type.

[0088] In an aspect, a Conditional Random Field (CRF) is used to recognize an entity in a clause, such as recognizing an entity in a clause as a building, a virtual item, virtual vegetation, or other scene object in a virtual scene. The entity in the clause is used to indicate a scene object that a virtual activity is directed to, such as the entity in the clause being a virtual building and the clause carrying a semantic indicating that an NPC initiates a virtual attack, then the virtual attack is performed on the virtual building, i.e., the virtual attack is initiated on the virtual building. On the other hand, a prediction network is called to predict a semantic type, a subject type, and an intent type in the clause, respectively. The semantic type includes whether or not the semantic indicates that the NPC initiates the virtual attack; the subject type is used to indicate that the subject of the clause is one or more NPCs or a user-controlled master virtual character; and the intent type is a desired virtual activity indicated by the clause, such as moving or initiating a virtual attack.

[0089] Stage three: query of spatial data;

[0090] For an entity in an input instruction, a spatial position of the entity in a virtual environment needs to be determined. In a case where multiple candidate entities in the virtual environment match the entity in the input instruction, the target entity matching the input instruction can be uniquely determined from the multiple candidate entities in combination with environmental perception information of a master virtual character and / or an NPC.

[0091] For example, the environmental perception information is the field of view information of the master virtual character, the spatial data in the virtual environment is a query range, and the real-time position and orientation of the master virtual character are query reference conditions. The spatial position of the target entity corresponding to the input instruction in the virtual environment is queried. This is so that the NPC can be subsequently controlled to move to the vicinity of the target entity or perform a behavior indicated by the input instruction on the target entity.

[0092] However, there can be multiple environmental perception information of the master virtual character and / or the NPC. The multiple environmental perception information can be fused to assist the determination of the target entity, or the multiple environmental perception information can be used according to a priority to assist the determination of the target entity, without limitation.

[0093] Stage four: voice feedback of the NPC;

[0094] The NPC will give voice feedback to the user. The types of voice feedback include at least one of immediate feedback, instruction execution feedback, and dynamic feedback. Among them, the immediate feedback refers to that the NPC gives voice feedback to the natural language command immediately after the user issues the natural language command to the NPC, such as "received" and the like; the instruction execution feedback refers to that the NPC gives voice feedback to the execution result of the control instruction after executing the control instruction indicated by the intention of the natural language command; and the dynamic feedback refers to that the NPC gives voice feedback according to real-time environmental perception when the user does not issue a natural language command to the NPC.

[0095] Among them, the voice feedback is obtained by converting the text content into speech. The text content is content inferred by a large language model in combination with the spatial data and dynamic environmental perception information. Using a large language model to dynamically generate the text content of the voice feedback can break through the monotony, mechanicalness, and easy repetition of fixed template voice feedback on the one hand; and on the other hand, it does not need to pre-save too many pre-produced template voices, which can avoid the problem of too large data volume of the client.

[0096] Stage five: behavior control of the NPC;

[0097] Taking the input instruction of the user as input and taking the spatial data of the virtual scene and the runtime information corresponding to the main control virtual role as reference, an AI model is used to determine a target entity in the virtual space and a control instruction for the NPC; the control instruction is an instruction that can be executed by a behavior tree of the NPC; the subsequent action indicated by the control instruction for the NPC to perform can be a single action or a sequence of actions. In the case where the control instruction is a sequence instruction of multiple actions, the sequence instruction of multiple actions is used to indicate the NPC to perform a sequence of actions, and the data structure for storing the sequence instruction of multiple actions is in the form of a list, and the to-be-executed instruction in the sequence instruction is buffered. When the behavior tree of the NPC executes the current instruction, the buffer is called back, and the to-be-executed instruction in the buffer is further issued for execution.

[0098] Hotword updating function:

[0099] For stage two, the present scheme further introduces a hotword system to improve the accuracy of voice recognition.

[0100] The hotword system is used to provide scene hotwords related to the scene involved in the natural language command in the voice recognition process, so that when a natural language command of voice input is received, the hotwords in the hotword library can be preferentially searched based on the currently running virtual scene as the instruction analysis result, so that the matching degree between the instruction analysis result and the virtual scene is higher.

[0101] The virtual scene includes at least one of a virtual battle scene, a virtual transaction scene, a virtual office scene, a virtual kitchen scene, and the like. Frequent words in the plurality of virtual scenes, specific virtual elements existing in the virtual scenes (such as virtual buildings and virtual characters existing only in a specific scene type), and the like are pre-analyzed as scene hot words to obtain a scene hot word library composed of a plurality of scene hot words, and the plurality of scene hot words correspond to the virtual scene. In addition, different scene language models are pre-trained for different virtual scenes, such as a scene language model in a battle scene paying more attention to battle-related words and a scene language model in a transaction scene paying more attention to transaction-related words. The scene language model can more specifically analyze the natural language command received in the virtual scene.

[0102] For example, a natural language command is received in the process of controlling a non-player character by the user. Based on the game state, the object position, and the current task completion condition, it is determined that the first virtual scene in which the target virtual character (at least one of the main virtual character and the non-player character) is currently located is a virtual battle scene, and a battle scene language model corresponding to the virtual battle scene is obtained. In addition, the first scene hot word in the virtual battle scene is obtained based on the field of view of the target virtual character, including "truck", "virtual grass", "enter", "defend", "desert", "attack", and the like. The natural language command is decoded and processed by the pre-trained natural language analysis model (including an acoustic network, a language network, and a preset dictionary) under the constraint of the first scene hot word, and an instruction analysis result is output to accurately control the non-player character. The instruction analysis result is illustrated as follows.

[0103] (1) The instruction analysis result is "No. 1, you defend in front of you". The instruction analysis result can be used to control No. 1 to move to the front to make a defensive posture, so as to avoid the enemy virtual character attacking the main virtual character first. If there is no constraint of the first scene hot word, the natural language command is easily identified as "No. 1, you make a sound in front of you", which affects the virtual combat plan of the player.

[0104] (2) The instruction analysis result is "No. 2, you explore the desert". The instruction analysis result can be used to control No. 2 to move to the desert to explore whether there is an enemy virtual character or a virtual element such as a virtual treasure chest. If there is no constraint of the first scene hot word, the natural language command is easily identified as "No. 2, you explore the mountain", so that No. 2 moves to the wrong position, and the human-computer interaction efficiency is low.

[0105] (3) The instruction analysis result is "No. 1 and No. 2, attack me", and the instruction analysis result can be used to control No. 1 and No. 2 to attack the virtual characters in the vicinity. Without the constraint of the first scene hotword, the natural language command is easily identified as "No. 1 and No. 2, supply me", so that No. 1 and No. 2 make an error behavior that does not meet the player's expectation, which cannot protect the main control virtual character and easily makes the enemy virtual character win the virtual game.

[0106] (4) The instruction analysis result is "No. 1, find an oak tree in the vicinity", and the instruction analysis result can be used to control No. 1 to search for an oak tree in the vicinity, so as to complete the game task or find an oak tree that can avoid attack. Without the constraint of the first scene hotword, the natural language command is easily identified as "No. 1, find a project book in the vicinity", so that No. 1 searches for a virtual element that does not meet the player's expectation in the vicinity, which cannot meet the player's virtual combat demand.

[0107] (5) The instruction analysis result is "No. 1 and No. 2, go to the stable over there", and the instruction analysis result can be used to control No. 1 and No. 2 to search for a stable in the vicinity and move to the position where the stable is located. Without the constraint of the first scene hotword, the natural language command is easily identified as "No. 1 and No. 2, go over there, it will be done soon", so that No. 1 and No. 2 mistakenly think that the main control virtual character currently wants to complete the virtual game by itself, and thus cannot provide better game assistance for the main control virtual character.

[0108] (6) The instruction analysis result is "No. 2, pick up the prop in front", and the instruction analysis result can be used to control No. 2 to search for a prop in front and pick up the prop. Without the constraint of the first scene hotword, the natural language command is easily identified as "No. 2, pick up the crab in front", so that No. 2 cannot accurately find the virtual element that the main control virtual character wants to pick up in front, which may prompt the player "cannot find the crab", or may pick up the wrong virtual element "crab" for the player, which interferes with the player's virtual combat process.

[0109] Environmental sound effect function:

[0110] In addition, for the environmental sound effect of the entire virtual scene, the scheme also provides a spatial audio enhancement scheme.

[0111] The audio played by the terminal includes environmental audio and NPC audio. The environmental audio is audio generated based on the scene characteristics of the virtual scene currently occupied by the main control virtual object, so as to give the user a sense of being there; the NPC audio is audio generated based on the characteristics of the NPC, so that the user can intuitively feel the emotion and body state of the role through the heard sound.

[0112] Ambient audio: identify scene elements in the virtual scene where the master virtual object is located, generate / choose appropriate element sound effects from a sound effect library in real time according to the scene elements, and synthesize the element sound effects to obtain the ambient audio. For example, the virtual scene is a forest at night, and the scene elements include trees, owls, insects, etc. The ambient audio includes rustling of leaves, hooting of owls, chirping of insects, etc.

[0113] NPC audio: determine the NPCs contained in the virtual scene, and generate corresponding character voices based on the personality type, current emotion, and current behavior of the NPCs. For example, when the NPC is a middle-aged man running, the generated character voice is a deep male voice with a breathing sound effect when running.

[0114] The terminal performs audio enhancement processing on the ambient audio and / or NPC audio to improve the realism of the audio when playing the ambient audio and / or NPC audio to the user.

[0115] In combination with the above description of the implementation environment, the method for broadcasting the non-player character provided by the embodiments of the present application is described, and the execution subject of the method is taken as the terminal shown in FIG. 1 for example.

[0116] FIG. 2 shows a flowchart of the method for broadcasting the non-player character provided by an example embodiment of the present application. The method can be executed by the terminal in FIG. 1 described above. The method includes:

[0117] Step 210: display at least one of the master virtual character and the non-player character in the three-dimensional virtual environment.

[0118] Optionally, the non-player character follows the master virtual character to move in the three-dimensional virtual environment, or the non-player character freely moves in a range indicated by the user account in the three-dimensional virtual environment, which is not limited herein.

[0119] For example, the method can be executed by a client installed on the terminal, the client is a client of an application program, and the application program is an application program supporting the three-dimensional virtual environment. For example, the application program can be an application program of a shooting game.

[0120] Optionally, the terminal displays a game session interface, and the game session interface includes a three-dimensional virtual environment picture, and the three-dimensional virtual environment picture displays the master virtual character and the non-player character. The three-dimensional virtual environment picture is a picture obtained by observing the three-dimensional virtual environment from the perspective of the master virtual character.

[0121] The master virtual role is a virtual role directly controlled by the terminal. The terminal can directly control the behavior of the master virtual role according to the received control instruction according to the preset control logic. For example, the master virtual role moves forward when the forward instruction is received; the master virtual role jumps up when the jump instruction is received.

[0122] For example, the terminal displays the control control (for example, a movement control, a shooting control, etc.) corresponding to the master virtual role, and the user controls the master virtual role to move in the three-dimensional virtual environment by triggering the control control. Alternatively, the terminal is connected with an input device (for example, a mouse, a keyboard, a sensor, etc.), and the user controls the master virtual role to move in the three-dimensional virtual environment by issuing a control instruction through the input device.

[0123] The non-player role is a virtual role controlled by the game application program, controlled by the server, or controlled by an artificial intelligence program in the client. In the embodiment of the application, the game application program can command the behavior of the non-player role according to the received natural language command. The game application program can understand the intention of the natural language command, generate a control instruction according to the intention, and command the non-player role to execute the intention indicated by the natural language command.

[0124] It should be noted that the non-player role has self-behavior ability, and the user can only command the non-player role and cannot directly control the non-player role. The non-player role understands and executes the natural language command based on its own self-behavior ability.

[0125] In an optional embodiment, the master virtual role is a virtual role controlled by the control instruction generated by triggering the UI control; and the non-player role is a virtual role controlled by the AI according to the natural language command. When the natural language command is not received, the AI can control the non-player role to move in the three-dimensional virtual environment; and when the natural language command is received, the AI can control the non-player role to move in the three-dimensional virtual environment according to the intention of the natural language command.

[0126] In some embodiments, the master virtual role and the non-player role belong to the same camp. For example, the non-player role is a teammate of the master virtual role, or the non-player role is a pet of the master virtual role, or the non-player role is a valet of the master virtual role, or the non-player role is an intelligent robot controlled by the master virtual role, or the non-player role is a duplicate of the master virtual role.

[0127] For example, in a team mode game session, a player needs to team up with at least one other player to play the game. When the player lacks a team partner, the game application can provide a non-player character as a team partner of the player to team up with the player to play the game session. In the game session, the player can command the non-player character to perform activities by issuing a voice command (natural language command), and make the non-player character broadcast feedback information according to environmental perception information.

[0128] For another example, in a game session, a player can carry a pet to explore a three-dimensional virtual environment through the pet. The player can command the pet to move in the three-dimensional virtual environment through a voice command or a text command to explore unknown areas in the three-dimensional virtual environment, and make the pet broadcast feedback information in real time according to environmental perception information during the exploration, so as to facilitate the master virtual character to make strategic decisions and deployment.

[0129] For another example, in a game session of a MOBA game, a master virtual character teams up with four virtual characters controlled by other players to play a game. When one of the team partners is offline and cannot control the virtual character, for example, a player of a first virtual character is offline and cannot control the first virtual character, the player can command the first virtual character during the offline period of the team partner through a voice command, so that the first virtual character assists other team partners to participate in team battles, or the first virtual character takes accurate actions to respond to changes in the game session; the first virtual character can broadcast feedback information according to environmental perception information, so as to facilitate the player to make battle decisions.

[0130] In some embodiments, a user controls a master virtual character to perform activities in a three-dimensional virtual environment through a client. Illustratively, the client receives a control operation on the master virtual character, and controls the master virtual character to perform activities in the three-dimensional virtual environment according to the control operation.

[0131] Step 220, receiving a natural language command for the non-player character.

[0132] In the embodiments of the present application, the client receives a natural language command for the non-player character from the user, and commands the non-player character to perform activities in the three-dimensional virtual environment according to an intention indicated by the natural language command.

[0133] In the embodiments of the present application, the natural language command is a command input by the user through voice received by the client. In some embodiments, a voice receiving control is displayed in the client, and in response to the voice receiving control receiving a trigger operation, the client starts to record voice audio of the user, and processes the recorded voice audio as the natural language command.

[0134] Step 230, broadcasting feedback information of the non-player character in response to the natural language command.

[0135] Illustratively, the feedback information corresponds to feedback text generated based on environmental perception information, and the environmental perception information includes information perceived by at least one of the main virtual character and the non-player character in the three-dimensional virtual environment.

[0136] Illustratively, after receiving the natural language command, the game application can generate corresponding feedback information in real time according to the environmental perception information of the non-player character, and broadcast the feedback information. For example, the game application can generate corresponding feedback information in real time according to the environmental perception information triggering the broadcast condition when the environmental perception information of the non-player character triggers the broadcast condition; or the game application can generate corresponding feedback information in real time according to the environmental perception information of the non-player character when the non-player character or the main virtual character triggers the broadcast condition.

[0137] Illustratively, the feedback text is obtained by calling a pre-trained large language model to infer based on static entity data of the three-dimensional virtual environment and dynamic environmental perception information of the non-player character. For different perception conditions of the non-player character, the large language model generates different feedback texts.

[0138] Illustratively, the feedback text can also be obtained by the large language model inferring based on static entity data of the three-dimensional virtual environment and dynamic environmental perception information of the non-player character, and the feedback text conforms to the personality characteristics of the non-player character. For example, when the non-player character is a robust uncle image, the feedback text can adopt a bold tone; when the non-player character is a reporter, the feedback text can adopt a news report tone.

[0139] It should be noted that due to the randomness of the generation result of the large language model, in the same scene, when the environmental perception information of the non-player character is the same, the generated feedback text can also be different.

[0140] Illustratively, the environmental perception information includes at least one of the following: visual information perceived within the visual range of the main virtual character; auditory information perceived within the auditory range of the main virtual character; information perceived by a perception skill or a perception device possessed by the main virtual character; visual information perceived within the visual range of the non-player character; auditory information perceived within the auditory range of the non-player character; information perceived by a perception skill or a perception device possessed by the non-player character.

[0141] For example, the environment perception information of the host virtual role can include: a picture or entity that the host virtual role can see when observing the three-dimensional virtual environment, an environmental sound that the host virtual role can hear, a source direction of the environmental sound, a type and a sound size of the environmental sound, perception information obtained by the host virtual role through use of a perception skill (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and perception information obtained by the host virtual role through use of a perception prop (for example, a sensing signal of a sensor, a position detection signal of a position detection prop, etc.).

[0142] The environment perception information of the non-player role can include: a picture or entity that the non-player role can see when observing the three-dimensional virtual environment, an environmental sound that the non-player role can hear, a source direction of the environmental sound, a type and a sound size of the environmental sound, perception information obtained by the non-player role through use of a perception skill, and perception information obtained by the non-player role through use of a perception prop.

[0143] It should be noted that the environment perception information can include information that is perceived by the non-player role in real time, or information that is perceived by the non-player role historically.

[0144] For example, the feedback information is inferred by the game application according to the environment perception information of the non-player role and static entity data of the three-dimensional virtual environment. The feedback information can be in the form of a voice broadcast or in the form of a text broadcast. The static entity data includes data of entities that are relatively constant in the three-dimensional virtual environment, for example, includes model information of various buildings, terrains, and vehicles in the three-dimensional virtual environment.

[0145] Optionally, the feedback information can include at least one of: immediate feedback and execution feedback. Both the immediate feedback and the execution feedback are feedback information generated for the natural language command.

[0146] 1. The immediate feedback is feedback information generated immediately upon receiving the natural language command. The immediate feedback can include a reply broadcast and an immediate broadcast. The reply broadcast is used to reply to a query in the natural language command; and the immediate broadcast is used to feed back a reception condition of the natural language command.

[0147] For example, the client can instruct the non-player role to move in the three-dimensional virtual environment according to an intent of the natural language command. Then, the client can generate immediate feedback corresponding to the natural language command upon receiving the natural language command, and the immediate feedback is used to indicate that the natural language command has been received, or the immediate feedback is used to immediately reply to the natural language command.

[0148] 2. The execution feedback is feedback generated when a task is performed according to an intent indicated by the natural language command after the natural language command is received. The execution feedback can also be referred to as an execution broadcast.

[0149] For example, when the natural language command is used to instruct the non-player character to perform a target task, the client can generate execution feedback during the performance of the target task by the non-player character, the execution feedback being used to indicate the execution status of the target task; or, the client can generate execution feedback after the non-player character completes the target task, the execution feedback being used to indicate the execution result of the target task.

[0150] In some other embodiments, the feedback information can further include dynamic feedback, the dynamic feedback being the feedback information generated spontaneously by the non-player character.

[0151] Illustratively, the dynamic feedback is the feedback information generated spontaneously according to the environmental perception information of the non-player character. The dynamic feedback can also be referred to as dynamic broadcasting.

[0152] For example, the client can spontaneously generate dynamic feedback based on the perception of the three-dimensional virtual environment without receiving the natural language command, the dynamic feedback being used to indicate an abnormal situation discovered by the non-player character in the three-dimensional virtual environment. For example, the abnormal situation can include: discovering an enemy virtual character, discovering a change in the state of a friendly virtual character, discovering a dangerous situation, discovering a battle trace or a looting trace, etc.

[0153] In an alternative embodiment, the feedback information is implemented as spatialized human voice audio of the non-player character, i.e., the feedback information played by the client is spatialized human voice audio processed based on the generated human voice audio. Illustratively, the human voice audio corresponding to the feedback text is obtained, the sound effect parameters corresponding to the non-player character are obtained based on the relative orientation relationship between the non-player character and the master virtual character in the three-dimensional virtual environment, the human voice audio is adjusted based on the sound effect parameters, and the spatialized human voice audio corresponding to the non-player character is generated and broadcasted. The spatialized human voice audio is used to represent the perception effect of the audio generated by the master virtual character on the non-player character under the relative orientation relationship.

[0154] Illustratively, the relative orientation relationship between the non-player character and the master virtual character in the three-dimensional virtual environment is used to indicate the relative relationship between a first position corresponding to the non-player character and a second position corresponding to the master virtual character in the three-dimensional virtual environment.

[0155] Optionally, the relative positional relationship comprises at least one of a distance, a direction, and a surrounding condition between the non-player character and the host virtual character, wherein the distance is used to indicate a distance relationship between the non-player character and the host virtual character, the direction is used to indicate an angle relationship between a first position of the non-player character and an orientation direction of the host virtual character, and the surrounding condition is used to indicate a surrounding condition of a space element forming a virtual space in the three-dimensional virtual environment, wherein the virtual space is used to indicate a propagation influence of the audio when the non-player character emits the audio.

[0156] In some embodiments, the determination of the relative positional relationship between the non-player character and the host virtual character can be implemented by: obtaining first position information of the non-player character in the three-dimensional virtual environment, obtaining second position information of the host virtual character in the three-dimensional virtual environment, and determining the relative positional relationship between the non-player character and the host virtual character based on the first position information and the second position information.

[0157] In some embodiments, the first position information of the non-player character and the second position information of the host virtual character are position information determined based on a same preset coordinate system. Optionally, the preset coordinate system can be a world coordinate system corresponding to the three-dimensional virtual environment, or the preset coordinate system can be a coordinate system established with the host virtual character as an origin.

[0158] In some embodiments, when the audio effect parameter corresponding to the non-player character is pre-generated and stored in a database, the audio effect parameter corresponding to the non-player character is obtained from the database. In other embodiments, the audio effect parameter corresponding to the non-player character can also be generated in real time, i.e., the audio effect parameter corresponding to the non-player character is generated based on the relative positional relationship.

[0159] Optionally, the audio effect parameter corresponding to the non-player character can comprise at least one of the following types:

[0160] Firstly, a distance type; illustratively, the audio effect parameter of the distance type is used to adjust a sound effect of the human voice audio based on a distance between the non-player character and the host virtual character. In some embodiments, the audio effect parameter of the distance type can adjust an audio volume of the human voice audio to simulate an influence of the distance on the sound size; and / or, the audio effect parameter of the distance type can adjust an audio time delay of the human voice audio to simulate an influence of the distance on a sound propagation time.

[0161] The second type is a direction type. Illustratively, the sound effect parameter of the direction type is used to indicate adjustment of the sound effect of the human audio based on the direction of the non-player character relative to the host virtual character. In some embodiments, the sound effect parameter of the direction type can adjust the audio volume of the human audio to simulate whether the host virtual character is facing the scene audio element; and / or, the sound effect parameter of the direction type can adjust the volume parameters of the human audio in different channels to simulate the direction of the sound source element in the scene emitting audio.

[0162] The third type is a spatial effect type. Illustratively, the sound effect parameter of the spatial effect type is used to indicate adjustment of the sound effect in the virtual space formed based on the spatial element, i.e., the sound effect performance of the non-player character emitting audio is affected by the virtual space through the sound effect parameter of the spatial effect type. In some embodiments, the sound effect parameter of the spatial effect type can increase the reverberation sound effect and the reverberation sound effect to simulate the effect of the sound in the space.

[0163] Optionally, the sound effect parameter of the spatial effect type includes at least one of a reverberation time parameter, a pre-delay parameter, a wetness parameter, a room width parameter, and a distance effect parameter. The reverberation time parameter is used to indicate the speed of audio decay in the virtual space; the pre-delay parameter is used to indicate the time difference between the direct arrival of audio to the host virtual character and the first reflection arrival to the host virtual character; the wetness parameter is used to indicate the proportion between directly propagated audio and reflected propagated audio; the room width parameter is used to indicate the degree of diffusion of audio on the horizontal plane in the virtual space; and the distance effect parameter is used to indicate the attenuation of audio with the propagation distance during propagation.

[0164] The fourth type is a custom type. Illustratively, the sound effect parameter of the custom type is a parameter for sound effect adjustment customized by the user. Optionally, the user can customize the overall volume of the scene audio, whether to add background music, customize the volume of different types of audio, etc.

[0165] In some embodiments, the client provides a custom interface for the sound effect parameter to the user, and the user configures the custom sound effect parameter through the custom interface.

[0166] In some embodiments, the server is configured with a parameter prediction model for personalized learning of the custom sound effect parameter for the user. Illustratively, the custom sound effect parameters of a plurality of candidate accounts are obtained, the plurality of custom sound effect parameters are input into the parameter prediction model to be trained, the parameter prediction model to be trained is iteratively trained to obtain the parameter prediction model, and the sound effect parameters generated by the system (e.g., the sound effect parameters of the distance type, the direction type, and the spatial effect type described above) are optimized through the parameter prediction model.

[0167] In the embodiments of the present application, the voice audio of the non-player character is adjusted according to the sound effect parameter corresponding to the non-player character, so as to generate the spatial voice audio corresponding to the non-player character. That is, the spatial voice audio is the audio data corresponding to the non-player character obtained by adjusting the voice audio through the sound effect parameter. In some embodiments, the spatial voice audio includes at least one of the audio adjusted by the sound effect parameter, audio time length information, a playback start timestamp of the audio, a playback end timestamp of the audio, playback condition information of the audio, and element identification information corresponding to the audio.

[0168] Optionally, when the sound effect parameter indicates that the volume size of the voice audio of the non-player character is adjusted, the volume size of the voice audio of the non-player character on at least one sound channel is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates that the playback delay of the voice audio of the non-player character is adjusted, the audio playback start time corresponding to the voice audio of the non-player character is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates that the sound speed size of the voice audio of the non-player character is adjusted, the playback speed of the voice audio of the non-player character is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates the tone of the voice audio of the non-player character, the frequency in the frequency spectrum of the voice audio of the non-player character is adjusted based on the sound effect parameter. When the sound effect parameter indicates the timbre of the voice audio of the non-player character, the filter corresponding to the non-player character is determined based on the sound effect parameter, and the timbre corresponding to the voice audio is adjusted through the filter.

[0169] In summary, when the user feeds back the relevant information of the three-dimensional virtual environment to the non-player character in the three-dimensional virtual environment through the natural language command, the feedback information corresponding to the non-player character is generated and broadcasted according to the information perceived by the non-player character and / or the host virtual character from the three-dimensional virtual environment, so that the feedback information can cover multiple scenes. For example, when the natural language command indicates that the non-player character asks whether there is an enemy in the three-dimensional virtual environment, when the non-player character perceives that the virtual character included in the three-dimensional virtual environment includes an enemy virtual character, the host virtual character is reminded of the existence of the enemy through the broadcast of the feedback information; when the natural language command instructs the non-player character to provide a virtual prop to the host virtual character, when the non-player character has a virtual prop required by the user, the non-player character submits the virtual prop to the host virtual character, and reminds the host virtual character that the virtual prop has been submitted through the broadcast of the feedback information. When the non-player character is controlled through the natural language command, the feedback information broadcast is performed according to the environmental perception information of the non-player character and / or the host virtual character, which can enrich the content that can be fed back by the feedback information, so that the broadcast is not limited to fixed triggering scenes. The game application program can flexibly broadcast the feedback information according to the environmental perception information in the game match, enrich the feedback information broadcast corresponding to the non-player character, and improve the human-computer interaction efficiency between the non-player character and the player.

[0170] And, since the user can instruct the non-player character to feedback information perceived from the three-dimensional virtual scene through a natural language command, the user can quickly acquire information in the three-dimensional virtual scene, and thus the user can perform a suitable operation according to the known three-dimensional virtual scene information, speed up the game session process, and reduce the game session duration, thereby reducing the cost of computing resources required to implement the game session and reducing the power consumption of the device.

[0171] The method provided by the embodiment can improve the immersion of the player during the game, increase the fluency of the interaction process between the player and the NPC, and make the feedback of the NPC tend to be the experience of the real player. In the session, the NPC can provide the player with immediate and effective voice feedback in the face of different environments, different battle situations, and different self-states, help the player acquire key information in real time, thereby speeding up the game session process and improving the completion efficiency of the virtual session. The player can also quickly respond to the instructions issued by the player, so that the player understands the current execution progress of the instructions, and at the same time, the surrounding environment of the instruction execution has new changes, so that the NPC is no longer in the interpretation mode of silently executing the behavior tree logic and playing fixed audio. The player is provided with a more human-like interaction experience with the NPC teammate, and the design of derivative gameplay is facilitated.

[0172] In an optional embodiment, as shown in FIG. 3, the feedback information corresponding to the non-player character includes at least one of dynamic feedback 301, immediate feedback 302, and execution feedback 303.

[0173] 1. The dynamic feedback 301 includes at least one of the following: fire scene feedback, environment perception feedback.

[0174] The fire scene feedback includes at least one of the following: attack feedback when the non-player character or the host virtual character is attacked; escape feedback when the enemy virtual character escapes; danger warning feedback when the non-player character discovers danger; kill feedback when the non-player character kills the enemy virtual character; death feedback when the friendly virtual character dies; bullet change feedback when the non-player character changes bullets; and thrown object attack feedback when the non-player character is attacked by a thrown object. The environment perception feedback includes at least one of the following: battle trace feedback when the non-player character discovers a battle trace in the three-dimensional virtual environment; and scavenging trace feedback when the non-player character discovers a scavenging trace in the three-dimensional virtual environment.

[0175] The fire exchange scene is an area in the game where the player fights with other players or NPCs, usually with rich terrain, cover, and tactical elements to increase the strategic and interesting nature of the battle. The throwing refers to the tactical behavior of the player using grenades, flash bombs, and other throwing items to attack or control the enemy in battle. The trace scavenging refers to the process in the game where the player can obtain information about the enemy's whereabouts, battle situation, etc. through trace scavenging, in order to better develop tactics and respond to the enemy. The battle trace is the trace left by the battle behavior in the game, such as bullet holes, damaged objects, etc., which can help the player judge the enemy's movement track and battle situation, mainly manifested as the reward box after defeating the enemy. For example, when the player does not issue an order to the NPC, the NPC automatically provides feedback to the player through its own perception of the game scene, providing real-time information about the battle situation and improving the gaming experience.

[0176] For fire exchange scene feedback, it can be divided into three types according to the reference for the enemy, the NPC itself, and the teammates. The fire exchange scene refers to a situation where there may be hostile virtual characters or potential dangers nearby.

[0177] Reference for the enemy: When the non-player character perceives the hostile virtual character, it will perform a spatial query based on the location of the hostile virtual character, obtain the location description of the location where the hostile virtual character is located, and then perform template splicing to obtain fire exchange scene feedback such as "XXX is attacked" and "XXX finds enemy". When the enemy is killed or escapes, the same logic of spatial query is performed, and fire exchange scene feedback such as "killed", "enemy down", and "enemy escape" is fed back according to the probability. Or splice the location description obtained by spatial query to feed back fire exchange scene feedback such as "XXX killed the enemy", "XXX one down", and "XXX enemy disappeared".

[0178] Reference for the NPC itself: When the non-player character is changing ammunition, it will feedback "changing ammunition"; when the non-player character is attacked by a throwing object, it will feedback "throwing object attack" (hand grenade, flash bomb, etc.); when the non-player character perceives a dangerous situation, it will feedback "nearby danger".

[0179] Reference for teammates: When the non-player character perceives the death of a friendly virtual character, it will feedback the broadcast of the death of the friendly virtual character, for example, "No. 3 has been killed". The purpose of this broadcast is to clearly inform the player which NPCs can currently receive natural language command instructions, avoiding the player issuing natural language commands to NPCs that have been killed.

[0180] For environmental awareness feedback, environmental awareness feedback covers the automatic feedback logic of NPC outside the fire-fighting scene, including the feedback of two types of intelligence, battle traces and scavenging traces. These two types of broadcasts can provide players with information about the historical situation of the three-dimensional virtual environment that has been released. For example, a certain area has been involved in a battle and someone has been defeated, or someone else has arrived in the area and has scavenged, at which point the non-player character needs to remind the player to pay attention to the surroundings to avoid being attacked through environmental awareness feedback.

[0181] Battle trace broadcast: a reward box will appear randomly in the three-dimensional virtual environment. The reward box is an item container in the game that provides additional rewards for players, usually containing valuable items, equipment or resources, which players can obtain by completing tasks, defeating enemies, etc. Since the reward box is dynamically generated after a kill in the game match, it cannot be found through static entity data in the three-dimensional virtual environment, so the game application will control the NPC to search for the reward box around it periodically, and after finding the reward box, it will determine whether it is killed by the team, whether it is due to the defeat of a teammate, and if neither of the above conditions is met and the non-player character first discovers the reward box, the non-player character needs to broadcast. For example, directly call the offline produced voice package to feedback "here discover battle trace". At the same time, the game application adds the reward box to the cache to avoid subsequent repeated broadcast.

[0182] Scavenging trace broadcast: Since scavenging traces are also generated by other virtual characters in the game match, they also need to be acquired, so the game application will also periodically search for scavengeable items around the NPC that are in the "opened" state, determine whether the past interaction record of the scavengeable item contains the host virtual character, and if not and the non-player character first discovers the scavengeable item, it can directly call the offline produced voice package to feedback "here discover scavenging trace". At the same time, add the scavenging item to the cache to avoid subsequent repeated broadcast.

[0183] 2. Instant feedback 302 includes at least one of the following: instant feedback of command reception, reply broadcast when responding to inquiry natural language commands.

[0184] The natural language command query can include at least one of battlefield information query, target position query, and casual conversation query. According to different query questions, corresponding reply broadcasts can be generated. For example, the instant feedback of receiving the natural language command can be the instant feedback of the NPC to the natural language command issued by the player, which can definitely make the player perceive that the natural language command just now has been successfully transmitted to the NPC. Instant feedback is required regardless of whether the natural language command is executable or not. For example, the instant feedback to fine-tuning commands and movement commands can be “received”, “understood”, and “No. 3 received”; the instant feedback to purposeful movement such as viewing and attack commands can be “advancing”; the instant feedback to set protection type movement commands can be “on the way” and “returning”; the instant feedback to player casual conversation type meaningless commands can be “I don't understand” and “battlefield, please don't chat”. For example, the instant feedback includes instant feedback to player inquiries. The player's inquiry in the game can include at least one of the following: information inquiry, position inquiry, and casual conversation inquiry.

[0185] For information inquiry: the player can inquire the NPC about the current situation, whether a certain position is safe, etc. At this time, the game application needs to query the area according to the position of the NPC accepting the question, and then generate a reply feedback according to whether the enemy virtual character is perceived in the NPC perception system and broadcast it. For example, broadcast “there are enemies on the second floor of the stable, be careful to hide” and “I am safe here”.

[0186] For position inquiry: the player can inquire the NPC about the current position of the NPC, the enemy virtual character, the attack sound, etc. At this time, the space query logic is used to generate reply feedback text, and the online text-to-speech service is used to generate reply feedback voice for broadcasting.

[0187] For casual conversation inquiry: the player can chat with the NPC at will, such as “what is the current weather like” and “have you eaten”. At this time, the corresponding reply feedback text is generated by the reasoning service of the large language model. When encountering a question not covered by the reasoning service, a reply is selected from the preset rejection template, such as broadcasting “don't chat, please fight seriously” and “conversation is irrelevant to the game”. After the game application receives the reply feedback text, it calls the online text-to-speech service to generate reply feedback voice for broadcasting.

[0188] 3. The execution feedback 303 includes at least one of entity query feedback, virtual prop use feedback, movement feedback, door hinge feedback, and item interaction feedback.

[0189] When the intent of the natural language command is to instruct the non-player character to perform a task, the task can be a task related to a target entity, the game application needs to first query the target entity referred to in the natural language command from the three-dimensional virtual environment, and then instruct the non-player character to perform the task related to the target entity, at this time, an entity query feedback can be generated based on the entity query result to indicate the result of the entity query. If the intent of the task natural language command is to instruct the non-player character to use a virtual prop, a virtual prop use feedback can be generated according to the use of the virtual prop. If the intent of the task natural language command is to instruct the non-player character to move, a movement feedback can be generated according to the movement result, for example, the movement feedback can include successfully moving to a target entity or a target position, or the movement feedback can include a movement failure. If the intent of the task natural language command is to instruct the non-player character to interact with a door (for example, open or close the door), a door feedback can be generated according to the interaction result of the non-player character and the door. If the intent of the task natural language command is to instruct the non-player character to perform an item interaction, for example, instructing the non-player character to exchange a virtual prop with a master virtual character, an item interaction feedback can be generated according to the item interaction result.

[0190] For example, for entity search feedback, entity search distinguishes between "findable" and "unfindable" two cases. The large language model directly returns feedback text according to the search result, for example, "folder found" and "safe nearby cannot be found", and the game application can directly convert the feedback text into feedback broadcast through text-to-speech. For example, for projectile (virtual prop) feedback, the projectile includes a hand grenade, a smoke bomb, a firecracker, a stun bomb, and a decoy bomb, and the non-player character will broadcast "XXX ready" in the preparation stage and "XXX has been thrown out" after the throwing is completed. Since the feedback text related to the projectile here is small and limited, the game application can directly use the offline packaged audio resources for playing. For example, for movement feedback, the movement instruction is divided into "move to entity position" and "move to space position" according to the movement target point, and the movement feedback is divided into "move successfully" and "move failed" according to the movement execution result.

[0191] In the case of successful movement, the feedback text of reaching the entity position is obtained by splicing the large language model, for example, "has reached XXX position" and "has reached XXX nearby", where XXX is the target entity extracted from the large language model. When reaching the building space position, the feedback text is designed to feedback according to the battle information, if there is an enemy virtual character in the building space position, it will broadcast "XXX has enemies, pay attention to concealment", if there is no enemy virtual character, it will broadcast "XXX is safe", where XXX is the name of the area.

[0192] In the case of a movement failure, the NPC directly broadcasts "Unable to continue approaching the target position." The situation in which the non-player character cannot reach the target position occurs when the center point of a large target position is too far from the navigation grid of the NPC, causing the NPC to only move to the edge of the target position. It is also possible that the target position point marked by the player on the mini-map fails to be mapped to the navigation grid.

[0193] For example, for door interaction feedback, the execution feedback of the door interaction in the scene is divided into two cases: "There is an interactive door" and "There is no interactive door." "There is an interactive door" means that the entity around the NPC can find the entity of the "door" through entity searching, and the current state of the door is opposite to the natural language command of the player; for example, the player issues an open door command, and the current door is in a closed state, at which time the NPC feedback content is "Opening / closing the door." "There is no interactive door" means that the entity searching around the NPC fails, or the state of the query result is consistent with the player's expectation; for example, the player issues an open door command, and the door itself is already in an open state, at which time the NPC can feedback "There is no door nearby that needs to be opened or closed."

[0194] For example, for item interaction feedback, item interaction refers to the player's instruction to the NPC to deliver an item, such as "Give me a medical kit" and "I have no bullets, give me some bullets." When the NPC can deliver the corresponding item, it will directly feedback "Give you XXX" and "XXX is placed here." When the NPC cannot deliver due to business logic restrictions or the absence of the item, it will fill in the entity word directly to feedback "I have no XXX" and "I cannot provide XXX."

[0195] In summary, the feedback types are divided into three parts: immediate feedback, instruction execution feedback, and dynamic feedback. The text content of the immediate feedback and the instruction execution feedback will make corresponding feedback according to the instruction inferred by the language model in combination with the static and dynamic game environment data, covering "immediate response", "start execution", "successful execution", "unsuccessful execution", etc. to generate; the text content of the dynamic feedback is the necessary feedback of the NPC according to the real-time environmental perception when the player has no instruction. That is, the non-player character provides feedback information to the user through diversified feedback methods, so that the user can obtain diversified information in the three-dimensional virtual scene through the non-player character, complete the non-player character interaction realized by the user through natural language commands through immediate feedback and instruction execution feedback, realize automatic feedback of three-dimensional virtual scene information through dynamic feedback, ensure the full coverage of diversified three-dimensional virtual environment information for the user when participating in the virtual game, ensure the real-time and effectiveness of the information, improve the human-computer interaction efficiency, and promote the game process of the virtual game, thereby improving the game completion efficiency of the virtual game and reducing the power consumption of the device.

[0196] In an alternative embodiment, the feedback information comprises a reply announcement. The reply announcement comprises a reply to the query natural language command generated in response to the query natural language command raised by the player. For example, when the query natural language command comprises a query for the location of a target entity, the reply announcement corresponding thereto should comprise the query result of the location of the target entity.

[0197] FIG. 4 shows a flowchart of a method for announcing a non-player character according to an example embodiment of the present application. The method can be performed by the terminal in FIG. 1. Based on the embodiment shown in FIG. 2, step 230 of the method can comprise step 231.

[0198] Step 231, in the case where the behavior intention of the natural language command is a query, announcing a reply announcement of the non-player character; the reply announcement comprises reply content to the natural language command generated according to the environmental perception information of the non-player character to the three-dimensional virtual environment.

[0199] The natural language command is used to query information from the non-player character. The natural language command can be referred to as a query natural language command. The query natural language command is a natural language command with an intention to query. The query natural language command contains at least one question.

[0200] The natural language command is a command indication conveyed using human natural language. The terminal receives the natural language command issued by the player and instructs the non-player character according to the intention corresponding to the natural language command. For example, the terminal receives the voice audio of the player, converts the voice audio into text, and obtains the natural language command. Alternatively, the terminal receives the text of the natural language command input by the player.

[0201] The player can use spoken language or written language to issue the natural language command. The game application program calls a pre-trained large language model to understand the natural language command, extracts the intention expressed in the natural language command, and generates a control instruction according to the intention to instruct the activities of the non-player character.

[0202] For example, the natural language command can be "come to me", and the game application program can parse the intention of the natural language command as "control the non-player character to move to the location of the master virtual character". The game application program obtains the location of the master virtual character, generates a navigation route for the non-player character to move to the location of the master virtual character, and controls the non-player character to move to the location of the master virtual character according to the navigation route. For another example, the natural language command can be "go and pick up the treasure chest", and the game application program can parse the intention of the natural language command as "control the non-player character to move to the location of the treasure chest and pick up the treasure chest". The game application program can obtain the location of the treasure chest, control the non-player character to move to the location of the treasure chest, and perform the operation of picking up the treasure chest after reaching the location of the treasure chest.

[0203] For example, in the case that the inquiry natural language command includes an intelligence inquiry to the target entity, a first reply broadcast is broadcast based on the perception of the non-player character to the target entity, and the first reply broadcast includes the intelligence perception result of the non-player character to the target entity. Or, in the case that the inquiry natural language command includes a location inquiry to the target entity, a second reply broadcast is broadcast based on the perception of the non-player character to the target entity, and the second reply broadcast includes a location description of the target entity.

[0204] For example, the pre-trained large language model is called to parse the received natural language command to obtain the intention of the natural language command. When the intention is an inquiry type, it can be determined that the natural language command is an inquiry natural language command. When the natural language command is an inquiry natural language command, the large language model can infer according to the intention of the natural language command, according to the static entity data and the environmental perception information, to obtain the reply feedback text. That is, the pre-trained large language model is called to parse the received natural language command, and the intention of the natural language command is output as an inquiry, and the reply feedback text of the natural language command. The game application program calls the text-to-speech service according to the processing logic of the inquiry intention, converts the reply feedback text into a reply broadcast, the reply broadcast is an audio form of broadcast, and sends the reply audio to the client for broadcast.

[0205] For example, when the large language model identifies that the intention of the natural language command is an inquiry, since there are thousands of questions that the user may ask, it is impossible to store the reply to each question locally on the client side. Therefore, the game application program calls the text-to-speech service to generate a reply broadcast in real time according to the reply feedback text returned by the large language model, so that the game application program can generate a corresponding reply broadcast in response to the player's question in real time.

[0206] For example, the game application program (client or server) calls the pre-trained large language model to parse the inquiry natural language command, infers the reply feedback text based on the static environment data and the environmental perception information of the three-dimensional virtual environment, generates human voice audio based on the reply feedback text, and obtains the reply broadcast.

[0207] For example, the client reports the received inquiry natural language command to the server, the server calls the pre-trained large language model to parse the inquiry natural language command, infers the reply feedback text based on the static environment data and the environmental perception information of the three-dimensional virtual environment, generates human voice audio based on the reply feedback text, and obtains the reply broadcast. The server sends the reply broadcast to the client, and the client broadcasts the reply broadcast.

[0208] As an example, the text-to-speech service is invoked to convert the reply text into human audio. The text-to-speech service obtains the reply playback by configuring parameters such as language, speech rate, and tone of the human audio.

[0209] Static environment data is used to describe relatively unchanging environmental features in a 3D virtual environment. Static environment data can include model data of stationary 3D models within the 3D virtual environment. Static environment data includes at least one of the following: the name of the 3D model, descriptive text, location information, orientation information, and spatial relationships (e.g., a cabin includes a kitchen and a bedroom). For example, static environment data may include entity information for all entities in the 3D virtual environment.

[0210] Dynamic reasoning can be performed based on static environmental data that remains unchanged, and dynamically changing environmental perception information. For example, if the environmental perception information includes the sound effect of flowing water from a river, then the river's position in the three-dimensional virtual environment can be determined from the static environmental data, based on the NPC's current location, the NPC's current orientation, the direction from which the sound effect is emitted, and the type of entity that can emit the sound effect (river or stream).

[0211] In summary, when a natural language command is received, the game application can combine the environmental awareness information of non-player characters to infer the corresponding response text and generate a response. By responding to the player's questions through this feedback, non-player characters can interact with the player in real time, reducing the response delay to player inquiries, improving information feedback efficiency, and providing the player with richer in-game information based on the non-player character's perception.

[0212] In one optional embodiment, the feedback information includes feedback broadcasts. These broadcasts include announcements of the task execution status and results generated in response to the player's natural language command for the task. For example, if the natural language command for the task includes moving to a target entity, the corresponding feedback broadcast might include: moving towards the target entity, or having reached the target entity.

[0213] Figure 5 shows a flowchart of a non-player character broadcasting method provided in an exemplary embodiment of this application. This method can be executed by the terminal shown in Figure 1 above. Based on the embodiment shown in Figure 2, step 230 of the method may include step 232.

[0214] Step 232: When the behavioral intent of the natural language command is to perform a task, broadcast a feedback broadcast; the feedback broadcast includes the execution status after executing the natural language command based on the non-player character's environmental perception information of the three-dimensional virtual environment.

[0215] The natural language command is used to instruct the non-player character to perform a task. The natural language command can be referred to as a task natural language command. The task natural language command includes a description text of the task, for example, the description text of the task can include at least one of the following: a task name, an action required to be performed by the non-player character, a method for performing the task, a task target, and a location of the task target.

[0216] For example, in a case where the intention of the task natural language command includes searching for a target entity, a first feedback broadcast is played, and the first feedback broadcast includes a search result of the target entity by the non-player character. Or, in a case where the intention of the task natural language command includes using a virtual prop, a second feedback broadcast is played, and the second feedback broadcast is used to indicate that the virtual prop is ready. Or, in a case where the intention of the task natural language command includes using a virtual prop, a third feedback broadcast is played, and the third feedback broadcast is used to indicate a use result of the virtual prop. Or, in a case where the intention of the task natural language command includes controlling the non-player character to move, a fourth feedback broadcast is played, and the fourth feedback broadcast is used to indicate a movement result of the movement. Or, in a case where the intention of the task natural language command includes interacting with a target entity, a fifth feedback broadcast is played, and the fifth feedback broadcast is used to indicate at least one of a search result of the target entity by the non-player character and an interaction result of the non-player character with the target entity. Or, in a case where the intention of the task natural language command includes an item interaction, a sixth feedback broadcast is played, and the sixth feedback broadcast is used to indicate an item interaction result of the non-player character performing the item interaction.

[0217] For example, a pre-trained large language model is called to parse the received natural language command to obtain an intention of the natural language command. When the intention is a task, it can be determined that the natural language command is a task natural language command. When the natural language command is a task natural language command, the large language model can infer a task instruction sequence according to the intention of the natural language command, the static entity data, and the environmental perception information, and the task instruction sequence is used to control the non-player character to perform the task. The game application program receives the task instruction sequence, controls the non-player character to perform the task according to the task instruction in the task instruction sequence, and controls the non-player character to perform the task according to the task instruction in the task instruction sequence. During the control of the non-player character according to the task instruction, the corresponding execution situation can be obtained according to the feedback broadcast logic corresponding to different task instructions to perform feedback broadcast.

[0218] For example, since the instructions that can be performed by the non-player character in the three-dimensional virtual environment are limited and traversable, the client can locally store feedback broadcasts corresponding to each character instruction. When the non-player character performs the corresponding instruction, the client can read the local feedback broadcast to perform voice broadcast.

[0219] Or, when there is variable content in the feedback report corresponding to a certain task instruction, for example, the target entity in the feedback report is variable, or the location information in the feedback report is variable; the client can also generate a report voice of variable content according to the target entity or location information returned by the large language model, and splice it with the report voice of the immutable content stored locally to obtain the final feedback report.

[0220] For example, when the task natural language command is "find the nearby treasure chest", calling the pre-trained large language model to analyze the task natural language command can obtain that its intention is a task, the task target is to find a target entity, and the target entity is a treasure chest; based on the static entity data and the environment perception information, it can be inferred that there is a treasure chest in the left front of the host virtual role. In the case where the target entity can be found, the feedback report corresponding to the task target is "find aaa in xxx", where xxx is the location of the target entity and aaa is the name of the target entity. The game application program can call the online text-to-speech technology to generate location voice according to the location information "left front" in the inference result; call the online text-to-speech technology to generate name voice according to the name "treasure chest" of the target entity, and then splice the location voice and the name voice into the feedback report to obtain the final feedback report.

[0221] For example, the game application program (client or server) calls the pre-trained large language model to analyze the task natural language command, infers the task execution instruction based on the static environment data and the environment perception information of the three-dimensional virtual environment, controls the non-player character to execute the task according to the task execution instruction, generates the state feedback text according to the execution state of the non-player character executing the task, and generates the human voice audio based on the state feedback text to obtain the feedback report.

[0222] For example, the client reports the task natural language command to the server; the server analyzes the task natural language command, infers the task execution instruction based on the static environment data and the environment perception information of the three-dimensional virtual environment, controls the non-player character to execute the task according to the task execution instruction, generates the state feedback text according to the execution state of the non-player character executing the task, generates the human voice audio based on the state feedback text to obtain the feedback report, and sends the feedback report to the client; the client receives and reports the feedback report.

[0223] In summary, when receiving a task natural language command, during the process of instructing the non-player character to perform the task, the execution feedback can be generated in real time according to the environmental perception information of the non-player character, and the state of the non-player character performing the task can be announced to the player. After the task is completed, the execution feedback can be generated according to the task completion result, and the result of the non-player character performing the task can be announced to the player. The player can grasp the state and progress of the non-player character performing the task in real time, so that the player can further indicate reasonable instructions to the non-player character, thereby improving the control efficiency of the non-player character, providing a smooth interactive experience for the player, and avoiding waste of computing resources for controlling the non-player character.

[0224] In an optional embodiment, the feedback information includes dynamic announcement. The dynamic announcement includes announcement of a perceived abnormal situation of the non-player character. For example, when the non-player character perceives that it is under attack, the announcement is that it is under attack; when the non-player character perceives that there is a dangerous situation in front, the announcement is that there is danger in front.

[0225] FIG. 6 shows a flowchart of a non-player character announcement method according to an example embodiment of the present application. The method can be performed by the terminal in FIG. 1. Based on the embodiment shown in FIG. 2, the method can further include step 240, which is a parallel step of steps 220-230.

[0226] Step 240: announcing a dynamic announcement when the environmental perception information of the non-player character in the three-dimensional virtual environment meets a dynamic announcement condition.

[0227] The dynamic announcement is generated by the game application based on the environmental perception information of the non-player character when it is determined that there is an abnormal situation that needs to be announced. The dynamic announcement is generated according to the abnormal situation. The dynamic announcement includes reminder information of the abnormal situation. For example, the abnormal situation can include at least one of the following: discovering an enemy virtual character, discovering a change in the state of a friendly virtual character, discovering new traces.

[0228] The dynamic announcement condition is used to trigger the dynamic announcement. The dynamic announcement condition is a trigger condition set based on the environmental perception information. The dynamic announcement condition can be set according to the abnormal situation. For example, the dynamic announcement condition can include at least one of the following: discovering an enemy virtual character according to the environmental perception information of the NPC (for example, the three-dimensional model of the enemy virtual character appears within the field of view of the NPC, the NPC receives the battle sound effect of the enemy virtual character, etc.), discovering a change in the state of a friendly virtual character according to the environmental perception information of the NPC (for example, the friendly virtual character dies within the field of view of the NPC, the NPC receives a notification message that the friendly virtual character has died, the NPC receives the hit sound effect of the friendly virtual character), discovering new traces according to the environmental perception information of the NPC (for example, new battle traces or looting traces appear within the field of view of the NPC).

[0229] For example, when a non-player character senses a hostile virtual character, a first dynamic announcement is broadcast; the first dynamic announcement includes the location of the hostile virtual character perceived by the non-player character. Alternatively, when a non-player character senses a dangerous situation, a second dynamic announcement is broadcast; the second dynamic announcement is used to alert the non-player character to the danger. Alternatively, when a non-player character senses a change in the status of a friendly virtual character, a third dynamic announcement is broadcast; the third dynamic announcement is used to alert the non-player character to the change in the friendly virtual character's status. Alternatively, when a non-player character discovers new traces in the 3D virtual environment, a fourth dynamic announcement is broadcast; the fourth dynamic announcement is used to alert the non-player character to the new traces. These new traces can be combat traces and / or looting traces.

[0230] For example, since the number of dynamic announcements that non-player characters can trigger in a 3D virtual environment is finite and traversable, the client can store dynamic announcement voices locally. When a corresponding dynamic announcement voice is triggered, the client can read the voice from the local storage and play it. Similarly, some dynamic announcement voices may contain variable content. The client can generate variable content in real time using text-to-speech technology and concatenate it with the dynamic announcement template to obtain the final dynamic announcement.

[0231] In summary, when a non-player character senses an anomaly, it can automatically generate dynamic feedback based on environmental perception information to alert the player to nearby anomalies. For example, when a non-player character detects an enemy virtual character, its location can be announced based on environmental perception information; when a non-player character detects danger, the player can be alerted to the danger; when a friendly virtual character's status changes, the player can be alerted to the status of their teammates, and so on. This allows the player to understand the non-player character's perception of the environment and make better combat decisions based on this perception, thereby improving the efficiency of game matches and reducing device power consumption during virtual matches.

[0232] In one optional embodiment, the feedback information includes a location description of the target entity. Indicatively, when generating the feedback information, it may be necessary to query the location of the target entity in the 3D virtual environment, or to determine the target entity indicated in the natural language command from the 3D virtual environment. In this case, an entity query method is needed to query the target entity.

[0233] Before executing a target entity query, static entity data for a 3D virtual environment needs to be pre-constructed. This static entity data construction serves both the reasoning of the large language model and the real-time feedback text generation on the game application side after the reasoning results are returned.

[0234] In an alternative embodiment, entities in a three-dimensional virtual environment are traversed, and entity information for each entity is exported from the game engine to obtain static entity data.

[0235] For example, for an in-game scene, an editor tool is developed to traverse the entire scene for StaticMesh, special Actors (such as interactive doors, which do not belong to StaticMesh), vegetation, and export the positions, orientations, and bounding box sizes of the entities as static entity data for subsequent tagging and inference by the large language model. For example, the vehicle 701, the box 702, the tree 703, and the interactive door of the building in FIG. 7 can all be collected as static entity data using this method.

[0236] Among them, StaticMesh is a type of static geometry resource in Unreal Engine 4 (UE4) game engine, used to represent unchangeable three-dimensional models such as buildings, props, etc. Actor is a basic object class in Unreal Engine 4 (UE4) game engine, representing an entity in the game world, such as characters, objects, light sources, etc. An Actor can contain multiple components to achieve different functions.

[0237] For example, spatial data can also be included in static entity data, which is obtained by manual labeling. Spatial data is more complex than static entities, as there are multiple floors and multiple layers of nesting within a floor in the space (such as a motel area containing a second floor, and the second floor area containing room 201). For such data, manual calibration is used to construct. Under the editor, add Volume to the scene, divide the required area of the entire map and label it accordingly. Volume is a special Actor class in Unreal Engine 4 (UE4) game engine, representing a three-dimensional area with specific functions, such as trigger areas, audio areas, etc. For example, the Volume area 704 in FIG. 8 is labeled as “Farm Map - Recycling Station Tower - 2nd Floor Area”.

[0238] After obtaining the static entity data, the entity or space can be queried based on the static entity data during game running.

[0239] For example, when the player issues a natural language command, the natural language command will be accompanied by the constructed static entity data and the real-time data captured at runtime (which includes the environmental perception information of non-player characters and / or the host virtual character) to request inference from the large language model. In addition, when the NPC needs to provide dynamic feedback based on runtime data, entity queries are also performed to generate feedback text based on the current state.

[0240] When the target entity is included in the natural language command, the game application program queries the target entity according to the current position and orientation of the host virtual character, the position and orientation of the teammates or enemies, and the target entity information of the target entity described in the natural language command, and filters the target entity matching the target entity information from at least one candidate entity within the perception range of the host virtual character and / or the non-player character based on the environmental perception information of the host virtual character and / or the non-player character.

[0241] Optionally, the natural language command is used to instruct the non-player character to perform an activity related to the target entity. For example, the target entity can be the moving destination of the non-player character, the target entity can be an object that needs to be observed by the non-player character, the target entity can be a target that needs to be attacked by the non-player character, or the target entity can be an object that needs to be interacted with by the non-player character.

[0242] The target entity is a virtual character existing in the three-dimensional virtual environment. The target entity is an entity described in the natural language command. The entity can refer to a virtual character that is fixed in the three-dimensional virtual environment, for example, the target entity can be a virtual building, a virtual terrain, a virtual vehicle, a virtual plant, a virtual prop, a virtual item, etc. The entity can also refer to a virtual character that can interact with a virtual character in the three-dimensional virtual environment, for example, the target entity can be an interaction point (e.g., a door, a window, a cabinet, a cellar, etc.), a virtual prop, a virtual character, a virtual light source, etc. in the three-dimensional virtual environment.

[0243] For example, the natural language command can be "please help me open the door", and the target entity can be "the door", and the non-player character needs to perform an activity related to the door "open the door". Or, the natural language command can be "is the kitchen safe", and the target entity can be "the kitchen", and the non-player character needs to perform an activity related to the kitchen "check whether there is an enemy virtual character or other dangerous situation in the kitchen".

[0244] Optionally, the natural language command includes target entity information of the target entity. The target entity information can be a description text of the target entity. For example, the target entity information can include at least one of the following: the name, type, location, and characteristics of the target entity. The game application program can find the target entity from a plurality of entities in the three-dimensional virtual environment based on the target entity information in the natural language.

[0245] The game application program analyzes the natural language command, extracts the target entity information therein, determines the target entity from a set of entities in the three-dimensional virtual environment based on the target entity information, and then controls the non-player character to perform an activity related to the target entity according to the intention of the natural language command.

[0246] The target entity is determined from the three-dimensional virtual environment in combination with environment perception information, and the environment perception information includes information perceived by at least one of the master virtual character or the non-player character from the three-dimensional virtual environment.

[0247] Since the target entity information in the natural language command is usually ambiguous, for example, the natural language command can be "move to the red truck", and the target entity information in the natural language command is "red truck", and there can be many red trucks in the three-dimensional virtual environment, and the target entity cannot be accurately determined from the three-dimensional virtual environment according to the target entity information in the natural language command.

[0248] Therefore, the embodiment of the present application provides a method for determining a target entity in combination with target entity information and environment perception information. For the above example, the "red truck" expressed by the player in the natural language command should be a red truck that the player can see, and therefore, in combination with the field of view of the master virtual character, the red truck located in the field of view of the master virtual character can be selected from the multiple red trucks in the three-dimensional virtual environment, and the red truck is the target entity indicated by the player in the natural language command.

[0249] Since the natural language command is issued by the player based on the perception of the three-dimensional virtual environment, in order to accurately identify the target entity indicated in the natural language command, the game application program will combine the environment perception information when the player issues the natural language command to identify the target entity in the natural language command. Based on the perception of the three-dimensional virtual environment by the player, the entity closest to the target entity information that the player can perceive in the three-dimensional virtual environment is inferred, which is the target entity.

[0250] For example, the three-dimensional virtual environment includes multiple candidate entities matching the natural language command, and the target entity is an entity selected from the multiple candidate entities based on the environment perception information of the master virtual character or the non-player character. For example, the game application program first selects multiple candidate entities matching the target entity information in the natural language command from the entity set, and then selects an entity that can be perceived by the master virtual character or the non-player character from the multiple candidate entities as the target entity according to the environment perception information.

[0251] For example, as shown in FIG. 9, the perception range of the master virtual character is determined according to the position and orientation 916 where the master virtual character is located. For example, the perception range is a fan-shaped area in front of the master virtual character with a certain field of view size; entities not within the fan-shaped range are filtered out. Then, the target entity is selected from the entities in the perception range according to the similarity between the target entity information and the entity information of the candidate entities. When there are two entities in the perception range with the highest similarity scores, for example, entity 2 and entity 3 both have a similarity score of 2, the entity closer to the master virtual character, entity 2, is selected as the target entity.

[0252] When it is necessary to generate feedback information according to the position description of the target entity, the game application can perform spatial query to generate the position description according to the positional relationship between the master virtual character and the target entity.

[0253] The main application scenario of spatial query is to obtain the positions of the master virtual character, non-player characters, enemy virtual characters, gunshots, etc. For example, a broadcast containing spatial information such as "enemy found on the second floor of the motel" can be made.

[0254] The spatial query logic defines the concepts of parent area and child area. The parent area refers to a large area in a three-dimensional virtual environment that covers multiple buildings or multiple isolated spaces, for example, a motel with a front yard and a back yard and containing multiple rooms and a basement. The child area refers to an independent space within the parent area, which can be a closed space or a special functional space, for example, a guest room or a kitchen in a motel.

[0255] In order to make the position description returned by the spatial query closer to human expression, when the master virtual character and the queried position are in different parent areas, the query result returns complete text information (for example, the master virtual character is outside the motel and the queried position is on the second floor of the motel, then the result is "the enemy is on the second floor of the motel"). When the master virtual character and the queried position are in the same parent area but different child areas, the query result only returns the child area information (for example, the master virtual character is on the first floor of the motel and the queried position is on the second floor, then the result is "the enemy is on the second floor"). When the master virtual character and the queried position are in the same parent area and the same child area, the query result only returns the direction relative to the master player's perspective (for example, the master virtual character and the queried position are both on the second floor of the motel, then the result is "the enemy is in the right front position").

[0256] Exemplarily, as shown in (1) of FIG. 10, in a case where the host virtual character 1011 and the target entity 1010 are located in different parent spaces, the location description includes the parent space and the child space in which the target entity is located. As shown in (2) of FIG. 10, in a case where the host virtual character 1011 and the target entity 1010 are located in different child spaces of the same parent space, the location description includes the child space in which the target entity is located. As shown in (3) of FIG. 10, in a case where the host virtual character 1011 and the target entity 1010 are located in the same child space of the same parent space, the location description includes the relative position of the target entity and the host virtual character; wherein one parent space includes at least one child space.

[0257] In some embodiments, the identification process of the target entity in the three-dimensional virtual environment can be determined according to the target entity information indicated by the natural language command and the environment perception information. Illustratively, the target entity is obtained from the entity set of the three-dimensional virtual environment according to the target entity information indicated by the natural language command and the environment perception information.

[0258] Optionally, the client infers the target entity according to the target entity information, the environment perception information and the preprocessed three-dimensional virtual environment data. The preprocessed three-dimensional virtual environment data includes an entity set of the three-dimensional virtual environment. The entity set includes entity information of each entity in the three-dimensional virtual environment. The entity information includes at least one of the following: name, type, location, feature, text label, embedding vector of the text label, image, embedding vector of the image. The text label can be at least one word obtained by tokenizing the feature (for example, a description text of the appearance of the entity), and the embedding vector corresponding to each text label is obtained by calling an embedding model. The image can include an image obtained by observing a three-dimensional model of the entity from at least one direction, for example, the image can include three views of the entity. The embedding vector of the image is an embedding vector of an image feature label, which is obtained by calling a multi-modal model based on the image and the text label of the entity.

[0259] In some embodiments, the natural language command is parsed to obtain target entity information, wherein the target entity information includes at least one of the following: entity type, entity name, entity position, entity feature; the similarity between the target entity information and the entity information of each entity in the entity set is calculated; and the target entity is determined from the entity set according to the similarity and the environment perception information.

[0260] Optionally, the client calls a pre-trained large language model to parse the natural language command to obtain the target entity information of the target entity. An embedding model is called to vectorize the target entity information to obtain a target embedding vector of the target entity information.

[0261] Exemplarily, a named entity recognition technique in the field of natural language processing can also be called to extract the target entity information from the natural language command. The named entity recognition technique is used to identify entities with specific meanings in the text, for example, to identify names, place names, adjectives, etc.

[0262] Further, the similarity includes a text similarity, and the entity set includes entity information of the first entity. Exemplarily, a calculation process of the similarity is implemented as follows: tokenizing the target entity information to obtain at least one target entity label; converting the at least one target entity label into at least one target embedding vector; obtaining the entity information of the first entity, the entity information including a text embedding vector, the text embedding vector being an embedding vector converted based on a text label of the first entity; respectively calculating text parent similarities of the at least one target embedding vector and the text embedding vector to obtain at least one text parent similarity corresponding to the at least one target embedding vector respectively; and determining a sum of the at least one text parent similarity as the text similarity of the target entity information and the entity information of the first entity.

[0263] In some embodiments, the client first performs a similarity search according to the target entity information to obtain at least one candidate entity with a higher similarity, and sorts the at least one candidate entity according to the similarity; then performs filtering and sorting based on position, orientation, direction, etc. of the master virtual character and / or the non-player character, and finally screens the target entity.

[0264] Exemplarily, the game application determines a target sensing range according to the target entity information and the environmental sensing information; and screens the target entity from entities sensed in the target sensing range according to the similarity.

[0265] Alternatively, the game application determines x entities with the highest similarity from the entity set according to the similarity to form a candidate entity list; determines a target sensing range according to the target entity information and the environmental sensing information, and determines entities in the target sensing range from the x entities as the target entity. x is a positive integer.

[0266] Exemplarily, the game application determines the entity with the highest similarity in the sensing range as the target entity.

[0267] It should be noted that when there are multiple entities with the same similarity in the target sensing range, the multiple entities can be sorted according to distances of the entities to the master virtual character. For example, in a case where the number of entities with the highest similarity in the sensing range is at least two, the entity with the highest similarity and closest to the master virtual character in the sensing range is determined as the target entity.

[0268] For example, the search range is determined according to the position and orientation of the host virtual character. For example, the target entity information indicates that the target entity is in front of the host virtual character, and the search range is a fan-shaped area in front of the host virtual character with a certain field of view size; entities not in the fan-shaped range are filtered out. Then, the entities are sorted according to the similarity scores, and when there are two entities with the highest and equal similarity scores in the search range, i.e., the similarity scores of entity 2 and entity 3 are both 2 points, the entity 2 closer to the host virtual character is selected as the target entity.

[0269] In an optional embodiment, the game application program determines the entity with the highest similarity in the perception range and above the threshold value as the target entity. For example, the number of perception ranges is at least one. For example, the perception range includes at least one of the following: the field of view range of the host virtual character, the hearing range of the host virtual character, the perception range of the host virtual character perceiving virtual props, the field of view range of the non-player character, the hearing range of the non-player character, and the perception range of the non-player character perceiving virtual props.

[0270] For example, the game application program can traverse the at least one perception range in sequence, or the game application program can traverse the at least one perception range determined according to the natural language command in sequence. For example, the natural language command is "Can you see the truck in front of you? Move to the truck", the game application program can traverse the field of view range of the host virtual character first, and then traverse the field of view range of the non-player character to match the entity with the highest similarity and above the threshold value.

[0271] For example, when there are at least two perception ranges, the traversal order of the at least two perception ranges can be preset, or the traversal order of the at least two perception ranges can be determined according to the intention indicated in the natural language command. For example, when the intention is to search for an object, the field of view range is traversed first. When the intention is to find a combat scene, the hearing range is traversed first.

[0272] If there are multiple levels of nested instructions in the natural language command, the first target entity is searched first, and then the next target entity is searched according to the position of the previous target entity. In the case where the natural language command includes a second target entity determined based on the position of the first target entity, the first target entity is queried from the entity set of the three-dimensional virtual environment according to the first target entity information of the first target entity indicated by the natural language command and the environmental perception information; then the second target entity is queried from the entity set of the three-dimensional virtual environment according to the second target entity information of the second target entity, the first target entity, and the environmental perception information.

[0273] It is worth noting that the above-mentioned target entity identification process can also be implemented by the server, which will not be described hereinafter.

[0274] In summary, the game application can generate feedback information closer to human expression. According to the current perception of the user, the game application can generate location description text that is more consistent with the current human perception. When the host virtual character and the target entity are located in different buildings, the complete location information of the target entity is broadcast to the player. When the host virtual character and the target entity are located in different spaces in the same building, the spatial position information is broadcast to the player. When the host virtual character and the target entity are located in the same space in the same building, the direction information is broadcast to the player. The player can intuitively know the spatial position relationship with the target entity through the broadcast content, and the information transmission efficiency of the broadcast is improved.

[0275] The method provided by the embodiment proposes a scheme for NPC to provide real-time voice feedback to players in the game process based on game static offline data and runtime data, combined with a large language model and text-to-speech technology, to bring an immersive experience to players during gameplay.

[0276] In an optional embodiment, the processing of the natural language command includes intent recognition of the natural language command. Illustratively, the non-player character has multiple behavior capabilities in the virtual environment, and the multiple behavior capabilities correspond to multiple classification labels. The command text corresponding to the natural language command is obtained, at least one hierarchical prediction network is called to perform intent recognition on the command text, and the classification label of the command text is obtained. The environment perception information in the virtual activity process corresponding to the intent of the non-player character is determined based on the intent indicated by the classification label, and the feedback text is generated based on the environment perception information. The hierarchical prediction network includes at least two sub-networks, and the at least two sub-networks are constructed based on a tree structure of the multiple classification labels. After the feedback text is generated, the non-player character performs the virtual activity according to the intent corresponding to the classification label, and the feedback information of the non-player character is broadcast. The feedback text corresponding to the feedback information is non-fixed text generated based on the environment perception information in the virtual activity process.

[0277] The hierarchical prediction network can predict the classification label corresponding to the command text. The hierarchical prediction network includes at least two sub-networks, an upper network and a lower network, which are cascaded. The lower network further performs prediction of the classification label according to the prediction result output by the upper network. The at least two sub-networks are constructed based on a tree structure of multiple classification labels. The hierarchical prediction network constructs at least two sub-networks corresponding to the hierarchical structure of the tree structure of the classification labels, divides the prediction task of the multiple classification labels into prediction sub-tasks performed by the at least two sub-networks, so as to reduce the complexity of classification prediction of each sub-network. Due to the richness of natural language expression, the classification labels corresponding to the natural language are various. By generalizing the classification labels with similar semantics into a high-level label, the various classification labels are constructed into a tree structure to clarify the similarity between the classification labels.

[0278] In an optional implementation, each of the at least one hierarchical prediction network is configured to predict a command intention of the command text for the non-player character in an intention dimension. The intention dimension includes at least one of a subject dimension, a semantic dimension, and a behavior dimension. In an example, the at least one hierarchical prediction network includes a subject prediction network, a behavior prediction network, and a semantic prediction network to perform parallel tactical instruction understanding of the command text. The command intention for the non-player character is described from three dimensions of the subject dimension, the semantic dimension, and the behavior dimension. It can be understood that the intention recognition of the command text is also referred to as the tactical instruction understanding of the command text. By invoking the at least one hierarchical prediction network to perform the tactical instruction understanding of the command text, the classification label of the command text for commanding the non-player character is obtained. The natural language has rich semantics, and the command text in the form of natural language includes rich semantic elements. Each hierarchical prediction network predicts the command intention in the command text from different dimensions, which reduces the number of classification labels required by the hierarchical prediction network for prediction, and reduces the complexity of the classification prediction task. In addition, the hierarchical prediction network only needs to understand the natural language semantics in the command text in one dimension, which reduces the difficulty of extracting the natural language semantics in the command text.

[0279] Exemplarily, in the at least one hierarchical prediction network, an i-th layer subnetwork in each hierarchical prediction network is configured to predict a primary behavior label of the command text in one intent dimension, an i+1-th layer subnetwork in the hierarchical prediction network is configured to predict a secondary behavior label corresponding to the primary behavior label, i is a positive integer; a high-level behavior label (e.g., a primary behavior label) predicted by a high-level subnetwork (e.g., an i-th layer subnetwork) in the hierarchical prediction network includes multiple sub-labels, or subordinate lower-level labels; and a corresponding lower-level subnetwork (e.g., an i+1-th layer subnetwork) needs to be invoked to further perform classification prediction (e.g., to predict a secondary behavior label). The prediction task of various classification labels is divided into prediction subtasks performed by at least two layers of subnetworks, so as to reduce the complexity of classification prediction of each layer of subnetworks.

[0280] In some embodiments, the recognition process for the command text incorporates hotword recognition. Exemplarily, a plurality of scene hotwords corresponding to a three-dimensional virtual environment are obtained, the three-dimensional virtual environment is a virtual environment in which at least one of a host virtual character or a non-player character is located, and the plurality of scene hotwords are scene-related words of the three-dimensional virtual environment; and the natural language command is converted into the command text based on the plurality of scene hotwords.

[0281] The plurality of scene hotwords are scene-related words of the three-dimensional virtual environment. That is, there is an association between the scene hotwords and the three-dimensional virtual environment, and the scene hotwords are words describing the three-dimensional virtual environment.

[0282] Optionally, the scene hotwords include a scene state of the three-dimensional virtual environment, such as a business state, a rest state, a closed state, and the like, when the three-dimensional virtual environment is a virtual restaurant. Optionally, the scene hotwords include an element name of a virtual element in the three-dimensional virtual environment, such as a virtual table, a virtual chair, a virtual kitchen utensil, a virtual dinner plate, and the like. Optionally, the scene hotwords include an interactive word for interacting with a virtual element in the three-dimensional virtual environment, such as attack, attack a virtual character, defense, and build a barrier, when the three-dimensional virtual environment is a virtual battlefield.

[0283] In some embodiments, the plurality of scene hotwords are collected based on the three-dimensional virtual environment after the natural language command is received, or the plurality of scene hotwords corresponding to the three-dimensional virtual environment are selected from a plurality of pre-obtained scene hotwords based on the three-dimensional virtual environment after the natural language command is received.

[0284] Illustratively, since the scene hot words are highly relevant words in the three-dimensional virtual environment, in order to better coordinate the non-player characters with the activities of the host virtual character through natural language commands, the scene hot words can be used as constraint conditions on the basis of the three-dimensional virtual environment, so that the command text is more suitable for the first virtual scene in the process of converting the natural language command into the command text, and the error rate of command analysis is reduced.

[0285] In some embodiments, the natural language command is analyzed by a pre-trained natural language analysis model, the natural language command and the plurality of scene hot words are taken as inputs of the model, instruction feature representations of the natural language command and hot word feature representations of the plurality of scene hot words are extracted by the natural language analysis model, the instruction feature representations and the hot word feature representations are fused to obtain fusion feature representations, and the command text is output based on the fusion feature representations.

[0286] In some embodiments, the natural language command is analyzed by a pre-trained natural language analysis model, the natural language command and the plurality of scene hot words are taken as inputs of the model, in the process of analyzing the natural language command by the natural language analysis model, an attention mechanism is used to determine the weight of the scene hot words in the instruction analysis process, so as to achieve the purpose of taking the plurality of scene hot words as analysis constraint conditions, and finally the output of the natural language analysis model is taken as the command text.

[0287] Optionally, the natural language analysis model includes a Gaussian Mixture Model—Hidden Markov Model (GMM-HMM), a deep learning model—Recurrent Neural Networks (RNN), a deep learning model—Transformer, an end-to-end model—deep learning speech recognition (End-to-End Deep Learning ASR), etc., wherein the end-to-end model—deep learning speech recognition includes a Connectionist Temporal Classification, a Sequence-to-Sequence, etc., and the selection of the natural language analysis model is not limited here.

[0288] In some embodiments, the process of obtaining the plurality of scene hotwords corresponding to the three-dimensional virtual environment can be implemented as at least one of the following: obtaining a first scene type corresponding to the three-dimensional virtual environment; obtaining a hotword set corresponding to the first scene type, taking the hotwords in the hotword set as the plurality of scene hotwords, the hotword set being a set of scene-related words collected based on the first scene type; or obtaining the plurality of scene hotwords based on the environment perception information; or obtaining a first scene type corresponding to the first virtual scene; obtaining a hotword set corresponding to the first scene type; and obtaining at least two words from the hotword set as the plurality of scene hotwords based on the environment perception information.

[0289] It is worth noting that the above-mentioned intent recognition and hotword recognition processes can also be implemented by the server, and will not be described again below.

[0290] It should be noted that the calculation in the above method embodiments can be completely performed by the client, for example, the game application program is an application program of a single-player game; or can be completely performed by the server, for example, in a cloud game scenario, all calculations of the application program are completed by the server, and the client is only responsible for display and collection of user operation behaviors; or can be completed by the client and the server in cooperation, for example, the client is responsible for part of the calculation, and the server is responsible for another part of the calculation.

[0291] FIG. 11 shows a flowchart of a non-player character broadcasting method according to an example embodiment of the present application. The method can be performed by the server in FIG. 1 described above. The method includes:

[0292] Step 510, receiving a natural language command for a non-player character sent by the client.

[0293] Illustratively, the client has control authority of the virtual character.

[0294] Illustratively, the method can be performed by a service end installed on the server, the service end being a service end of an application program, and the application program being an application program supporting a three-dimensional virtual environment. For example, the application program can be an application program of a shooting game. Illustratively, the server controls the non-player character to move in the three-dimensional virtual environment.

[0295] Step 520, obtaining environment perception information of at least one of the virtual character and the non-player character.

[0296] For example, the environmental perception information includes at least one of the following: visual information perceived within a visual field of the virtual character; auditory information perceived within an auditory field of the virtual character; information perceived by a perception skill or a perception device possessed by the virtual character; visual information perceived within a visual field of the non-player character; auditory information perceived within an auditory field of the non-player character; information perceived by a perception skill or a perception device possessed by the non-player character.

[0297] The environmental perception information of the non-player character can include: a picture or entity that can be seen by the non-player character when observing the three-dimensional virtual environment, an environmental sound that can be heard by the non-player character, a source direction of the environmental sound, a type and sound size of the environmental sound, perception information obtained by the non-player character by using a perception skill (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and perception information obtained by the non-player character by using a perception device (for example, a sensing signal of a sensor, a position detection signal of a position detection prop, etc.).

[0298] It should be noted that the environmental perception information can include information perceived by the non-player character in real time, or information perceived by the non-player character historically.

[0299] Step 530, determining the feedback text of the non-player character according to the natural language command and the environmental perception information.

[0300] The feedback text is a non-fixed text generated based on the environmental perception information of the non-player character.

[0301] In some embodiments, the feedback text can be generated by calling a pre-trained large language model. Illustratively, the pre-trained large language model is called to infer the natural language command based on static environmental data of the three-dimensional virtual environment and the environmental perception information, to obtain an inference result of the natural language command; and the feedback text is determined according to the inference result.

[0302] In some embodiments, the static environmental data includes data of environmental elements that exist in the three-dimensional virtual environment, such as virtual buildings, virtual terrains (for example, grasslands, caves, mountains, canyons, etc.), virtual animals, virtual plants, virtual streams, etc.

[0303] In some embodiments, the feedback text includes a reply feedback text, and in a case where the behavior intention of the natural language command is a query, the reply feedback text of the non-player character is determined according to the environmental perception information; the reply feedback text includes reply content generated according to the environmental perception information of the non-player character in the three-dimensional virtual environment and directed to the natural language command.

[0304] Illustratively, in the case that the natural language command includes an intelligence inquiry to the target entity, the first reply feedback text is determined according to the environmental perception information based on the perception of the non-player character to the target entity; the first reply feedback text includes the intelligence perception result of the non-player character to the target entity; in the case that the natural language command includes a location inquiry to the target entity, the second reply feedback text is determined according to the environmental perception information based on the perception of the non-player character to the target entity; the second reply feedback text includes the location description of the target entity.

[0305] In some other embodiments, the feedback text includes a state feedback text, in the case that the behavior intention of the natural language command is to execute a task, the state feedback text is determined according to the environmental perception information; the state feedback text includes the execution state after the non-player character executes the natural language command according to the environmental perception information of the three-dimensional virtual environment.

[0306] Illustratively, in the case that the behavior intention of the natural language command includes finding a target entity, the first state feedback text is determined according to the environmental perception information; the state feedback text includes the finding result of the non-player character to the target entity; in the case that the behavior intention of the natural language command includes using a virtual prop, the second state feedback text is determined according to the environmental perception information; the second state feedback text is used to indicate that the virtual prop is ready; in the case that the behavior intention of the task natural language command includes using a virtual prop, the third state feedback text is determined according to the environmental perception information; the third state feedback text is used to indicate the use result of the virtual prop; in the case that the behavior intention of the natural language command includes controlling the non-player character to move, the fourth state feedback text is determined according to the environmental perception information; the fourth state feedback text is used to indicate the movement result of the movement; in the case that the behavior intention of the natural language command includes executing an interaction with a target entity, the fifth state feedback text is determined according to the environmental perception information; the fifth state feedback text is used to indicate at least one of the finding result of the non-player character to the target entity, the interaction result of the non-player character and the target entity; in the case that the behavior intention of the natural language command includes an article interaction, the sixth state feedback text is determined according to the environmental perception information; the sixth state feedback text is used to indicate the article interaction result of the non-player character executing the article interaction.

[0307] Illustratively, in the case that the behavior intention of the natural language command is to execute a task, a pre-trained large language model is called to parse the natural language command, and a task execution instruction is inferred based on the static environmental data and the environmental perception information of the three-dimensional virtual environment; the non-player character is controlled to execute the task according to the task execution instruction; and the state feedback text is determined according to the execution state of the non-player character executing the task.

[0308] In step 540, the relevant information for playing the feedback information is sent to the client based on the feedback text.

[0309] In the embodiments of the present application, the feedback information is human voice audio corresponding to the feedback text, that is, the client broadcasts the feedback information corresponding to the feedback text.

[0310] For example, the feedback information can be determined in at least one of the following ways: directly broadcasting the offline voice stored locally by the client as the feedback information, splicing the offline voice stored locally by the client with the entity voice of the target entity to obtain the feedback information, and converting the inference result of the large language model into voice to obtain the feedback information.

[0311] In combination with the above determination manner of the feedback information, the related information sent by the server to the client includes at least one of the following:

[0312] First, the feedback text

[0313] Illustratively, when the feedback information is implemented as offline voice stored locally by the client or splicing of offline voice and entity voice, the server directly sends the generated feedback text to the client, and the client generates feedback information according to the feedback text, and then broadcasts the feedback information.

[0314] Second, the feedback information

[0315] Illustratively, when the feedback information is obtained by converting the inference result of the large language model into voice, the server generates the feedback information based on the feedback text, sends the generated feedback information to the client, and the client broadcasts the received feedback information.

[0316] Third, the voice identifier corresponding to the feedback information

[0317] Illustratively, when the feedback information is implemented as offline voice stored locally by the client or splicing of offline voice and entity voice, the server determines the feedback information that needs to be broadcast by the client based on the feedback text, sends the voice identifier corresponding to the feedback information to the client, and the client reads the corresponding feedback information locally according to the voice identifier for broadcasting.

[0318] In some embodiments, the server determines the feedback information according to the first offline voice corresponding to the inference result in the case that the client stores the first offline voice. Illustratively, the feedback text corresponding to the first offline voice or the first voice identifier corresponding to the first offline voice is sent to the client, and the first voice identifier is used to instruct the client to generate the feedback information based on the first offline voice, that is, in the case that the client stores the first offline voice, the client queries the corresponding first offline voice according to the received feedback text to generate the feedback text, or queries the corresponding first offline voice according to the received first voice identifier to generate the feedback text.

[0319] Optionally, after receiving the feedback text, the client determines at least one first offline voice corresponding to the feedback text, and when the number of the determined first offline voice is one, directly plays the first offline voice as the feedback information; when the number of the determined first offline voice is multiple, splices the multiple first offline voices to obtain the feedback information, and plays the feedback information.

[0320] Optionally, when the client receives one first voice identifier, queries the corresponding first offline voice according to the first voice identifier to play as the feedback information; when the client receives multiple first voice identifiers, splices the first offline voices corresponding to the multiple first voice identifiers to obtain the feedback information, and plays the feedback information.

[0321] In some embodiments, when the entity information of the target entity is included in the inference result, the server calls a text-to-speech service to generate an entity voice corresponding to the entity information, and sends the entity voice and the feedback text corresponding to the first offline voice to the client; or, sends the first voice identifier corresponding to the entity voice and the first offline voice to the client; the client is used to splice the entity voice and the first offline voice to obtain the feedback information; wherein the entity information includes at least one of the following: the entity name of the target entity, the position description of the target entity.

[0322] In other embodiments, when the client does not store the offline voice corresponding to the behavior intention and the inference result, the text-to-speech service is called to generate the feedback information based on the feedback text, and the feedback information is sent to the client.

[0323] In some embodiments, when the behavior intention of the natural language command is inquiry, the reply feedback text of the non-player character is determined according to the environmental perception information; the reply feedback text includes the reply content to the natural language command generated according to the environmental perception information of the non-player character in the three-dimensional virtual environment. For example, the natural language command is "where are you", and the reply feedback text returned by the large language model is "I am in the small wooden house", and the reply broadcast (feedback information) is "I am in the small wooden house".

[0324] Alternatively, when the behavior intention is inquiry, and the inference result provided by the large language model includes failure information such as failure to generate reply feedback text, failure to find, and failure to identify behavior intention, the first offline voice corresponding to the inquiry failure stored locally by the client is determined as the feedback information. For example, the natural language command is "help me find a nearby reward box", and the inference result fed back by the large language model is "no reward box found near the NPC", and the reply broadcast can be the offline voice broadcast "not found" stored locally by the client.

[0325] Or, in the case of the behavior intention is a query, and the reasoning result provided by the large language model includes: yes, confirm, no problem, and confirmation information, the first offline voice corresponding to the confirmation stored locally by the client is determined as the feedback information. For example, the natural language command is "are you on the second floor", and the reasoning result feedback by the large language model is "yes, the NPC is located on the second floor", then the reply broadcast can be the offline voice broadcast "yes" stored locally by the client.

[0326] Or, in the case of the behavior intention is a query, and the reasoning result provided by the large language model includes the entity information of the target entity, a text-to-speech service is called to generate entity voice based on the entity information; according to the query behavior intention and the specific content provided in the reasoning result, the corresponding first offline voice is determined, and the entity voice is filled into the position reserved for the entity information in the first offline voice to obtain the feedback information. For example, the natural language command is "is there any danger ahead", and the reasoning result returned by the large language model is "there is a first enemy virtual character at the first position", then according to the "xxx discovers enemy" offline voice template stored locally by the client, and the entity voice of "first position" obtained by using the text-to-speech technology, the feedback information of "first position discovers enemy" can be spliced.

[0327] In some embodiments, in the case of the behavior intention of the natural language command is to execute a task, a state feedback text is determined according to the environment perception information; the state feedback text includes the execution state after the non-player character executes the natural language command according to the environment perception information of the three-dimensional virtual environment. Illustratively, in the case of the behavior intention of the natural language command is to execute a task, a pre-trained large language model is called to analyze the natural language command, and a task execution instruction is inferred based on the static environment data and the environment perception information of the three-dimensional virtual environment; the non-player character is controlled to execute the task according to the task execution instruction; and the state feedback text is determined according to the execution state of the non-player character executing the task.

[0328] For example, the natural language command is "throw the first virtual prop to the first virtual character", then the reasoning result can include the control instruction of controlling the NPC to throw the first virtual prop to the first virtual character; when the NPC fails to throw the first virtual prop to the first virtual character, the server inputs each execution result of the NPC executing the control instruction into the large language model to generate feedback text corresponding to the execution result, for example, the execution result can be that the NPC does not have the first virtual prop, then the feedback text can be "currently does not have the first virtual prop"; or the execution result can be that the first virtual character is not hit, then the feedback text can be "sorry, I missed it"; a text-to-speech service is called to generate human voice audio corresponding to the feedback text to obtain the feedback information.

[0329] Or, in the case of the behavior intention is a task, according to the execution state of the task, the first offline voice corresponding to the execution state stored locally by the client is determined as the feedback information. For example, the natural language command is "throw the first virtual prop to the first virtual character", and the inference result can include the control instruction of controlling the NPC to throw the first virtual prop to the first virtual character; the server controls the NPC according to the control instruction, and when the execution result of the control instruction includes successfully hitting the first virtual character, the corresponding first offline voice "hit" is determined as the feedback information.

[0330] Or, in the case of the behavior intention is a task, and the task is related to the target entity, and the first offline voice corresponding to the execution state of the task includes the entity information of the target entity, the text-to-speech service is called to generate an entity voice based on the entity information; the entity voice is filled into the position reserved for the entity information in the first offline voice to obtain the feedback information.

[0331] Optionally, the server can also directly determine the determination method of the feedback information according to the behavior intention. For example, in the case of the behavior intention being the first type of behavior intention, the text-to-speech service is called to generate feedback information based on the inference result output by the large language model. In the case of the behavior intention being the second type of behavior intention, the offline voice stored by the client corresponding to the second type of behavior intention is taken as the feedback information.

[0332] In some embodiments, the client also broadcasts dynamic broadcast information. Illustratively, the server determines dynamic broadcast text according to the environmental perception information of the non-player character in the case that the environmental perception information of the non-player character in the three-dimensional virtual environment satisfies the dynamic broadcast condition; the dynamic broadcast text includes an abnormal situation identified according to the environmental perception information of the non-player character; based on the dynamic broadcast text, the server sends relevant information for playing dynamic broadcast information to the client, and the dynamic broadcast information is a human voice audio corresponding to the dynamic broadcast text.

[0333] Illustratively, in the case that the non-player character perceives an enemy virtual character, a first dynamic broadcast text is generated; the first dynamic broadcast text includes the position of the enemy virtual character perceived by the non-player character; in the case that the non-player character perceives a dangerous situation, a second dynamic broadcast text is generated; the second dynamic broadcast text is used to prompt the dangerous situation; in the case that the non-player character perceives a state change of an ally virtual character, a third dynamic broadcast text is generated; the third dynamic broadcast text is used to prompt the state change of the ally virtual character; in the case that the non-player character discovers a new trace in the three-dimensional virtual environment, a fourth dynamic broadcast text is generated; the fourth dynamic broadcast text is used to prompt the new trace.

[0334] In some embodiments, the server determines the dynamic broadcast information according to the second offline voice corresponding to the dynamic broadcast condition, in the case that the environment perception information of the non-player role satisfies the dynamic broadcast condition.

[0335] For example, the server can determine the second offline voice corresponding to the dynamic broadcast condition as the dynamic broadcast information. The second offline voice is an offline voice broadcast stored locally by the client. For example, the server determines the offline voice "danger ahead" as the feedback information when an enemy virtual role appears within the field of view of the non-player role.

[0336] For example, the server queries entity information of the target entity in the case that the environment perception information includes the target entity; calls a text-to-speech service to generate an entity voice corresponding to the entity information; splices the entity voice and the second offline voice to obtain the dynamic broadcast information; wherein the entity information includes at least one of the following: an entity name of the target entity, a position description of the target entity. For example, the server generates an entity voice according to the nickname "Robot No. 1" of the enemy virtual role when an enemy virtual role appears within the field of view of the non-player role, and splices the entity voice with the offline voice template "xxx found ahead" to obtain the feedback information "Robot No. 1 found ahead".

[0337] In summary, when a user feeds back relevant information of a three-dimensional virtual environment to a non-player role in the three-dimensional virtual environment through a natural language command, feedback information corresponding to the non-player role is generated and broadcast according to the information perceived by the non-player role from the three-dimensional virtual environment, so that the feedback information can cover multiple scenarios. When the natural language command instructs the non-player role to inquire whether there is an enemy in the three-dimensional virtual environment, and the non-player role perceives that the virtual role included in the three-dimensional virtual environment contains an enemy virtual role, the host virtual role is reminded of the existence of the enemy through the broadcast of the feedback information; when the natural language command instructs the non-player role to provide a virtual prop to the host virtual role, and the non-player role has a virtual prop required by the user, the non-player role submits the virtual prop to the host virtual role, and reminds the host virtual role that the virtual prop has been submitted through the broadcast of the feedback information. When the non-player role is controlled through a natural language command, the feedback information is broadcast according to the environment perception information of the non-player role, which can enrich the content that can be fed back by the feedback information, so that the broadcast is not limited to fixed triggering scenarios. The game application program can flexibly broadcast the feedback information according to the environment perception information in the game match, enrich the feedback information broadcast corresponding to the non-player role, and improve the human-computer interaction efficiency between the non-player role and the player.

[0338] In an optional embodiment, the server performs spatial sound effect processing on the human voice audio of the non-player role to obtain a spatial human voice audio as the feedback information.

[0339] FIG. 12 shows a flowchart of a method for broadcasting a non-player character according to an example embodiment of the present application. The method can be performed by the server in FIG. 1. Based on the example embodiment shown in FIG. 11, the method can further include steps 541-544, which are sub-steps of step 540.

[0340] In step 541, the human voice audio corresponding to the feedback text is obtained.

[0341] In step 542, the sound effect parameter corresponding to the non-player character is obtained based on the relative positional relationship between the non-player character and the master virtual character in the three-dimensional virtual environment.

[0342] Illustratively, the relative positional relationship between the non-player character and the master virtual character in the three-dimensional virtual environment indicates the relative relationship between a first position corresponding to the non-player character and a second position corresponding to the master virtual character in the three-dimensional virtual environment.

[0343] Optionally, the relative positional relationship includes at least one of a positional distance, a positional direction, and a spatial surrounding condition between the non-player character and the master virtual character, wherein the positional distance indicates the proximity relationship between the non-player character and the master virtual character, the positional direction indicates the angle size relationship between the first position corresponding to the non-player character and the direction of the master virtual character, and the spatial surrounding condition indicates the surrounding condition of the master virtual character by the spatial elements forming a virtual space in the three-dimensional virtual environment, wherein the virtual space has a propagation influence on the audio emitted by the non-player character.

[0344] In some embodiments, the determination of the relative positional relationship between the non-player character and the master virtual character can be implemented by obtaining first position information of the non-player character in the three-dimensional virtual environment, obtaining second position information of the master virtual character in the three-dimensional virtual environment, and determining the relative positional relationship between the non-player character and the master virtual character based on the first position information and the second position information.

[0345] In some embodiments, the first position information of the non-player character and the second position information of the master virtual character are position information determined based on a same preset coordinate system. Optionally, the preset coordinate system can be a world coordinate system corresponding to the three-dimensional virtual environment, or the preset coordinate system can be a coordinate system established with the master virtual character as the origin.

[0346] In some embodiments, when the sound effect parameter corresponding to the non-player character is pre-generated and stored in a database, the sound effect parameter corresponding to the non-player character is obtained from the database. In other embodiments, the sound effect parameter corresponding to the non-player character can also be generated in real time, i.e., the sound effect parameter corresponding to the non-player character is generated based on the relative positional relationship.

[0347] Optionally, the sound effect parameter corresponding to the non-player character can include at least one of the following types: a distance type, a direction type, a spatial effect type, and a custom type.

[0348] At step 543, the human voice audio is adjusted based on the sound effect parameter to generate spatial human voice audio corresponding to the non-player character.

[0349] The spatial human voice audio is used to represent the perception effect of the audio generated by the host virtual character on the non-player character in the relative orientation relationship.

[0350] In the embodiments of the present application, the human voice audio is adjusted according to the sound effect parameter corresponding to the non-player character, thereby generating spatial human voice audio corresponding to the non-player character. That is, the spatial human voice audio is the audio data corresponding to the non-player character obtained by adjusting the human voice audio through the sound effect parameter. In some embodiments, the spatial human voice audio includes at least one of the following data: the audio adjusted by the sound effect parameter, audio duration information, a playback start timestamp of the audio, a playback end timestamp of the audio, playback condition information of the audio, and element identification information corresponding to the audio.

[0351] Optionally, when the sound effect parameter indicates that the volume of the human voice audio of the non-player character is adjusted, the volume of the human voice audio of the non-player character in at least one sound channel is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates that the playback delay of the human voice audio of the non-player character is adjusted, the audio playback start time corresponding to the human voice audio of the non-player character is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates that the speed of the human voice audio of the non-player character is adjusted, the playback speed of the human voice audio of the non-player character is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates the pitch of the human voice audio of the non-player character, the frequency in the frequency spectrum of the human voice audio of the non-player character is adjusted based on the sound effect parameter. When the sound effect parameter indicates the timbre of the human voice audio of the non-player character, a filter corresponding to the non-player character is determined based on the sound effect parameter, and the timbre of the human voice audio is adjusted through the filter.

[0352] At step 544, the spatial human voice audio is sent to the client.

[0353] Illustratively, the server sends the obtained spatial human voice audio to the client, and the client broadcasts the spatial human voice audio as feedback information in response to the natural language command.

[0354] In summary, when the player seeks voice feedback of the non-player character through the natural language command, the sound effect parameter of the non-player character is acquired according to the relative positional relationship between the host virtual character and the non-player character, and the spatial voice frequency of the non-player character is obtained by adjusting the voice frequency of the non-player character through the sound effect parameter. Since the spatial voice frequency is generated in real time based on the relative positional relationship between the host virtual character and the non-player character, the application program providing the three-dimensional virtual environment does not need to pre-configure too many application resources in order to improve the diversity of sound effect performance, thereby reducing the waste of storage space of the computer device running the above-mentioned application program.

[0355] Meanwhile, the voice frequency of the non-player character after being adjusted by the sound effect parameter can make the spatial voice frequency played by the client simulate the perception effect of the audio generated by the host virtual character to the non-player character under the current relative positional relationship, so that the spatial voice frequency played by the client can make the player perceive the voice of the non-player character in the three-dimensional virtual environment as the host virtual character, thereby improving the diversity of the sound effect perceived by the player in the game process, and improving the immersion of the player in the three-dimensional virtual environment and the user experience.

[0356] In an optional embodiment, when the sound effect parameter is generated based on the relative positional relationship between the non-player character and the host virtual character, the first position of the non-player character in the three-dimensional virtual environment and the second position of the host virtual character in the three-dimensional virtual environment need to be combined.

[0357] FIG. 13 shows a flowchart of a non-player character broadcasting method according to an example embodiment of the present application. The method can be performed by the server in FIG. 1. Based on the embodiment shown in FIG. 12, the method can further include steps 5421-5423, which are sub-steps of step 542.

[0358] Step 5421: Obtain first position information of the non-player character in the three-dimensional virtual environment; and obtain second position information of the host virtual character in the three-dimensional virtual environment.

[0359] In the embodiments of the present application, the non-player character is a virtual character that moves following the host virtual character in the three-dimensional virtual environment, and thus the position of the non-player character in the three-dimensional virtual environment is determined by monitoring the position of the non-player character in real time.

[0360] In some embodiments, the first position information is coordinate point data determined based on a preset coordinate system.

[0361] In some embodiments, the three-dimensional virtual environment is divided into a grid structure, and the position information of the scene element in the three-dimensional virtual environment is indicated by a grid point in the grid structure. Illustratively, a first grid point in which the non-player character is located in the grid structure is obtained, and a position of the first grid point in the grid structure is determined as the first position information.

[0362] As shown in FIG. 14, which shows a schematic diagram of a grid structure according to an example embodiment of the present application, a three-dimensional virtual environment 1400 is divided into a grid structure composed of multiple grids, and a virtual character 1401 in the three-dimensional virtual environment 1400 occupies a certain number of grid points in the grid structure.

[0363] The user can control the host virtual character to move in the three-dimensional virtual environment, and thus the position of the host virtual character in the three-dimensional virtual environment is not fixed. Illustratively, the position of the host virtual character in the three-dimensional virtual environment is detected to obtain second position information.

[0364] In some embodiments, the second position information is coordinate point data determined based on a preset coordinate system.

[0365] In some embodiments, the three-dimensional virtual environment is divided into a grid structure, and illustratively, a second grid point corresponding to the position of the host virtual character in the three-dimensional virtual environment in the grid structure is determined, and a position of the second grid point in the grid structure is determined as the second position information.

[0366] At step 5422, the relative orientation relationship between the non-player character and the host virtual character is determined based on the first position information and the second position information.

[0367] In some embodiments, when the first position information and the second position information are coordinate point data determined based on a preset coordinate system, the relative orientation relationship between the non-player character and the host virtual character is determined based on a first coordinate point indicated by the first position information and a second coordinate point indicated by the second position information.

[0368] In some embodiments, when the first position information and the second position information are grid points indicated based on the grid structure, the relative orientation relationship between the non-player character and the host virtual character is determined based on a first grid point corresponding to the first position information and a second grid point indicated by the second position information.

[0369] Optionally, the relative positional relationship comprises at least one of a positional distance, a positional direction, and a spatial surrounding condition between the non-player character and the host virtual character, wherein the positional distance is used to indicate a distance relationship between the non-player character and the host virtual character, the positional direction is used to indicate an angle size relationship between a first position of the non-player character and an orientation direction of the host virtual character, and the spatial surrounding condition is used to indicate a surrounding condition of the host virtual character in a virtual space formed by spatial elements in the three-dimensional virtual environment, and the virtual space is used to indicate a propagation influence of the spatial elements on the audio when the non-player character emits the audio.

[0370] At step 5423, the sound effect parameter corresponding to the non-player character is generated based on the relative positional relationship.

[0371] Optionally, the generation of the sound effect parameter corresponding to the non-player character can be implemented in at least one of the following manners:

[0372] Firstly, the sound effect parameter is generated based on a sound propagation path of the audio generated by the non-player character under the relative positional relationship.

[0373] The sound propagation path is used to indicate a propagation path of sound waves when the non-player character acts as a sound source. It is worth noting that one non-player character can correspond to multiple sound propagation paths.

[0374] Illustratively, the sound propagation path of the audio generated by the non-player character between the non-player character and the host virtual character is determined based on the relative positional relationship, and the sound effect parameter corresponding to the non-player character is generated based on the sound propagation path. That is, since the non-player character acts as a sound source in the three-dimensional virtual environment, the audio emitted by the non-player character propagates in the three-dimensional virtual environment, and the path pointing to the host virtual character is the sound propagation path of the audio between the non-player character and the host virtual character.

[0375] In some embodiments, the sound propagation path between the non-player character and the host virtual character can be implemented as a ray path from the non-player character to the host virtual character, that is, the sound directly propagates from the non-player character to the host virtual character.

[0376] In some embodiments, since the audio emitted by the non-player character also has other direction propagation paths, and after reflection by other scene elements in the three-dimensional virtual environment, the host virtual character also has a sound propagation path to receive the audio, that is, the sound propagation path is a path of the audio emitted by the non-player character after reflection in the three-dimensional virtual environment.

[0377] In one example, the sound propagation path of the audio received by the host virtual character includes a direct propagation path and a reflected propagation path. For example, based on the relative positional relationship, a first propagation path of the audio generated by the non-player character directly propagating from the non-player character to the host virtual character is determined, and based on the relative positional relationship, a second propagation path of the audio generated by the non-player character propagating to the host virtual character after being reflected in the three-dimensional virtual environment is determined. Based on the first propagation path and the second propagation path, the sound effect parameter corresponding to the non-player character is determined.

[0378] Specifically, please refer to Formula One and Formula Two, which show the radiation process of sound in the three-dimensional virtual environment. Formula One and Formula Two collectively describe the change of sound energy in the process of audio propagating from the non-player character to the host virtual character. Formula One is the change of total energy, and Formula Two is the energy change of audio in the transmission process due to the scattering effect of the non-uniformity of the medium or other sound sources.

[0379] where t is time, x' is the first position of the non-player character, x is the second position of the host virtual character, L(x'→x, t) represents the sound field contribution from x' to x at time t, L0(x'→x, t) represents the primary sound field directly radiated from x' to x, L s (x'→x, t) represents the scattering effect of the sound wave in the process of propagating from x' to x due to the non-uniformity of the medium or other sound sources, (x'→x, t) represents the scattering effect of the sound wave in the process of propagating from x' to x due to the non-uniformity of the medium or other sound sources, (x'→x, t) represents the scattering effect of the sound wave in the process of propagating from x' to x due to the non-uniformity of the medium or other sound sources,

[0380] In some embodiments, the detection of the sound propagation path can be realized by Ray Tracing or Image Source Method to simulate the direct propagation or reflected propagation of sound.

[0381] In some embodiments, in the process of determining the sound effect parameter, the sound effect generated in the reflection propagation process of the sound can also be integrated into the reverberation calculation. Specifically, the early reflections of the audio emitted by the non-player character in the three-dimensional virtual environment are calculated, which usually arrive at the master virtual character within a specified time length (e.g., 50 milliseconds), and then the First In First Out (FIFO) sound queue and the Schroeder algorithm are used to simulate the environmental tail sound, thereby adding the audio effect of the reverberation tail. As shown in FIG. 15, which shows a schematic diagram of a reverberation sound wave provided by an exemplary embodiment of the present application, in the reverberation sound wave schematic diagram 1500, the horizontal coordinate is used to represent the time when the sound arrives at the master virtual character, and the vertical coordinate is used to represent the volume of the received sound, wherein the volume of the first sound 1501 received at t1 through the direct propagation path to the master virtual character is the largest, the volume of the second sound 1502 received through the early reflection propagation path to the master virtual character within the t2 time period is reduced to a certain extent, and the volume of the third sound 1503 received as the reverberation tail within the t3 time period is low.

[0382] That is, by simulating the propagation path of the sound in the three-dimensional virtual environment to generate the corresponding sound effect parameter, the scene audio can embody the result that the audio emitted by each non-player character arrives at the master virtual character after propagating in the three-dimensional virtual environment, improve the authenticity of the scene audio, make the user more immersive, and improve the user experience. Further, when simulating the propagation path of the sound, in addition to the direct propagation path to the master virtual character, the result of the reflection of the sound in the three-dimensional virtual environment is also simulated, which can improve the spatial sense of the spatial audio corresponding to the non-player character, that is, the reflected sound can increase the depth and width sense of the space, making the sound sound more rich and stereoscopic, thereby improving the audio effect of the scene audio and improving the user experience.

[0383] Secondly, whether the non-player character is within the field of view of the master virtual character is used to determine the corresponding sound effect parameter.

[0384] Specifically, the field of view range of the master virtual character in the three-dimensional virtual environment is determined based on the orientation of the master virtual character, in the case that the non-player character is within the field of view range, the sound effect parameter is determined based on the first parameter, and in the case that the non-player character is outside the field of view range, the sound effect parameter is determined based on the second parameter, wherein the audio volume corresponding to the first parameter is greater than the audio volume corresponding to the second parameter. That is, when the master virtual character faces the non-player character, the user is made aware that the sound source is in front of the master virtual character by increasing the audio volume of the played non-player character.

[0385] In some embodiments, the field of view of the host virtual character is the range of the three-dimensional virtual environment displayed by the client when providing the user with a three-dimensional virtual environment picture, i.e., the three-dimensional virtual environment picture is a picture of the three-dimensional virtual environment observed from the perspective of the host virtual character, and the range of the three-dimensional virtual environment corresponding to the picture is the field of view.

[0386] In some embodiments, the field of view of the host virtual character is determined by a virtual camera (Camera), which is a crucial component in the three-dimensional virtual environment that determines the content and perspective of what the user sees on the screen. The position of the virtual camera in the three-dimensional space determines the observation point for observing the three-dimensional virtual environment, the orientation of the virtual camera in the three-dimensional space determines the observation direction for observing the three-dimensional virtual environment, and the field of view (FOV) of the virtual camera determines the range of the three-dimensional virtual environment that can be displayed on the screen. Typically, the user cannot directly observe the virtual camera in the three-dimensional virtual environment.

[0387] Optionally, when the user-controlled host virtual character observes the three-dimensional virtual environment in the first-person perspective, the virtual camera is bound to the head position of the host virtual character; when the user-controlled host virtual character observes the three-dimensional virtual environment in the third-person perspective, the virtual camera is bound to the position behind the host virtual character.

[0388] In one example, the field of view is determined by the frustum of the virtual camera, which is a spatial region defined by the position of the virtual camera and the FOV. Illustratively, after determining the loading range corresponding to the three-dimensional virtual environment, frustum culling is performed to determine which scene elements are within the frustum, thereby determining which scene elements need to be rendered, and then determining the non-player character that emits the audio from the rendered scene elements, i.e., the non-player character within the field of view, which corresponds to the sound effect parameter configured with the first parameter.

[0389] Specifically, six planes (top, bottom, left, right, front, back) are extracted from the frustum of the camera to extract the frustum planes. For each scene element, its bounding box is tested with the frustum planes, and if the bounding box is completely outside the frustum, the scene element is culled from the field of view of the host virtual character.

[0390] When the sound effect parameters of different volumes are configured according to the field of view of the master virtual role, the user can distinguish the positions of different non-player roles in the three-dimensional virtual environment through scene audio, determine whether the non-player role is within the field of view of the user, and more accurately determine the direction of the sound source while enhancing the sense of reality. At the same time, when the user rotates the view angle of the master virtual role, the volume of the audio emitted by the non-player role at different positions changes, thereby realizing interactive feedback of view angle change at the sound experience level.

[0391] Thirdly, the corresponding sound effect parameters are determined based on the space type of the virtual space formed by the space elements.

[0392] In some embodiments, the three-dimensional virtual environment includes space elements forming a virtual space in the three-dimensional virtual environment, and the master virtual role is in the virtual space. The surrounding state information of the space elements to the master virtual role is determined based on the relative positional relationship between the space elements and the master virtual role. The space type of the virtual space in which the master virtual role is located is determined based on the surrounding state information. The space sound effect parameters matching the space type are obtained. The sound effect parameters are determined based on the space audio parameters. The surrounding state information is used to indicate the relative relationship between the position of the master virtual role and the distribution of the space elements in the three-dimensional virtual environment.

[0393] Optionally, the space type can be implemented as at least one of an open type, a closed type, and a semi-closed type. When the space type is the open type, the three-dimensional virtual environment can be implemented as a scene such as a prairie, a plain, or a sea surface. When the space type is the closed type, the three-dimensional virtual environment can be implemented as a scene such as a cave or a room. When the space type is the semi-closed type, the three-dimensional virtual environment can be implemented as a scene such as a canyon or a forest.

[0394] In some embodiments, different space types correspond to different sound effect parameters. For example, when the space type is the open type, the corresponding sound effect parameters make the audio more transparent and clear, correspond to lower reverberation, and faster sound attenuation. Illustratively, the sound effect parameters corresponding to different space types can be pre-configured, that is, the database pre-stores the pre-configured sound effect parameter strategies corresponding to different space types respectively. When the space type of the virtual space in which the master virtual role is located is determined based on the surrounding state information, the sound effect parameter strategy corresponding to the space type is found from the database based on the type identifier of the space type, and the sound effect parameters matching the space type are obtained from the sound effect parameter strategy. That is, the sound effect parameters matching the space type are obtained according to the space type formed by the space elements in the three-dimensional virtual environment, so that the user can perceive the virtual space of the three-dimensional virtual environment from the auditory perception level, enhance the immersion of the user, and make the user feel more immersed in the three-dimensional virtual environment.

[0395] Fourthly, the sound propagation volume attenuation and delay based on the distance between the master virtual character and the non-player character are used to determine the corresponding sound effect parameters.

[0396] In some embodiments, since there is a certain distance between the master virtual character and the non-player character, the sound effect parameters are also used to adjust the attenuation effect and / or delay effect of the sound. Illustratively, the distance information between the master virtual character and the non-player character is determined based on the relative positional relationship, the attenuation parameters and delay parameters corresponding to the audio generated by the non-player character are determined based on the distance information, and the sound effect parameters are determined based on the attenuation parameters and delay parameters.

[0397] In some embodiments, the above-mentioned attenuation parameters are preset sound attenuation functions. In some embodiments, the above-mentioned delay parameters are delay durations determined based on the distance information and the sound propagation speed.

[0398] That is, the attenuation effect of the sound after propagation in the three-dimensional virtual environment is simulated through the attenuation parameters, and the delay effect of the sound in the three-dimensional virtual environment is simulated through the delay parameters, thereby improving the authenticity of the scene audio.

[0399] In summary, when the player seeks voice feedback of the non-player character through the natural language command, the sound effect parameters of the non-player character are obtained according to the relative positional relationship between the master virtual character and the non-player character, and the spatial human voice audio of the non-player character is obtained through the adjustment of the sound effect parameters. Since the spatial human voice audio is generated in real time based on the relative positional relationship between the master virtual character and the non-player character, the application program providing the three-dimensional virtual environment does not need to pre-configure too many application resources in order to improve the diversity of sound effect performance, thereby reducing the waste of storage space of the computer device running the above-mentioned application program.

[0400] Meanwhile, after the adjustment of the sound effect parameters, the spatial human voice audio played by the client can simulate the perception effect of the audio generated by the master virtual character to the non-player character under the current relative positional relationship, so that the spatial human voice audio played by the client can enable the player to perceive the human voice generated by the non-player character in the three-dimensional virtual environment as the master virtual character, thereby improving the diversity of the sound effect perceived by the player during the game, thereby improving the immersion of the player in the three-dimensional virtual environment and improving the user experience.

[0401] In the embodiments of the present application, the sound effect parameters conforming to the relative positional relationship between the non-player character and the master virtual character are generated in real time according to the first position information of the non-player character in the three-dimensional virtual environment and the second position information of the master virtual character in the three-dimensional virtual environment, so that the spatial human voice audio finally played by the client can truly simulate the perception of the sound of the non-player character by the master virtual character at the current position in the three-dimensional virtual environment, thereby improving the user experience.

[0402] In an optional embodiment, the role characteristic information of the non-player character can be incorporated into the generation process of the human voice audio for the non-player character.

[0403] FIG. 16 shows a flowchart of a method for broadcasting a non-player character according to an example embodiment of the present application. The method can be performed by the server in FIG. 1. Based on the embodiment shown in FIG. 1, the method can further include steps 5411-5412, which are sub-steps of step 541.

[0404] In step 5411, the feedback text corresponding to the non-player character and the role characteristic information are obtained.

[0405] In the embodiments of the present application, the role characteristic information is used to describe the attributes of the non-player character, i.e., the role characteristic information is information used to describe the current state and characteristics of the non-player character.

[0406] In some embodiments, the attributes of the non-player character include at least one of basic attributes of the non-player character and scenario performance attributes of the non-player character. The basic attributes are attributes of the non-player character that are pre-configured for the non-player character, i.e., attributes of the non-player character that do not change due to the three-dimensional virtual environment or the current scenario, such as age information of the non-player character, male / female information of the non-player character, faction information of the non-player character, etc. The scenario performance attributes are attributes of the non-player character that are determined in real time in the scenario of the three-dimensional virtual environment, i.e., attributes associated with the current scenario, such as emotional information of the non-player character in the current scenario, behavior information of the non-player character, etc.

[0407] Optionally, the role characteristic information includes at least one of role basic information, role emotional information, and role behavior information. The role basic information is used to indicate the basic situation of the non-player character, i.e., the information of the basic attributes of the non-player character, such as age information of the non-player character, male / female information of the non-player character, faction information of the non-player character, etc. The role emotional information is used to indicate the emotional state of the non-player character in the dialogue scenario corresponding to the feedback text, i.e., the scenario performance attribute of the non-player character, such as indicating that the non-player character is in an excited state, an angry state, a sad state, etc. The role behavior information is used to indicate the action performed by the non-player character in the dialogue scenario, i.e., the scenario performance attribute of the non-player character, such as being in a running state, being in an attacking state, being in a state of receiving treatment, etc.

[0408] It is worth noting that in the embodiments of the present application, different role characteristic information corresponds to different performances of the human voice audio. In one example, when the role characteristic information indicates that the non-player character is a young man and is currently in a battle state and is excited, the timbre of the corresponding human voice audio is the timbre of a young man. Since the non-player character is in a battle state, the human voice audio is mixed with the panting sound of the battle process, and since the non-player character is excited when speaking, the pitch of the human voice audio is high. That is, generating human voice audio with different performances according to different role characteristic information can improve the authenticity of the human voice audio of the non-player character, making the user more immersive when interacting with the non-player character.

[0409] In step 5412, the human voice audio matching the role characteristic information is generated according to the feedback text.

[0410] In the embodiments of the present application, the human voice audio of the non-player character is generated based on the feedback text and the role characteristic information, wherein the speech content of the human voice audio is determined by the feedback text, and the speech performance of the human voice audio is determined by the role characteristic information.

[0411] In some embodiments, the generation of the human voice audio is implemented by a pre-trained speech generation model, which is used to generate speech matching the role characteristic information of the non-player character. Illustratively, the feedback text and the role characteristic information are input into the pre-trained speech generation model to obtain the human voice audio as the human voice audio.

[0412] Optionally, the above-mentioned speech generation model can be implemented by a neural network model such as a convolutional neural network, a feedforward neural network, a residual network, a Transformer, a multimodal large language model (MLLM), etc., which is not specifically limited here.

[0413] In some embodiments, the above-mentioned speech generation model includes a text encoder, a role encoder, and a decoder. The text encoder is used to perform text encoding on the input feedback text, i.e., the feedback text is input into the text encoder to obtain a text encoding representation. The role encoder is used to perform feature encoding on the input role characteristic information, i.e., the role characteristic information is input into the role encoder to obtain a role encoding representation.

[0414] Illustratively, after obtaining the text encoding representation and the role encoding representation, the text encoding representation and the role encoding representation are fused to obtain an encoding representation input into the decoder, i.e., the text encoding representation and the role encoding representation are fused to obtain a joint encoding representation, and the joint encoding representation is input into the decoder to generate the human voice audio.

[0415] In some embodiments, the role encoder comprises at least one of a first sub-encoder, a second sub-encoder, and a third sub-encoder. The first sub-encoder is configured to perform feature encoding based on the basic attribute of the non-player role, i.e., in the case where the role feature information comprises role basic information, the role basic information is input into the first sub-encoder to obtain a first role encoding representation; the second sub-encoder is configured to perform feature encoding based on the emotional state of the non-player role, i.e., in the case where the role feature information comprises role emotional information, the role emotional information is input into the second sub-encoder to obtain a second role encoding representation; and the third sub-encoder is configured to perform feature encoding based on the behavior state of the non-player role, i.e., in the case where the role feature information comprises role behavior information, the role behavior information is input into the third sub-encoder to obtain a third role encoding representation.

[0416] As shown in FIG. 17, which shows a schematic diagram of a speech generation model 1700 provided by an example embodiment of the present application, the speech generation model 1700 comprises an encoder part 1710 and a decoder part 1720, wherein the encoder part 1710 comprises a text encoder 1711, a first sub-encoder 1712, a second sub-encoder 1713, and a third sub-encoder 1714. The feedback text 1701 is input into the text encoder 1711, and a text encoding representation is output. The role basic information 1702 of the non-player role is input into the first sub-encoder 1712, and a first role encoding representation is output. The role emotional information 1703 of the non-player role is input into the second sub-encoder 1713, and a second role encoding representation is output. The role behavior information 1704 of the non-player role is input into the third sub-encoder 1714, and a third role encoding representation is output. The text encoding representation, the first role encoding representation, the second role encoding representation, and the third role encoding representation are fused by a fusion layer 1730 to obtain a joint encoding representation, which is input into the decoder part 1720. The human voice audio is output by the decoder.

[0417] In the embodiments of the present application, the method for generating the sound effect parameter of the human voice audio can be implemented as at least one of the following methods in addition to the method provided in the foregoing embodiments:

[0418] Firstly, an adjustment parameter is generated based on the distance between the non-player role and the master virtual role, and the adjustment parameter is used to determine the sound effect parameter.

[0419] Illustratively, the first position information of the non-player role in the three-dimensional virtual environment is obtained, and the second position information of the master virtual role in the three-dimensional virtual environment is obtained. In the case where the positional distance between the first position information and the second position information reaches a first threshold value, a tone adjustment parameter is generated, the tone adjustment parameter is used to adjust the tone of the human voice audio, and the sound effect parameter is determined based on the tone adjustment parameter.

[0420] In some embodiments, the first position information of the non-player character and the second position information of the host virtual character are position information determined based on a same preset coordinate system. Optionally, the preset coordinate system can be implemented as a world coordinate system corresponding to the three-dimensional virtual environment, or the preset coordinate system can be implemented as a coordinate system established with the host virtual character as an origin.

[0421] In some embodiments, the first tone adjustment parameter is generated when the position distance between the first position information and the second position information reaches a first threshold, and the second tone adjustment parameter is generated when the position distance between the first position information and the second position information does not reach the first threshold, wherein the tone level corresponding to the first tone adjustment parameter is higher than the tone level corresponding to the second tone adjustment parameter.

[0422] In some embodiments, the first tone adjustment parameter and the second tone adjustment parameter can be preset parameters; in other embodiments, the first tone adjustment parameter can be determined based on a first control function, and the second tone adjustment parameter can be determined based on a second control function. Optionally, the first control function and the second control function can be implemented as the same function or different functions.

[0423] In other embodiments, the adjustment parameter can also be implemented as a volume control parameter. Illustratively, the first position information of the non-player character in the three-dimensional virtual environment is obtained, and the second position information of the host virtual character in the three-dimensional virtual environment is obtained. When the position distance between the first position information and the second position information reaches a first threshold, a volume control parameter is generated, and an audio effect parameter is determined based on the volume control parameter, wherein the volume control parameter is used to adjust the volume of the human voice audio, so that the volume of the human voice audio is positively correlated with the position distance.

[0424] In some embodiments, the first volume control parameter is generated when the position distance between the first position information and the second position information reaches a first threshold, and the second volume control parameter is generated when the position distance between the first position information and the second position information does not reach the first threshold, wherein the volume corresponding to the first volume control parameter is greater than the volume corresponding to the second volume control parameter.

[0425] In some embodiments, the first volume control parameter and the second volume control parameter can be preset parameters; in other embodiments, the first volume control parameter can be determined based on a third control function, and the second volume control parameter can be determined based on a fourth control function. Optionally, the third control function and the fourth control function can be implemented as the same function or different functions.

[0426] In one example, when the host virtual character is located outside the virtual building and the non-player character is located on the second floor of the virtual building, the speech tone and / or the speech volume of the human voice audio of the non-player character is boosted by the volume control parameter, so that the user can receive clearer feedback voice of the non-player character.

[0427] Secondly, when the non-player character is a continuously moving virtual character, the scene change of the three-dimensional virtual environment during the movement is detected, and the corresponding scene sound effect parameter is generated based on the scene change, so as to determine the sound effect parameter.

[0428] Illustratively, the movement path of the non-player character in the three-dimensional virtual environment is obtained; in the case that the movement path indicates that the non-player character moves between a plurality of virtual sub-scenes, the scene sound effect parameters corresponding to the plurality of virtual sub-scenes are obtained respectively; and the sound effect parameter is determined based on the scene sound effect parameters.

[0429] In some embodiments, the three-dimensional virtual environment is a scene composed of a plurality of virtual sub-scenes, and the three-dimensional virtual environment is connected between the virtual sub-scenes, for example, a virtual character gets off a virtual vehicle, walks along a path and enters a virtual building, and the plurality of virtual sub-scenes include a vehicle interior scene, a road scene and a building interior scene.

[0430] In some embodiments, different virtual sub-scenes are pre-configured with different scene sound effect parameters, for example, when the virtual sub-scene is a vehicle interior scene, since the vehicle is a small closed space, the corresponding scene sound effect parameter will process the human voice audio as a small volume and a low and dull tone effect.

[0431] In some embodiments, after obtaining the scene sound effect parameters corresponding to different virtual sub-scenes, the scene change nodes in the movement path are determined based on the movement path, the time points of the scene changes in the speech time axis are determined based on the speech time axis of the human voice audio, the movement time axis corresponding to the movement path and the scene change nodes, the different time periods in the human voice audio are assigned with the scene sound effect parameters corresponding to the corresponding virtual sub-scenes based on the time points of the scene changes, and the parameter effective time stamps of the different scene sound effect parameters are recorded in the process of determining the sound effect parameter based on the scene sound effect parameters.

[0432] Thirdly, the tone adjustment parameter is generated based on the environmental sound information corresponding to the natural language command, so as to determine the sound effect parameter.

[0433] Illustratively, the environmental sound information corresponding to the natural language command is obtained, the environmental sound information is used to indicate the information related to the background environmental sound in the voice when the voice including the natural language command is collected; the tone adjustment parameter of the human voice audio is generated based on the environmental sound information; and the sound effect parameter is determined based on the tone adjustment parameter.

[0434] The natural language command is a terminal-recorded user voice. After receiving the natural language command, the server detects the environmental sound information corresponding to the natural language command. Illustratively, the user's voice content is filtered out from the sound spectrum corresponding to the natural language command to obtain environmental sound data. The environmental sound data is detected to determine the corresponding environmental sound information.

[0435] In some embodiments, the environmental sound information includes an environmental sound volume, and the tone adjustment parameter corresponds to a tone height that is positively correlated with the size of the environmental sound volume. Illustratively, the environmental sound volume corresponding to the environmental sound data is detected. In a case where the environmental sound volume is less than a preset volume threshold, a third tone adjustment parameter is generated. In a case where the environmental sound volume is greater than or equal to the preset volume threshold, a fourth tone adjustment parameter is generated. The tone of the third tone adjustment parameter is lower than the tone of the fourth tone adjustment parameter.

[0436] In other embodiments, the environmental sound information includes an environmental recognition type corresponding to a background environmental sound, and a fifth tone adjustment parameter is generated based on the environmental recognition type. Illustratively, different environmental recognition types are preconfigured with different tone adjustment parameters.

[0437] In one example, when the environmental recognition type is a cybercafe type, the tone adjustment parameter corresponds to a first tone. When the environmental recognition type is a daytime bedroom, the tone adjustment parameter corresponds to a second tone. When the environmental recognition type is a nighttime bedroom, the tone adjustment parameter corresponds to a third tone. The first tone is higher than the second tone, and the second tone is higher than the third tone. That is, in a relatively noisy environment, the tone of the dialogue voice replied by the non-player character is higher, and in a relatively quiet environment, the tone of the dialogue voice replied by the non-player character is lower.

[0438] In other embodiments, the environmental sound information also affects the volume of the human voice audio. Illustratively, the environmental sound information corresponding to the natural language command is obtained, and a volume adjustment parameter of the human voice audio is generated based on the environmental sound information. The sound effect parameter is determined based on the volume adjustment parameter. Illustratively, the environmental sound volume corresponding to the environmental sound data is detected. In a case where the environmental sound volume is less than a preset volume threshold, it is determined that the natural language command corresponds to a first environmental sound state. In a case where the environmental sound volume is greater than or equal to the preset volume threshold, it is determined that the natural language command corresponds to a second environmental sound state. A first volume adjustment parameter is obtained based on the first environmental sound state, and a second volume adjustment parameter is obtained based on the second environmental sound state. The volume size corresponding to the first volume adjustment parameter is smaller than the volume size corresponding to the second volume adjustment parameter.

[0439] That is, when the environmental sound state corresponding to the natural language command indicates that the environmental sound is relatively noisy, the volume adjustment parameter is increased to increase the volume of the human voice audio, so that the user can better receive the reply result of the non-player character.

[0440] In some embodiments, in order to ensure the accuracy of the output human voice audio, the speech duration of the human voice audio is synchronized in a static synchronization manner. Illustratively, before generating the human voice audio, the speech speed setting information corresponding to the non-player character is obtained, wherein the speech speed setting information is used to indicate the speech speed setting condition of the non-player character, the estimated playing duration corresponding to the feedback text is determined based on the speech speed setting information, the speech duration corresponding to the human voice audio is obtained, and the playing speed of the human voice audio is adjusted based on the difference between the speech duration and the estimated duration.

[0441] That is, in the static synchronization execution process, the time length of each text segment in the feedback text is calculated in advance before generating the human voice audio, and it is ensured that the generated speech segment matches it. Specifically, the input segment text is analyzed, and the speech length of each segment text is estimated. In an example, the estimation can be based on a preset speech speed (such as the number of words per second), thereby assigning an estimated time length to each segment text. After generating the human voice audio, the difference between the actual generated human voice audio and the estimated time length is corrected to ensure the accurate synchronization of the speech and the text, so that the speech performance of the final human voice audio is more natural, and the readability of the human voice audio is ensured to be synchronized with the text.

[0442] In some embodiments, when the human voice audio is a reply voice generated based on a natural language command of a first account for a non-player character, the real-time and continuity of the human voice audio need to be ensured when the generated human voice audio is played to feedback the natural language command, and therefore, dynamic speech data synchronization needs to be ensured.

[0443] Illustratively, a natural language command of a first account for a non-player character is received, wherein the first account is an account for controlling a virtual character. A plurality of text segments of the non-player character are continuously generated based on the natural language command to form a feedback text, and a first start timestamp and a first end timestamp of each text segment are recorded. In a case where an i-th text segment is generated, an i-th character segment voice matching the character feature information is generated based on the i-th text segment as the human voice audio, and a second start timestamp and a second end timestamp of the i-th character segment voice are recorded. i is a positive integer. A first difference between the first start timestamp and the second start timestamp is determined, and a second difference between the first end timestamp and the second end timestamp is determined.

[0444] Optionally, the first generation speed of the i+1th text segment is adjusted based on the difference between the first difference and the second difference; or the second generation speed of the i+1th character segment voice is adjusted based on the difference between the first difference and the second difference; or the transmission speed of the scene audio data corresponding to the i th text segment to the terminal is adjusted based on the difference between the first difference and the second difference.

[0445] That is, dynamic synchronization is achieved based on real-time adjustment of voice generation speed and text processing speed in the process of generating human voice audio. In the process of voice generation, the length of the generated voice and the progress of text processing are detected in real time, and based on the real-time detection data, the voice generation speed and / or the text processing speed can be dynamically adjusted. Through the feedback mechanism, it is ensured that the voice generation and text processing are consistent in time.

[0446] Specifically, a timer or clock is used to record the time of voice generation and text processing in real time, and based on the real-time detection data, the voice generation speed or text processing speed is dynamically adjusted. The voice generation speed is adjusted by changing the speech rate parameter of the Text-to-Speech (TTS) engine.

[0447] In other embodiments, the buffer can also be managed to control the transmission speed of the generated human voice audio, wherein the buffer is used to store the generated character segment voice and text segment to ensure that they can be output on time. The generated character segment voice and text segment are stored in the buffer. The character segment voice and text segment are output from the buffer on time. In one example, a circular buffer is used to manage data storage and output to ensure that data is output on time. A precise timestamp is added to each text segment and character segment voice during the generation of the text segment and character segment voice, recording its start and end time, and the timestamp is recorded in the data structure. The text processing module for generating text segments and the voice generation module for generating character segment voices exchange time information in real time. The server detects the state of text processing and voice generation in real time during communication interaction to ensure that they are consistent in time, and detects the time error of text processing and voice generation in real time. The transmission speed of the human voice audio corresponding to the scene audio data from the circular buffer to the client is controlled by calculating the error based on the timestamp information in the circular buffer and adjusting the system parameters in real time, to ensure time synchronization.

[0448] By recording the corresponding timestamp of text generation and the corresponding timestamp of voice generation to dynamically adjust the generation speed of text / voice and the transmission speed of generated voice, the real-time and continuity of output voice can be ensured when the non-player character issues a natural language command to the user for feedback, avoiding the problem of stuttering of output human voice audio, and improving the time accuracy of data output.

[0449] In summary, when the player seeks voice feedback of the non-player character through the natural language command, the sound effect parameter of the non-player character is acquired according to the relative positional relationship between the host virtual character and the non-player character, and the spatial voice frequency of the non-player character is adjusted through the sound effect parameter. Since the spatial voice frequency is generated in real time based on the relative positional relationship between the host virtual character and the non-player character, the application program providing the three-dimensional virtual environment does not need to pre-configure too many application resources in order to improve the diversity of sound effect performance, thereby reducing the waste of storage space of the computer device running the above application program.

[0450] Meanwhile, the voice frequency corresponding to the non-player character after being adjusted through the sound effect parameter can make the spatial voice frequency played by the client simulate the perception effect of the audio generated by the host virtual character to the non-player character under the current relative positional relationship, so that the spatial voice frequency played by the client can make the player perceive the voice of the non-player character in the three-dimensional virtual environment as the host virtual character, thereby improving the diversity of the sound effect perceived by the player in the game process, thereby improving the immersion of the player in the three-dimensional virtual environment and improving the user experience.

[0451] In the embodiments of the present application, the sound effect parameter is applied to the voice frequency of the non-player character in the three-dimensional virtual environment, so that the user receives the voice frequency with the effect after propagation in the three-dimensional virtual environment, thereby improving the sound effect authenticity of the voice frequency of the non-player character when the non-player character dialogues with the host virtual character, and improving the user experience when the user perceives the voice frequency of the non-player character.

[0452] An exemplary embodiment of the present application provides a broadcasting method of a non-player character. The method comprises:

[0453] Step 1, displaying at least one of a host virtual character and an NPC in a virtual environment;

[0454] Exemplarily, the host virtual character is a virtual character directly controlled by the user in the virtual environment. There is one or more non-player characters (NPC) in the virtual environment; the NPC and the host virtual character belong to the same virtual camp and are teammates of the host virtual character, and the NPC follows the host virtual character to move in the virtual environment. In addition, the NPC can also be a follower character, a pet character, etc. controlled by the host virtual character. In addition, the NPC can also be a neutral character and only performs cooperative action with the host virtual character under certain conditions.

[0455] Step 2, acquiring a natural language command;

[0456] The natural language command includes at least one of a behavior intention and a scene entity. The behavior intention is used to indicate a type of virtual activity performed by the NPC, such as indicating a type of virtual activity performed in the virtual environment. The scene entity includes a first scene object, such as indicating a scene object in the virtual environment to which the virtual activity is performed.

[0457] In one example, the natural language command controls the NPC from two aspects of a type of virtual activity performed and a scene object in the virtual environment to which the virtual activity is performed. The natural language command is a complex instruction for controlling the NPC.

[0458] Step 3, in response to the natural language command, controlling the NPC to perform the virtual activity.

[0459] In some examples, the virtual activity is determined based on environment perception information of the host virtual character and / or the NPC. In some examples, the host virtual character and / or the NPC obtains the environment perception information within a perception range of a visual, an auditory, a movement trajectory, and the like. In some examples, the virtual activity is close to a current perception situation of the virtual environment by the host virtual character and / or the NPC. In some examples, the NPC is controlled based on the current perception situation of the virtual environment by the host virtual character and / or the NPC. In some examples, the NPC has autonomous behavior capability, and the user only commands the NPC. The user gives a command, and the NPC understands and executes the command based on its autonomous behavior capability.

[0460] In summary, the complex instruction for controlling the NPC from two aspects of a type of virtual activity performed and a scene object in the virtual environment to which the virtual activity is performed is provided. The virtual activity to be performed is determined based on environment perception information of the host virtual character and / or the NPC. The NPC is guaranteed to perform the virtual activity in cooperation with the host virtual character based on the environment perception information of the host virtual character and / or the NPC.

[0461] In an optional embodiment, the server is further provided with a command analysis function for the natural language command. The command analysis function includes entity recognition, intention recognition, and hot word detection in the natural language command.

[0462] FIG. 18 shows a flowchart of a non-player character broadcasting method according to an example embodiment of the present application. The method can be performed by the terminal in FIG. 1. The method includes:

[0463] Step 201: obtaining a spatial dataset of a virtual environment;

[0464] The spatial dataset of the virtual environment includes visual labels of scene objects, the visual labels being used to describe visual features of the scene objects in at least one dimension, such as in the dimensions of material, transparency, color, shape, etc.

[0465] Step 202: displaying at least one of a master virtual character and an NPC in the virtual environment;

[0466] Illustratively, the master virtual character is a virtual character directly controlled by a user in the virtual environment. There is also one or more NPCs in the virtual environment; the NPC and the master virtual character belong to the same virtual camp and are teammates of the master virtual character.

[0467] Step 203: obtaining a natural language command in the form of speech;

[0468] The natural language command includes natural semantics for directing the NPC, and the command information includes a behavior intention and a scene entity; the behavior intention is used to indicate a type of virtual activity performed by the NPC, and the scene entity is used to indicate a target entity to which the virtual activity is directed;

[0469] Step 204: converting the natural language command in the form of speech into a natural language command in the form of text;

[0470] Illustratively, Automatic Speech Recognition (ASR) processing is performed on the natural language command in the form of speech to determine the natural language command in the form of text. The ASR processing generally includes invoking acoustic models, language models, etc. components to realize recognition of pronunciation, vocabulary and syntax structure, etc. information of the natural language command in the form of speech, and conversion into the natural language command in the form of text.

[0471] Step 205: performing intention recognition on the natural language command to obtain a first classification label;

[0472] Illustratively, the first intention corresponding to the first classification label indicates a command intention for a non-player character in one intention dimension; such as at least one of the following: indicating which non-player character of a plurality of non-player characters is directed by the command text, indicating whether a virtual activity performed by the non-player character is related to a virtual attack, and indicating what kind of virtual activity is performed by the non-player character.

[0473] Step 206: performing entity recognition on the natural language command to obtain a target entity;

[0474] The target entity is determined from the virtual environment in combination with environmental perception information, the environmental perception information including information perceived by at least one of the virtual role controlled by the player and the non-player role from the virtual environment. The target entity can be an entity currently seen by the player, or an entity historically seen by the player.

[0475] Step 207: in response to the natural language command, controlling the NPC to perform a virtual activity according to a first intent corresponding to the first classification label, or controlling the NPC to perform a virtual activity associated with the target entity, or controlling the NPC to perform a virtual activity associated with the target entity according to a first intent corresponding to the first classification label;

[0476] For example, the NPC is controlled to perform a virtual activity according to the first classification label and / or the indication of the target entity.

[0477] Step 208: broadcasting feedback information of the NPC;

[0478] For example, the game application can generate corresponding feedback information in real time according to the environmental perception information of the non-player role, and broadcast it. For example, the game application can generate corresponding feedback information in real time according to the environmental perception information triggering the broadcast condition when the environmental perception information of the non-player role triggers the broadcast condition; or the game application can generate corresponding feedback information in real time according to the environmental perception information of the non-player role when the non-player role or the virtual role controlled by the player triggers the broadcast condition.

[0479] The preprocessing stage of step 201 can be implemented as:

[0480] Substep 1: obtaining attribute text of a scene object in a virtual scene;

[0481] For example, the attribute text is used to introduce the inherent attributes of the scene object in the virtual scene; on the one hand, the attribute text realizes the description of the scene object in the form of text, which describes the scene object and provides semantic information for the prediction of the visual label of the scene object. On the other hand, the attribute text is the description of the scene object in the virtual scene, and in the case of a large number of scene objects in the virtual scene and the reuse of object models, it can more accurately describe the inherent attributes of the scene object in the virtual scene. In an optional implementation, the attribute text includes at least one of the name of the scene object in the virtual scene and the size of the scene object in the virtual scene.

[0482] Substep 2: obtaining appearance images of the scene object in the virtual scene;

[0483] Exemplarily, the appearance image is used to describe the style of the scene object; the appearance image carries the appearance style of the scene object such as color, texture, shape and mutual positional relationship between each subpart on the picture modality, and can comprehensively describe the scene object from the picture modality (or visual modality). In an optional implementation manner, the appearance image of the scene object includes images obtained by observing the scene object from at least two perspectives.

[0484] Substep 3: calling the multi-modal model to perform prediction on the attribute text and the appearance image of the scene object to obtain a visual label of the scene object;

[0485] Exemplarily, the multi-modal model has the capability of performing model prediction on text information and picture information of different modalities. In the embodiment, the input parameter of the multi-modal model is the attribute text and the appearance image of the scene object; the multi-modal model predicts the visual label of the scene object from two modalities of the text modality and the picture modality, and the visual label is used to describe the visual feature of the scene object in at least one dimension.

[0486] Optionally, the multi-modal model includes a visual question and answer model; the attribute text of the scene object is carried in the question sentence; the question sentence is used to guide the visual question and answer model to convert the appearance image into the visual label of the scene object, and at the same time, the question sentence provides supplementary information of the scene object in a text manner.

[0487] Obtaining expected information of the scene object;

[0488] Exemplarily, the expected information is used to indicate the description dimension of the scene object expected in the visual label, and / or the expected format of the visual label; in an example, the expected information used to indicate the description dimension of the scene object expected in the visual label includes but is not limited to at least one of type description, material, transparency, color, surface feature and shape of the scene object. In another example, the expected information used to indicate the expected format of the visual label is at least one of Comma Separated Values (CSV), JavaScript Object Notation (JSON) and eXtensible Markup Language (XML).

[0489] Constructing a question sentence of the appearance image according to the expected information and the attribute text;

[0490] Exemplarily, the first subpart in the question sentence is supplementary introduction information of the scene object, carrying the attribute text of the scene object; the second subpart in the question sentence is an answer guiding sentence for the visual question and answer model, carrying the expected information.

[0491] Optionally, the spatial data set comprises a matching label of the scene object in the virtual scene; accordingly: performing paralinguistic rewriting on the visual label of the scene object to obtain a matching label conforming to the natural language spoken expression.

[0492] For example, as introduced above, the visual label is used to describe the visual features of the scene object in at least one dimension, and the appearance image of the scene object presents rich visual features of the scene object. However, the description of the scene object in the natural language spoken expression cannot cover the visual features of the scene object in each dimension. The purpose of performing paralinguistic rewriting on the visual label of the scene object is to obtain a matching label that is more similar to the spoken expression; it can be understood that performing paralinguistic rewriting can be deleting part of the content in the visual label, or changing the visual label to a label with the same semantics but different text expression.

[0493] In an optional implementation, the paralinguistic rewriting is performed by calling a natural language model; the visual label of the scene object is input into the natural language model to predict a matching label conforming to the natural language spoken expression. The paralinguistic rewriting on the visual label is realized based on calling the natural language model. For example, the natural language model carries prior knowledge of the natural language spoken expression. For example, the natural language model is implemented as a large language model (LLM).

[0494] Optionally, the spatial data set further comprises spatial information of the scene object, and accordingly further comprises: obtaining the spatial position of the scene object in the virtual scene, and determining the spatial position of the scene object as auxiliary information of the visual label of the scene object.

[0495] For example, the spatial position of the scene object in the virtual scene is used to indicate the deployment situation of the scene object in the virtual scene, and the spatial information is used to indicate the position, size, etc. of the scene object in the virtual scene after the scene object is deployed in the virtual scene.

[0496] Further, at least one of the coordinate position, orientation information, bounding box information, and shelter point information of the scene object in the virtual scene is obtained.

[0497] The coordinate position (Location) is used to indicate the position of the scene object in the virtual scene, such as the coordinate information of the center point or the preset point of the scene object in the virtual scene; the orientation (Rotation) is used to indicate the direction faced by the scene object in the virtual scene, such as the direction faced by the front of the scene object in the virtual scene; the bounding box (Bounding Box) is used to indicate the size of the scene object in the virtual scene; and the cover points (Cover points) are used to indicate the recommended position points of the virtual character when the virtual character approaches the scene object, so that the scene object can mask the virtual character.

[0498] The intention recognition stage for step 205 can be implemented as follows:

[0499] Substep 4: calling at least one hierarchical prediction network to perform intention recognition on the natural language command to obtain a first classification label of the natural language command commanding the NPC;

[0500] For example, the hierarchical prediction network has the ability to predict the classification label corresponding to the command text. For example, the hierarchical prediction network includes at least two subnetworks, an upper network and a lower network, which are cascaded, and the lower network further performs classification label prediction according to the prediction result output by the upper network. The at least two subnetworks are constructed based on a tree structure of multiple classification labels, and the hierarchical prediction network constructs at least two subnetworks of corresponding hierarchical structure for the tree structure of classification labels, and divides the prediction task of a variety of classification labels into prediction subtasks performed by the at least two subnetworks, so as to reduce the complexity of classification prediction of each subnetwork.

[0501] In an optional implementation, each hierarchical prediction network in the at least one hierarchical prediction network is used to predict the commanding intention of the command text for the non-player character in one intention dimension; the intention dimension includes at least one of the subject dimension, the semantic dimension, and the behavior dimension. For example, in the at least one hierarchical prediction network, the i-th subnetwork in each hierarchical prediction network is used to predict a first-level behavior label of the command text in one intention dimension, and the i+1-th subnetwork in the hierarchical prediction network is used to predict a second-level behavior label corresponding to the first-level behavior label in the command text, where i is a positive integer; the high-level behavior label (such as the first-level behavior label) predicted by the high-level subnetwork (such as the i-th subnetwork) in the hierarchical prediction network includes multiple sub-labels, or subordinate lower-level labels; and a corresponding lower-level subnetwork (such as the i+1-th subnetwork) needs to be called to further perform classification prediction (such as to predict the second-level behavior label). The prediction task of a variety of classification labels is divided into prediction subtasks performed by the at least two subnetworks, so as to reduce the complexity of classification prediction of each subnetwork.

[0502] Further, the intent dimension comprises a subject dimension, and the hierarchical prediction network comprises a hierarchical subject prediction network, the subject prediction network having the capability of predicting a subject type in the natural language command; the subject type is used to indicate the identity of the NPC commanded by the natural language command; for example, the subject prediction network is used to predict which of the plurality of non-player characters is commanded by the command text, and the number of non-player characters commanded by the command text can be one or more.

[0503] Further, the intent dimension comprises a semantic dimension, and the hierarchical prediction network comprises a hierarchical semantic prediction network, the semantic prediction network having the capability of predicting a semantic type in the natural language command; the semantic type is used to indicate the control manner of the NPC for initiating a virtual attack; for example, the subject prediction network is used to predict whether the virtual activity performed by the non-player character commanded by the command text is related to initiating a virtual attack.

[0504] Further, the intent dimension comprises a behavior dimension, and the hierarchical prediction network comprises a hierarchical behavior prediction network, the behavior prediction network having the capability of predicting a behavior intent type in the natural language command; the behavior intent type is used to indicate the behavior manner of the NPC for performing a virtual activity. For example, the behavior prediction network is used to predict what virtual activity is performed by the non-player character.

[0505] Correspondingly, in step 207, in response to the natural language command, the NPC is controlled to perform a virtual activity according to a first intent corresponding to the first classification label;

[0506] For example, the first intent corresponding to the first classification label indicates a command intent for the non-player character in one intent dimension; such as at least one of the following: indicating which of the plurality of non-player characters is commanded by the command text, indicating whether the virtual activity performed by the non-player character is related to initiating a virtual attack, and indicating what virtual activity is performed by the non-player character. The NPC is controlled to perform a virtual activity according to the indication of the first classification label.

[0507] The target entity recognition process for step 206 can be implemented as follows:

[0508] Sub-step 5: querying the target entity from the entity set of the virtual environment according to the target entity information of the target entity indicated by the natural language command and the environment perception information;

[0509] The target entity is determined from the virtual environment in combination with the environment perception information, and the environment perception information includes information perceived by the master virtual character and / or the non-player character from the virtual environment.

[0510] Since the target entity information in the natural language command is usually ambiguous, for example, the natural language command can be "move to the red truck", and the target entity information therein is "red truck", and there can be many red trucks in the virtual environment, and the target entity cannot be accurately determined from the virtual environment according to the target entity information in the natural language command.

[0511] Therefore, the embodiments of the present application provide a method of determining a target entity in combination with target entity information and environment perception information. For the above example, the "red truck" expressed by the player in the natural language command should be a red truck that the player can see, and therefore, in combination with the visual range of the virtual character controlled by the player, the red truck located in the visual range of the virtual character controlled by the player can be filtered from the multiple red trucks in the virtual environment, and the red truck is the target entity indicated by the player in the natural language command.

[0512] Since the natural language command is issued by the player based on the perception of the virtual environment, in order to accurately identify the target entity indicated in the natural language command, the game application program will combine the environment perception information of the player when issuing the natural language command to identify the target entity in the natural language command. Based on the perception of the virtual environment by the player, the entity closest to the target entity information that the player can perceive in the virtual environment is inferred as the target entity.

[0513] For example, the virtual environment includes multiple candidate entities matching the natural language command, and the target entity is an entity filtered from the multiple candidate entities based on the environment perception information of the virtual character controlled by the player or the non-player character. For example, the game application program first selects multiple candidate entities matching the target entity information in the natural language command from the entity set, and then selects an entity that can be perceived by the virtual character controlled by the player or the non-player character from the multiple candidate entities as the target entity according to the environment perception information.

[0514] For example, the environment perception information of the virtual character controlled by the player can include: a picture or entity that can be seen by the virtual character controlled by the player when observing the virtual environment, an environmental sound that can be heard by the virtual character controlled by the player, a source direction of the environmental sound, a type and sound size of the environmental sound, perception information obtained by the virtual character controlled by the player by using a perception skill (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and perception information obtained by the virtual character controlled by the player by using a sensing device (for example, a sensing signal of a sensor, a position detection signal of a position detection prop, etc.).

[0515] The environment perception information of the non-player character can include: a picture or entity that the non-player character can see when observing the virtual environment, an environmental sound that the non-player character can hear, a source direction of the environmental sound, a type and sound size of the environmental sound, perception information (for example, visual information, auditory information, sound wave information, light reflection information, etc.) obtained by the non-player character using a perception skill, and perception information (for example, a sensing signal of a sensor, a position detection signal of a position detection prop, etc.) obtained by the non-player character using a perception device.

[0516] It should be noted that the environment perception information can include information that the host virtual character and the non-player character perceive in real time when the natural language command is received, or information that the host virtual character and the non-player character perceive historically before the natural language command is received. That is, the target entity can be an entity that the player currently sees, or an entity that the player has historically seen. Therefore, the game application needs to combine the real-time environment perception information and the historical environment perception information of the host virtual character and / or the non-player character to identify the target entity.

[0517] For example, when the natural language command is “Let’s go back to the hotel we just passed by”, the game application needs to query the hotel that the host virtual character has passed by according to the historical environment perception information of the host virtual character.

[0518] In an optional embodiment, the client queries the target entity from the entity set of the virtual environment according to the target entity information indicated by the natural language command and the environment perception information. In another optional embodiment, the client reports the natural language command to the server, and the server queries the target entity from the entity set of the virtual environment according to the target entity information indicated by the natural language command and the environment perception information.

[0519] Optionally, the game application infers the target entity according to the target entity information, the environment perception information, and preprocessed virtual environment data. The preprocessed virtual environment data includes the entity set of the virtual environment.

[0520] The entity set includes entity information of each entity in the virtual environment. The entity information includes at least one of the following: a name, a type, a position, a feature, a text label, an embedding vector of the text label, an image, and an embedding vector of the image.

[0521] The text label can be at least one word obtained by tokenizing the feature (for example, a description text of the appearance of the entity), and the embedding model is called to obtain an embedding vector corresponding to each text label.

[0522] The image can include an image obtained by observing a three-dimensional model of the entity from at least one direction, for example, the image can include a three-view image of the entity. The embedding vector of the image is an embedding vector of an image feature label. The image feature label is obtained by calling a multi-modal model to perform feature recognition based on the image and the text label of the entity.

[0523] For example, the entity set includes first-class entities and second-class entities, the first-class entities include at least one of a region and a building, and the second-class entities include entity objects. The text label of the first-class entity is manually labeled, for example, the text label of a certain room of a certain building is manually labeled as: a certain region, a certain building, a certain floor, and a room. The text label of the second-class entity is automatically generated by calling a pre-trained large language model.

[0524] The method for generating the text label of the second-class entity can include: obtaining at least one image of the entity, inputting the at least one image into the large language model to obtain a description text of the entity; performing word segmentation on the description text to obtain at least one text label of the entity, each text label including one word obtained after word segmentation; and performing a vectorization operation on the text label to obtain an embedding vector corresponding to the text label.

[0525] The embedding vector is a high-dimensional vector data of the text label, and the embedding vector of the text label is generated in advance, so that it is not necessary to repeatedly generate the vector of the text label of each entity when matching the target entity based on the target entity information from the entity set, and the matching efficiency of the target entity information and the entity information can be improved.

[0526] It should be noted that the method does not directly use the description text as the text label, but uses the word segmentation result of the description text as the text label, because the description text is usually of different lengths, and for a longer description text, the embedding vector thereof has a poor effect when actually searching and comparing the similarity. For example, when searching for a “truck” using “a blue small rusty truck”, the “a blue small rusty truck” is probably ranked after “a car” in the recall result, but actually, for the search of “truck”, no matter how complex the additional description is, the “truck” should be searched first, and therefore, when performing the similarity search, the comparison should be performed on the feature words one by one, rather than on the description text.

[0527] In addition, in order to further extract the image features of the entity, the method provided in the embodiments of the present application can further extract hidden visual features in the image of the entity based on the text label and the image of the entity. Taking the first entity as an example, the method includes: obtaining at least one perspective image of the first entity; obtaining at least one text label of the first entity; calling a multi-modal model to extract image features of the first entity based on the at least one perspective image and the at least one text label of the first entity, and obtaining an image embedding vector of the first entity.

[0528] Since the text label is manually annotated or generated by a large language model based on human language characteristics, the text label extracted based on human language habits may ignore some features of the entity. For example, when manually annotating the text label for an oil drum, the more noticeable text labels such as "metal" and "rust" may be annotated, and the detailed features such as "rust", "blue", "yellow", and "right lower corner paint drop" may be ignored. Therefore, the method provided in the embodiments of the present application also uses a multi-modal model to extract more image features from the text label and the entity image of the entity based on the multi-modal model, and performs image similarity matching with the target entity based on the image features to improve the recognition accuracy of the target entity.

[0529] The training method of the multi-modal model can be: inputting the sample entity image and the sample label into the pre-trained multi-modal model, fine-tuning the pre-trained multi-modal model according to the loss of the predicted label output by the pre-trained multi-modal model and the sample label, obtaining a multi-modal model capable of outputting image feature labels based on input text labels and images, and the image feature labels include more detailed description texts extracted from the image. For example, the text description of "a metal oil drum" can only be associated with the two text labels "metal" and "oil drum", but the Clip (multi-modal) model can also identify hidden visual information in the image of the oil drum, such as the image feature labels "rust", "blue", and "yellow". These image features do not appear in the text label, so combining the Clip model for image feature search can further improve the accuracy of entity search.

[0530] In an optional embodiment, the game application program can query the target entity by using the following method.

[0531] (1) Analyzing the natural language command to obtain target entity information; the target entity information includes at least one of the following: entity type, entity name, entity position, and entity feature.

[0532] Optionally, the game application program calls a pre-trained large language model to analyze the natural language command to obtain the target entity information of the target entity. The embedding model is called to perform vectorization processing on the target entity information to obtain a target embedding vector of the target entity information.

[0533] For example, the natural language command is "come here in front of the blue truck", and the target entity information analyzed by the large language model includes the position information "in front of" and the entity name "blue truck".

[0534] Exemplarily, a named entity recognition technique in the field of natural language processing can also be called to extract target entity information from the natural language command. The named entity recognition technique is used to identify entities with specific meanings in text, for example, to identify names, place names, adjectives, and the like.

[0535] For example, the natural language command is "find a paper box behind the red sofa on the first floor of the motel". Using the named entity recognition technique, the building name "motel first floor", the article name "sofa" and "paper box", the adverb "behind", and the adjective "red" can be identified. After logical construction, a hierarchical scene query call form with search type, content, adverb, and constraint information is formed. Among them, the search type is used to narrow the data retrieval range based on similarity matching, the search content is the specific entity description, and the adjective is combined with the description text for query. The floor constraint is determined by the player's location, which needs to narrow the range up and down in the indoor scene, thereby avoiding finding entities that are not visible across floors.

[0536] (2) Calculate the similarity of the target entity information and the entity information of each entity in the entity set.

[0537] Exemplarily, the entity set includes entity information of at least one entity, and the entity information of each entity can include at least one of a text embedding vector of a text label and an embedding vector of an image feature. The target entity information can be calculated with the text embedding vector and the image embedding vector, respectively, and the entity with higher similarity is determined as the target entity.

[0538] The methods of similarity matching with the text embedding vector and the image embedding vector are given below, respectively.

[0539] 1) The similarity includes the text similarity of the target entity information and the text embedding vector.

[0540] Taking the first entity in the entity set as an example, the game application program (client or server) tokenizes the target entity information to obtain at least one target entity label; converts the at least one target entity label into at least one target embedding vector; obtains the entity information of the first entity, which includes a text embedding vector converted based on the text label of the first entity; calculates the text parent similarity of the at least one target embedding vector and the text embedding vector, respectively, to obtain at least one text parent similarity corresponding to the at least one target embedding vector, respectively; and determines the sum of the at least one text parent similarity as the text similarity of the target entity information and the entity information of the first entity.

[0541] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to text embedding vector 1, and entity 2 corresponds to text embedding vector 2. Then, similarity 1 of target entity label 1 and text embedding vector 1 is calculated, similarity 2 of target entity label 2 and text embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the text similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and text embedding vector 2 is calculated, similarity 4 of target entity label 2 and text embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the text similarity of target entity information and entity 2.

[0542] For example, the entity information of the first entity includes at least one text embedding vector, and the at least one target embedding vector includes a first target embedding vector. The game application calculates text sub-similarities of the first target embedding vector and the at least one text embedding vector respectively, and obtains at least one text sub-similarity. The highest value in the at least one text sub-similarity is determined as a text parent-similarity corresponding to the first target embedding vector.

[0543] For example, the number of target entity labels is one: target entity label 1. The number of entities in the entity set is also one: entity 1. Entity 1 corresponds to text embedding vector 1 and text embedding vector 3. Then, similarity 1 of target entity label 1 and text embedding vector 1 is calculated, similarity 5 of target entity label 1 and text embedding vector 3 is calculated, and the greater value of similarity 1 and similarity 5 is determined as the similarity of target entity information (target entity label 1) and entity 1.

[0544] For example, for the input target entity information "blue car", first perform word segmentation processing on the target entity information to decompose into target entity labels "blue" and "car", representing two features of the query target entity. Then, the similarity of each feature with each entity in the entity set is calculated, the maximum value of the similarity of each feature is taken, and finally the sum of the similarity scores corresponding to all query features is taken, which represents the text similarity of the target entity information and the queried entity.

[0545] For example, the target entity information contains two target entity labels "blue" and "car", which are converted into two target embedding vectors. Then, the entity information in the entity set is obtained, for example, the entity set includes entity 1 and entity 2, the text labels of entity 1 include "metal", "old", "car", "truck", "blue", and "damaged", and the text labels of entity 2 include "metal", "scratch", "car", "damaged", "money truck", and "black".

[0546] The similarity between the target entity information and the text of entity 1 is calculated: first, the similarity between the target entity label "blue" and each text label of entity 1 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "blue" and the text label "blue" of entity 1 is 1 (1 is the maximum value); then, the similarity between the target entity label "car" and each text label of entity 1 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "car" and the text label "car" of entity 1 is 1, and then the sum of the two similarities of the target entity labels "car" and "blue" is taken to obtain the final similarity between the target entity information and the entity 1 as 2.

[0547] Similarly, the similarity between the target entity information and the text of entity 2 is calculated: first, the similarity between the target entity label "blue" and each text label of entity 2 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "blue" and the text label "black" of entity 1 is 0.91; then, the similarity between the target entity label "car" and each text label of entity 1 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "car" and the text label "car" of entity 2 is 1 (1 is the maximum value), and then the sum of the two similarities of the target entity labels "car" and "blue" is taken to obtain the final similarity between the target entity information and the entity 2 as 1.91.

[0548] It can be seen that the similarity between the target entity information and entity 1 is 2, which is higher than the similarity between the target entity information and entity 2, which is 1.91.

[0549] 2) The similarity includes an image similarity between the target entity information and the image embedding vector.

[0550] Taking the first entity in the entity set as an example, the game application (client or server) tokenizes the target entity information to obtain at least one target entity label; converts the at least one target entity label into at least one target embedding vector; obtains the entity information of the first entity, and the entity information includes an image embedding vector, which is an embedding vector extracted based on the image of the first entity; calculates the image parent similarity between the at least one target embedding vector and the text embedding vector respectively to obtain at least one image parent similarity corresponding to the at least one target embedding vector respectively; and determines the sum of the at least one image parent similarity as the image similarity between the target entity information and the entity information of the first entity.

[0551] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to image embedding vector 1, and entity 2 corresponds to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the image similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the image similarity of target entity information and entity 2.

[0552] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to image embedding vector 1, and entity 2 corresponds to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the image similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the image similarity of target entity information and entity 2.

[0553] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to image embedding vector 1, and entity 2 corresponds to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the image similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the image similarity of target entity information and entity 2.

[0554] In an optional embodiment, the calculation of the image similarity can use a target image embedding vector, that is, the target image embedding vector is used instead of the target embedding vector in the above method. The target image embedding vector can be obtained by the following method: tokenizing the target entity information to obtain at least one target entity label; inputting the at least one target entity label into a multi-modal model to obtain the target image embedding vector.

[0555] 3) The similarity includes a text similarity and an image similarity.

[0556] In the case where the first entity and the target entity correspond to a text similarity and an image similarity, the game application determines the average of the text similarity and the image similarity as the similarity of the first entity and the target entity.

[0557] Alternatively, in the case where the first entity and the target entity correspond to a text similarity and an image similarity, the greater value of the text similarity and the image similarity is determined as the similarity of the first entity and the target entity.

[0558] Alternatively, in a case where the first entity and the target entity correspond to a text similarity and an image similarity, a sum of the text similarity and the image similarity is determined as the similarity of the first entity and the target entity.

[0559] (3) determining the target entity from the entity set according to the similarity and the environmental perception information.

[0560] In an optional embodiment, the game application first performs similarity searching according to the target entity information to obtain at least one candidate entity with a higher similarity, and sorts the candidate entities according to the similarity; then, based on the position, orientation, direction, etc. of the master virtual character and / or the non-player character, the game application performs filtering and sorting of the perception, and finally screens the target entity.

[0561] For example, the game application determines a target perception range according to the target entity information and the environmental perception information; and screens the target entity from the entities perceived within the target perception range according to the similarity.

[0562] Alternatively, the game application determines x entities with the highest similarity from the entity set according to the similarity to form a candidate entity list; and determines the target entity from the entities within the target perception range according to the target entity information and the environmental perception information. x is a positive integer.

[0563] For example, the game application determines the entity with the highest similarity within the perception range as the target entity.

[0564] It should be noted that when there are multiple entities with the same similarity within the target perception range, the multiple entities can be sorted according to the distance between the entities and the master virtual character. For example, in a case where the number of entities with the highest similarity within the perception range is at least two, the entity with the highest similarity within the perception range and closest to the master virtual character is determined as the target entity.

[0565] In an optional embodiment, the game application determines the entity with the highest similarity within the perception range and higher than a threshold value as the target entity. For example, the number of perception ranges is at least one. For example, the perception range includes at least one of the following: the visual range of the master virtual character, the auditory range of the master virtual character, the perception range of the master virtual character for perceiving virtual props, the visual range of the non-player character, the auditory range of the non-player character, the perception range of the non-player character for perceiving virtual props.

[0566] For example, the game application can sequentially traverse the at least one perception range described above, or the game application can sequentially traverse the at least one perception range determined according to the natural language command. For example, the natural language command is "Can you see the truck in front of you? Move to the truck", the game application can traverse the visual range of the master virtual role first, and then traverse the visual range of the non-player character, and match the entity with the highest similarity and higher than the threshold.

[0567] For example, when there are at least two perception ranges, the traversal order of the at least two perception ranges can be preset, or the traversal order of the at least two perception ranges can be determined according to the intention indicated in the natural language command. For example, when the intention is to perform a search for an item, the visual range is traversed first. When the intention is to find a combat scene, the auditory range is traversed first.

[0568] If there are multiple levels of nested instructions in the natural language command, the first target entity is searched first, and then the next target entity is searched according to the position of the previous target entity. In the case where the natural language command includes a second target entity determined based on the position of the first target entity, the first target entity is queried from the entity set of the virtual environment according to the first target entity information of the first target entity indicated by the natural language command and the environmental perception information; and then the second target entity is queried from the entity set of the virtual environment according to the second target entity information of the second target entity, the first target entity and the environmental perception information.

[0569] In response to the natural language command, the NPC performs a virtual activity associated with the target entity.

[0570] For example, a pre-trained large language model is called to identify the behavior intention of the natural language command, a control instruction or a control instruction sequence is generated according to the behavior intention, and the non-player character is controlled to complete an activity based on the target entity according to the control instruction or the control instruction sequence.

[0571] The NPC feedback information broadcast in step 208 can be implemented as:

[0572] Sub-step 6: Broadcast the feedback information of the non-player character, and the feedback text corresponding to the feedback information is non-fixed text generated based on the environmental perception information of the non-player character;

[0573] Exemplarily, the game application can generate corresponding feedback information in real time according to the environmental perception information of the non-player character, and play the feedback information. For example, the game application can generate corresponding feedback information in real time according to the environmental perception information triggering the play condition when the environmental perception information of the non-player character triggers the play condition; or the game application can generate corresponding feedback information in real time according to the environmental perception information of the non-player character when the non-player character or the master virtual character triggers the play condition.

[0574] Exemplarily, the feedback text is obtained by calling a pre-trained large language model to infer based on the static entity data of the three-dimensional virtual environment and the dynamic environmental perception information of the non-player character. For different perception situations of the non-player character, the large language model generates different feedback texts.

[0575] Exemplarily, the feedback text can also be obtained by the large language model inferring based on the static entity data of the three-dimensional virtual environment and the dynamic environmental perception information of the non-player character, and the feedback text conforms to the personality characteristics of the non-player character. For example, when the non-player character is a robust uncle image, the feedback text can adopt a bold tone; when the non-player character is a reporter, the feedback text can adopt a news report tone.

[0576] Among them, the environmental perception information can include information perceived by the non-player character in real time, and can also include information perceived by the non-player character historically.

[0577] Exemplarily, the feedback information is obtained by the game application inferring based on the environmental perception information of the non-player character and the static entity data of the three-dimensional virtual environment. The feedback information can be a voice form of play, or a text form of play.

[0578] Exemplarily, the feedback information can include at least one of the following: immediate feedback, execution feedback, and dynamic feedback. Among them, the immediate feedback and the execution feedback are feedback information generated for natural language commands, and the dynamic feedback is feedback information generated spontaneously by the non-player character.

[0579] 1. The immediate feedback is feedback information generated immediately when the natural language command is received. The immediate feedback can include reply play and immediate play. The reply play is used to recover the query in the natural language command; the immediate play is used to feed back the receiving situation of the natural language command.

[0580] In an optional embodiment, the feedback information includes reply play. The reply play includes a reply to the query natural language command generated for the query natural language command proposed by the player. For example, when the query natural language command includes a query of the position of the target entity, the corresponding reply play should include the query result of the position of the target entity.

[0581] In a case where the behavior intention of the received natural language command is inquiry, the terminal broadcasts a non-player character reply broadcast; the reply broadcast includes reply content to the natural language command generated according to environmental perception information of the non-player character to the three-dimensional virtual environment.

[0582] For example, a pre-trained large language model is called to parse the received natural language command to obtain the intention of the natural language command, and when the intention is an inquiry type, it can be determined that the natural language command is an inquiry natural language command. When the natural language command is an inquiry natural language command, the large language model can infer according to the static entity data and the environmental perception information according to the intention of the natural language command to obtain reply text. For example, when the large language model identifies that the intention of the natural language command is inquiry, since there are thousands of questions that the user may ask, it is impossible to store the reply to each question locally on the client side, so the game application program calls a text-to-speech service to generate a reply broadcast in real time according to the reply text returned by the large language model, so that the game application program can generate a corresponding reply broadcast in real time to reply to the player's question.

[0583] For example, the game application program (client or server) calls a pre-trained large language model to parse the inquiry natural language command, infers to obtain reply text based on the static environmental data and environmental perception information of the three-dimensional virtual environment, and generates human voice based on the reply text to obtain a reply broadcast.

[0584] For example, the client reports the received inquiry natural language command to the server, the server calls a pre-trained large language model to parse the inquiry natural language command, infers to obtain reply text based on the static environmental data and environmental perception information of the three-dimensional virtual environment, generates human voice based on the reply text to obtain a reply broadcast, the server sends the reply broadcast to the client, and the client broadcasts the reply broadcast.

[0585] 2. The execution feedback is generated when a task indicated by the natural language command is executed according to the intention of the natural language command after the natural language command is received. The execution feedback can also be referred to as execution broadcast.

[0586] In an optional embodiment, the feedback information includes feedback broadcast (which can also be referred to as execution feedback). The feedback broadcast includes a broadcast of the task execution state and the task execution result generated for the task natural language command proposed by the player. For example, when the task natural language command includes moving to a target entity, the corresponding feedback broadcast can include that the target entity is being moved to, or that the target entity has been reached.

[0587] The terminal broadcasts a feedback broadcast in a case where a behavior intention of the received natural language command is to perform a task; the feedback broadcast includes an execution state after the non-player character executes the natural language command according to the environmental perception information of the three-dimensional virtual environment.

[0588] The natural language command is used to instruct the non-player character to perform a task.

[0589] For example, a pre-trained large language model is called to parse the received natural language command to obtain an intention of the natural language command, and when the intention is a task, it can be determined that the natural language command is a task natural language command. When the natural language command is a task natural language command, the large language model can infer a task instruction sequence according to the static entity data and the environmental perception information according to the intention of the natural language command, and the task instruction sequence is used to control the non-player character to perform the task. The game application program controls the non-player character to perform the task according to the task instruction in the task instruction sequence when the task instruction sequence is received. During the control of the non-player character according to the task instruction, the corresponding execution state can be obtained according to the feedback broadcast logic corresponding to different task instructions to perform feedback broadcast.

[0590] For example, since the instructions that can be performed by the non-player character in the three-dimensional virtual environment are limited and traversable, the client can locally store the feedback broadcast corresponding to each character instruction, and the client can read the local feedback broadcast to perform voice broadcast when the non-player character performs the corresponding instruction.

[0591] Alternatively, when there is variable content in the feedback broadcast corresponding to a task instruction, for example, the target entity in the feedback broadcast is variable, or the position information in the feedback broadcast is variable; the client can also generate a broadcast voice of the variable content according to the target entity or the position information returned by the large language model, and splice the broadcast voice of the variable content with the broadcast voice of the non-variable content stored locally to obtain the final feedback broadcast.

[0592] For example, the game application program (client or server) calls a pre-trained large language model to parse a task natural language command, infers a task execution instruction based on the static environment data and the environmental perception information of the three-dimensional virtual environment, controls the non-player character to perform the task according to the task execution instruction, generates a feedback text according to the execution state of the non-player character performing the task, and generates a human voice audio based on the feedback text to obtain a feedback broadcast.

[0593] 3. Dynamic feedback is feedback information generated spontaneously according to the environmental perception information of the non-player character. Dynamic feedback can also be referred to as dynamic broadcast.

[0594] For example, the client can generate dynamic feedback based on the perception of the three-dimensional virtual environment without receiving a natural language command, which is used to indicate an abnormal situation discovered by the non-player character in the three-dimensional virtual environment. For example, the abnormal situation can include: discovering an enemy virtual character, discovering a change in the state of a friendly virtual character, discovering a dangerous situation, discovering a battle trace or a looting trace, etc.

[0595] In an optional embodiment, the feedback information includes a dynamic announcement (also referred to as "dynamic feedback"). The dynamic announcement includes an announcement of a perceived abnormal situation of the non-player character. For example, when the non-player character perceives that it is under attack, the announcement is that it is under attack; when the non-player character perceives that there is a dangerous situation ahead, the announcement is that there is a danger ahead.

[0596] The terminal announces the dynamic announcement when the environmental perception information of the non-player character in the three-dimensional virtual environment meets ...

Claims

A non-player character broadcasting method, the method is executed by a computer device, the method comprises: Displaying at least one of a master virtual character and a non-player character located in a three-dimensional virtual environment; Receiving a natural language command for the non-player character; In response to the natural language command, broadcasting feedback information of the non-player character, the corresponding feedback text of the feedback information is a non-fixed text generated based on environment perception information, and the environment perception information includes information perceived by at least one of the master virtual character and the non-player character in the three-dimensional virtual environment. The method of claim 1, wherein, The feedback information includes a reply broadcast; The response to the natural language command, broadcasting the feedback information of the non-player character, comprises: In the case where the behavior intention of the natural language command is inquiry, the reply broadcast of the non-player character is broadcasted; The reply broadcast includes reply content for the natural language command generated according to the environment perception information of the non-player character in the three-dimensional virtual environment. The method according to claim 1 or 2, wherein The reply broadcast of the non-player character in the case where the behavior intention of the natural language command is inquiry comprises at least one of: In the case where the natural language command includes an intelligence inquiry to a target entity, a first reply broadcast is broadcasted based on the perception of the non-player character to the target entity; the first reply broadcast includes the intelligence perception result of the non-player character to the target entity; In the case where the natural language command includes a position inquiry to a target entity, a second reply broadcast is broadcasted based on the perception of the non-player character to the target entity, and the second reply broadcast includes a position description of the target entity. The method according to any one of claims 1 to 3, wherein The method further comprises: Calling a pre-trained large language model to analyze the natural language command, and inferring a reply feedback text based on static environment data of the three-dimensional virtual environment and the environment perception information; Generating human voice audio based on the reply feedback text to obtain the reply broadcast. The method according to any one of claims 1 to 4, wherein The feedback information includes a feedback broadcast; The response to the natural language command, broadcasting the feedback information of the non-player character, comprises: In the case where the behavior intention of the natural language command is to execute a task, the feedback broadcast is broadcasted; the feedback broadcast includes an execution state after the non-player character executes the natural language command according to the environment perception information of the non-player character in the three-dimensional virtual environment; Wherein, the natural language command is used to instruct the non-player character to execute a task. The method according to any one of claims 1 to 5, wherein The feedback broadcast is broadcasted in the case where the behavior intention of the natural language command is to execute a task, which comprises at least one of: In the case where the behavior intention of the natural language command includes searching for a target entity, a first feedback broadcast is broadcasted, and the first feedback broadcast includes a search result of the non-player character to the target entity; In the case where the behavior intention of the natural language command includes using a virtual prop, a second feedback broadcast is broadcasted, and the second feedback broadcast is used to indicate that the virtual prop is ready; In a case where the behavior intention of the task natural language command includes using a virtual prop, a third feedback broadcast is broadcasted, the third feedback broadcast being used to indicate a use result of the virtual prop; In a case where the behavior intention of the natural language command includes controlling the non-player character to move, a fourth feedback broadcast is broadcasted, the fourth feedback broadcast being used to indicate a movement result of the movement; In a case where the behavior intention of the natural language command includes performing an interaction with a target entity, a fifth feedback broadcast is broadcasted, the fifth feedback broadcast being used to indicate at least one of a search result of the non-player character on the target entity, an interaction result of the non-player character and the target entity; In a case where the behavior intention of the natural language command includes an item interaction, a sixth feedback broadcast is broadcasted, the sixth feedback broadcast being used to indicate an item interaction result of the non-player character performing the item interaction. The method according to any one of claims 1 to 6, wherein The method further comprises: calling a pre-trained large language model to analyze the natural language command, and reasoning a task execution instruction based on the static environment data of the three-dimensional virtual environment and the environment perception information; controlling the non-player character to perform a task according to the task execution instruction; generating a state feedback text according to an execution state of the non-player character performing the task; generating a human voice audio based on the state feedback text to obtain the feedback broadcast. The method according to any one of claims 1 to 7, wherein The method further comprises: in a case where the environment perception information of the non-player character on the three-dimensional virtual environment meets a dynamic broadcast condition, a dynamic broadcast is broadcasted; the dynamic broadcast includes an abnormal situation identified according to the environment perception information of the non-player character. The method according to any one of claims 1 to 8, wherein The in a case where the environment perception information of the non-player character on the three-dimensional virtual environment meets a dynamic broadcast condition, a dynamic broadcast is broadcasted, comprises at least one of: in a case where the non-player character perceives an enemy virtual character, a first dynamic broadcast is broadcasted; the first dynamic broadcast includes a position of the enemy virtual character perceived by the non-player character; in a case where the non-player character perceives a dangerous situation, a second dynamic broadcast is broadcasted; the second dynamic broadcast is used to prompt the dangerous situation; in a case where the non-player character perceives a state change of a friendly virtual character, a third dynamic broadcast is broadcasted; the third dynamic broadcast is used to prompt that the state of the friendly virtual character has changed; in a case where the non-player character discovers a new trace in the three-dimensional virtual environment, a fourth dynamic broadcast is broadcasted; the fourth dynamic broadcast is used to prompt the new trace. The method according to any one of claims 1 to 9, wherein The feedback information includes a position description of a target entity; the method further comprises: generating the position description according to a position relationship between the main control virtual character and the target entity. The method according to any one of claims 1 to 10, wherein Before the generating the position description according to the position relationship between the main control virtual character and the target entity, the method further comprises: querying the target entity from an entity set of the three-dimensional virtual environment according to target entity information indicated by the natural language command and the environment perception information. The method according to any one of claims 1 to 11, wherein The feedback information is a spatial human voice audio of the non-player character; The broadcast of the feedback information of the non-player character includes: Obtaining the human voice audio corresponding to the feedback text; Based on the relative positional relationship between the non-player character and the master virtual character in the three-dimensional virtual environment, the sound effect parameter corresponding to the non-player character is obtained; Based on the sound effect parameter, the spatial human voice audio corresponding to the non-player character is generated and broadcasted, and the spatial human voice audio is used to represent the perception effect of the audio generated by the master virtual character to the non-player character under the relative positional relationship. A non-player character broadcasting method, the method is executed by a computer device, and the method includes: Receiving a natural language command for a non-player character sent by a client, the client has control authority of a master virtual character; Obtaining environmental perception information of at least one of the master virtual character and the non-player character; According to the natural language command and the environmental perception information, the feedback text of the non-player character is determined, and the feedback text is a non-fixed text generated based on the environmental perception information of the non-player character; Based on the feedback text, relevant information for playing feedback information is sent to the client, and the feedback information is human voice audio corresponding to the feedback text. The method of claim 13, wherein, The feedback text of the non-player character is determined according to the natural language command and the environmental perception information, including: Calling a pre-trained large language model, inferring the natural language command based on the static environmental data of the three-dimensional virtual environment and the environmental perception information, and obtaining an inference result of the natural language command; According to the inference result, the feedback text is determined. The method according to claim 13 or 14, wherein Based on the feedback text, relevant information for playing feedback information is sent to the client, including: In the case that the client stores a first offline voice corresponding to the inference result, the feedback text corresponding to the first offline voice or a first voice identifier corresponding to the first offline voice is sent to the client, and the first voice identifier is used to instruct the client to generate the feedback information based on the first offline voice; In the case that the client does not store offline voice corresponding to the inference result, calling a text-to-speech service to generate the feedback information based on the feedback text; the feedback information is sent to the client. The method according to any one of claims 13 to 15, wherein The feedback text includes a reply feedback text; According to the natural language command and the environmental perception information, the feedback text of the non-player character is determined, including: In the case that the behavior intention of the natural language command is inquiry, the reply feedback text of the non-player character is determined according to the environmental perception information; The reply feedback text includes reply content generated according to the environmental perception information of the non-player character to the three-dimensional virtual environment for the natural language command. The method according to any one of claims 13 to 16, wherein In the case that the behavior intention of the natural language command is inquiry, the reply feedback text of the non-player character is determined according to the environmental perception information, including at least one of: In a case where the natural language command includes an intelligence inquiry on a target entity, a first reply feedback text is determined according to the environmental perception information based on the perception of the non-player character on the target entity; the first reply feedback text includes the intelligence perception result of the non-player character on the target entity; In a case where the natural language command includes a position inquiry on a target entity, a second reply feedback text is determined according to the environmental perception information based on the perception of the non-player character on the target entity; the second reply feedback text includes a position description of the target entity. The method according to any one of claims 13 to 17, wherein The feedback text includes a state feedback text; The determination of the feedback text of the non-player character according to the natural language command and the environmental perception information includes: In a case where the behavior intention of the natural language command is to execute a task, the state feedback text is determined according to the environmental perception information; The state feedback text includes an execution state after the execution of the natural language command according to the environmental perception information of the non-player character on the three-dimensional virtual environment. The method according to any one of claims 13 to 18, wherein The determination of the state feedback text according to the environmental perception information in a case where the behavior intention of the natural language command is to execute a task includes at least one of the following: In a case where the behavior intention of the natural language command includes searching for a target entity, a first state feedback text is determined according to the environmental perception information; the state feedback text includes a search result of the non-player character on the target entity; In a case where the behavior intention of the natural language command includes using a virtual prop, a second state feedback text is determined according to the environmental perception information; the second state feedback text is used to indicate that the virtual prop is ready; In a case where the behavior intention of the natural language command includes using a virtual prop, a third state feedback text is determined according to the environmental perception information; the third state feedback text is used to indicate a use result of the virtual prop; In a case where the behavior intention of the natural language command includes controlling the movement of the non-player character, a fourth state feedback text is determined according to the environmental perception information; the fourth state feedback text is used to indicate a movement result of the movement; In a case where the behavior intention of the natural language command includes executing an interaction with a target entity, a fifth state feedback text is determined according to the environmental perception information; the fifth state feedback text is used to indicate at least one of a search result of the non-player character on the target entity, an interaction result of the non-player character and the target entity; In a case where the behavior intention of the natural language command includes an article interaction, a sixth state feedback text is determined according to the environmental perception information; the sixth state feedback text is used to indicate an article interaction result of the execution of the article interaction by the non-player character. The method according to any one of claims 13 to 19, wherein The determination of the state feedback text according to the environmental perception information in a case where the behavior intention of the natural language command is to execute a task includes: In a case where the behavior intention of the natural language command is to perform a task, a pre-trained large language model is invoked to analyze the natural language command, and task execution instructions are inferred based on static environment data of the three-dimensional virtual environment and the environment perception information; The non-player character is controlled to perform the task according to the task execution instructions; The state feedback text is determined according to the execution state of the non-player character performing the task. The method according to any one of claims 13 to 20, wherein The method further comprises: In a case where the environment perception information of the non-player character in the three-dimensional virtual environment meets a dynamic reporting condition, dynamic reporting text is determined according to the environment perception information; the dynamic reporting text includes an abnormal situation identified according to the environment perception information of the non-player character; Based on the dynamic reporting text, relevant information for playing dynamic reporting information is sent to the client, and the dynamic reporting information is human voice audio corresponding to the dynamic reporting text. The method according to any one of claims 13 to 21, wherein In a case where the environment perception information of the non-player character in the three-dimensional virtual environment meets a dynamic reporting condition, dynamic reporting text is determined according to the environment perception information, including at least one of the following: In a case where the non-player character perceives an enemy virtual character, a first dynamic reporting text is generated; the first dynamic reporting text includes the position of the enemy virtual character perceived by the non-player character; In a case where the non-player character perceives a dangerous situation, a second dynamic reporting text is generated; the second dynamic reporting text is used to prompt the dangerous situation; In a case where the non-player character perceives a state change of a friendly virtual character, a third dynamic reporting text is generated; the third dynamic reporting text is used to prompt the state change of the friendly virtual character; In a case where the non-player character discovers new traces in the three-dimensional virtual environment, a fourth dynamic reporting text is generated; the fourth dynamic reporting text is used to prompt the new traces. The method according to any one of claims 13 to 22, wherein The feedback text includes a position description of a target entity; the method further comprises: The position description is generated according to the position relationship between the main control virtual character and the target entity. The method according to any one of claims 13 to 23, wherein Before the position description is generated according to the position relationship between the main control virtual character and the target entity, the method further comprises: According to the target entity information indicated by the natural language command and the environment perception information, the target entity is queried from an entity set of the three-dimensional virtual environment. The method according to any one of claims 13 to 24, wherein The feedback information is spatial human voice audio of the non-player character; The method further comprises: Obtaining human voice audio corresponding to the feedback text; Based on the relative positional relationship between the non-player character and the main control virtual character in the three-dimensional virtual environment, obtaining sound effect parameters corresponding to the non-player character; Based on the sound effect parameters, the human voice audio is adjusted to generate spatial human voice audio corresponding to the non-player character, and the spatial human voice audio is used to represent the perception effect of the audio generated by the main control virtual character on the non-player character in the relative positional relationship. The method according to any one of claims 13 to 25, wherein The sound effect parameter corresponding to the non-player character is obtained based on the relative positional relationship between the non-player character and the host virtual character in the three-dimensional virtual environment, including: obtaining first position information of the non-player character in the three-dimensional virtual environment; and obtaining second position information of the host virtual character in the three-dimensional virtual environment; determining the relative positional relationship between the non-player character and the host virtual character based on the first position information and the second position information; generating the sound effect parameter corresponding to the non-player character based on the relative positional relationship. The method according to any one of claims 13 to 26, wherein The sound effect parameter corresponding to the non-player character is generated based on the relative positional relationship, including at least one of the following: determining a sound propagation path of the audio generated by the non-player character between the non-player character and the host virtual character based on the relative positional relationship; generating the sound effect parameter corresponding to the non-player character based on the sound propagation path; or, determining a field of view range corresponding to the host virtual character based on the orientation of the host virtual character in the three-dimensional virtual environment; in the case that the non-player character is in the field of view range, determining the sound effect parameter based on a first parameter; in the case that the non-player character is out of the field of view range, determining the sound effect parameter based on a second parameter; wherein the audio volume corresponding to the first parameter is greater than the audio volume corresponding to the second parameter; or, determining distance information between the host virtual character and the non-player character based on the relative positional relationship; determining an attenuation parameter and a delay parameter corresponding to the audio generated by the non-player character based on the distance information; determining the sound effect parameter based on the attenuation parameter and the delay parameter; or, determining the surrounding state information of the space element to the host virtual character based on the relative positional relationship between the space element and the host virtual character, the three-dimensional virtual environment including the space element forming a virtual space in the three-dimensional virtual environment, and the host virtual character being in the virtual space; determining the space type of the virtual space where the host virtual character is located based on the surrounding state information; obtaining a space sound effect parameter matched with the space type; determining the sound effect parameter based on the space audio parameter. The method according to any one of claims 13 to 27, wherein The human voice audio corresponding to the feedback text is obtained, including: obtaining the feedback text and character feature information corresponding to the non-player character, the character feature information being used to describe the attributes of the non-player character, wherein the attributes of the non-player character include at least one of a basic attribute and a scene performance attribute of the non-player character, the basic attribute being an attribute pre-configured for the non-player character, and the scene performance attribute being an attribute determined in real time by the non-player character in the scene of the three-dimensional virtual environment; generating the human voice audio matched with the character feature information according to the feedback text. The method according to any one of claims 13 to 28, wherein The non-player character has a plurality of behavior capabilities in the virtual environment, and the plurality of behavior capabilities correspond to a plurality of classification tags; The determining the feedback text of the non-player character according to the natural language command and the environment perception information comprises: obtaining a command text corresponding to the natural language command; calling at least one hierarchical prediction network to perform intent recognition on the command text to obtain a classification label of an intent of the command text instructing the non-player character; wherein the hierarchical prediction network comprises at least two sub-networks, and the at least two sub-networks are constructed based on a tree structure of the multiple classification labels; determining environment perception information of the non-player character in a virtual activity process corresponding to the intent based on the classification label; generating the feedback text based on the environment perception information. The method according to any one of claims 13 to 29, wherein The obtaining the command text corresponding to the natural language command comprises: obtaining a plurality of scene hot words corresponding to a three-dimensional virtual environment, the three-dimensional virtual environment being a virtual environment in which at least one of the main control virtual character or the non-player character is located, and the plurality of scene hot words being scene-related vocabularies of the three-dimensional virtual environment; converting the natural language command into the command text based on the plurality of scene hot words. A non-player character broadcasting device, the device comprising: a display module configured to display at least one of a main control virtual character and a non-player character located in a three-dimensional virtual environment; a receiving module configured to receive a natural language command for the non-player character; a broadcasting module configured to broadcast feedback information of the non-player character in response to the natural language command, the feedback information corresponding to feedback text being non-fixed text generated based on environment perception information, the environment perception information comprising information perceived by at least one of the main control virtual character and the non-player character in the three-dimensional virtual environment. A non-player character broadcasting device, the device comprising: a receiving module configured to receive a natural language command for a non-player character sent by a client, the client having control authority of a main control virtual character; an obtaining module configured to obtain environment perception information of at least one of the main control virtual character and the non-player character; a determining module configured to determine feedback text of the non-player character according to the natural language command and the environment perception information, the feedback text being non-fixed text generated based on the environment perception information of the non-player character; a sending module configured to send, to the client, information related to playing feedback information based on the feedback text, the feedback information being human voice audio corresponding to the feedback text. A computer device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the non-player character broadcasting method according to any one of claims 1 to 30. A computer readable storage medium, the readable storage medium storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by a processor to implement the non-player role broadcasting method according to any one of claims 1 to 30. A computer program product or a computer program, the computer program product or the computer program comprising computer instructions stored in a computer readable storage medium; a processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device is caused to execute to implement the non-player role broadcasting method according to any one of claims 1 to 30.

Citation Information

Patent Citations

  • Virtual object prompting method and device thereof, storage medium and computer equipment

    CN113018864A

  • Non-player character interaction processing method and device, electronic equipment and storage medium

    CN116966574A

  • Non-player character interaction method and system

    CN118079400A

  • Non-player character broadcasting method and device, equipment, medium and product

    CN118987618A

  • Information processing device, information processing method, and program

    JP2020137819A