Method and apparatus for controlling virtual character, device, medium, and program product

By combining environmental perception information to identify behavioral intentions in natural language commands, the problem of insufficient accuracy of speech recognition systems in NPC control is solved, and more efficient human-computer interaction is achieved.

WO2026031888A1PCT designated stage Publication Date: 2026-02-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/105722
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-06-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

In existing technologies, speech recognition systems have low accuracy when recognizing text words with similar pronunciations, resulting in low precision in NPC control and affecting the efficiency of human-computer interaction.

Method used

By acquiring natural language commands in the form of speech and combining them with the environmental perception information of the main virtual character and NPCs regarding the virtual environment, behavioral intentions can be identified to control NPCs to perform virtual activities.

Benefits of technology

It improves the accuracy of identifying behavioral intentions, reduces the rate of player errors, and enhances the efficiency of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025105722_12022026_PF_FP_ABST
    Figure CN2025105722_12022026_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for controlling a virtual character, a device, a medium, and a program product. The method comprises: displaying at least one of a virtual main player character and an NPC located in a virtual environment (310); acquiring a natural language command in the form of speech (320); and, by means of a behavior intent expressed by the natural language command, controlling the NPC to execute a virtual activity (330).
Need to check novelty before this filing date? Find Prior Art

Description

Virtual role control method, device, apparatus, medium and program product

[0001] The present application claims priority to the Chinese patent application No. 202411096374.8, filed on August 9, 2024, and entitled "Virtual role control method, device, apparatus, medium and program product", the entire content of which is incorporated herein by reference. TECHNICAL FIELD

[0002] Embodiments of the present application relate to the field of human-computer interaction, and in particular to a virtual role control method, device, apparatus, medium and program product. BACKGROUND

[0003] In current game applications, a player can control a master virtual role to participate in a virtual game in a personal or team mode.

[0004] In related technologies, in order to improve the interest of the game, in the case that a player controls a master virtual role to participate in a virtual game, the player can select an NPC to participate in the virtual game together with the master virtual role; if the player controls the NPC by using a natural language command in a voice form, a pre-trained voice recognition system is usually used to analyze the natural language command in text, so as to control the NPC by using the text content obtained by the analysis.

[0005] However, the voice recognition system is trained by using a large number of text words, and thus it is difficult to flexibly distinguish text words with similar pronunciation in the text analysis process, so that the meaning of the output text content is usually broad and the accuracy is low, and thus the control accuracy of the NPC is low, which affects the human-computer interaction efficiency. SUMMARY

[0006] Embodiments of the present application provide a virtual role control method, device, apparatus, medium and program product, which can better understand a natural language command by using environmental perception information, improve the accuracy of the behavior intention obtained by recognition, make the behavior intention more matched with a virtual environment, reduce the error operation rate of a player, and improve the human-computer interaction efficiency. The technical solution is as follows.

[0007] In one aspect, a virtual role control method is provided, which is executed by a computer device, and the method comprises:

[0008] displaying at least one of a master virtual role and an NPC located in a virtual environment;

[0009] obtaining a natural language command in a voice form, the natural language command comprising a natural semantic used for controlling the NPC;

[0010] control the NPC to perform a virtual activity according to a behavior intention expressed by the natural language command.

[0011] The behavior intention is identified from the natural language command according to environment perception information, and the environment perception information includes information perceived by at least one of the host virtual character and the NPC in the virtual environment.

[0012] In another aspect, a method for controlling a virtual character is provided, and the method includes:

[0013] acquiring a natural language command in a voice form, the natural language command including natural semantics for controlling an NPC;

[0014] converting the natural language command into a command analysis result in a text form, the command analysis result being a result of performing voice-to-text on the natural language command based on environment perception information, and the environment perception information including information perceived by at least one of a host virtual character and the NPC in a virtual environment;

[0015] performing behavior intention identification on the command analysis result to obtain a behavior intention, the behavior intention being used to control the NPC to perform a virtual activity.

[0016] In another aspect, a device for controlling a virtual character is provided, and the device includes:

[0017] a display module configured to display at least one of a host virtual character and an NPC in a virtual environment;

[0018] an acquisition module configured to acquire a natural language command in a voice form, the natural language command including natural semantics for controlling the NPC;

[0019] a control module configured to control the NPC to perform a virtual activity according to a behavior intention expressed by the natural language command, wherein the behavior intention is identified from the natural language command according to environment perception information, and the environment perception information includes information perceived by at least one of the host virtual character and the NPC in the virtual environment.

[0020] In another aspect, a device for controlling a virtual character is provided, and the device includes:

[0021] an acquisition module configured to acquire a natural language command in a voice form, the natural language command including natural semantics for controlling an NPC;

[0022] a conversion module configured to convert the natural language command into a command analysis result in text form, the command analysis result being a result of performing speech-to-text on the natural language command based on environment perception information, the environment perception information including information perceived by at least one of the host virtual character and the NPC in the virtual environment;

[0023] a recognition module configured to perform behavior intention recognition on the command analysis result to obtain a behavior intention, the behavior intention being used to control the NPC to perform the virtual activity.

[0024] In another aspect, a computer device is provided, the computer device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the method for controlling a virtual character according to any one of the above embodiments of the present application.

[0025] In another aspect, a computer readable storage medium is provided, the storage medium storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by a processor to implement the method for controlling a virtual character according to any one of the above embodiments of the present application.

[0026] In another aspect, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to cause the computer device to perform the method for controlling a virtual character according to any one of the above embodiments.

[0027] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:

[0028] After obtaining the natural language command in the form of voice for commanding the NPC, the environmental perception information perceived by at least one of the host virtual character and the NPC to the virtual environment is used to identify the behavior intention from the natural language command, so as to control the NPC to perform the virtual activity through the behavior intention. With the environmental perception information of the host virtual character and / or the NPC, the natural language command can be better understood, the analysis accuracy of the natural language command can be improved, and thus the accuracy of the identified behavior intention can be improved, so that the behavior intention is more matched with the virtual environment, and the invalid or incorrect command to the NPC is greatly avoided, which helps to reduce the incorrect operation rate of the player and also helps the player to reduce the operation times in the virtual game process, reduce the data processing amount related to the operation, improve the utilization rate of the computing resource and the human-computer interaction efficiency. BRIEF DESCRIPTION OF DRAWINGS

[0029] Fig. 1 is a flowchart of commanding the NPC in the virtual environment with the natural language command (natural language instruction) in the form of voice according to an example embodiment of the present application;

[0030] Fig. 2 is a structural block diagram of a control system according to an example embodiment of the present application;

[0031] Fig. 3 is a flowchart of a control method of a virtual character according to an example embodiment of the present application;

[0032] Fig. 4 is a flowchart of a control method of a virtual character according to another example embodiment of the present application;

[0033] Fig. 5 is a flowchart of a control method of a virtual character according to still another example embodiment of the present application;

[0034] Fig. 6 is an interface diagram for determining the object perception range according to an example embodiment of the present application;

[0035] Fig. 7 is a flowchart of behavior intention recognition according to an example embodiment of the present application;

[0036] Fig. 8 is a diagram for performing the behavior intention recognition on the analysis result of the command according to an example embodiment of the present application;

[0037] Fig. 9 is a diagram of a first-level candidate label according to an example embodiment of the present application;

[0038] Fig. 10 is a diagram of a first-level candidate label according to still another example embodiment of the present application;

[0039] Fig. 11 is an analysis flowchart of behavior intention recognition according to an example embodiment of the present application;

[0040] Fig. 12 is a diagram of a command recognition result obtained by segmentation according to an example embodiment of the present application;

[0041] FIG. 13 is a schematic diagram of a method for training a prediction network according to an example embodiment of the present application;

[0042] FIG. 14 is a flowchart of a method for analyzing a behavior intention based on a natural language command according to an example embodiment of the present application;

[0043] FIG. 15 is a technical flowchart of a method for controlling a virtual role according to an example embodiment of the present application;

[0044] FIG. 16 is a schematic diagram of an interface of a method for controlling a virtual role according to an example embodiment of the present application;

[0045] FIG. 17 is a flowchart of a method for controlling a virtual role according to an example embodiment of the present application;

[0046] FIG. 18 is a structural block diagram of a device for controlling a virtual role according to an example embodiment of the present application;

[0047] FIG. 19 is a structural block diagram of a device for controlling a virtual role according to another example embodiment of the present application;

[0048] FIG. 20 is a structural schematic diagram of a terminal according to an example embodiment of the present application. DETAILED DESCRIPTION

[0049] First, the terms involved in the embodiments of the present application are briefly introduced.

[0050] Virtual scene: a virtual scene displayed (or provided) by an application when running on a terminal.

[0051] Virtual model: a model used to simulate a real scene in a virtual scene.

[0052] Virtual role / virtual object: a movable object in a virtual scene.

[0053] In the embodiments of the present application, a method for controlling a virtual role is introduced, which can better understand a natural language command in a voice form with the aid of environmental perception information in a virtual scene, so as to obtain a more accurate command analysis result in a text form, and further improve the precision of a behavior intention recognized based on the command analysis result, so that the behavior intention is more matched with the virtual environment, and the accuracy of controlling an NPC in the virtual environment is improved. The method for controlling a virtual role provided by the embodiments of the present application can be applied to a terminal game scene, a virtual reality scene, an augmented reality scene, a virtual travel scene, a virtual conference scene, a virtual shopping scene, a virtual fitness scene, and other virtual scenes requiring a voice conversion process, which are not limited by the embodiments of the present application.

[0054] In an optional embodiment, the method of controlling a virtual role is applied to a terminal game.

[0055] Illustratively, a game application is installed in the terminal, and the game application displays a virtual environment in which a player plays a virtual game during running. The virtual environment includes at least one of a host virtual role and an NPC, and the NPC in the virtual scene can be commanded / controlled by a voice control mode. The player issues a natural language command in the form of voice, and the terminal acquires the natural language command. The natural language command includes a natural semantic for controlling the NPC. In the case of receiving the natural language command, environment perception information is obtained according to the perception of the virtual environment by the host virtual role and / or the NPC, such as perceiving a virtual stable, a virtual obstacle, and the like around. Then, the environment perception information is used to assist in recognizing the natural language command to obtain a behavior intention, and the NPC is controlled to perform a virtual activity according to the behavior intention. Due to the constraint of the environment perception information, the behavior intention is more matched with the virtual environment, which helps to improve the recognition accuracy of the behavior intention, and further more accurately controls the NPC to perform the virtual activity, and improves the efficiency and accuracy of controlling the NPC.

[0056] In an optional embodiment, the method of controlling a virtual role is applied to a virtual reality scene.

[0057] Illustratively, a virtual environment received from a computer device can be transmitted and presented by a VR head-mounted display. The player observes the virtual environment by wearing the VR head-mounted display, and the player can also move to other positions in the virtual environment by moving, adjusting body parts, and the like to explore the virtual scene. Taking the player himself as the host virtual role when the player wears the VR head-mounted display as an example, there can also be virtual roles controlled and / or commanded by other players and NPCs in the virtual environment presented by the VR head-mounted display. Optionally, the player can control the NPC by a natural language command in the form of voice; the VR head-mounted display can receive the natural language command in the form of voice through a deployed microphone array, and autonomously analyze the natural language command or send the natural language command to a connected computer device for analysis. The natural language command includes a natural semantic for controlling the NPC. In the case of receiving the natural language command, environment perception information is obtained according to the perception of the virtual environment by the host virtual role and / or the NPC, such as perceiving an enemy virtual role, a teammate virtual role, a virtual kitchen knife, and the like around. Then, the environment perception information is used to assist in recognizing the natural language command to obtain a behavior intention, and the NPC is controlled to perform a virtual activity according to the behavior intention. Due to the constraint of the environment perception information, the behavior intention is more matched with the virtual environment, which helps to improve the recognition accuracy of the behavior intention, and further more accurately controls the NPC to perform the virtual activity, and enriches the interest of the virtual reality scene, so that the player can obtain a more real and in-depth virtual experience.

[0058] It should be noted that, before collecting the relevant data of the user and in the process of collecting the relevant data of the user, the application can display a prompt interface, a pop-up window or output voice prompt information, which prompts the user that the relevant data of the user is currently being collected, so that the application only starts to perform the relevant steps of obtaining the relevant data of the user after obtaining the confirmation operation of the user to the prompt interface or the pop-up window, otherwise (i.e. without obtaining the confirmation operation of the user to the prompt interface or the pop-up window), ending the relevant steps of obtaining the relevant data of the user, that is, not obtaining the relevant data of the user. In other words, all the user data collected by the application is collected with the consent and authorization of the user, and the collection, use and processing of the relevant user data need to comply with the relevant laws, regulations and standards of the relevant region. It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant region. For example, the natural language commands and other contents involved in the application are obtained with full authorization.

[0059] The embodiment of the application provides a scheme for a user to use a natural language voice command in a natural language voice form to direct and control an NPC (Non Player Character, NPC) in a virtual environment. In the scheme, the user usually has one or more NPCs as teammates in addition to a self-controlled master virtual character. The user can use a human conversation mode of relatively arbitrary and non-mechanical commands to direct the NPC to make feedback on scene objects in the virtual environment as desired by the user, so as to direct the NPC to complete a task in cooperation with the master virtual character. Referring to an exemplary reference figure 1, the user controls a master virtual character 10 to play a game, and the master virtual character 10 has an NPC teammate 20. There is a car 30 in the visual range of the master virtual character 10, and the user speaks a natural language voice command instruction in a natural language voice form: "No. 2, go to hide behind the car 30", and then the NPC teammate 20 will move to hide behind the car 30. It should be noted that there can be multiple cars in the scene shown in figure 1, and the NPC teammate 20 will accurately understand the car said by the user as the car 30 in the visual range of the user. That is, the NPC teammate 20 has relatively intelligent natural language understanding capability.

[0060] Compared with the mechanical fixed instruction used to direct the NPC in the conventional technology, the relatively arbitrary natural language command used to direct the NPC in the embodiment of the application can provide the user with more natural, more complex and more flexible language directing capability.

[0061] In another aspect, the natural language understanding capability of the intelligent NPC team member is embodied in that the NPC team member considers not only the literal information of the natural language command when understanding the natural language command, but also combines the environmental perception information of the virtual environment of the host virtual character and / or the NPC itself to assist in understanding the semantics of the natural language command. That is, the NPC considers not only the information of the natural language command in one modality, but also considers the perception information of other modalities such as visual field, hearing, radar, etc. to assist in understanding and executing the natural language command. Since the natural language command in the form of voice is a command in the form of spoken language, not a command in the form of written language, considering only the literal information of the natural language command may have multiple candidate understanding ways or unclear places. The NPC team member combines the environmental perception information of the virtual environment of the host virtual character and / or the NPC itself to determine a reasonable understanding way among the multiple candidate understanding ways or eliminate doubts about unclear places, thereby realizing a more intelligent natural language understanding capability. The environmental perception information mentioned above includes at least one of the following:

[0062] · visual information perceived within the visual field range of the host virtual character;

[0063] · hearing information perceived within the hearing range of the host virtual character;

[0064] · information perceived by the perception skill or perception device possessed by the host virtual character;

[0065] · visual information perceived within the visual field range of the NPC;

[0066] · hearing information perceived within the hearing range of the NPC;

[0067] · information perceived by the perception skill or perception device possessed by the NPC.

[0068] In combination with reference to FIG. 1, the scheme includes at least one of the following five stages:

[0069] Stage one: preprocessing of spatial data;

[0070] The server preprocesses the scene objects in the virtual scene to construct spatial data of the virtual scene.

[0071] The spatial data of the virtual scene includes matching labels of the scene objects and spatial positions of the scene objects, and in some examples, the spatial data further includes appearance images of the scene objects. The scene objects are any objects appearing in the virtual scene, such as the car, wall, box, etc. in FIG. 1.

[0072] Firstly, the pre-acquisition process of the matching label visual label of the scene object is introduced. Firstly, the attribute information of the scene object in the virtual scene is acquired, such as at least one of the name, size and the like of the scene object; and the appearance image of the scene object is acquired, the appearance image including images obtained by observing the scene object from at least two perspectives, and observing the scene object from multiple perspectives can ensure that the appearance image carries comprehensive appearance information of the scene object.

[0073] The server calls the multi-modal model to perform prediction on the attribute information and the appearance image of the scene object, extracts hidden layer features from the attribute information in the text dimension and the appearance image in the image dimension based on the ability of the multi-modal model to process information in the text dimension and the image dimension, and performs decoding on the hidden layer features to predict the visual object label of the scene object, the visual object label being used to describe the scene object in at least one dimension in a text manner; such as describing the scene object in the dimensions of material, transparency, color, shape, etc. The natural language model is called to perform label optimization on the visual object label to obtain the matching label visual label of the scene object; the natural language model has a text generation capability, and performs rewriting on the visual object label input into the natural language model in a manner conforming to the spoken expression of natural language, so as to realize label optimization of the visual object label and obtain the optimized matching label visual label of the scene object. The optimized visual label conforms to the spoken expression of natural language, and describes the scene object in at least one dimension in a manner close to the spoken expression. Taking a virtual bed in a virtual scene as an example, usually the color, material, and placement position of the virtual bed are concerned, while usually the surface features of the head and the side plate of the bed, such as carving and embossing, are ignored. The purpose of performing spoken language-like rewriting on the visual label of the scene object is to obtain a matching label closer to the spoken expression; such as enriching the labels of the dimensions of color, material, and placement position concerned in the spoken expression with a plurality of words with the same semantics, and deleting the label of the surface feature dimension of the head and the side plate of the bed ignored in the spoken expression.

[0074] Next, the spatial position of the scene object is introduced. The spatial position includes at least one of the following of the scene object: coordinate position (Location), such as coordinate information of a center point or a preset point of the scene object in the virtual scene; orientation (Rotation), such as a direction in which a front surface of the scene object faces in the virtual scene; bounding box (Bounding Box), used to indicate the size of the scene object in the virtual scene; cover point (Cover points), used to indicate a recommended virtual character position point when the virtual character approaches the scene object, so as to realize that the scene object can mask the virtual character.

[0075] Next, the pre-acquisition process of the appearance image of the scene object is introduced. The appearance image is acquired in the process of predicting the matching label visual label of the scene object, and the appearance image includes images obtained by observing the scene object from at least two perspectives. Optionally, the spatial data of each scene object in the entire virtual environment is pre-acquired and arranged into a data set for use in the subsequent query process.

[0076] Stage two: speech recognition and behavior intention recognition;

[0077] In the process of controlling the master virtual role by the user, the terminal receives the natural language command input by the user in the form of speech; the server performs speech recognition and behavior intention recognition on the natural language command to analyze the input instruction of the user;

[0078] The speech recognition is text conversion on the natural language instruction input by the user in the form of speech to obtain the instruction text corresponding to the natural language instruction, and converts the speech information into text information. The encoding network is called to perform feature encoding on the instruction text to obtain the feature representation of the natural language instruction in the hidden layer space; the text segmentation network is called to perform segmentation on the feature representation to obtain multiple clauses in the instruction text, and then the behavior intention recognition is performed on each of the multiple clauses. In the case that the natural language instruction input by the user in the form of speech is a long sentence, the segmentation of the instruction text corresponding to the natural language instruction can be realized, and the demand for analyzing the input instruction of the user in the long sentence scenario can be met.

[0079] The behavior intention recognition of each clause is introduced as follows: the input instruction of the user is obtained through the behavior intention recognition, and the input instruction of the user includes the entity, semantic type, subject type and intention type in at least one clause.

[0080] On the one hand, the Conditional Random Field (CRF) method is used to recognize the entity in the clause, such as recognizing the entity in the clause as a scene object in the virtual scene, such as a virtual building, a virtual article, virtual vegetation, etc. The entity in the clause is used to indicate the scene object to which the virtual activity is directed, such as the entity in the clause being a virtual building, and the clause carrying the semantics indicating that the NPC initiates a virtual attack, then the virtual attack is performed on the virtual building, that is, the virtual attack is initiated on the virtual building. On the other hand, the prediction network is called to predict the semantic type, subject type and intention type in the clause respectively. Among them, the semantic type includes whether to indicate that the NPC initiates a virtual attack; the subject type is used to represent that the subject of the clause is one or more NPCs or a master virtual role controlled by the user; and the intention type is the expected virtual activity indicated by the clause, such as moving, initiating a virtual attack, etc.

[0081] Stage three: query of spatial data;

[0082] The input instruction of the user is input, the spatial data of the virtual scene and the runtime information corresponding to the master virtual character are referenced, and an AI model is used to determine the target entity in the virtual space and the control instruction of the NPC.

[0083] For the entity in the input instruction, the spatial position of the entity in the virtual environment needs to be determined. For example, first, the input instruction in the form of text is analyzed to determine the entity in the text instruction. In the case where multiple candidate entities in the virtual environment match the entity in the input instruction, the target entity matching the input instruction can be uniquely determined from the multiple candidate entities in combination with the environmental perception information of the master virtual character and / or the NPC. For example, the environmental perception information is the field of view information of the master virtual character, and then the spatial data of the scene in the virtual environment is used as the query range, and the real-time position and orientation of the master virtual character are used as the query reference conditions to query the spatial position of the target entity corresponding to the input instruction in the virtual environment. In order to subsequently control the NPC to move to the vicinity of the target entity position, or perform the behavior indicated by the input instruction on the target entity position.

[0084] However, there can be multiple environmental perception information of the master virtual character and / or the NPC, and multiple environmental perception information can be fused to assist the determination process of the target entity, or multiple environmental perception information can be used according to priority to assist the determination process of the target entity, which is not limited.

[0085] Stage four: voice feedback of NPC;

[0086] The NPC will make voice feedback to the user. The types of voice feedback include at least one of immediate feedback, instruction execution feedback, and dynamic feedback. The immediate feedback refers to the NPC making voice feedback to the natural language instruction immediately after the user issues the natural language instruction to the NPC, such as "received", "start execution", "good" and the like. The execution instruction feedback refers to the NPC making voice feedback to the execution result of the control instruction after executing the control instruction indicated by the natural language instruction, such as "successfully executed" and "failed to successfully execute". The dynamic feedback refers to the NPC making voice feedback according to real-time environmental perception when the user does not issue a natural language instruction to the NPC.

[0087] The voice feedback is obtained by converting the text content into voice. The text content is inferred by a large language model in combination with the spatial data and dynamic runtime information environmental perception information. Using a large language model to dynamically generate the text content of the voice feedback can break the monotony, mechanicalness, and easy repetition of fixed template voice feedback on the one hand, and can avoid the problem of excessive data volume of the client on the other hand without pre-storing too many pre-produced template voices.

[0088] Stage five: behavior control of NPC;

[0089] The input instruction of the user is input, the spatial data of the virtual scene and the runtime information corresponding to the master virtual role are referenced, and an AI model is used to determine the target entity in the virtual space and the control instruction of the NPC;

[0090] The natural language command usually further includes a control instruction, which is an instruction that can be executed by the behavior tree of the NPC; the subsequent action that the NPC needs to perform indicated by the control instruction can be a single action or a sequence of actions. In the case of a sequence instruction of multiple actions, the sequence instruction of multiple actions is used to indicate that the NPC performs a sequence of actions, and the data structure for storing the sequence instruction of multiple actions is in the form of a list. The instructions to be executed in the cache are executed when the behavior tree of the NPC calls the cache after executing the current instruction.

[0091] Hotword updating function:

[0092] For stage two, the hotword system is also introduced to assist in improving the accuracy of voice recognition.

[0093] The hotword system is used to provide scene hotwords related to the scene involved in the natural language instruction in the voice recognition process, so that when the natural language instruction of the voice input is received, the vocabulary in the hotword library can be preferentially searched as the voice recognition result instruction analysis result based on the currently running virtual scene, so that the matching degree between the instruction analysis result voice recognition result and the virtual scene is higher.

[0094] The virtual scene includes at least one of a plurality of scene types, such as at least one of a plurality of virtual scenes including a virtual battle scene type, a virtual transaction scene type, a virtual office scene type, a virtual kitchen scene type, etc. The frequently used vocabulary under a plurality of virtual scene types, specific virtual elements existing in the virtual scene type (such as virtual buildings, virtual role objects, etc. that exist only in a specific scene type) are analyzed in advance as scene hotwords, and a scene hotword library composed of a plurality of scene hotwords is obtained, and the plurality of scene hotwords correspond to the virtual scene type. In addition, different scene language models are trained in advance for different virtual scene types, such as a scene language model under a battle scene that pays more attention to words related to battles and a scene language model under a transaction scene that pays more attention to words related to transactions; the scene language model can more specifically analyze the natural language instruction received in the virtual scene type.

[0095] For example, in the process of controlling a non-player-controlled virtual object by a user, a natural language instruction is received, based on the game state, object position, and current task completion, it is determined that the first virtual scene in which the virtual character and the NPC are currently located is a virtual battle scene, the target scene type is a battle scene type, and a battle scene language model corresponding to the virtual battle scene type is obtained. Based on the field of view of the target virtual object controlled by the user, the target first scene hot words in the virtual battle scene type are obtained, including "truck", "virtual grass", "virtual house", "enter", "open truck", "defend", "desert", "attack", "oak tree", "stable", and "enemy A" and "attack" within the field of view, and the natural language instruction is decoded by a pre-trained natural language analysis model (including an acoustic model network, a target scene language network model, a preset dictionary, etc.) under the constraint of the first target scene hot words. The output text content is "attack me", instead of the result obtained by voice recognition in the usual case - "supply me", which improves the accuracy of the user's input instruction and outputs the instruction analysis result to accurately control the NPC. The instruction analysis result is illustrated as follows.

[0096] (1) The instruction analysis result is "No. 1, go to the front and defend", which can be used to control No. 1 to move to the front and make a defensive posture, avoiding the enemy virtual character attacking the user-controlled virtual character first. Without the constraint of the first scene hot words, the natural language instruction is easily recognized as "No. 1, go to the front and counterattack", which affects the virtual combat plan of the player.

[0097] (2) The instruction analysis result is "No. 2, go to the desert and explore", which can be used to control No. 2 to move to the desert and explore whether there are enemy virtual characters or virtual treasure chests; without the constraint of the first scene hot words, the natural language instruction is easily recognized as "No. 2, go to the mountain and explore", which makes No. 2 move to the wrong position, and the human-computer interaction efficiency is low.

[0098] (3) The instruction analysis result is "No. 1 and No. 2, attack me", which can be used to control No. 1 and No. 2 to attack the virtual characters within the attack range; without the constraint of the first scene hot words, the natural language instruction is easily recognized as "No. 1 and No. 2, supply me", which makes No. 1 and No. 2 make an error behavior that does not meet the player's expectation, which cannot protect the user-controlled virtual character and easily makes the enemy virtual character win the virtual game.

[0099] (4) The instruction analysis result is "No. 1, go find an oak tree nearby", which can be used to control No. 1 to search for an oak tree in the nearby area so as to complete the game task or find an oak tree that can be used to avoid attack. Without the constraint of the first scene hotword, the natural language command is easily identified as "No. 1, go find a project book nearby", so that No. 1 searches for a virtual element that does not meet the player's expectation to find, and cannot meet the player's virtual combat demand.

[0100] (5) The instruction analysis result is "No. 1 and No. 2, go to the stable over there", which can be used to control No. 1 and No. 2 to search for a stable nearby and move to the position where the stable is located. Without the constraint of the first scene hotword, the natural language command is easily identified as "No. 1 and No. 2, go over there, it will be done soon", so that No. 1 and No. 2 mistakenly think that the host virtual character currently wants to complete the virtual game by himself, and thus cannot provide better game assistance for the host virtual character.

[0101] Ambient sound effect function:

[0102] In addition, for the ambient sound effect of the entire virtual scene, the scheme also provides a spatial audio enhancement scheme.

[0103] The audio played by the terminal includes ambient audio and NPC audio. The ambient audio is audio generated based on the scene characteristics of the virtual scene currently occupied by the host virtual object, so as to give the user a sense of being there; the NPC audio is audio generated based on the characteristics of the NPC, so that the user can intuitively feel the emotion and physical state of the role through the heard sound.

[0104] Ambient audio: identify the scene elements in the virtual scene where the host virtual object is located, generate / choose appropriate element sound effects from the sound effect library in real time according to the scene elements, and synthesize the element sound effects to obtain the ambient audio. For example, the virtual scene is a forest at night, and the scene elements include trees, owls, insects, etc., and the ambient audio includes rustling of leaves, owl calls, insect calls, etc.

[0105] NPC audio: determine the NPC contained in the virtual scene, generate the corresponding role voice based on the character type, current emotion and current behavior of the NPC. For example, when the NPC is a middle-aged man who is running, the generated role voice is a deep male voice with a breathing sound effect when running. When the terminal plays ambient audio and / or NPC audio to the user, the ambient audio and / or NPC audio are subjected to audio enhancement processing to improve the realism of the audio.

[0106] In some embodiments, the control system involved in the embodiments of the present application is described, and the control method of the virtual role provided by the embodiments of the present application can be implemented by the terminal alone, or by the server, or by the terminal and the server through data interaction, and the embodiments of the present application do not limit this. Optionally, the control method of the virtual role is taken as an example for description.

[0107] Illustratively, referring to FIG. 2, the control system involves a terminal 210 and a server 220, and the terminal 210 and the server 220 are connected through a communication network 230.

[0108] In some embodiments, the terminal 210 is installed with a game application, and at least one of a master virtual role and an NPC in a virtual environment is displayed during the running of the game application. The NPC includes at least one of a virtual soldier, a virtual hero, a virtual monster, a virtual mount and the like. The player can control the NPC through a voice control mode.

[0109] Optionally, the player is taken as an example for commanding the NPC through voice; the terminal 210 is deployed with a microphone array, and can obtain a natural language command in the form of voice. The natural language command includes a natural semantic used for controlling the NPC.

[0110] In some embodiments, the terminal 210 analyzes by itself based on the natural language command, or the terminal 210 sends the natural language command to the server 220 through the communication network 230, and the server 220 analyzes the natural language command.

[0111] Optionally, the terminal 210 is taken as an example for sending the natural language command to the server 220 and analyzing the natural language command by the server 220.

[0112] The server 220 determines environment perception information according to the information perceived by at least one of the master virtual role and the NPC in the virtual environment, then identifies a behavior intention from the natural language command according to the environment perception information, and sends the behavior intention to the terminal 210 through the communication network 230.

[0113] In some embodiments, the terminal 210 controls the NPC to perform a virtual activity through the behavior intention expressed by the natural language command, and the terminal 210 renders and displays an activity animation corresponding to the activity animation data on a screen interface.

[0114] It is worth noting that the terminal described above includes but is not limited to mobile terminals such as mobile phones, tablets, portable laptop computers, smart voice interaction devices, smart home appliances, vehicle-mounted terminals, etc., and can also be implemented as a desktop computer, etc. The server described above can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server.

[0115] In combination with the above-mentioned name introduction and application scenarios, the control method of the virtual role provided by the present application is described. For example, the method is applied to a terminal, as shown in FIG. 3, which includes the following steps 310 to 330.

[0116] Step 310: Display at least one of a master virtual role and an NPC in a virtual environment.

[0117] Optionally, the virtual environment is a virtual scene displayed when an application runs on the terminal; or, the virtual environment is a virtual scene projected and displayed. The virtual environment can be a simulated scene of a real scene, a semi-simulated and semi-fictional scene, or a purely fictional scene. The virtual environment can be any one of a two-dimensional virtual scene, a 2.5-dimensional virtual scene, and a three-dimensional virtual scene.

[0118] Illustratively, the master virtual role is a virtual role manually controlled by a player, such as a virtual hero, a virtual mount, etc. The NPC is a virtual role displayed in a game, or a virtual role displayed in other virtual environments such as a virtual reality scene, an augmented display scene, etc. The NPC has autonomous action capability and can move in the virtual environment on its own.

[0119] In some embodiments, the NPC is a virtual role available for control by the player. When the NPC is not controlled by the player, it moves in the virtual environment according to a preset trajectory, or attacks an enemy virtual role in the virtual environment according to a preset game logic, assists the master virtual role in performing a game task, etc. When the NPC is controlled by the player, the NPC will perform virtual activities in the virtual environment according to the command in a targeted manner to assist the master virtual role more purposefully.

[0120] Generally, a virtual match includes a master virtual role and at least one NPC. There is a certain difference between the master virtual role and the NPC.

[0121] Optionally, the master virtual role is a virtual role mainly controlled by the player, and the NPC is a virtual role selected for command by the player; or, the master virtual role is a virtual role autonomously selected by the player, and the NPC is a virtual role configured by default by the system; or, both the master virtual role and the NPC are virtual roles autonomously selected by the player, wherein the master virtual role is a virtual role with the highest priority, and the NPC is a virtual role with a lower priority, etc.

[0122] For example, when the NPC is a virtual character selected by the player, a plurality of candidate control characters with different personalities and / or autonomous behavior capabilities can be provided. The personality can represent the attack strength of the candidate control character in the virtual game, such as a candidate control character with a violent personality having a larger attack strength, a candidate control character with a gentle personality having a smaller attack strength, etc. The autonomous behavior capability represents the skill situation of the candidate control character in the virtual game, such as candidate control character 1 being able to use skill A to attack enemy virtual characters in a circular area, candidate control character 2 being able to use skill B to defend against attacks from nearby enemy virtual characters, etc. The player's selection operation on the characters in the plurality of candidate control characters can select at least one candidate control character as an NPC that can be commanded in the game, etc.

[0123] At step 320, a natural language command in the form of speech is obtained.

[0124] Speech form is a form expressed in the form of speech, i.e., through speech, sound, information or intent. In the field of computer science, especially in human-computer interaction and natural language processing, speech form usually refers to the way in which an object inputs information using spoken language.

[0125] Illustratively, the natural language command as a command in the form of speech is usually implemented as audio data. For example, the player issues a voice during the operation of the terminal, and the terminal automatically collects the voice based on the deployment of a microphone array and obtains audio data as the natural language command in the form of speech; or the player issues a voice during the operation of the terminal based on the triggering operation on the audio collection control, and the terminal receives the voice and obtains audio data as the natural language command in the form of speech, etc.

[0126] The natural language command includes a natural semantic for controlling the NPC.

[0127] Illustratively, the natural semantic is semantic information expressed by the natural language command for controlling the NPC, and the NPC can be controlled to perform virtual activities based on the natural semantic.

[0128] Optionally, the natural semantic includes an identification semantic for indicating the NPC and an intent semantic for indicating the activity of the NPC. Illustratively, the identification semantic includes at least one of a name of the NPC (such as "A hero"), a number of the NPC (such as "No. 1"), and location information of the NPC (such as "next to the house").

[0129] Optionally, the NPC is a virtual character commanded by the player through voice, including at least one of a virtual soldier, a virtual hero, a virtual monster, a virtual mount, and the like. Based on the natural language command in the form of voice issued by the player during the game, the NPC in the virtual environment can be commanded to act according to the natural language command.

[0130] In some embodiments, the process of commanding the NPC is a process that is unlocked after the game account reaches a preset level. The game account is an account logged in by the player. For example, the preset level is level 10. When the game account has not reached the preset level, the NPC is an uncontrollable virtual character, that is, the player cannot command the NPC through voice. When the game account reaches the preset level, the NPC is converted into a controllable virtual character for the player, that is, the player can currently command the NPC through voice.

[0131] In step 330, the NPC is controlled to perform a virtual activity according to the behavior intention expressed by the natural language command.

[0132] Illustratively, the behavior intention is an operation intention or purpose expressed by the natural language command. Accurate understanding of the behavior intention helps to control the NPC to perform the virtual activity purposefully.

[0133] In some embodiments, the behavior intention is identified from the natural language command according to environment perception information. The environment perception information includes information perceived by at least one of the main virtual character and the NPC in the virtual environment.

[0134] Illustratively, since the natural language command is voice information received during the display of the virtual environment, the virtual environment can be associated with the natural language command for analysis when the natural language command is analyzed.

[0135] In some embodiments, the environment perception information is obtained based on information perceived by at least one of the main virtual character and the NPC in the virtual environment. Illustratively, the environment perception information is obtained based on information perceived by the main virtual character in the virtual environment; or, the environment perception information is obtained based on information perceived by the NPC in the virtual environment; or, the environment perception information is obtained based on information perceived by the main virtual character and the NPC in the virtual environment.

[0136] Optionally, the environment perception information is used to represent environment information perceived and obtained by the main virtual character and / or the NPC within an object perception range. The object perception range is a three-dimensional space range in which the main virtual character and / or the NPC has perception of other scene elements.

[0137] Optionally, the environment perception information includes a virtual element determined by emitting a ray, and can also include region structure information determined by analyzing the geometric structure and layout information of the virtual environment, and the like.

[0138] Illustratively, the virtual environment further includes a plurality of virtual elements, and in determining the environment perception information, virtual elements perceived by the host virtual character and / or the NPC in the virtual environment are determined as the environment perception information; or in determining the environment perception information, an object perception range of the host virtual character and / or the NPC in the virtual environment is determined, and virtual elements within the object perception range are determined as the environment perception information, and the like.

[0139] In some embodiments, the intent recognition is performed on the natural language command under the constraint of the environment perception information, a behavior intent is obtained, and the NPC is controlled to perform a virtual activity through the behavior intent.

[0140] Illustratively, since the environment perception information is information obtained by the host virtual character and / or the NPC, the information is closely related not only to the virtual environment but also to the host virtual character and / or the NPC in the virtual environment, and the analysis process of the natural language command is constrained through the environment perception information, which helps to finely obtain the intent information expressed in the natural language command.

[0141] Optionally, the behavior intent recognition process is performed through a pre-trained machine learning model, the environment perception information and the natural language command are taken as inputs of the machine learning model, and an encoding process and a decoding process are performed through the machine learning model to output the behavior intent.

[0142] Illustratively, the machine learning model includes at least one of a plurality of networks such as a recurrent neural network (RNN), a long short-term memory (LSTM), a gated recurrent unit (GRU), a convolutional neural network (CNN), and the like, which are not limited here.

[0143] To sum up, with the help of the environment perception information of the host virtual character and / or the NPC, the natural language command can be better understood, the analysis accuracy of the natural language command can be improved, and the accuracy of the behavior intent obtained through recognition can be improved, so that the behavior intent is more matched with the virtual environment, the invalid or incorrect command to the NPC is greatly avoided, the error operation rate of the player is reduced, the player's operation times in the virtual game are reduced to a certain extent, the data processing amount related to the operation is reduced, and the utilization rate of the computing resources and the human-computer interaction efficiency are improved.

[0144] In an optional embodiment, the control method of the terminal performing the virtual role is taken as an example for illustration. The terminal performs speech-to-text processing on the natural language command to obtain a command analysis result in the form of text, which can be displayed on the terminal interface. The behavior intention is recognized through the command analysis result in the form of text to control the NPC to perform the virtual activity. Illustratively, as shown in FIG. 4, the step 330 shown in FIG. 3 can also be implemented as steps 410 to 420.

[0145] The step 410 converts the natural language command into a command analysis result in the form of text.

[0146] Illustratively, the terminal receives the natural language command in the form of speech, and first converts the natural language command in the form of speech into a command analysis result in the form of text through the speech-to-text technology.

[0147] The command analysis result is obtained by performing speech-to-text on the natural language command based on the environment perception information.

[0148] Illustratively, in the process of performing speech-to-text on the natural language command, the environment perception information of the host virtual role and / or the NPC is used to constrain the conversion of the natural language command into the form of text, i.e., to obtain the command analysis result.

[0149] The command analysis result is used for behavior intention recognition.

[0150] Illustratively, the terminal performs the behavior intention recognition process based on the command analysis result in the form of text. The form of text is more intuitive and explicit than the form of speech, and is convenient for the accuracy of the behavior intention recognition.

[0151] The behavior intention is the intention information expressed in the command analysis result. For example, the command analysis result is “No. 1, move to A”, and the behavior intention is determined to be “No. 1” and “move to A” based on the behavior intention recognition. That is, the intention information for controlling the NPC to perform the virtual activity is obtained from the command analysis result through the behavior intention recognition.

[0152] In an optional embodiment, as shown in FIG. 5, the step 410 can be implemented as steps 411 to 412.

[0153] The step 411 obtains at least one first scene hotword based on the first virtual scene in which the host virtual role and / or the NPC is located.

[0154] The at least one first scene hotword is a scene-related vocabulary of the first virtual scene. That is, there is an association relationship between the first scene hotword and the first virtual scene, and the first scene hotword is a vocabulary describing the first virtual scene.

[0155] Optionally, the first scene hotword comprises a scene state of the first virtual scene, such as: the first virtual scene is a virtual restaurant, and the first scene hotword comprises a business state, a rest state, a closed state, and the like.

[0156] Optionally, the first scene hotword comprises an element name of a virtual element in the first virtual scene, such as: the first virtual scene is a virtual restaurant, and the first scene hotword comprises a virtual table, a virtual chair, a virtual kitchen utensil, a virtual dinner plate, and the like.

[0157] Optionally, the first scene hotword comprises an interactive vocabulary for interacting with a virtual element in the first virtual scene, such as: the first virtual scene is a virtual battlefield, and the first scene hotword comprises attack, attack a virtual character, defense, and build a barrier.

[0158] In some embodiments, at least one first scene hotword is collected based on the first virtual scene after receiving the natural language command; or, at least one first scene hotword corresponding to the first virtual scene is filtered from a plurality of scene hotwords obtained in advance based on the first virtual scene after receiving the natural language command, and the like.

[0159] In an optional embodiment, a first scene type corresponding to the first virtual scene in which the host virtual character and / or the NPC is located is obtained.

[0160] Illustratively, the first scene type is a scene type corresponding to the first virtual scene, and the scene type is used to represent a scene state of the virtual scene, such as: the scene type comprises an office type, a battle type, a kitchen type, a bedroom type, a classroom type, a store type, a factory type, and the like, which are various types of describing scene states.

[0161] Optionally, the first virtual scene corresponds to the first scene type, and the first scene type is used to describe a scene state represented by the first virtual scene, such as: the first virtual scene is a virtual office, and the first scene type is an office type; or, the first virtual scene is a virtual battle area, and the first scene type is a battle type; or, the first virtual scene is a virtual store, and the first scene type is a store type, and the like.

[0162] In an optional embodiment, a hotword set corresponding to the first scene type is obtained, and at least one vocabulary in the hotword set is taken as at least one first scene hotword, and the hotword set is a set of scene-related vocabularies collected based on the first scene type.

[0163] Illustratively, a plurality of scene types respectively correspond to a hotword set, and the hotword set represents a set of scene-related vocabularies collected based on the scene type, and also represents a set of scene-related vocabularies collected based on at least one virtual scene under the scene type.

[0164] For example, the office type corresponds to hotword set 1, and the hotword set 1 includes multiple scene-related words collected based on at least one virtual office scene under the office type, such as scene-related words collected based on virtual office 1 and scene-related words collected based on virtual office 2, where virtual office 1 and virtual office 2 are two virtual offices under the office type.

[0165] For example, the battle type corresponds to hotword set 2, and the hotword set 2 includes multiple scene-related words collected based on at least one virtual battle scene under the battle type, such as scene-related words collected based on virtual battle scene 1 and scene-related words collected based on virtual battle scene 2, where virtual battle scene 1 and virtual battle scene 2 are two virtual battle scenes under the battle type.

[0166] Optionally, after determining the first scene type corresponding to the first virtual scene, a hotword set corresponding to the first scene type is obtained. The hotword set is a set of scene-related words collected based on the first scene type, that is, the hotword set corresponding to the first scene type is a set of scene-related words collected based on the first scene type. For example, the first scene type of the first virtual scene is the battle type, and the hotword set 2 corresponding to the battle type is taken as the hotword set corresponding to the first scene type.

[0167] In some embodiments, the words in the hotword set are referred to as scene-related words, or as hotword words, or as words. After obtaining the hotword set corresponding to the first scene type, the words in the hotword set are taken as at least one first scene hotword, that is, when obtaining at least one first scene hotword based on the first virtual scene, multiple words in the hotword set corresponding to the first scene type to which the first virtual scene belongs are taken as at least one first scene hotword corresponding to the first virtual scene.

[0168] Illustratively, the first virtual scene is virtual kitchen 1, the scene type corresponding to virtual kitchen 1 is determined to be the kitchen type, a hotword set corresponding to the kitchen type is obtained, and words in the hotword set are taken as at least one first scene hotword. The hotword set corresponding to the kitchen type includes words obtained based on virtual kitchen 1 and words obtained based on virtual kitchen 2, so the at least one first scene hotword obtained in this way has a more extensive characteristic.

[0169] In an optional embodiment, an object perception range of the main virtual character and / or NPC in the first virtual scene is obtained.

[0170] Optionally, the object perception range of the master virtual role in the first virtual scene is acquired; or, the object perception range of the NPC in the first virtual scene is acquired; or, the object perception range of the master virtual role in the first virtual scene is acquired and the object perception range of the NPC in the first virtual scene is acquired.

[0171] The object perception range is a three-dimensional space range in which the master virtual role and / or the NPC perceives other scene elements.

[0172] Illustratively, the master virtual role and the NPC are virtual roles in the first virtual scene, which can be regarded as a kind of virtual element in the first virtual scene, and the first virtual scene can also include other scene elements in addition to the master virtual role and the NPC; the object perception of the master virtual role and / or the NPC represents the perception of the master virtual role and / or the NPC to other scene elements, and the object perception range is the range of the object perception, which is realized as a three-dimensional space range in the first virtual scene.

[0173] In some embodiments, the object perception range includes at least one of the following.

[0174] (1) The visual perception range of the master virtual role and / or the NPC.

[0175] It is generally used to describe the size of the area that can be covered or perceived visually by a living organism or a technical device. The visual perception range of the master virtual role and / or the NPC is a three-dimensional space range perceived by the master virtual role and / or the NPC as an observer through observation of the master virtual role and / or the NPC.

[0176] Optionally, the visual perception range is acquired based on the position of the master virtual role and / or the NPC in the first virtual scene and the visual observation capability.

[0177] Illustratively, the position is used to represent the position information of the master virtual role and / or the NPC relative to the first virtual scene, such as represented by position coordinates; and the visual observation capability is used to represent the ability of the master virtual role and / or the NPC to observe the first virtual scene, such as the visual observation capability including a visual distance (i.e., the farthest distance that can be observed and perceived) and a visual angle (i.e., the angle range that can be observed and perceived). By integrating the position and the visual observation capability, the visual perception range of the master virtual role and / or the NPC can be roughly determined from the first virtual scene.

[0178] In some embodiments, in the case where the object perception range includes the visual perception range, a preset visual cone is generated as the visual perception range, with the position of the master virtual role and / or the NPC as the starting point and the orientation of the master virtual role and / or the NPC as the central line of the cone.

[0179] The preset visual cone is a three-dimensional region for measuring the visual perception range.

[0180] Illustratively, when the object perception range includes the visual perception range, the position where the master virtual character and / or NPC is located and the orientation of the master virtual character and / or NPC are determined, and the preset visual cone is taken as the visual perception range determined based on the visual perception capability, with the position as the starting point and the orientation as the cone center line.

[0181] Optionally, the preset visual cone is a tetrahedron; or the preset visual cone is a prism (such as a large tetrahedron from which a small tetrahedron is cut to obtain the preset visual cone); or the preset visual cone is a circular cone; or the preset visual cone is a polygonal prism, etc.

[0182] The above describes the content that the visual perception range is a preset visual cone determined based on the position and orientation of the master virtual character and / or NPC. By generating the preset visual cone as the visual perception range based on the current position and orientation of the master virtual character and / or NPC, the visual field perception of the virtual character can be accurately simulated, and the perception capability of the virtual character in the virtual environment can be ensured. The preset visual cone is determined by considering the orientation and position of the virtual character, so that the visual perception range of the virtual character can be dynamically defined, the visual behavior of a human or an animal in the physical world can be simulated, and thus the behavior of the virtual character in the virtual scene is more consistent with the real logic. Especially in a complex environment, the virtual character can identify potential threats through visual perception and make corresponding behavior adjustments. Therefore, the preset visual cone as the visual perception range greatly enhances the immersion of the virtual environment, ensures the operation experience of the player, and also helps to provide more intelligent, flexible and realistic behavior control of the virtual character, improve the utilization rate of computing resources and the efficiency of human-computer interaction.

[0183] (2) The auditory perception range of the master virtual character and / or NPC.

[0184] Illustratively, the auditory perception range refers to the spatial range that a listener can perceive or understand in the auditory sense, and is usually used to describe the range of sounds that a biological body or a technical device can cover or perceive in the auditory sense. The auditory perception range of the master virtual character and / or NPC refers to the three-dimensional spatial range perceived by the master virtual character and / or NPC as a listener through the auditory perception capability of the master virtual character and / or NPC.

[0185] Optionally, the auditory perception range is obtained based on the position of the master virtual character and / or NPC in the first virtual scene and the auditory perception capability.

[0186] Illustratively, the position is used to represent the position information of the host virtual character and / or NPC relative to the first virtual scene, such as represented by position coordinates; the auditory perception capability is used to represent the ability of the host virtual character and / or NPC to perceive sound changes in the first virtual scene, such as the auditory perception capability including an auditory distance (i.e., the farthest distance of a sound that can be heard), an auditory angle (i.e., the direction of a sound source or the coverage angle range of a sound that can be heard), a frequency range (i.e., the frequency range of a sound that can be perceived), and the like. In combination with the position and the auditory perception capability, the auditory perception range of the host virtual character and / or NPC can be roughly determined from the first virtual scene.

[0187] (3) The olfactory perception range of the host virtual character and / or NPC.

[0188] Illustratively, the olfactory perception range refers to the spatial range of a smell that can be perceived or recognized in an olfactory manner. The olfactory perception range of the host virtual character and / or NPC refers to the three-dimensional spatial range that can be perceived by the host virtual character and / or NPC in an olfactory manner.

[0189] Optionally, based on the position of the host virtual character and / or NPC in the first virtual scene and the olfactory perception capability, the olfactory perception range is obtained.

[0190] Illustratively, the position is used to represent the position information of the host virtual character and / or NPC relative to the first virtual scene, such as represented by position coordinates; the olfactory perception range is used to represent the ability of the host virtual character and / or NPC to perceive smell changes in the first virtual scene, such as the olfactory perception range including an olfactory sensitivity (i.e., the ability to react to odor molecules to different degrees), an olfactory perception distance (i.e., the distance from the source of the smell to the host virtual character and / or NPC), and the like. In combination with the position and the olfactory perception capability, the olfactory perception range of the host virtual character and / or NPC can be roughly determined from the first virtual scene.

[0191] (4) The perception skill possessed by the host virtual character and / or NPC or the range of perception of the perception skill.

[0192] Illustratively, the perception skill is the ability of the host virtual character and / or NPC to perform scene perception on the first virtual scene, and the perception skill includes at least one of a perception acquisition form and a perception enhancement form.

[0193] Optionally, the perception acquisition form is used to represent a process of acquiring a perception skill that is not possessed. After the perception skill is acquired through the perception acquisition form, the host virtual character and / or NPC can perform targeted scene perception on the first virtual scene through the perception skill. The perception enhancement form is used to represent a process of enhancing a perception capability through the possession of a perception skill in the case of possessing the perception capability.

[0194] Illustratively, the perception prop is a virtual prop for perceiving the first virtual scene, and the perception prop comprises at least one of a perception obtaining form and a perception enhancing form; the perception prop in the perception obtaining form is a first perception prop, and the perception prop in the perception enhancing form is a second perception prop.

[0195] Optionally, the host virtual character and / or the NPC can obtain the perception prop based on completion of a specific game task or reaching a certain game level. Optionally, different perception skills or perception props each correspond to a preset effective range, representing a range interval that can be perceived when the perception skill or the perception prop is applied, and the preset effective range represents the object perception range.

[0196] The above describes an illustrative implementation of the object perception range. The visual perception range, the auditory perception range, and the olfactory perception range are real perception capabilities simulated by the host virtual character and / or the NPC, and diversified perception channels help to more comprehensively understand and respond to dynamic changes in the virtual environment, and improve the naturalness of interaction and immersion of the virtual environment. In addition, the host virtual character and / or the NPC can also extend its perception capabilities to fields beyond traditional senses through perception skills or perception props, and setting diversified perception methods can further improve the intelligent response capability of the virtual character, and also make its behavior in a complex scene more flexible and more consistent with actual operation logic. Therefore, multiple object perception ranges not only enhance the adaptability of the virtual character to the virtual environment, but also improve the depth of interaction between the player and the virtual environment, and improve the efficiency of human-computer interaction.

[0197] In some embodiments, the object perception range is determined based on a first position of the host virtual character in the first virtual scene.

[0198] In some embodiments, the object perception range is determined based on a second position of the NPC in the first virtual scene.

[0199] In some embodiments, when the host virtual character and the NPC are in the same first virtual scene, a first object perception range of the host virtual character in the first virtual scene is obtained, and a second object perception range of the NPC in the first virtual scene is obtained, and the first object perception range and the second object perception range are collectively taken as the object perception range.

[0200] It is worth noting that the above object perception range is only an illustrative example, and the embodiments of the present application are not limited thereto.

[0201] In an optional embodiment, at least one first scene hotword is obtained based on environmental perception information in the object perception range.

[0202] The environmental perception information is used to represent the environmental information perceived and obtained by the host virtual character and / or the NPC within the object perception range. Optionally, the environmental perception information includes virtual elements determined by ray casting, and can also include region structure information determined by analyzing the geometric structure and layout information of the virtual environment, etc.

[0203] Optionally, based on the range of the object perception range in the first virtual scene, and the environmental perception information determined by the host virtual character and / or the NPC within the object perception range based on perception, at least one first scene hotword representing the state of the host virtual character and / or the NPC is obtained.

[0204] In an optional embodiment, the environmental perception information includes virtual elements; and the virtual elements within the object perception range in the first virtual scene are determined.

[0205] Optionally, the object perception range is a part of the three-dimensional space range in the first virtual scene; and the first virtual scene includes a large number of virtual elements, and the virtual elements are elements constituting the virtual scene, such as virtual ground, virtual buildings, virtual trees, virtual characters, virtual stones, etc.

[0206] Optionally, in the case of determining the object perception range, the virtual elements within the object perception range are determined from a plurality of virtual elements corresponding to the first virtual scene.

[0207] Optionally, as shown in FIG. 6, taking the visual perception range as an example, if the first virtual scene in which the host virtual character is located is displayed in the first-person perspective of the host virtual character, a frustum as shown in FIG. 6 can be determined as the visual perception range based on the position and orientation of the host virtual object.

[0208] Optionally, the position and orientation of the host virtual object are represented based on the position and orientation of the virtual camera 610; and a tetrahedron frustum as shown in FIG. 6 is obtained as the object perception range with the position of the virtual camera 610 as the starting point and the orientation of the virtual camera as the median line of the frustum.

[0209] Alternatively, two limiting planes in the preset camera rendering, i.e., the near clipping plane 621 and the far clipping plane 622 in FIG. 6, are parallel to the XY plane of the virtual camera, and are separated by a certain distance along the center line. Any virtual element closer to the virtual camera than the near clipping plane 621 and any virtual element farther from the virtual camera than the far clipping plane 622 will not be rendered, i.e., a small tetrahedron 632 is cut from a large tetrahedron 631 as shown in FIG. 6, and a prism is obtained as the frustum.

[0210] Illustratively, the distance between the far clipping plane 622 and the virtual camera 610 is the farthest distance observed by the master virtual character, and the frustum is taken as the object perception range, which includes virtual grass, virtual buildings, virtual cars and other virtual elements in the three-dimensional scene, i.e., the virtual car, the virtual grass, the virtual building and other virtual elements are virtual elements in the object perception range, as shown in FIG. 6, and the multiple virtual elements are presented by projecting and displaying the multiple virtual elements on the far clipping plane 622. In fact, the multiple virtual elements can be scattered and displayed at multiple positions in the visual perception range (such as the cuboid surrounded by the near clipping plane 621 and the far clipping plane 622).

[0211] Optionally, if the visual perception range is determined as the object perception range in the third-person perspective, the virtual camera is located in a region outside the master virtual character, and can not only observe the multiple virtual elements, but also observe the master virtual character moving in the virtual environment. At this time, if the visual perception range is determined, the range captured by the virtual camera can be taken as the visual perception range (wherein the virtual camera moves with the master virtual character, and the visual perception range is still related to the master virtual character); or, a frustum determined according to the position and orientation of the master virtual character is taken as the visual perception range (refer to FIG. 6, which takes the observation perspective of the master virtual character as an example, and the observation perspective of the virtual camera 610 is consistent). The determination process of the visual perception range is not limited here.

[0212] In some embodiments, the virtual elements in the first virtual scene within the object perception range are highlighted.

[0213] Optionally, after the virtual elements within the object perception range are determined from the first virtual scene, the virtual elements within the object perception range are highlighted, so as to distinguish and present different virtual elements within and outside the object perception range, and also help to prompt the player to analyze the situation of the virtual elements related to the natural language command.

[0214] In an optional embodiment, the element name of the virtual element is taken as the first scene hotword; or, the action name of the interactive action corresponding to the virtual element is taken as the first scene hotword.

[0215] Illustratively, the virtual element has an element name, and the element name is used to represent the virtual element itself. The element name can be one or multiple, i.e., multiple element names can all represent the virtual element. For example, the element name of the virtual element 621 is truck or car; and the element name of the virtual element 622 is grass or grass cluster.

[0216] Optionally, after determining the virtual elements within the object perception range, the element names of the virtual elements are summarized, and the multiple element names are all taken as the first scene hot words corresponding to the first virtual scene, to obtain at least one first scene hot word.

[0217] Illustratively, during the game process, the player can command the master virtual character or the NPC to interact with the virtual elements, i.e., there are interaction actions to implement the interaction process, and the action names of the interaction actions are taken as the first scene hot words.

[0218] For example, the element name of the virtual element 621 is a truck or a car, the player can control the master virtual character or command the NPC to interact with the virtual element 621, such as taking a car, getting off a car, hitting a car, climbing onto a car, etc., which are the action names of the interaction actions; or the element name of the virtual element 622 is grass or a grass clump, the player can control the master virtual character or command the NPC to interact with the virtual element 622, such as hiding in a grass clump, trimming a grass clump, treading on a grass clump, etc., which are the action names of the interaction actions, etc.

[0219] Optionally, the action names that can interact with the virtual elements are summarized, and the at least one action name corresponding to each of the multiple virtual elements is taken as the first scene hot word corresponding to the first virtual scene, to obtain at least one first scene hot word.

[0220] In some embodiments, the first scene hot word corresponding to the virtual element is highlighted beside the virtual element.

[0221] Illustratively, after determining the first scene hot words corresponding to the multiple virtual elements, the first scene hot words associated with the virtual elements can also be displayed beside the multiple virtual elements in the interface. For example, the text word "stable" is displayed beside the virtual element "stable", and the text word "attack him" is displayed beside the virtual element "enemy virtual character", etc.

[0222] The above describes a method of highlighting virtual elements within the object perception range of the virtual character and using the element name or action name as the first scene hotword. This process can significantly enhance the intelligent interaction experience in the virtual environment. Highlighting virtual elements helps the virtual character quickly identify and focus on key elements related to the current task or scene, reducing the complexity of information processing and improving decision-making efficiency. Secondly, using the element name of the virtual element or the action name of its interactive action as the first scene hotword helps to create a stronger association between the virtual character and the subsequent operation behavior through natural language commands. This allows the virtual character to flexibly respond to changes in natural language commands and even scene changes, providing more accurate feedback and responses. This process also makes the behavior connection between virtual elements and virtual characters in the virtual scene more closely, greatly enhancing the authenticity and smoothness of the interaction, making the virtual character more flexible, natural, and accurate in executing virtual activities in a constantly changing virtual environment, ensuring human-computer interaction efficiency.

[0223] In an optional embodiment, a first scene type corresponding to the first virtual scene is obtained; and a hotword set corresponding to the first scene type is obtained.

[0224] Illustratively, the first scene type is a scene type corresponding to the first virtual scene, and the scene type is used to represent the scene state of the virtual scene. For example, the scene type includes office type, battle type, kitchen type, bedroom type, classroom type, store type, factory type, and other types of scene state descriptions.

[0225] Illustratively, a plurality of scene types correspond to a hotword set respectively, and the hotword set represents a set of scene-related words collected based on the scene type, and also represents a set of scene-related words collected based on at least one virtual scene under the scene type.

[0226] Optionally, after determining the first scene type corresponding to the first virtual scene, the hotword set corresponding to the first scene type is obtained. The hotword set is a set of scene-related words collected based on the first scene type. For example, the first scene type of the first virtual scene is the battle type, and the hotword set 2 corresponding to the battle type is used as the hotword set corresponding to the first scene type.

[0227] In an optional embodiment, the object perception range of the main virtual character and / or NPC in the first virtual scene is obtained.

[0228] Illustratively, the master virtual role and / or the NPC includes at least one of the master virtual role and the NPC. When the master virtual role and / or the NPC is obtained based on the master virtual role and / or the NPC, when the master virtual role and / or the NPC is the master virtual role, the object perception range of the master virtual role in the first virtual scene is obtained; when the master virtual role and / or the NPC is the NPC, the object perception range of the NPC in the first virtual scene is obtained; or, when the master virtual role and / or the NPC is the master virtual role and the NPC, the object perception range of the master virtual role in the first virtual scene is obtained and the object perception range of the NPC in the first virtual scene is obtained.

[0229] Optionally, the object perception range includes at least one of a visual perception range, an auditory perception range, an olfactory perception range of the master virtual role and / or the NPC, and a range perceived by a perception skill or a perception device possessed by the master virtual role and / or the NPC.

[0230] In an optional embodiment, at least one word is obtained from the hot word set as at least one first scene hot word based on the environmental perception information in the object perception range.

[0231] Optionally, after the object perception range is determined, a virtual element in the object perception range is determined, and a word having an element association relationship with the virtual element is obtained from the hot word set as at least one first scene hot word.

[0232] Illustratively, the element association relationship is used to represent a relationship embodying a state of the virtual element. For example, a word including an element name of the virtual element is obtained from the hot word set as the first scene hot word; or, a word used to trigger the virtual element is obtained from the hot word set as the first scene hot word; or, an interactive word used to interact with the virtual element is obtained from the hot word set as the first scene hot word, etc.

[0233] The above describes determining the first scene hotword based on at least one of the first scene type and the object perception range. Among them, by obtaining the first scene type in which at least one of the host virtual role and the NPC virtual role is located and the corresponding hotword set, the interactive content range can be dynamically adjusted according to the characteristics of the scene and the position of the role, so as to accurately understand the key elements and potential information in the first virtual scene, so that the virtual role can accurately identify the high-frequency words related to the first virtual scene, improve the intelligent perception and interaction ability of the virtual environment, and enhance the interactive experience of the virtual role. In addition, based on the object perception range, the environment perception information can be dynamically obtained and identified according to the field of view and the interaction range of the virtual role, so as to further improve the reaction ability and task execution efficiency of the virtual role in the virtual environment through the screened first scene hotword. In addition, the first scene hotword is selected by comprehensively selecting the hotword set and the object perception range, which can more specifically realize the purpose of hotword screening in the first virtual scene, improve the intelligent level of the virtual role in the virtual environment, and ensure the accuracy of the subsequent command analysis result based on the first scene hotword.

[0234] It is worth noting that the above is only an illustrative example, and the embodiments of the present application do not limit this.

[0235] In an optional embodiment, based on the first virtual game participated by the host virtual role and / or the NPC, a first virtual scene corresponding to the first virtual game is determined. Among them, a plurality of virtual games respectively correspond to a virtual scene, and the plurality of virtual games are simulation battle environments provided for the host virtual role and / or the NPC.

[0236] Illustratively, in addition to determining the first virtual scene based on the position of the host virtual role and / or the NPC in the virtual environment, the virtual scene corresponding to the first virtual game participated by the host virtual role and / or the NPC can also be taken as the first virtual scene. For example: the virtual environment includes a plurality of virtual games, each virtual game corresponds to a virtual scene (such as virtual game 1 is a virtual snow scene, virtual game 2 is a virtual forest scene, etc.), if the virtual game participated by the host virtual role and / or the NPC is the first virtual game, the virtual scene corresponding to the first virtual game is taken as the first virtual scene (if the host virtual role participates in the virtual game 1 with the enemy virtual role in the virtual environment, the virtual scene is taken as the first virtual scene, etc.); or, a plurality of virtual games are provided before the game starts, each virtual game corresponds to a virtual scene, if the player selects the host virtual role to participate in the first virtual game, the NPC participates in the first virtual game with the host virtual role, and the virtual scene corresponding to the first virtual game is taken as the first virtual scene, etc.

[0237] Optionally, the virtual match includes various forms such as game tasks (e.g., task 1 corresponds to virtual scene 1, task 2 corresponds to virtual scene 2, etc.), game levels (e.g., game level 1 corresponds to virtual scene 1, game level 2 corresponds to virtual scene 2, etc.), and the like, which are not limited here.

[0238] In an optional embodiment, a plurality of scene hotwords are acquired, and the plurality of scene hotwords correspond to at least two virtual scenes.

[0239] Illustratively, a plurality of scene hotwords corresponding to a plurality of virtual scenes in a virtual environment are acquired in advance. Optionally, based on virtual elements in each virtual scene, at least one scene hotword corresponding to each virtual scene is acquired, and the plurality of scene hotwords are stored.

[0240] Optionally, an element name of a virtual element in a virtual scene is acquired as a scene hotword; or, an action name of an interactive action interacting with the virtual element in the virtual scene is acquired as a scene hotword; or, a state name describing a state of the virtual element in the virtual scene is acquired as a scene hotword, and the like.

[0241] In an optional embodiment, based on a first virtual scene in which the host virtual character and / or NPC is located, at least one first scene hotword corresponding to the first virtual scene is acquired from the plurality of scene hotwords.

[0242] Optionally, the plurality of scene hotwords correspond to scene identifiers, and the scene identifiers are used to represent virtual scenes when the scene hotwords are acquired. For example, if scene hotword 1 and scene hotword 2 are acquired based on virtual scene A, scene identifier a corresponding to virtual scene A is marked for scene hotword 1, and scene identifier a is also marked for scene hotword 2; if scene hotword 3 is acquired based on virtual scene B, scene identifier b corresponding to virtual scene B is marked for scene hotword 3, and the like.

[0243] Optionally, based on a first virtual scene in which the host virtual character and / or NPC is located, a plurality of candidate scene hotwords having a first scene identifier are acquired from the plurality of scene hotwords, and the first scene identifier is a scene identifier corresponding to the first virtual scene.

[0244] Illustratively, after the first virtual scene in which the host virtual character and / or NPC is located is determined, each of the plurality of scene hotwords is marked with a scene identifier, so that scene hotwords marked with the first scene identifier can be screened from the plurality of scene hotwords as candidate scene hotwords.

[0245] The above describes the content of obtaining at least one first scene hotword based on the first virtual scene from the obtained multiple scene hotwords. This process can significantly improve the dynamic adaptability of the virtual environment and improve the interaction quality of the virtual character in the virtual environment. By associating the specific hotword of each virtual scene, the key elements or activities related to the current scene can be accurately identified and located. When the host virtual character or NPC is located in a specific virtual scene, the system can automatically filter out the most relevant hotword from the preset multiple scene hotwords, further accurately adjust the character behavior or dialogue, and ensure the consistency between the behavior of the virtual character and the virtual scene. In addition, dynamically obtaining scene hotwords also helps to adjust the actions and reactions of virtual characters in real time through the changes of virtual scenes, providing strong support for intelligent interaction in complex virtual environments, ensuring interaction accuracy and human-computer interaction efficiency.

[0246] In some embodiments, based on the object perception range of the host virtual character and / or NPC in the first virtual scene, at least one candidate scene hotword is selected from the multiple candidate scene hotwords as the first scene hotword.

[0247] Illustratively, after filtering multiple candidate scene hotwords based on the first virtual scene, the object perception range is part of the first virtual scene, so the multiple candidate scene hotwords can be further filtered based on the object perception range, and at least one candidate scene hotword related to the object perception range is selected as the first scene hotword.

[0248] Optionally, a virtual element within the object perception range is determined as a perceived virtual element, i.e. the perceived virtual element is a virtual element that the host virtual character and / or NPC can perceive in the first virtual scene; the element name of the perceived virtual element is filtered from the multiple candidate scene hotwords as the first scene hotword; or the action name of the interaction action corresponding to the perceived virtual element is filtered from the multiple candidate scene hotwords as the first scene hotword, etc.

[0249] The above describes the content of further obtaining the first scene hotword from the filtered multiple candidate scene hotwords by integrating the first virtual scene and the object perception range. Among them, the introduction of the scene identifier establishes a clear association between each hotword and the virtual scene, so that the relevant hotword can be quickly identified based on the first scene identifier of the first virtual scene in the virtual environment, avoiding confusion between the virtual scene and the hotword, ensuring the accuracy of the selection of the hotword, and preliminarily ensuring the interaction accuracy between the virtual character and the virtual environment. Secondly, further filtering the candidate scene hotword determined based on the first scene identifier based on the object perception range of the character can dynamically adjust the priority of the hotword, so that the virtual character can interact more effectively with the virtual elements in the current virtual scene, improve the flexibility of the virtual character interaction and the player operation experience, and make full use of computing resources to achieve more accurate and complete interaction purposes.

[0250] At step 412, the natural language command is converted into a command analysis result in text form based on the at least one first scene hotword.

[0251] In an optional embodiment, the first virtual scene is one of a plurality of virtual scenes in a virtual environment; the virtual environment is a three-dimensional simulation environment in which the host virtual character and the NPC jointly participate in a virtual game.

[0252] Illustratively, a plurality of virtual characters participating in the virtual game are in a virtual environment, and the virtual environment includes a plurality of virtual scenes, which collectively constitute the virtual environment, for example, the virtual environment includes virtual kitchen 1, virtual kitchen 2, virtual office, and virtual battle scene, etc.

[0253] Optionally, based on the positions of the host virtual character and / or the NPC in the virtual environment, the first virtual scene in which the host virtual character and / or the NPC is located is determined from the plurality of virtual scenes.

[0254] Optionally, the virtual scene in which the host virtual character is located is determined as the first virtual scene from the plurality of virtual scenes; or the virtual scene in which the NPC is located is determined as the first virtual scene from the plurality of virtual scenes; or the virtual scene in which the host virtual character and the NPC jointly are located is determined as the first virtual scene from the plurality of virtual scenes.

[0255] The above describes the determination of the first virtual scene from the plurality of virtual scenes in the virtual environment. Taking the virtual scene in which the host virtual object or the NPC is located as the first virtual scene can ensure the uniqueness of the first virtual scene and facilitate targeted analysis based on the first virtual scene; taking the virtual scene in which the host virtual object is located and the virtual scene in which the NPC is located as the first virtual scene can ensure the comprehensiveness of the determination of the first virtual scene and facilitate more comprehensive control of the virtual object based on the first virtual scene; taking the virtual scene in which the host virtual object and the NPC jointly are located can help achieve flexibility in tactical strategies in the virtual environment, so that the host virtual object and the NPC cooperate to produce more rich operations.

[0256] In addition, automatically identifying the virtual scene in which the virtual character is located based on the position of the virtual character can achieve precise environment adaptation and scene switching. When the virtual character moves to the corresponding scene, the information of the scene in which it is located can be updated in real time, ensuring that the interaction between the virtual character and the virtual environment remains consistent. Secondly, this kind of dynamic scene selection mechanism also helps the virtual character to make reasonable decisions and behavioral responses according to the characteristics of the scene in which it is located, improving the naturalness and logic of the interaction. In addition, this process also helps to avoid the abruptness caused by scene switching, making the switching between virtual scenes smoother and more natural, improving the operation experience of the player and enhancing the dynamicity and intelligence of the virtual environment.

[0257] In some embodiments, in response to the host virtual character and / or the NPC moving from the first virtual scene to a second virtual scene without obtaining the command analysis result, at least one second scene hotword corresponding to the second virtual scene is obtained.

[0258] Illustratively, the virtual environment includes a plurality of virtual scenes, and the second virtual scene is another virtual scene different from the first virtual scene. The host virtual character and / or the NPC are virtual characters participating in the virtual game, and their positions can change in the virtual environment, such as moving from A to B. If the host virtual character and / or the NPC move from the first virtual scene to the second virtual scene during the process of receiving the natural language command but before obtaining the command analysis result, the scene hotword used to constrain the generation of the command analysis result needs to be determined based on the second virtual scene that the host virtual character and / or the NPC recently moved to.

[0259] For example, during the process of receiving the natural language command but before generating the command analysis result, if the host virtual character and / or the NPC move from the first virtual scene where they are currently located to the second virtual scene, i.e., the positions of the host virtual character and / or the NPC move from the first virtual scene to the second virtual scene, in order to more flexibly analyze the situation of the virtual scene after the movement, at least one second scene hotword corresponding to the second virtual scene is obtained.

[0260] In some embodiments, the natural language command is converted into a command analysis result in text form based on the at least one second scene hotword.

[0261] Illustratively, after obtaining the at least one second scene hotword, the natural language command is converted into a command analysis result in text form by a pre-trained natural language analysis model with the at least one second scene hotword as a constraint.

[0262] Optionally, the natural language command and the at least one second scene hotword are taken as inputs of the model, the instruction feature representation of the natural language command and the hotword feature representation of the at least one second scene hotword are extracted by the natural language analysis model, the instruction feature representation and the hotword feature representation are fused to obtain a fusion feature representation, and the command analysis result in text form is output based on the fusion feature representation.

[0263] In an optional embodiment, a pre-trained natural language analysis model is obtained.

[0264] The pre-trained natural language analysis model includes an acoustic network, a language network, and a preset dictionary, and the preset dictionary includes at least one first scene hotword, other scene hotwords, and general vocabulary.

[0265] Optionally, the natural language analysis model can also be referred to as a natural language analysis system, and the acoustic network, the language network and the preset dictionary in the natural language analysis model can also be referred to as an acoustic model, a language model and a preset dictionary in the natural language analysis system. Here, the natural language analysis model and the related network are not limited in terms.

[0266] Illustratively, the acoustic network can also be referred to as an acoustic model (Acoustic Model) for processing input in the form of speech to convert the sound signal into potential speech units such as phonemes or segments. A phoneme is the smallest unit of speech in a language and is a phonetic unit. The acoustic network can analyze natural language commands in the form of speech based on the acoustic features of the natural language instructions (such as frequency spectrum, tone, speech rate, etc.) to identify possible sequences of speech units. The acoustic network can learn the mapping from acoustic features to speech units through machine learning techniques such as deep neural networks (Deep Neural Network, DNN) or recurrent neural networks (Recurrent Neural Network, RNN).

[0267] Illustratively, the language network can also be referred to as a language model (Language Model) for improving the probability score of a language unit sequence with high possibility when understanding text input. The language network predicts the rationality and fluency of a given sequence of words or speech units based on the grammatical rules and statistical information of the language. The language network includes n-gram model (n-gram), recurrent neural network language model (Recurrent Neural Network Language Model, RNNLM) or transformer model (Transformer) for modeling and generating text sequences.

[0268] Illustratively, the preset dictionary contains words and their corresponding pronunciation or language feature information that may be encountered in speech recognition or text understanding. It provides a way to map text words to their pronunciation or language features. In speech recognition, the preset dictionary helps to reduce possible pronunciation ambiguity and improve system recognition accuracy.

[0269] Among them, the preset dictionary includes at least one first scene hot word, other scene hot word and general vocabulary.

[0270] Illustratively, the first-scene hotword is a hotword corresponding to the first virtual scene, the other hotwords are hotwords corresponding to other virtual scenes except the first virtual scene, and the general words are words frequently used in daily life, such as “we”, “I”, “here”, “there”, “big”, “small”, and the like. The preset dictionary is a collection of a plurality of words pre-collected, which includes the first-scene hotword, the other hotwords, and the general words.

[0271] In an optional embodiment, the plurality of phonetic units corresponding to the natural language command are obtained through an acoustic network.

[0272] The phonetic unit is a basic constituent unit of pronunciation of a word.

[0273] In some embodiments, an instruction feature representation corresponding to the natural language command is extracted.

[0274] Illustratively, the feature extraction process is to divide the continuous sound signal, i.e., the natural language command, into a plurality of short time periods, and to calculate acoustic features such as frequency spectrum, pitch, speech rate, and the like in the plurality of time periods, each time period corresponding to an acoustic feature. The instruction feature representation can represent an acoustic feature sequence obtained by arranging the plurality of acoustic features according to the time periods, and can also represent a feature representation obtained by fusing the plurality of acoustic features, which is not limited here.

[0275] In some embodiments, the instruction feature representation is analyzed through an acoustic network.

[0276] Illustratively, taking the instruction feature representation representing the plurality of acoustic features as an example, the acoustic network aims to map the continuous acoustic feature sequence to a phonetic unit sequence, such as a phoneme or a segment. It learns to predict the possible phonetic unit (phoneme or segment) sequence given the acoustic features.

[0277] Optionally, the acoustic network outputs one or more possible phonetic unit sequences, and the phonetic unit sequence includes a plurality of phonetic units. The different phonetic unit sequences can have different phonetic units, or can only have different orders of the phonetic units. The plurality of phonetic unit sequences represent the arrangement of the phonetic units considered by the language network to be the most possible when understanding the input natural language command.

[0278] In some embodiments, in the process of analyzing the instruction feature representation through the acoustic network, a context analysis network is additionally introduced, which is used to analyze the context and is helpful to obtain a more reliable phonetic unit sequence by considering the time sequence dependency between the plurality of acoustic features.

[0279] In an optional embodiment, the sequence matching relationship between the plurality of speech units and the selected vocabulary in the preset vocabulary dictionary is analyzed by a language network to obtain a command analysis result in a text form.

[0280] Illustratively, the preset vocabulary dictionary includes a plurality of vocabularies, each of which corresponds to at least one speech unit. After obtaining the sequence of at least one speech unit, the sequence of at least one speech unit can be matched with the vocabularies in the preset vocabulary dictionary to analyze at least one vocabulary represented by the sequence of speech units.

[0281] Optionally, considering that the at least one first scene hotword is obtained based on the first virtual scene in advance, the first scene hotword is used to constrain the command analysis result, so that at least the at least one first scene hotword can be used as a selected vocabulary to be applied in the preset vocabulary dictionary matching process when matching the vocabulary by the sequence of speech units; in addition, considering that some common general vocabularies are also involved in the speech recognition process, the general vocabularies can also be used as selected vocabularies in the vocabulary matching, that is, the sequence of at least one speech unit is matched with the first scene hotword and the general vocabulary in the selected vocabulary to analyze at least one vocabulary represented by the sequence of speech units, and the at least one vocabulary is more likely to include the first scene hotword, avoiding the interference of other scene hotwords. That is, the selected vocabulary at least includes the at least one first scene hotword and the general vocabulary.

[0282] The above describes the content of obtaining the command analysis result by the pre-trained natural language analysis model. With the acoustic network, the language network and the preset vocabulary dictionary in the natural language analysis model, the understanding accuracy and response efficiency of the natural voice command in the virtual environment can be greatly improved. The acoustic network extracts the basic pronunciation unit through the analysis of the natural voice command, so as to accurately recognize the voice input of the player and provide high-quality voice data support for subsequent language processing. The language network realizes deep analysis at the semantic level by matching the sequence between the speech unit and the vocabulary in the preset vocabulary dictionary. The preset vocabulary dictionary not only includes the hotword related to the current virtual scene, but also includes the general vocabulary, so that the natural voice command of the player can be recognized and understood in a wider context. This multi-level analysis method makes the analysis of the natural voice command more detailed and intelligent. Therefore, the combination of natural language processing technology can provide a more natural and smooth human-computer interaction experience for the field of virtual reality, intelligent assistants and the like, and enhance the immersion and operation convenience of the user.

[0283] In an optional embodiment, the selected vocabulary also includes other scene hotwords.

[0284] Illustratively, if the selected vocabulary is the first-scene hotword and the general vocabulary, the at least one phonetic unit sequence can be matched with the first-scene hotword and the general vocabulary; if the selected vocabulary is the first-scene hotword, the other-scene hotword and the general vocabulary, the at least one phonetic unit sequence can be matched with the first-scene hotword, the other-scene hotword and the general vocabulary.

[0285] The vocabulary in the preset dictionary includes an analysis weight, a first analysis weight of the at least one first-scene hotword is higher than a second analysis weight of the other-scene hotword, and the analysis weight refers to a degree of attention paid to the vocabulary when the vocabulary participates in sequence matching.

[0286] Illustratively, the vocabulary in the preset dictionary corresponds to an analysis weight respectively, the greater the analysis weight, the greater the degree of attention paid to the vocabulary when the vocabulary participates in sequence matching, i.e., the vocabulary is more likely to be paid attention to; the smaller the analysis weight, the smaller the degree of attention paid to the vocabulary when the vocabulary participates in sequence matching, i.e., the vocabulary is less likely to be paid attention to; the degree of attention paid to different vocabularies in the sequence matching process can be adjusted through the analysis weight, so that the vocabulary with a greater analysis weight is preferentially matched with the vocabulary sequence, and the vocabulary with a greater analysis weight is also more likely to be selected as a vocabulary constituting the command analysis result.

[0287] Optionally, at least one first-scene hotword is obtained based on a first virtual scene in which the host virtual character and / or the NPC is located, the first-scene hotword can better reflect the situation of the first virtual scene and has a greater probability of having an association relationship with the natural language command, so that a first analysis weight of the first-scene hotword is set to be higher than a second analysis weight of the other-scene hotword, so that the first-scene hotword is more likely to be paid attention to when the vocabulary participates in sequence matching, and the first-scene hotword is also more likely to be selected as a vocabulary constituting the command analysis result.

[0288] Optionally, the general vocabulary corresponds to a third analysis weight, the third analysis weight is less than the first analysis weight and greater than the second analysis weight; or the third analysis weight is equal to the first analysis weight and greater than the second analysis weight; or the third analysis weight is less than the first analysis weight and equal to the second analysis weight; or the third analysis weight is less than the first analysis weight and less than the second analysis weight, and the value of the third analysis weight is not limited here.

[0289] The first analysis weight, the second analysis weight and the third analysis weight are usually values between 0 and 1, and the specific value of the analysis weight is not limited here as long as the first analysis weight is greater than the second analysis weight.

[0290] The above introduces the way that the first analysis weight of the first scene hot word is higher than the second analysis weight of other scene hot words, and focuses more on the content of the first scene hot word during voice analysis. The introduction of the vocabulary analysis weight mechanism enables the natural language processing model to more accurately process the scene hot words corresponding to different virtual scenes respectively, thereby improving the intelligent level of voice recognition and semantic understanding. By assigning different analysis weights to the words in the preset dictionary, the attention of the words in the sequence matching process can be automatically adjusted according to the importance of the words, optimizing the voice recognition process and reducing the possibility of misrecognition. In addition, this process also enables the system to better cope with the diversity of scene changes and user needs, ensuring the smoothness and naturalness of voice interaction, promoting the continuous development of intelligent voice technology, and improving the efficiency of human-computer interaction.

[0291] In an optional embodiment, the matching relationship between the plurality of voice units and the selected words in the preset dictionary is analyzed to obtain a plurality of candidate word sequences.

[0292] Illustratively, the process of analyzing the matching relationship represents the process of analyzing the matching between the selected words and the voice units. Illustratively, each selected word corresponds to at least one voice unit, and by analyzing the matching between the plurality of voice units and the selected words, the conversion probability of the voice units converting to the selected words is determined. The conversion probability represents the probability of matching between the selected words and the voice units. The greater the conversion probability, the closer the voice unit to the selected word, i.e., the greater the possibility that the voice unit may express the selected word, and the stronger the matching relationship between the selected word and the voice unit.

[0293] Optionally, the plurality of voice units form at least one voice unit sequence, and based on the matching analysis process between the voice units and the selected words, at least one candidate word sequence corresponding to at least one voice unit sequence is determined, and a plurality of candidate word sequences are obtained.

[0294] For example, based on the voice network, the natural language command is analyzed to obtain a phonetic unit sequence 1 and a phonetic unit sequence 2. The phonetic unit sequence 1 is composed of a phonetic unit a, a phonetic unit b and a phonetic unit c. The phonetic unit sequence 2 is composed of the phonetic unit a, a phonetic unit d and the phonetic unit c. It can be considered that the phonetic units include the phonetic unit a, the phonetic unit b, the phonetic unit c and the phonetic unit d. After the matching analysis of the phonetic units and the selected vocabulary in the preset dictionary, the candidate vocabulary sequence based on the phonetic unit sequence 1 includes a candidate vocabulary sequence 11 and a candidate vocabulary sequence 12. The candidate vocabulary sequence 11 is “help me attack the place”, and the candidate vocabulary sequence 12 is “help me attack the enemy”. The candidate vocabulary sequence based on the phonetic unit sequence 2 includes a candidate vocabulary sequence 21 and a candidate vocabulary sequence 22. The candidate vocabulary sequence 21 is “help me supply the place”, and the candidate vocabulary sequence 22 is “help me attack the enemy”. That is, the multiple candidate vocabulary sequences include the candidate vocabulary sequence 11, the candidate vocabulary sequence 12, the candidate vocabulary sequence 21 and the candidate vocabulary sequence 22.

[0295] The multiple candidate vocabulary sequences include at least one vocabulary in the multiple selected vocabularies. The candidate vocabulary sequence is a vocabulary sequence obtained based on the change relationship between the multiple phonetic units.

[0296] For example, the candidate vocabulary sequence is a vocabulary sequence obtained based on the selected vocabulary in the preset dictionary. Therefore, the vocabulary in the candidate vocabulary sequence is a vocabulary in the selected vocabulary. The candidate vocabulary sequence can be a sequence obtained by a single vocabulary or a sequence obtained by multiple vocabularies, which is not limited here.

[0297] The candidate vocabulary sequence is a sequence obtained after the matching analysis of the multiple phonetic units and the selected vocabulary. Therefore, the phonetic unit change of at least one vocabulary in the candidate vocabulary sequence matches the change between the multiple phonetic units. That is, the candidate vocabulary sequence is a vocabulary combination sequence predicted based on the change between the multiple phonetic units.

[0298] In an optional embodiment, the sequence semantics of the multiple candidate vocabulary sequences are analyzed by the language network, and at least one candidate vocabulary sequence is obtained from the multiple candidate vocabulary sequences as the command analysis result in the text form.

[0299] For example, the sequence semantics are used to represent the semantic information expressed by the candidate vocabulary sequence. The sequence semantics can not only reflect the semantic change of at least one vocabulary in the candidate vocabulary sequence, but also reflect the sentence semantics of the entire candidate vocabulary sequence.

[0300] For example, the language network evaluates the multiple candidate vocabulary sequences respectively, so that at least one candidate vocabulary sequence more consistent with the current semantic situation is obtained as the command analysis result.

[0301] Optionally, the plurality of candidate word sequences are probabilistically evaluated by the language network to obtain a plurality of predicted probabilities respectively corresponding to the plurality of candidate word sequences; and at least one candidate word sequence with the largest predicted probability is taken as the command analysis result.

[0302] Illustratively, the predicted probability is used to represent a probability estimation result of the correctness of the candidate word sequence under the constraint of the plurality of speech units. For example, the predicted probability of the candidate word sequence 11 is 0.32, the predicted probability of the candidate word sequence 12 is 0.8, the predicted probability of the candidate word sequence 13 is 0.69, and the predicted probability of the candidate word sequence 22 is 0.45; if the candidate word sequence with the largest predicted probability is taken as the command analysis result, the candidate word sequence 12, i.e., “help me attack the enemy” is output as the command analysis result.

[0303] In some embodiments, after the sequence semantics of the plurality of candidate word sequences are analyzed by the language network, a plurality of probability distributions respectively corresponding to the plurality of candidate word sequences are generated; the plurality of probability distributions and the plurality of speech units are input into a decoding network, and at least one candidate word sequence is output as the command analysis result based on the constraint of the plurality of speech units.

[0304] Illustratively, the probability distribution is used to describe the naturalness between at least one word in the candidate word sequence; and the decoding network (or decoder) is used to select the best candidate word sequence as the command analysis result by combining the plurality of candidate word sequences, the corresponding probability distributions, and the plurality of language units.

[0305] Optionally, the decoding network can use beam search or other optimization algorithms to search in the plurality of candidate word sequences, taking into account the probability distribution output by the language network and the constraint of the plurality of speech units (phoneme or phonetic segment sequence) output by the acoustic network; the beam search can find the most likely candidate word sequence in the potential sequence space as the command analysis result, thereby improving the accuracy of speech recognition.

[0306] It is worth noting that the decoding network can be configured outside the acoustic network and the language network, or inside the language network; and the preset dictionary can be configured inside the language network, or outside the language network as an independent neural network layer. Here, the network structure of the natural language analysis model is not limited.

[0307] In some embodiments, the language network includes language sub-networks respectively corresponding to at least two virtual scenes.

[0308] Illustratively, the plurality of virtual scenes respectively correspond to a language sub-network, and different language sub-networks are used for analyzing the virtual scenes to which they correspond. For example, virtual scene A corresponds to language sub-network 1, and virtual scene B corresponds to language sub-network 2. When analysis based on virtual scene A is required, language sub-network 1 is used for language analysis, and when analysis based on virtual scene B is required, language sub-network 2 is used for language analysis.

[0309] Alternatively, if the natural language analysis model is referred to as a natural language analysis system, language models corresponding to at least two virtual scenes can be configured, and the plurality of language models correspond one-to-one to the plurality of virtual scenes. Here, the model structure or system architecture is not limited, and the above content is only a schematic description of the analysis process.

[0310] In some embodiments, based on the first virtual scene in which the virtual character and / or NPC is located, a first language sub-network corresponding to the first virtual scene is determined from at least two language sub-networks.

[0311] Illustratively, the first virtual scene is a virtual scene in the at least two virtual scenes. In addition to determining the first scene hotword based on the first virtual scene in which the virtual character and / or NPC is located, a first language sub-network corresponding to the first virtual scene is also determined from the at least two language sub-networks based on the first virtual scene. The first language sub-network is a sub-network layer based on the first virtual scene for analyzing the candidate word sequence.

[0312] In some embodiments, the sequence semantics of the plurality of candidate word sequences are analyzed by the first language sub-network, and at least one candidate word sequence is obtained from the plurality of candidate word sequences as a command analysis result in text form.

[0313] Alternatively, different language sub-networks are language sub-networks pre-trained based on corresponding virtual scenes, and the first language sub-network is a language sub-network trained based on the first virtual scene.

[0314] Illustratively, taking the first language sub-network trained based on the first virtual scene as an example, the first scene hotword corresponding to the first virtual scene is obtained, and the first scene hotword and the general word are used as sample data to train the first language sub-network. Alternatively, a plurality of scene description words used to describe the first virtual scene are obtained, and the scene description words are used as sample data to train the first language sub-network. The scene description words include a plurality of words describing the state of the virtual scene and describing virtual elements therein. Other language sub-networks are also trained based on the words (scene hotword, general word, scene description word, etc.) related to the corresponding virtual scene, which will not be described here.

[0315] Optionally, after obtaining the plurality of candidate word sequences under the constraint of the first scene hotword, the first language subnetwork corresponding to the first virtual scene is used to perform semantic analysis on the plurality of candidate word sequences respectively to obtain predicted probabilities corresponding to the plurality of candidate word sequences respectively, and at least one candidate word sequence with the largest predicted probability is taken as the command analysis result in the text form.

[0316] It is worth noting that the process of determining the first language subnetwork based on the first virtual scene in which the host virtual character and / or NPC is located is only an illustrative example, and a plurality of virtual scenes can also be configured to correspond to one sound network (for example, the sound network is a network trained based on the corresponding virtual scene), and the corresponding first sound subnetwork is selected based on the first virtual scene for analysis, or the corresponding first sound subnetwork and the first language subnetwork are selected based on the first virtual scene for analysis, which is not limited here.

[0317] In an optional embodiment, the command analysis result in the text form is displayed.

[0318] Illustratively, after the speech-to-text process on the natural language command in the speech form, the command analysis result in the text form is obtained, and the terminal can display the command analysis result on the terminal interface to prompt the player that the terminal has successfully received and analyzed the natural language command.

[0319] Optionally, the command analysis result is displayed in a preset area on the terminal interface, such as a preset dialog box area, and the dialog box area is used to represent the simulated dialogue process between the host virtual character and the NPC; or the command analysis result is displayed in the central area of the terminal interface, such as being displayed in the center of the terminal interface with a preset font size and a preset color.

[0320] In an optional embodiment, a reply statement of the NPC in reply to the command analysis result is displayed.

[0321] The reply statement is a text statement generated based on the semantic analysis of the command analysis result. Illustratively, after the terminal analyzes the command analysis result, a pre-trained semantic analysis model is called to analyze the command analysis result, and a reply statement for replying to the command analysis result is generated. For example, the command analysis result is "No. 1, go to A place to see", and the reply statement generated based on the semantic analysis model for this command analysis result is "No. 1 received" or "good" and the like, and the form of the reply statement is not limited here.

[0322] Optionally, the reply statement is displayed below the command analysis result; or the reply statement is displayed after the command analysis result is no longer displayed.

[0323] The above describes the display of the command analysis result in the form of text and the content of the reply sentence. The display of the text can greatly improve the interaction intelligence and transparency between the virtual character and the virtual environment, facilitate the player to intuitively understand, interpret and process the process of natural language commands, and enhance the intelligibility in the interaction process. In addition, the reply sentence of the NPC is generated after semantic analysis, which enables the NPC to not only respond based on the literal meaning, but also to understand the intention and context information behind the command analysis result, thereby generating a reply text that is more in line with the player's expectations, improving the semantic relevance and the quality of the dialogue between the virtual object and the NPC, ensuring the player's interactive immersion and interactive experience, and improving the interaction richness and diversity between the player and the virtual environment, ensuring the efficiency and accuracy of human-computer interaction.

[0324] In step 420, the NPC is controlled to perform a virtual activity based on the behavior intention.

[0325] Illustratively, after the command analysis result in the form of text is obtained, the behavior intention is identified based on the command analysis result, and the behavior intention is obtained, and then the NPC is controlled to perform the virtual activity indicated by the behavior intention based on the behavior intention.

[0326] Illustratively, the command analysis result is a text content in the form of text, which has more intuitive characteristics than natural language commands in the form of voice, and can control the NPC more specifically.

[0327] Optionally, the form of commanding the NPC to perform the activity based on the behavior intention includes at least one of the following: (1) commanding the NPC to move in the virtual environment; (2) commanding the NPC to search for a specified item; (3) commanding the NPC to cooperate with the main virtual character to defend against the attack of the enemy virtual character; (4) commanding the NPC to attack the enemy virtual character; (5) commanding the NPC to change the prop accessory, clothing accessory, etc. The form of the activity is not limited here.

[0328] For example, the command analysis result obtained after analyzing the natural language command is "attack me", the behavior intention of the natural language command is determined based on the behavior intention recognition to be: commanding the NPC to perform virtual attack on the virtual character that can be attacked in the virtual environment, so the NPC can be controlled to perform virtual attack on other virtual characters based on the behavior intention.

[0329] It is worth noting that the above is only an illustrative example, and the embodiments of the present application are not limited in this regard.

[0330] In the embodiments of the present application, the content of obtaining the behavior intention by analyzing the command analysis result after voice-to-text is introduced to control the NPC to perform virtual activities. The process can make the player realize seamless interaction with the virtual environment through voice instructions, greatly simplify the input method, reduce the dependence on traditional input devices, and improve the operation experience of the player. By converting the natural language command into a text form command analysis result, the voice information can be quantified, and the noise interference in the voice input can be eliminated. The environmental perception information is also considered in the process of obtaining the command analysis result, which can improve the understanding ability of the complex context and guarantee the accuracy of obtaining the command analysis result. Further processing of the command analysis result can identify the behavior intention therein, so as to accurately understand the user's demand or instruction, so that the NPC can autonomously and specifically perform virtual activities, enhance the interactivity and intelligence of the virtual world, guarantee the operation accuracy, and improve the human-computer interaction efficiency.

[0331] In the embodiments of the present application, the content of determining the first virtual scene by the position of the main control virtual character and / or NPC, and then determining the object perception range from the first virtual scene based on the position and orientation and obtaining the first scene hotword is introduced. If multiple virtual characters are active in a virtual environment including multiple virtual scenes, the first virtual scene can be determined from the virtual environment according to the position of the main control virtual character and / or NPC, so as to determine the object perception range in the first virtual scene in combination with the orientation of the main control virtual character and / or NPC, thereby overcoming the problem that only a command analysis result with broad meaning can be obtained through a voice recognition system. Through the environmental perception information represented by the object perception range, the text vocabulary used in the voice-to-text process is constrained (such as preferentially using the scene hotword), thereby helping to improve the text accuracy of the command analysis result.

[0332] In addition, even in the case where the command analysis result is not generated, the second scene hotword can be obtained based on the process of moving the main control virtual character and / or NPC to the second virtual scene to constrain the analysis process of the command analysis result, which not only improves the flexibility of obtaining the scene hotword for constraining the analysis process, but also helps to further improve the flexibility of obtaining the command analysis result, so as to more accurately and flexibly command the NPC through the text form command analysis result.

[0333] In the embodiments of the present application, the process of determining the corresponding first language sub-network for targeted analysis based on the first virtual scene in addition to the first scene hotword based on the first virtual scene where the host virtual character and / or NPC is located is also introduced. Among them, the corresponding language sub-network is trained for each virtual scene, so that the language sub-network can analyze the candidate word sequence more purposefully under the constraint of the virtual scene, avoiding the blindness problem when obtaining the command analysis result; the process of obtaining the candidate word sequence by constraining the first scene hotword improves the accuracy of the candidate word sequence, and then combining the first language sub-network to further constrain the analysis process of the candidate word sequence, which helps to further improve the relevance between the command analysis result and the first virtual scene, improve the analysis accuracy of the natural language command, and facilitate to improve the command accuracy of the NPC through more accurate command analysis results.

[0334] In an optional embodiment, the control method of the terminal executing the virtual role is taken as an example for illustration. The NPC has multiple behavior capabilities in the virtual scene, and the multiple behavior capabilities correspond to multiple classification labels. After the terminal obtains the command analysis result in the form of text based on voice-to-text, the hierarchical prediction network is used to perform behavior intention recognition on the command analysis result to obtain the behavior intention and control the NPC to execute the virtual activity. Illustratively, as shown in FIG. 7, the step 420 shown in FIG. 4 can also be implemented as steps 710 to 720.

[0335] In step 710, at least one hierarchical prediction network is called to perform behavior intention recognition on the command analysis result, and a first classification label of the command analysis result for commanding the NPC is obtained as the behavior intention.

[0336] Illustratively, the behavior intention recognition is a recognition process for recognizing the behavior intention from the command analysis result. The behavior intention is the intention information expressed in the command analysis result, and includes at least one of subject information, semantic information and behavior information. The subject information is used to represent the identity information of the NPC pointed to by the command analysis result (such as the name, nickname, etc. of the NPC included); the semantic information is used to represent the meaning and thought conveyed by the command analysis result (such as whether the command analysis result commands the NPC to initiate a virtual attack, etc.); and the behavior information is used to represent the behavior mode of the virtual activity that the NPC pointed to by the command analysis result needs to execute (such as whether the command analysis result commands the NPC to move, jump, etc. in the virtual environment).

[0337] Among them, the hierarchical prediction network includes at least two sub-networks, and the at least two sub-networks are constructed based on the tree structure of the multiple classification labels.

[0338] The hierarchical prediction network is capable of predicting the classification label corresponding to the analysis result of the command. The hierarchical prediction network includes at least two sub-networks, an upper network and a lower network, which are cascaded, and the lower network further performs the prediction of the first classification label according to the prediction result output by the upper network.

[0339] In this embodiment, the hierarchical prediction network is called to perform the behavior intention recognition on the fusion feature of the analysis result of the command, to obtain the first classification label of the analysis result of the command.

[0340] FIG. 8 shows a schematic diagram of performing the behavior intention recognition on the analysis result of the command according to an example embodiment of the present application. The encoding network 460 based on the attention mechanism is called to perform the feature encoding on the analysis result 450 of the command, to obtain the first feature vector 451 of the analysis result of the command in the hidden layer space; the first feature vector 451 can indicate the natural semantics of the characters in the analysis result 450 of the command.

[0341] In this embodiment, the information structure prediction network includes a first full connection layer 462 and a first activation function 464; the first full connection layer 462 is used to perform the dimension encoding on the first feature vector 451, to obtain the second feature vector 452 in the control dimension; the second feature vector 452 can describe the natural semantics of the analysis result 450 of the command in the dimension of commanding the NPC. The first activation function 464 is used to perform the activation processing on the second feature vector 452, to obtain the first output 465, which is used to indicate whether the natural semantics of commanding the NPC is carried in the analysis result 450 of the command.

[0342] In the case that the first output 465 is used to indicate that the natural semantics of commanding the NPC is carried in the analysis result 450 of the command, the first bilinear transformation layer 466 is called to perform the bilinear transformation on the first feature vector 451 and the second feature vector 452, to fuse the natural language semantics of the analysis result 450 of the command in the two dimensions, to obtain the third feature vector 453.

[0343] Next, the hierarchical prediction network is called to perform the behavior intention recognition on the analysis result of the command, and the hierarchical prediction network includes two sub-networks. First, for the first layer sub-network of the upper layer, the first layer sub-network includes a second full connection layer 472 and a second activation function 474; the second full connection layer 472 performs the encoding on the third feature vector 453 from the dimension of the intention feature, to obtain the fourth feature vector 454; the intention dimension corresponding to the fourth feature vector 454 includes any one of the subject dimension, the semantic dimension and the behavior dimension. The second activation function 474 is used to perform the activation processing on the fourth feature vector 454, to obtain the second output 475, which indicates the first level behavior label of the analysis result 450 of the command in one intention dimension.

[0344] The second bilinear transformation layer 476 is called to perform bilinear transformation on the third feature vector 453 and the fourth feature vector 454, to fuse the natural language semantics of the command analysis result 450 in two dimensions, and obtain a fifth feature vector 455.

[0345] Next, the second layer network of the hierarchical prediction network is called to predict the corresponding second-level behavior label of the command analysis result 450 in the first-level behavior label. For the second layer network of the low layer: the second layer network includes a third full connection layer 482 and a third activation function 484; the third full connection layer 482 performs encoding on the fifth feature vector 455 from the dimension of the intention feature, and performs activation processing on the encoding result based on the second activation function 474 to obtain a third output 485, which indicates the corresponding second-level behavior label of the command analysis result 450 in the first-level behavior label.

[0346] In an optional embodiment, the behavior intention includes at least one intention dimension; each of the at least one hierarchical prediction network is used to predict the command intention of the command analysis result for the NPC in one intention dimension; the intention dimension includes at least one of: subject dimension, semantic dimension, and behavior dimension.

[0347] In an example, the at least one hierarchical prediction network includes a subject prediction network, a behavior prediction network, and a semantic prediction network to implement parallel execution of tactical instruction understanding of the command analysis result; the command intention for the NPC is described from three dimensions of the subject dimension, the semantic dimension, and the behavior dimension. It can be understood that the behavior intention recognition of the command analysis result is also called the tactical instruction understanding of the command analysis result; the first classification label of the command analysis result commanding the NPC is obtained by calling the at least one hierarchical prediction network to perform the tactical instruction understanding of the command analysis result. For example, the natural language has rich semantics, and the command analysis result in the form of natural language has rich semantic elements; each hierarchical prediction network respectively predicts the command intention in the command analysis result from different dimensions, which on the one hand reduces the number of label types that need to be classified by the hierarchical prediction network for prediction task, and reduces the complexity of the classification prediction task; on the other hand, the hierarchical prediction network only needs to understand the natural language semantics in the command analysis result in one dimension, which reduces the difficulty of extracting the natural language semantics in the command analysis result.

[0348] Exemplarily, in the at least one hierarchical prediction network, an i-th subnetwork in each hierarchical prediction network is configured to predict a primary behavior label of the command analysis result in one intention dimension, and an i+1-th subnetwork in the hierarchical prediction network is configured to predict a secondary behavior label corresponding to the primary behavior label of the command analysis result, where i is a positive integer; a high-level behavior label (e.g., the primary behavior label) predicted by a high-level subnetwork (e.g., the i-th subnetwork) in the hierarchical prediction network includes a plurality of sub-labels or subordinate low-level labels; and a corresponding low-level subnetwork (e.g., the i+1-th subnetwork) needs to be invoked to further perform classification prediction (e.g., to predict the secondary behavior label). The prediction task of various classification labels is divided into prediction subtasks performed by at least two subnetworks, so as to reduce the complexity of classification prediction of each subnetwork.

[0349] In an optional embodiment, the behavior intention includes at least one intention dimension in a subject dimension, a semantic dimension, and a behavior dimension; the subject dimension analysis process is referred to as a subject prediction process; the semantic dimension analysis process is referred to as a semantic prediction process; and the behavior dimension analysis process is referred to as a behavior prediction process; and the prediction process for each dimension is described as follows.

[0350] (1) Subject prediction process

[0351] In some embodiments, an i-th subnetwork in the subject prediction network is invoked to perform subject prediction on the command analysis result, and a primary behavior label corresponding to the command analysis result in a plurality of primary candidate labels is predicted.

[0352] In this embodiment, the intention dimension includes the subject dimension, the hierarchical prediction network includes a hierarchical subject prediction network, and the subject prediction network has the ability to predict a subject type in the command analysis result; the subject type is used to represent the identity of the NPC commanded by the command analysis result, that is, to represent the subject information in the command analysis result. Exemplarily, the subject prediction network is used to predict which role of a plurality of NPCs is commanded by the command analysis result, and the number of NPCs commanded by the command analysis result can be one or more.

[0353] Exemplarily, the i-th subnetwork is used to predict a primary behavior label corresponding to the command analysis result in a plurality of primary candidate labels. The plurality of primary candidate labels are used to indicate whether there is a subject in the command analysis result; and the i-th subnetwork is used to perform a binary classification task of whether there is a subject. In an optional example, the i-th subnetwork includes an encoding layer, which is used to extract a feature representation of the command analysis result in the subject dimension, and then perform classification prediction on the feature representation of the subject dimension to obtain the primary behavior label corresponding to the command analysis result.

[0354] In some embodiments, the subject-prediction network corresponding to the i+1th sub-network of the primary action label is invoked to perform subject prediction on the command analysis result, to obtain a secondary action label corresponding to the command analysis result from the plurality of secondary candidate labels.

[0355] Exemplarily, in the case where the primary action label is used to indicate that there is a subject in the command analysis result, the plurality of secondary candidate labels are used to indicate that the subject of the command analysis result is one or more NPCs controlled by the sender of the command analysis result and / or in the virtual camp where the master virtual character is located. The plurality of secondary candidate labels are used to indicate the identity of the virtual character commanded by the command analysis result. In one example, in the case where the secondary action label indicates that the subject of the command analysis result is the master virtual character, although there is a natural semantic of controlling the virtual character in the command analysis result, it is not a control command to the NPC, but only the user's self-talk or the exchange text among the plurality of users for the purpose of informing each other of the activity of the master virtual character. In the case where the secondary action label indicates that the subject of the command analysis result is one or more NPCs, the secondary action label can determine the identity of the NPC that needs to perform the virtual activity according to the command analysis result.

[0356] In one example, FIG. 9 shows a schematic diagram of the primary candidate labels provided by one exemplary embodiment of the present application. The i-th sub-network in the subject-prediction network is invoked to perform subject prediction on the command analysis result, and the prediction result indicates that the command analysis result corresponds to a primary action label from the plurality of primary candidate labels; the plurality of primary candidate labels include a positive label 902 and a negative label 904; the positive label 902 is used to indicate that there is a subject in the command analysis result, and the negative label 904 is used to indicate that there is no subject in the command analysis result. The positive label 902 includes a plurality of secondary candidate labels: an all-roles label 905, a master virtual character label 906, and an NPC label 907; the all-roles label 905 is used to indicate that the subject in the command analysis result is all virtual characters in the virtual camp; the master virtual character label 906 is used to indicate that the subject in the command analysis result is the master virtual character controlled by the user; and the NPC label 907 is used to indicate that the subject in the command analysis result is an NPC.

[0357] In some embodiments, in the case where the primary action label is used to indicate that there is no subject in the command analysis result, the first classification label of the command analysis result is determined to be the first label.

[0358] Exemplarily, in a case where the primary behavior label is used to indicate that there is no subject in the command analysis result, the command analysis result does not explicitly show the identity of the NPC that needs to perform the virtual activity. The first classification label of the command analysis result is determined as the first label, and the first label is used to indicate that the subject of the command analysis result is a historical subject in a historical command, the historical command being an adjacent command input before the command analysis result; the historical subject in the historical command input before is continued, and continuous control of one or more NPCs can be realized with the historical subject as a reference.

[0359] Optionally, in a case where the primary behavior label is used to indicate that there is no subject in the command analysis result, and the input time interval between the command analysis result and the historical command does not exceed the time length threshold, the first classification label of the command analysis result is determined as the first label; and in a case where the input time interval between the command analysis result and the historical command does not exceed the time length threshold, the command analysis result and the historical command have continuity of sentences, which is a continuous control command for one or more NPCs.

[0360] In some embodiments, in a case where the primary behavior label is used to indicate that there is no subject in the command analysis result, the first classification label of the command analysis result is determined as the second label.

[0361] Similarly, in a case where the primary behavior label is used to indicate that there is no subject in the command analysis result, the command analysis result does not explicitly show the identity of the NPC that needs to perform the virtual activity. The first classification label of the command analysis result is determined as the second label, and the second label is used to indicate that the subject of the command analysis result is all NPCs in the virtual camp where the virtual character is located. In natural language spoken expression, the subject is usually omitted when spoken expression is made for a group object, and the subject of the command analysis result is determined as all NPCs in the virtual camp where the virtual character is located, which conforms to the natural language spoken expression.

[0362] Optionally, in a case where the primary behavior label is used to indicate that there is no subject in the command analysis result, and there is no historical command in a time length threshold before the input time of the command analysis result, the first classification label of the command analysis result is determined as the second label; and in a case where there is no historical command in the time length threshold before the input time of the command analysis result, the command analysis result is isolated command information, and is not a continuous expression of previous information, and according to the spoken expression habit, the subject of the command analysis result is determined as all NPCs.

[0363] (2) Semantic prediction process

[0364] In an optional embodiment, an i-th layer subnetwork in the semantic prediction network is called to perform semantic prediction on the command analysis result, and the command analysis result is predicted to correspond to a primary behavior label in the plurality of primary candidate labels.

[0365] In the embodiment, the intention dimension includes a semantic dimension, the hierarchical prediction network includes a hierarchical semantic prediction network, and the semantic prediction network has the ability to predict semantic types in the command analysis result; the semantic types are used to represent the control mode of the NPC for initiating a virtual attack, i.e., to represent semantic information in the command analysis result. For example, the subject prediction network is used to predict whether a virtual activity performed by the NPC under the command of the command analysis result is related to initiating a virtual attack. It should be noted that the virtual scene provides a space for virtual competition between virtual characters; in the process of virtual competition, it is necessary to maintain the survival of virtual characters. Initiating a virtual attack is highly related to maintaining the survival of virtual characters, such as actively initiating a virtual attack to obtain (or pick up) additional virtual materials by defeating an enemy virtual character, but the enemy virtual character will counterattack, and there is a risk of virtual life value loss; therefore, whether to initiate a virtual attack is highly related to maintaining the survival of virtual characters, and in this embodiment, for the scenario of virtual competition in a virtual scene, the semantic dimension of whether to initiate a virtual attack is predicted as a separate dimension.

[0366] For example, the i-th layer subnetwork is used to predict a primary behavior label corresponding to the command analysis result from a plurality of primary candidate labels. The plurality of primary candidate labels are used to indicate whether there is a semantic of initiating a virtual attack in the command analysis result; the i-th layer subnetwork is used to perform a binary classification task of whether there is a semantic of initiating a virtual attack. In an optional example, the i-th layer subnetwork includes an encoding layer for extracting a feature representation of the command analysis result in the semantic dimension, and then performing classification prediction on the feature representation of the semantic dimension to obtain the primary behavior label corresponding to the command analysis result.

[0367] In an optional embodiment, the i+1-th layer subnetwork corresponding to the primary behavior label in the semantic prediction network is called to perform semantic prediction on the command analysis result to obtain a secondary behavior label corresponding to a plurality of secondary candidate labels of the command analysis result.

[0368] For example, in the case where the primary behavior label is used to indicate that there is no natural semantic of initiating a virtual attack in the command analysis result, the plurality of secondary candidate labels are used to indicate that the natural semantic in the command analysis result is to avoid initiating a virtual attack or other semantics unrelated to a virtual attack.

[0369] For example, the primary behavior label is used to indicate that the command analysis result does not indicate actively initiating a virtual attack, but how to face a virtual attack is predicted by the i+1-th layer subnetwork. For example, avoiding initiating a virtual attack is used to indicate that the avoidance and evasion mode is adopted for a virtual attack, and the number of virtual attacks is intentionally reduced. The other semantics unrelated to a virtual attack are used to indicate that a neutral mode is adopted for a virtual attack, and do not indicate actively initiating or intentionally avoiding a virtual attack.

[0370] In one example, FIG. 9 shows a schematic diagram of the primary candidate labels provided by one example embodiment of the present application. The i-th layer subnetwork in the semantic prediction network is invoked to perform subject prediction on the command analysis result, and the primary action label is predicted to correspond to the command analysis result among the plurality of primary candidate labels. The plurality of primary candidate labels includes the positive semantic label 912 and the negative semantic label 914. The positive semantic label 912 is used to indicate the active initiation of virtual attacks, and the negative semantic label 914 is used to indicate the absence of active initiation of virtual attacks. The negative semantic label 914 includes a plurality of secondary candidate labels: the stop label 915 and the other label 916. The stop label 915 is used to indicate that the virtual attacks are taken in the avoidance and evasion manner, and the number of virtual attacks is intentionally reduced. The other label 916 is used to indicate that the virtual attacks are taken in the neutral manner, and neither active initiation nor intentional avoidance of virtual attacks is indicated.

[0371] (3) Action prediction process

[0372] In one optional embodiment, the i-th layer subnetwork in the action prediction network is invoked to perform action prediction on the command analysis result, and the primary action label is predicted to correspond to the command analysis result among the plurality of primary candidate labels.

[0373] In the present embodiment, the intent dimension includes the action dimension, and the hierarchical prediction network includes the hierarchical action prediction network. The action prediction network has the ability to predict the action type in the command analysis result, i.e., to predict the behavior of the NPC in performing the virtual activity. For example, the action prediction network is used to predict the virtual activity performed by the NPC.

[0374] For example, the i-th layer subnetwork is used to predict the primary action label corresponding to the command analysis result among the plurality of primary candidate labels. The plurality of primary candidate labels is used to indicate the virtual activity type. The virtual activity type predicted by the i-th layer subnetwork includes a plurality of sub-labels or subordinate lower-level labels. The i+1-th layer subnetwork corresponding to the primary action label needs to be invoked for further prediction.

[0375] In one optional embodiment, the i+1-th layer subnetwork corresponding to the primary action label in the action prediction network is invoked to perform action prediction on the command analysis result, and the secondary action label is predicted to correspond to the command analysis result among the plurality of secondary candidate labels.

[0376] For example, the plurality of secondary candidate labels is the plurality of sub-labels or subordinate lower-level labels of the primary action label predicted by the i-th layer subnetwork. The plurality of secondary candidate labels is used to indicate a plurality of action items in the virtual activity type corresponding to the primary action label. The action item is used to indicate the manner in which the NPC performs the virtual activity.

[0377] In an alternative implementation, the following is introduced for the secondary candidate labels.

[0378] FIG. 10 shows a schematic diagram of the primary candidate labels according to an example embodiment of the present application. In an example, the i-th layer subnetwork in the action prediction network is invoked to perform subject prediction on the command analysis result, and the prediction result indicates that the command analysis result corresponds to a primary action label among a plurality of primary candidate labels. The plurality of primary candidate labels include throwing a virtual prop 921, triggering a virtual skill 922, adjusting a motion feature 923, interacting with a virtual item in a virtual scene 924, moving in the virtual scene 925, and answering an inquiry 926.

[0379] In the case where the primary action label is used to indicate that the virtual activity type is throwing a virtual prop, the plurality of secondary candidate labels are used to describe the type of virtual prop to be thrown. For example, in a shooting game, the plurality of secondary candidate labels include, but are not limited to, at least one of a virtual explosive label 931, a virtual smoke bomb label 932, and a virtual flash bomb label 933. The type of virtual prop to be thrown is not limited in the present application. The virtual prop can be a virtual prop equipped before entering the virtual scene, or a virtual prop picked up in the virtual scene. It should be noted that the process of taking a virtual prop out of the storage space (e.g., a virtual backpack) of an NPC virtual character and triggering the use of the virtual prop by giving up holding the virtual prop, such as taking a virtual artillery out of the virtual backpack and arranging the virtual artillery in the virtual scene without holding the virtual artillery, is also considered as throwing a virtual prop. Accordingly, throwing a virtual prop is also referred to as deploying a virtual prop.

[0380] In the case where the primary action label is used to indicate that the virtual activity type is triggering a virtual skill, the plurality of secondary candidate labels are used to describe the type of virtual skill to be triggered. For example, in a MOBA game, the plurality of secondary candidate labels include at least one of an attack-type virtual skill label 934, a transfer-type virtual skill label 935, and a shield-type virtual skill label 936. The virtual skill can be triggered for one or more virtual objects, or triggered for a certain range / direction of effect. For example, an NPC has three virtual skills, such as a first skill, a second skill, and a third skill (or a big move). The type of virtual skill indicated by the secondary candidate label can be the name of the virtual skill, or the serial number of the virtual skill among the three virtual skills possessed by the NPC.

[0381] And / or, in the case that the first-level behavior label is used to indicate that the virtual activity type is adjusting the motion feature, the plurality of second-level candidate labels include an adjusting NPC body posture label 937 or a moving speed label 938; for example, the body posture includes, but is not limited to, at least one of jumping, standing, squatting, lying down, aiming, etc. The moving speed includes, but is not limited to, at least one of running speed, walking speed, slow walking speed, crawling speed, and crouching speed. The adjusting motion feature is used to indicate that the appearance style (such as different projection observation areas for different body postures) and / or the moving feature of the NPC are adjusted.

[0382] And / or, in the case that the first-level behavior label is used to indicate that the virtual activity type is interacting with a virtual object in the virtual scene, the plurality of second-level candidate labels include an opening label 939 or a closing label 940; the virtual object in the embodiment is an object with different switch states, such as at least one of a virtual door, a virtual window, a virtual safe, a virtual box, a virtual computer, a virtual machine, etc. Taking the virtual object as a virtual door in the virtual scene as an example, the virtual door has a switch state, and has different appearance styles (different door panel angles in the open state and the closed state) and / or different human-computer interaction modes (different door panel angles in the open state and the closed state) in different switch states.

[0383] And / or, in the case that the first-level behavior label is used to indicate that the virtual activity type is moving in the virtual scene, the plurality of second-level candidate labels include a moving path label 941 or a companion activity label 942 in the moving process; for example, the second-level candidate label describing the moving path includes, but is not limited to, at least one of the following: moving between two position points based on the line connecting the two position points, moving in a way of bypassing a virtual object in the virtual environment, and a plurality of NPCs moving for the purpose of surrounding a virtual object (surrounding the virtual object). The second-level candidate label describing the companion activity includes, but is not limited to, at least one of the following: accompanying initiating a virtual attack, accompanying avoiding a virtual attack, and accompanying looking for a masking object in the virtual environment.

[0384] And / or, in the case that the first-level behavior label is used to indicate that the virtual activity type is answering an inquiry, the plurality of second-level candidate labels include a location inquiry label 943 including location information in the virtual scene in the answer statement or an information inquiry label 944 for description information of the virtual scene. In the case that the second-level candidate label is used to describe that the answer statement includes the location information in the virtual scene, the inquiry purpose of the inquiry statement is to inquire the location of the NPC in the virtual environment; in the case that the second-level candidate label is used to describe that the answer statement includes the description information in the virtual scene, the inquiry purpose of the inquiry statement is to inquire the observation result of the NPC in the virtual environment.

[0385] It is worth noting that the above is only an illustrative example, and the embodiments of the present application are not limited in this regard.

[0386] In an optional embodiment, the natural language command is converted into a command analysis sentence group including a plurality of command analysis sentences based on speech-to-text technology, wherein each command analysis sentence can be a command analysis result subjected to behavior intention recognition.

[0387] In some embodiments, a process of performing intention analysis based on the command analysis sentence group is described. As shown in FIG. 11, a schematic diagram of a processing method of a command analysis sentence provided by an example embodiment of the present application is shown.

[0388] A command analysis sentence group 1102 is obtained, and the text content of the command analysis sentence group 1102 includes: “You put the gun on the first floor, and I go to the front to get the treasure box”; a text segmentation network 1110 is called to perform semantic segmentation on the command analysis sentence group 1102, and it is obtained that the command analysis sentence group 1102 includes two command analysis sentences: a first command analysis sentence 1104 and a second command analysis sentence 1106; wherein the text content of the first command analysis sentence 1104 includes: “You put the gun on the first floor”; and the text content of the second command analysis sentence 1106 includes: “I go to the front to get the treasure box”.

[0389] Next, the process of behavior intention recognition and entity recognition of the first command analysis sentence 1104 is described. First, an information structure prediction network 1120 is called to perform classification prediction on the first command analysis sentence 1104, and a control label 1125 is obtained, which is used to indicate whether the first command analysis sentence 1104 carries natural semantics for commanding NPCs. In the case that the control label 1125 indicates that the first command analysis sentence 1104 carries natural semantics for commanding NPCs, a hierarchical prediction network 1100 and an entity prediction network 1160 are respectively called to perform behavior intention recognition and entity recognition on the first command analysis sentence 1104. NPCs are servers or clients or controlled by artificial intelligence (AI), and NPCs have autonomous behavior ability. Command is a command given by a player, and NPCs understand and execute the command according to their autonomous behavior ability.

[0390] For the behavior intention recognition of the first command analysis sentence 1104, the hierarchical prediction network 1100 includes three parts, which respectively predict the command intention for the NPC from three intention dimensions. Among them, the subject prediction network 1130 is used to predict the subject label 1135 of the first command analysis sentence 1104, the subject label 1135 is used to indicate the subject type, and the subject type is used to characterize the identity of the NPC commanded by the command analysis sentence. In the first command analysis sentence 1104, the subject label 1135 indicates that the command analysis sentence commands the teammate role. The semantic prediction network 1140 is used to predict the semantic label 1145 of the first command analysis sentence 1104, the semantic label 1145 is used to indicate the semantic type, and the semantic type is used to characterize the control manner of the NPC for initiating a virtual attack. In the first command analysis sentence 1104, the semantic label 1145 indicates that there is a semantic of initiating a virtual attack. The behavior prediction network 1150 is used to predict the behavior label 1155 of the first command analysis sentence 1104, the behavior label 1155 is used to indicate the intention type, and the intention type is the behavior manner of the NPC for executing a virtual activity. In the first command analysis sentence 1104, the behavior label 1155 indicates moving in the virtual environment to go to a specified location.

[0391] For example, each of the subject prediction network 1130, the semantic prediction network 1140, and the behavior prediction network 1150 includes at least two layers of sub-networks, which are constructed based on a tree structure of multiple classification labels. Taking the subject prediction network 1130 as an example, the high-level sub-network of the subject prediction network 1130 is used to predict whether there is a subject in the first command analysis sentence 1104, and in the case that there is a subject in the first command analysis sentence 1104, the low-level sub-network is used to predict that the subject of the first command analysis sentence 1104 is a master virtual role (also referred to as a player role) and / or one or more NPCs (also referred to as a teammate role) in the virtual camp where the master virtual role is located.

[0392] For the entity recognition of the first command analysis sentence 1104, the entity prediction network 1160 can be implemented as a conditional random field to predict the entity label 1165 of the first command analysis sentence 1104. In the first command analysis sentence 1104, the word "you" corresponds to a person entity label, the word "first floor" corresponds to a building entity label, and there is no word corresponding to an adverb label.

[0393] Similarly, the process of behavior intent recognition, entity recognition and the first sentence are similar for the second sentence 1106. In the second sentence 1106, the subject label 1135 indicates that the command analysis sentence instructs the player character, the semantic label 1145 indicates that there is a semantic of initiating a virtual attack, and the behavior label 1155 indicates moving in the virtual environment to go to the specified location. In the second sentence 1106, the word "I" corresponds to the person entity label, there is no word corresponding to the building entity label, and the word "front" corresponds to the direction label in the modifier.

[0394] In the process of behavior intent recognition, by sequentially calling two-layer sub-networks of each part in the hierarchical prediction network 1100, the classification labels of the command analysis sentence in the subject dimension, the semantic dimension, and the behavior dimension are predicted, and the type of the control object carried in the command analysis sentence is determined; at least two layers of sub-networks in the hierarchical prediction network are constructed based on the tree structure of multiple classification labels, which reduces the classification prediction complexity of each layer of sub-networks; in the process of entity recognition prediction, the words corresponding to the person entity label, the building entity label, and the modifier label are realized based on the conditional random field, which improves the accuracy of the classification prediction of the natural semantic execution in the command analysis sentence.

[0395] It is worth noting that the above is only an illustrative example, and the embodiments of the present application are not limited thereto.

[0396] In an optional embodiment, the content of the command analysis sentence obtained from the command analysis sentence group is described.

[0397] In some embodiments, an encoder is called to perform feature encoding on the command analysis sentence group to obtain a first feature representation of the command analysis result in the hidden layer space.

[0398] Illustratively, the encoder has the ability to extract the feature representation of the command analysis result in the hidden layer space, and encodes the command analysis result in the natural language form into at least one of a feature value, a feature vector, and a feature matrix. Illustratively, the encoder can be implemented as a natural language model, such as a Bidirectional Encoder Representations from Transformers (BERT).

[0399] In some embodiments, a segmenter is called to perform semantic segmentation on the first feature representation to obtain at least two sub-feature representations.

[0400] Illustratively, the segmenter determines whether the characters in the command analysis sentence group are the starting positions of the sentences according to the first feature representation obtained by encoding, and further splits the first feature representation into at least two sub-feature representations according to the characters that are the starting positions of the sentences.

[0401] In some embodiments, the decoder is invoked to perform feature decoding on the at least two sub-feature representations respectively, to obtain at least two command analysis results, the at least two command analysis results corresponding to the at least two sub-feature representations one-to-one.

[0402] For example, the decoder has the capability of decoding the hidden layer features of the natural language into the natural language, and performs feature decoding on the at least two sub-feature representations respectively, to obtain at least two command analysis results, the at least two command analysis results corresponding to the at least two sub-feature representations one-to-one, so as to realize the splitting of the command analysis result group.

[0403] FIG. 12 shows a schematic diagram of splitting command analysis results according to an example embodiment of the present application. The command analysis result group 1202 is used to continuously instruct the NPC to perform virtual activities; the command analysis result group 1202 includes a non-sentence starting symbol [CLS] and N tokens (Tok). The BERT 1210 is invoked to perform feature encoding on the command analysis result group 1202, wherein the BERT 1210 performs linear transformation on the starting symbol [CLS] and the N tokens, to obtain linear encoding: E CLS and E1 to E N ; the BERT 1210 performs feature encoding on the linear encoding, to obtain feature representation of the command analysis result in the hidden layer space: T CLS and T1 to T N The text splitter 1220 is invoked to perform semantic segmentation on the first feature representation, to obtain two sub-feature representations; the two sub-feature representations include E1 to E3, E’1 to E’3; the above two sub-feature representations correspond to two command analysis sentences one-to-one. The text splitter 1220 has the capability of determining whether the N tokens of the command analysis result group 1202 are related to the starting position of the sentence, for example, in the case that the segmentation label of the token is S B , the token is related to the starting position of the sentence; while in the case that the segmentation label of the token is S O , the token is not related to the starting position of the sentence. The decoder is invoked to perform feature decoding on the at least two sub-feature representations respectively, to obtain at least two command analysis results 1242.

[0404] In some embodiments, the first command analysis result in the at least two command analysis results is determined as the command analysis result.

[0405] In the present embodiment, each step in the following is introduced by taking the first classification label prediction of a command analysis result as an example, and by repeatedly performing the behavior intention recognition on each command analysis sentence in the at least two command analysis results, to obtain the first classification label of each command analysis sentence instructing the NPC.

[0406] It is worth noting that the above is only an illustrative example, and the embodiments of the present application do not limit this.

[0407] In some embodiments, a feature encoding network based on an attention mechanism is called to perform feature encoding on the command analysis result, to obtain a hidden layer feature representation of the command analysis result in a hidden layer space.

[0408] For example, the network structure of the encoding network in this step and the above-mentioned encoder can be the same or different. In the above-mentioned content, the hidden layer feature extracted by the encoder can indicate whether the character in the command analysis result group is the starting position of the sentence; the encoding network based on the attention mechanism in this step can indicate the natural semantics of the character in the command analysis result; it can be seen that the encoding purposes of the two are different, and the network parameters of the two are usually different as well. The difference in the above-mentioned network parameters is caused by the use of different loss functions in the network training process. For example, the hidden layer feature representation includes at least one of a feature value, a feature vector, and a feature matrix.

[0409] It should be noted that in the above-mentioned new embodiments, the classification prediction performed by calling the hierarchical prediction network is performed on the hidden layer feature representation of the command analysis result.

[0410] In an optional embodiment, an information structure prediction network is called to perform classification prediction on the command analysis result, and it is predicted that the command analysis result corresponds to the first control label.

[0411] The information structure prediction network has the ability to predict whether the command analysis result carries the natural semantics of commanding the NPC, and the first control label is used to indicate that the command analysis result carries the natural semantics of commanding the NPC. In the case where it is predicted that the command analysis result corresponds to the first control label, it indicates that the command analysis result is used to control the NPC, rather than casual conversation content, and thus the behavior intention recognition needs to be performed to obtain the first classification label of the intention of commanding the NPC. Further, the structure prediction network is used to perform a binary classification task of whether it belongs to casual conversation content.

[0412] In an optional embodiment, a feature encoding network is called to perform dimension encoding on the command analysis result, to obtain a dimension feature representation of the command analysis result.

[0413] For example, the dimension feature representation is used to indicate the feature representation of the command analysis result in at least one of the subject dimension, the semantic dimension, the behavior dimension, and the control dimension; wherein the control dimension is used to indicate whether the command analysis result carries the natural semantics of commanding the NPC. In some examples, the corresponding dimension of the dimension encoding performed by the feature encoding network is different from the intention dimension of the behavior intention recognition, and by performing encoding in different dimensions, more information carrying feature information can be obtained, and the information source basis for performing the behavior intention recognition is increased.

[0414] In an optional embodiment, a linear transformation is performed on the hidden layer feature representation and the dimensional feature representation of the command analysis result, to obtain a fused feature of the command analysis result.

[0415] By way of example, the natural semantic dimension and the dimensional feature (at least one of the subject dimension, the semantic dimension, the behavior dimension, and the control dimension as introduced above) are fused by performing a linear transformation on the hidden layer feature representation and the dimensional feature representation, to provide more information sources for performing the behavior intention recognition. Accordingly, in this embodiment, the classification prediction performed by invoking the hierarchical prediction network is performed on the fused feature of the command analysis result.

[0416] It should be noted that, in the above new embodiment, the classification prediction performed by invoking the hierarchical prediction network is performed on the fused feature of the command analysis result.

[0417] It should be noted that the above is only by way of example, and the embodiments of the present application are not limited in this regard.

[0418] In an optional embodiment, the training process of the hierarchical prediction network is described.

[0419] In some embodiments, a sample command and a sample label of the sample command are obtained, the sample label being a type label of the sample command; at least one hierarchical prediction network is invoked to perform behavior intention recognition on the sample command, to obtain a predicted label of the sample command indicating the command of the NPC.

[0420] By way of example, the sample command is a command analysis result that has been labeled with a sample label, and the sample label is a type label of the sample command; the sample command and the sample label can be obtained by manual labeling, and are used to train the hierarchical prediction network.

[0421] By way of example, the hierarchical prediction network has the ability to predict a classification label corresponding to the command analysis result. By way of example, the hierarchical prediction network includes at least two sub-networks, and the at least two sub-networks are constructed based on a tree structure of multiple classification labels. The hierarchical prediction network constructs at least two sub-networks of a corresponding hierarchical structure for the tree structure of the classification labels, and divides the prediction task of a variety of classification labels into prediction sub-tasks performed by the at least two sub-networks, to reduce the complexity of the classification prediction of each sub-network.

[0422] In some embodiments, backward error propagation training is performed on the hierarchical prediction network according to the difference between the predicted label and the sample label, to obtain a trained hierarchical prediction network.

[0423] The difference between the predicted label and the sample label is used to indicate the accuracy of the hierarchical prediction network in performing the behavior intention recognition. The hierarchical prediction network is trained by back-propagation of errors according to the difference between the predicted label and the sample label, and the network parameters of the hierarchical prediction network are adjusted to minimize the difference between the predicted label and the sample label as a training target, to obtain the trained hierarchical prediction network.

[0424] It should be noted that the behavior intention includes at least one intention dimension, and each hierarchical prediction network in the at least one hierarchical prediction network in the embodiment is used to predict the command analysis result in one intention dimension for the command intention of the NPC. The intention dimension includes at least one of a subject dimension, a semantic dimension, and a behavior dimension. For the subject dimension, the semantic dimension, and the behavior dimension, please refer to the respective embodiments in the foregoing description, which will not be repeated here.

[0425] For example, the back-propagation of errors of the hierarchical prediction network in the step can be training of only one hierarchical prediction network in one intention dimension, or can be training of multiple hierarchical prediction networks together, to obtain the trained multiple hierarchical prediction networks.

[0426] Optionally, a difference loss between the predicted label and the sample label is obtained, a first number is obtained, and a weighting parameter is determined according to the first number, the first number and the weighting parameter are in a negative correlation relationship, the first number is a number of training commands, the training commands and the sample command correspond to the same sample label, and the training commands are used to train the hierarchical prediction network. The hierarchical prediction network is trained by back-propagation of errors according to the product of the difference loss and the weighting parameter, to obtain the trained hierarchical prediction network.

[0427] In some embodiments, in the case that the first predicted label does not exist corresponding (i+1)th subnetwork, the hierarchical prediction network is trained by back-propagation of errors according to the difference between the first predicted label and the sample label, to obtain the trained hierarchical prediction network.

[0428] For example, in the case that the first predicted label does not exist corresponding (i+1)th subnetwork, only one subnetwork in the hierarchical prediction network is needed to realize the prediction of the predicted label, and a loss function used to train the hierarchical prediction network is determined according to the difference between the first predicted label and the sample label.

[0429] As shown in FIG. 13, the training method of the prediction network provided by an example embodiment of the present application is shown in the figure. The long sentence set 1302 includes long sentences for continuously commanding NPCs in a virtual scene, also known as sample command groups; the long sentence set 1302 corresponds to the clause set 1304 and the label set 1306; taking one long sentence in the long sentence set 1302 as an example, the clause set 1304 is used to indicate the clauses after the long sentence is segmented. The label set 1306 is used to indicate the classification labels of each word in the long sentence, and the classification labels include labels in multiple dimensions, which will be described below.

[0430] The long sentence in the long sentence set 1302 is input into the text segmentation network 1310 to obtain the segmentation result 1315 of the long sentence. For example, the text segmentation network 1310 is introduced in the encoder, the segmenter, and the decoder in the foregoing description; here it will not be repeated one by one. The segmentation result 1315 can be the predicted segmentation label of each word in the long sentence, indicating whether each word in the long sentence is the starting position of a sentence; the first loss function 1316 is obtained by comparing the segmentation result 1315 and the real segmentation label of each word in the long sentence in the label set 1306; for example, the first loss function 1316 is a cross-entropy loss function.

[0431] Next, the entity prediction and intent prediction process of a clause is introduced. To achieve entity prediction and intent prediction of each clause in the long sentence, it is only necessary to repeat the execution, and perform entity prediction and intent prediction for each clause in the long sentence.

[0432] Firstly, the clause in the clause set 1304 is input into the information structure prediction network 1320 to obtain a control label 1325; the control label 1325 is used to indicate whether the clause carries natural semantics of commanding the NPC. In the case that the control label 1325 is used to indicate that the clause carries the natural semantics of commanding the NPC, the subject prediction network 1330, the semantic prediction network 1340, the behavior prediction network 1350, and the entity prediction network 1360 are called to perform prediction on the clause respectively. The subject label 1335 is predicted based on the subject prediction network 1330, the semantic label 1345 is predicted based on the semantic prediction network 1340, and the behavior label 1355 is predicted based on the behavior prediction network 1350; the above-mentioned three networks are behavior intention recognition in different intention dimensions. It should be noted that each of the above-mentioned three networks includes at least two sub-networks; for the above-mentioned three networks, please refer to the processing method of the command analysis result in each of the embodiments described above, which will not be repeated here. The entity label 1365 is predicted based on the entity prediction network 1360, and the entity label 1365 corresponds to an environmental element in the virtual scene. The entity label 1365 is used to describe which environmental element in the virtual environment the intention command of the NPC performs the virtual activity. For example, the entity prediction network 1360 is implemented as a conditional random field; for the unfinished introduction of the entity prediction network 1360, please refer to step 530 in the above, which will not be repeated here.

[0433] The true label of each word in the clause in the label set 1306 in the subject dimension, the semantic dimension, the behavior dimension, the control dimension, and the entity dimension is included. By comparing the difference between the true label and the predicted label, the second loss function 1326 is obtained. According to the sum of the first loss function 1316 and the second loss function 1326, the text segmentation network 1310, the information structure prediction network 1320, the subject prediction network 1330, the semantic prediction network 1340, the behavior prediction network 1350, and the entity prediction network 1360 are trained by backward error propagation to realize model training.

[0434] It should be noted that the above is only an illustrative example, and the embodiments of the present application are not limited in this regard.

[0435] Step 720, controlling the NPC to perform a virtual activity according to a first intention corresponding to the first classification label.

[0436] Illustratively, the first classification label is determined based on the above process, and the classification label and the intention are corresponding, so the NPC can be controlled to perform a virtual activity by a corresponding first intention in the case of determining the first classification label.

[0437] Optionally, the first classification label is a classification label determined based on at least one of the subject dimension, the semantic dimension, and the behavior dimension, so that the intention information can be more comprehensively presented, and control of the virtual activity performed by the NPC is more accurate.

[0438] It is worth noting that the above is only an illustrative example, and the embodiments of the present application do not limit this.

[0439] In the embodiments of the present application, if the subject prediction network is implemented by two sub-networks, the subject label corresponding to the command analysis result is predicted, the identity of the NPC commanded by the command analysis result is determined; at least two sub-networks in the subject prediction network are constructed based on the tree structure of multiple subject labels, which reduces the classification prediction complexity of each sub-network; whether the subject exists in the command analysis result is predicted in the i th sub-network, and the specific role identity commanded by the command analysis result is predicted in the i+1 th sub-network; in the scene of controlling the NPC, the subject prediction network is constructed according to the classification label of the tree structure of the subject label, and the accuracy of classification prediction of the natural semantics in the command analysis result is improved.

[0440] In the embodiments of the present application, if the semantic prediction network is implemented by two sub-networks, the semantic label corresponding to the command analysis result is predicted, and the control mode of initiating a virtual attack is determined; at least two sub-networks in the semantic prediction network are constructed based on the tree structure of multiple semantic labels, which reduces the classification prediction complexity of each sub-network; whether the semantic of initiating a virtual attack exists in the command analysis result is predicted in the i th sub-network, and the command analysis result specifically indicates that the virtual attack is avoided or is irrelevant to the virtual attack in the i+1 th sub-network; in the scene of controlling the NPC, the semantic prediction network is constructed according to the classification label of the tree structure of the semantic label, and the accuracy of classification prediction of the natural semantics in the command analysis result is improved.

[0441] In the embodiments of the present application, if the behavior prediction network is implemented by two sub-networks, the behavior label corresponding to the command analysis result is predicted, and the virtual activity type is determined; at least two sub-networks in the behavior prediction network are constructed based on the tree structure of multiple behavior labels, which reduces the classification prediction complexity of each sub-network; the virtual activity type of the command analysis result is predicted in the i th sub-network, and the virtual activity item of the command analysis result is predicted in the i+1 th sub-network; in the scene of controlling the NPC, the behavior prediction network is constructed according to the classification label of the tree structure of the behavior label, and the accuracy of classification prediction of the natural behavior in the command analysis result is improved.

[0442] In addition, at least two sub-networks in the hierarchical prediction network are constructed based on a tree structure of multiple classification labels, reducing the classification prediction complexity of each sub-network; wherein the classification labels are label information representing the behavior ability of the NPC, and the tree structure constructed by multiple classification labels performs behavior intention recognition on the command analysis result, which can effectively utilize the hierarchical relationship to achieve efficient overall recognition process, helping to extract more matched behavior intention from the command analysis result, improving the influence of environmental perception information in the virtual environment on behavior intention recognition, overcoming the problem that related technologies cannot extract behavior intention from text content, and helping to improve the recognition accuracy of behavior intention; in the scene of controlling the NPC, the hierarchical prediction network is constructed according to the classification labels of the tree structure of the NPC, improving the accuracy of classification prediction of the natural semantics in the command analysis result, obtaining more accurate behavior intention, and further helping to better control the NPC to perform virtual activities through more accurate behavior intention; in addition, the encoder, the splitter and the decoder can also be used to split long sentences to adapt to the way of continuously commanding the NPC.

[0443] In addition, the intention dimensions include subject dimension, semantic dimension, behavior dimension, etc. By introducing the hierarchical prediction network based on multi-dimensional intention analysis, more complex instructions can be accurately understood and executed, which can significantly improve the intelligent performance of virtual characters or NPCs under complex commands. The analysis of the subject dimension determines the subject type in the command analysis result through the subject prediction network, accurately identifies the NPC under command, and ensures that the execution of the command has strong pertinence. The analysis of the semantic dimension predicts the semantic type in the command through the semantic prediction network, thereby representing the attack control mode of the NPC in the virtual environment and improving the rationality and accuracy of the attack. The introduction of the behavior dimension predicts the behavior type of the NPC in the command through the behavior prediction network, further enhancing the natural performance of the virtual character when executing tasks. In summary, multi-dimensional analysis makes the voice command in the virtual environment more accurate and smooth, especially in complex interactions or situations, which can effectively improve the response ability of the virtual character and the efficiency of command execution.

[0444] In an optional embodiment, the control method of the virtual character executed by the server is taken as an example for illustration. After receiving the natural language command in the form of voice, the terminal sends the natural language command to the server, and the server performs voice-to-text processing on the natural language command to obtain a command analysis result in the form of text. The command analysis result can be sent to the terminal and displayed on the terminal interface. In addition, the server can also perform behavior intention recognition on the command analysis result in the form of text to obtain behavior intention and control the NPC to perform virtual activities. Illustratively, as shown in FIG. 14, the step 330 shown in FIG. 3 can also be implemented as steps 1410 to 1430.

[0445] At step 1410, the natural language command in the form of voice is obtained.

[0446] Illustratively, the natural language command is implemented as a voice command in the form of audio data. For example, the player speaks during the operation of the terminal, and the terminal automatically collects the voice based on the deployment of a microphone array and obtains audio data as the natural language command in the form of voice; or the player speaks after triggering the audio collection control during the operation of the terminal, and the terminal receives the voice and obtains audio data as the natural language command in the form of voice.

[0447] The natural language command includes a natural semantic for controlling the NPC.

[0448] Illustratively, the natural semantic is semantic information expressed by the natural language command for controlling the NPC, and the NPC can be controlled to perform virtual activities based on the natural semantic.

[0449] Optionally, the natural semantic includes an identification semantic for indicating the NPC and an intention semantic for indicating the activity of the NPC. Illustratively, the identification semantic includes at least one of a name of the NPC (such as "A hero"), a number of the NPC (such as "No. 1"), and location information of the NPC (such as "next to the house").

[0450] At step 1420, the natural language command is converted into a command analysis result in the form of text.

[0451] The command analysis result is obtained by performing voice-to-text conversion on the natural language command based on environment perception information, and the environment perception information includes information perceived by at least one of the main virtual character and the NPC in the virtual environment.

[0452] Optionally, the environment perception information is used to represent environment information perceived and obtained by the main virtual character and / or the NPC within an object perception range, and the object perception range is a three-dimensional space range in which the main virtual character and / or the NPC has perception of other scene elements.

[0453] In some embodiments, behavior intention recognition is performed on the natural language command under the constraint of the environment perception information, a behavior intention is obtained, and the NPC is controlled to perform virtual activities through the behavior intention.

[0454] Optionally, the behavior intention recognition process is performed through a pre-trained machine learning model, the environment perception information and the natural language command are taken as inputs of the machine learning model, and the machine learning model is used to perform an encoding process and a decoding process to output the behavior intention.

[0455] In an optional embodiment, at least one first scene hotword is obtained based on a first virtual scene in which the host virtual role and / or the NPC is located.

[0456] The at least one first scene hotword is a scene-related vocabulary of the first virtual scene.

[0457] Optionally, the first virtual scene is one of a plurality of virtual scenes in a virtual environment; and the first virtual scene in which the host virtual role and / or the NPC is located is determined from the plurality of virtual scenes based on a position of the host virtual role and / or the NPC in the virtual environment.

[0458] In some embodiments, a position coordinate of the host virtual role and / or the NPC in the virtual environment is obtained; a plurality of scene coordinate intervals respectively corresponding to the plurality of virtual scenes are obtained; and the first virtual scene in which the host virtual role and / or the NPC is located is determined from the plurality of virtual scenes based on a positional relative relationship between the position coordinate and the plurality of scene coordinate intervals.

[0459] The scene coordinate interval is used to represent a position range of the virtual scene in the virtual environment.

[0460] Illustratively, the position coordinate of the host virtual role and / or the NPC is content representing position information in coordinate data. Optionally, the position coordinate of the host virtual role and / or the NPC is determined based on a world coordinate system corresponding to the virtual environment.

[0461] Illustratively, the plurality of virtual scenes in the virtual environment respectively correspond to a scene coordinate interval, and the scene coordinate interval is used to limit a three-dimensional space range corresponding to the virtual scene from three dimensions, i.e., a horizontal direction, a vertical direction, and a depth direction in the three-dimensional space. The scene coordinate intervals respectively corresponding to different virtual scenes are independent and do not overlap with each other. Optionally, the scene coordinate intervals respectively corresponding to the plurality of virtual scenes are determined based on a world coordinate system corresponding to the virtual environment.

[0462] Illustratively, after obtaining the position coordinate, the position coordinate is matched with the plurality of scene coordinate intervals, and a virtual scene corresponding to a scene coordinate interval including the position coordinate is taken as the first virtual scene in which the host virtual role and / or the NPC is located. For example, the scene coordinate interval including the position coordinate is referred to as a first scene coordinate interval, the first virtual scene corresponds to the first scene coordinate interval, and the object coordinate is located in the first scene coordinate interval.

[0463] Optionally, if the host virtual role is the host virtual role and / or the NPC, a first position coordinate of the host virtual role in the virtual environment is obtained; and the first virtual scene in which the host virtual role and / or the NPC is located is determined from the plurality of virtual scenes based on a positional relative relationship between the first position coordinate and the plurality of scene coordinate intervals.

[0464] Optionally, if the NPC is a master virtual character and / or NPC, a first position coordinate of the NPC in the virtual environment is obtained; based on a positional relative relationship between the first position coordinate and a plurality of scene coordinate intervals, a first virtual scene in which the master virtual character and / or NPC is located is determined from the plurality of virtual scenes.

[0465] Optionally, in the absence of the command analysis result, in response to the master virtual character and / or NPC moving from the first virtual scene to a second virtual scene, at least one second scene hotword corresponding to the second virtual scene is obtained; and the natural language command is converted into a text form command analysis result based on the at least one second scene hotword.

[0466] In an optional embodiment, based on a first virtual game in which the master virtual character and / or NPC participates, a first virtual scene corresponding to the first virtual game is determined.

[0467] In the plurality of virtual games, each virtual game corresponds to a virtual scene, and the plurality of virtual games are simulated battle environments provided for the master virtual character and / or NPC.

[0468] In some embodiments, a first scene type corresponding to the first virtual scene in which the master virtual character and / or NPC is located is obtained.

[0469] Illustratively, the first scene type is a scene type corresponding to the first virtual scene, and the scene type is used to represent a scene state of the virtual scene, such as the scene type including an office type, a battle type, a kitchen type, a bedroom type, a classroom type, a store type, a factory type, and a plurality of types describing scene states.

[0470] Optionally, the first virtual scene corresponds to a first scene type, and the first scene type is used to describe a scene state represented by the first virtual scene, such as the first virtual scene being a virtual office, the first scene type being an office type; or the first virtual scene being a virtual battle area, the first scene type being a battle type; or the first virtual scene being a virtual store, the first scene type being a store type, and the like.

[0471] In some embodiments, a hotword set corresponding to the first scene type is obtained, and a word in the hotword set is taken as at least one first scene hotword, and the hotword set is a set of scene-related words collected based on the first scene type.

[0472] Illustratively, a plurality of scene types respectively correspond to a hotword set, and the hotword set represents a set of scene-related words collected based on a scene type, and also represents a set of scene-related words collected based on at least one virtual scene under the scene type.

[0473] Optionally, a plurality of scene hotwords corresponding to at least two virtual scenes are acquired.

[0474] Illustratively, a plurality of scene hotwords corresponding to a plurality of virtual scenes in a virtual environment are acquired in advance. Optionally, at least one scene hotword corresponding to each virtual scene is acquired based on a virtual element in the virtual scene, and the plurality of scene hotwords are stored. Optionally, an element name of a virtual element in a virtual scene is acquired as a scene hotword; or an action name of an interactive action interacting with the virtual element in the virtual scene is acquired as a scene hotword; or a state name describing a state of the virtual element in the virtual scene is acquired as a scene hotword, etc.

[0475] Optionally, at least one first scene hotword corresponding to a first virtual scene in which the virtual character and / or NPC is located is acquired from the plurality of scene hotwords based on the first virtual scene.

[0476] Optionally, a plurality of candidate scene hotwords having a first scene identifier are acquired from the plurality of scene hotwords based on the first virtual scene in which the virtual character and / or NPC is located, wherein the first scene identifier is a scene identifier corresponding to the first virtual scene; and at least two candidate scene hotwords are selected as the first scene hotword from the plurality of candidate scene hotwords based on an object perception range of the virtual character and / or NPC in the first virtual scene.

[0477] For example, if scene hotword 1 and scene hotword 2 are acquired based on virtual scene A, scene identifier a corresponding to virtual scene A is marked for scene hotword 1, and scene identifier a is also marked for scene hotword 2; if scene hotword 3 is acquired based on virtual scene B, scene identifier b corresponding to virtual scene B is marked for scene hotword 3, etc.

[0478] In some embodiments, an object perception range of the virtual character and / or NPC in the first virtual scene is acquired, the object perception range being a three-dimensional spatial range in which the virtual character and / or NPC has perception of other scene elements; and at least one first scene hotword is acquired based on environmental perception information in the object perception range.

[0479] Illustratively, after the object perception range is determined, at least one first scene hotword representing a state of the virtual character and / or NPC is acquired based on a range of the object perception range in the first virtual scene and environmental perception information determined by the virtual character and / or NPC based on perception in the object perception range.

[0480] Optionally, the object perception range comprises at least one of the following: a visual perception range of the master virtual character and / or the NPC; an auditory perception range of the master virtual character and / or the NPC; an olfactory perception range of the master virtual character and / or the NPC; a perception skill or a perception device owned by the master virtual character and / or the NPC.

[0481] Optionally, in the case where the object perception range comprises a visual perception range, a preset visual cone is generated as the visual perception range, with the position of the master virtual character and / or the NPC as the starting point and the orientation of the master virtual character and / or the NPC as the cone central line, the preset visual cone being a three-dimensional region for measuring the visual perception range.

[0482] Illustratively, in the case where the object perception range comprises a visual perception range, the position of the master virtual character and / or the NPC and the orientation of the master virtual character and / or the NPC are determined, and a preset visual cone is generated as the visual perception range determined based on the visual perception capability, with the position as the starting point and the orientation as the cone central line.

[0483] In some embodiments, a first scene type corresponding to the first virtual scene is obtained; a hotword set corresponding to the first scene type is obtained; an object perception range of the master virtual character and / or the NPC in the first virtual scene is obtained; at least one first scene hotword is obtained from the hotword set based on environmental perception information in the object perception range.

[0484] Optionally, the environmental perception information comprises a virtual element; a virtual element in the first virtual scene within the object perception range is determined; an element name of the virtual element is taken as the first scene hotword; or an action name of an interactive action corresponding to the virtual element is taken as the first scene hotword.

[0485] Illustratively, the element association relationship is used to represent a relationship reflecting the state of the virtual element. For example: a word comprising an element name of a virtual element is obtained from the hotword set as the first scene hotword; or a word for triggering a virtual element is obtained from the hotword set as the first scene hotword; or an interactive word for interacting with a virtual element is obtained from the hotword set as the first scene hotword, etc.

[0486] In an optional embodiment, the natural language command is converted into a command analysis result in text form based on the at least one first scene hotword.

[0487] In some embodiments, a pre-trained natural language analysis model is obtained, the pre-trained natural language analysis model comprising an acoustic network, a language network, and a preset dictionary, the preset dictionary comprising at least one first scene hotword, other scene hotwords, and general words.

[0488] In some embodiments, the natural language command corresponds to a plurality of phonetic units obtained through an acoustic network, the phonetic unit being a basic unit of a lexical pronunciation; and a sequence matching relationship between the plurality of phonetic units and selected lexicons in a preset lexicon dictionary is analyzed through a language network to obtain a command analysis result in a text form.

[0489] The selected lexicons include at least one first-scene hotword and a general lexicon.

[0490] Optionally, the selected lexicons further include other scene hotwords.

[0491] Illustratively, if the selected lexicons are the first-scene hotword and the general lexicon, the at least one phonetic unit sequence can be matched with the first-scene hotword and the general lexicon; and if the selected lexicons are the first-scene hotword, the other scene hotwords and the general lexicon, the at least one phonetic unit sequence can be matched with the first-scene hotword, the other scene hotwords and the general lexicon.

[0492] The lexicons in the preset lexicon dictionary include an analysis weight, a first analysis weight of the at least one first-scene hotword being higher than a second analysis weight of the other scene hotwords, the analysis weight representing a degree of attention of the lexicon when participating in sequence matching.

[0493] Illustratively, the greater the analysis weight, the greater the degree of attention of the lexicon when participating in sequence matching, i.e., the lexicon is more likely to be paid attention to; the smaller the analysis weight, the smaller the degree of attention of the lexicon when participating in sequence matching, i.e., the lexicon is less likely to be paid attention to; the analysis weight can be used to adjust the degree of attention to different lexicons in the sequence matching process, so that the lexicon with a greater analysis weight is preferentially matched with the lexicon sequence, and the lexicon with a greater analysis weight is also more likely to be selected as a lexicon constituting the command analysis result.

[0494] Optionally, a matching relationship between the plurality of phonetic units and the selected lexicons in the preset lexicon dictionary is analyzed to obtain a plurality of candidate lexicon sequences, the plurality of candidate lexicon sequences including at least one lexicon of the plurality of selected lexicons, the candidate lexicon sequence being a lexicon sequence obtained based on a variation relationship between the plurality of phonetic units.

[0495] Illustratively, the process of analyzing the matching relationship represents a process of analyzing the matching between the selected lexicons and the phonetic units.

[0496] Optionally, a sequence semantics of the plurality of candidate lexicon sequences is analyzed through the language network to obtain at least one candidate lexicon sequence from the plurality of candidate lexicon sequences as the command analysis result in the text form.

[0497] Illustratively, the sequence semantics is used to represent the semantic information expressed by the candidate vocabulary sequence, which can not only reflect the semantic change of at least one vocabulary in the candidate vocabulary sequence, but also reflect the sentence semantics of the entire candidate vocabulary sequence. Optionally, the plurality of candidate vocabulary sequences are probabilistically evaluated by the language network to obtain a plurality of predicted probabilities respectively corresponding to the plurality of candidate vocabulary sequences; and at least one candidate vocabulary sequence with the maximum predicted probability is taken as the command analysis result.

[0498] The above describes the content of analyzing the sequence semantics of the candidate vocabulary sequence to obtain the command analysis result from the plurality of candidate vocabulary sequences. By analyzing the sequence matching relationship between the plurality of voice units and the selected vocabulary in the preset vocabulary dictionary, the accuracy and flexibility of the voice recognition process can be significantly improved. Among them, the language network can deeply analyze the change relationship between the voice unit and the selected vocabulary, effectively weaken the noise in the voice, and provide the plurality of candidate vocabulary sequences to help select the most suitable vocabulary sequence in different situations, achieve the purpose of deeply understanding the overall meaning and context information of the sentence, improve the intelligence of voice recognition, and protect the naturalness and efficiency of the interaction process.

[0499] Optionally, the language network includes a language sub-network corresponding to at least two virtual scenes respectively.

[0500] Illustratively, the plurality of virtual scenes respectively correspond to one language sub-network, and different language sub-networks are used for targeted analysis of the virtual scenes corresponding thereto. For example, virtual scene A corresponds to language sub-network 1, virtual scene B corresponds to language sub-network 2, when analysis based on virtual scene A is needed, language sub-network 1 is used for language analysis, when analysis based on virtual scene B is needed, language sub-network 2 is used for language analysis, etc.

[0501] Illustratively, based on the first virtual scene in which the host virtual character and / or NPC is located, a first language sub-network corresponding to the first virtual scene is determined from at least two language sub-networks; and the sequence semantics of the plurality of candidate vocabulary sequences is analyzed by the first language sub-network to obtain at least one candidate vocabulary sequence from the plurality of candidate vocabulary sequences as the command analysis result in text form.

[0502] Optionally, different language sub-networks are language sub-networks pre-trained based on corresponding virtual scenes, and the first language sub-network is a language sub-network trained based on the first virtual scene.

[0503] Optionally, after obtaining the plurality of candidate vocabulary sequences under the constraint of the first scene hotword, the plurality of candidate vocabulary sequences are respectively subjected to semantic analysis by the first language sub-network corresponding to the first virtual scene to obtain a plurality of predicted probabilities respectively corresponding to the plurality of candidate vocabulary sequences, and at least one candidate vocabulary sequence with the maximum predicted probability is taken as the command analysis result in text form.

[0504] The above introduces the content of analyzing the sequence semantics of the candidate vocabulary sequence through the language sub-network corresponding to the virtual scene to obtain the command analysis result. Each virtual scene corresponds to a special language sub-network, so that the language understanding can be optimized according to the context characteristics of different virtual scenes, ensuring a high match between the voice input and the virtual scene, and reducing the possibility of ambiguity and misanalysis. Among them, the candidate vocabulary sequence is analyzed through the first language sub-network in the first virtual scene, so that in-depth semantic understanding can be carried out according to the scene-specific language rules and context, and a more accurate command analysis result is selected from multiple candidate sequences. The accuracy of voice interaction between the virtual character and the player and the sense of immersion are improved, and a more smooth and personalized user experience is provided. The selection of accurate voice sub-networks also helps to achieve more accurate analysis purposes with less computing resources, improving data processing efficiency.

[0505] Step 1430, performing behavior intention recognition on the command analysis result to obtain a behavior intention.

[0506] In some embodiments, at least one hierarchical prediction network is called to perform behavior intention recognition on the command analysis result, and a first classification label of the command analysis result intention commanding the NPC is obtained as the behavior intention.

[0507] Among them, the hierarchical prediction network includes at least two sub-networks, and the at least two sub-networks are constructed based on a tree structure of multiple classification labels.

[0508] In some embodiments, the behavior intention includes at least one intention dimension; each hierarchical prediction network in the at least one hierarchical prediction network is used to predict the command analysis result for the command intention of the NPC in one intention dimension; the i-th layer sub-network in each hierarchical prediction network is used to predict a first-level behavior label of the command analysis result in one intention dimension, and the i+1-th layer sub-network in the hierarchical prediction network is used to predict a second-level behavior label corresponding to the first-level behavior label in the command analysis result, i is a positive integer; wherein the intention dimension includes at least one of: subject dimension, semantic dimension, and behavior dimension.

[0509] In some embodiments, the intention dimension includes a subject dimension, the hierarchical prediction network includes a hierarchical subject prediction network, the subject prediction network has the ability to predict a subject type in the command analysis result; the subject type is used to represent the identity of the NPC commanded by the command analysis result; and / or, the intention dimension includes a semantic dimension, the hierarchical prediction network includes a hierarchical semantic prediction network, the semantic prediction network has the ability to predict a semantic type in the command analysis result; the semantic type is used to represent the control manner of the NPC for initiating a virtual attack; and / or, the intention dimension includes a behavior dimension, the hierarchical prediction network includes a hierarchical behavior prediction network, the behavior prediction network has the ability to predict a behavior type in the command analysis result; the behavior type is used to represent the behavior manner of the NPC for performing a virtual activity.

[0510] It is worth noting that the content of performing the behavior intention recognition in step 1430 can refer to the above-mentioned step 710, which will not be repeated here.

[0511] The behavior intention is used to control the NPC to perform a virtual activity. Illustratively, the behavior type indicates the behavior manner of the NPC for performing a virtual activity. For example: indicating the NPC to perform a throwing action; or, indicating the NPC to adjust the motion characteristics; or, indicating the NPC to interact with a virtual object, etc.

[0512] In some embodiments, after the server controls the NPC to perform a virtual activity based on the behavior intention, the server sends the activity animation data to the terminal, so as to render and display the activity animation of the NPC performing the virtual activity on the terminal; or, the server sends the behavior intention to the terminal, and the terminal controls the NPC to perform a virtual activity based on the behavior intention, and then renders and displays the corresponding activity animation, etc.

[0513] It is worth noting that the above is only an illustrative example, and the embodiments of the present application are not limited thereto.

[0514] In the embodiments of the present application, the process of performing speech-to-text processing on the natural language command by the server is introduced, and the process of performing behavior intention recognition by the server is introduced. The natural language command is analyzed by the server to obtain the behavior intention for controlling the NPC to perform a virtual activity, which takes advantage of the powerful computing capability of the server, not only reduces the pressure of the terminal for data analysis, but also helps to continuously improve the analysis process and analysis efficiency by means of the algorithm integration capability of the server, and improves the privacy of data analysis.

[0515] In an optional embodiment, the control method of the virtual character described above can be applied in a game scene, and the player issues a natural language command in the form of voice during the game process, so that the terminal realizes the purpose of voice command of the NPC based on the behavior intention in the natural language command. Illustratively, as shown in FIG. 15, a technical flowchart of a voice conversion method based on a virtual scene is shown, which includes the following steps 1510 to step 1590.

[0516] Step 1510, the player enters a game level.

[0517] Illustratively, the game level is a specific part or stage in an electronic game, and the player needs to complete specific goals or tasks in this part to continue to the next level. The level is usually designed to gradually increase the difficulty to challenge the player's skills and strategies. Each level may contain different environments, enemies, puzzles and rewards, aiming to provide a diverse gaming experience.

[0518] Optionally, the player can control the host virtual character to participate in a virtual game during the game process, and the player can also command other NPCs to cooperate with the host virtual character to participate in the virtual game. The NPC is a virtual character other than the host virtual character, such as a virtual soldier or a virtual hero that is deployed in the virtual game by the system by default. For example, the player can manually control the virtual character to participate in a virtual battle during the game process, and can also command NPCs to participate in a virtual battle through voice.

[0519] Step 1520, prepare a hotword list based on virtual elements in the virtual scene of the game level.

[0520] Illustratively, a game usually provides multiple different game levels, and each game level corresponds to a virtual environment. A virtual environment includes multiple virtual scenes, and a game scene refers to the three-dimensional space range in which the host virtual character and / or NPC are located in a specific game level. The virtual scene is not only a visual presentation, but also includes virtual elements, game sound effects and other details, which together constitute the player's gaming experience. A well-designed virtual scene can greatly enhance the immersion and interest of the game. Multiple virtual scenes correspond to a scene type, and the scene types of multiple virtual scenes may be the same or different, which are not limited here.

[0521] Optionally, each scene type has its unique virtual elements and atmosphere, and these virtual elements can enhance the story and interactivity of the game. For example, in a forest scene, the player may need to avoid wild animals or find hidden paths; and in a futuristic city scene, the player may need to use high-tech equipment to complete the task, etc.

[0522] Optionally, each game level can have some specific virtual elements, such as keys for unlocking specific doors or treasure chests, i.e. specific keys or completion of specific tasks are required to open, doors leading to new areas or game levels. Virtual props obtained in a specific game level can have unique attack effects or abilities, such as providing additional defense or special resistance, helping the player to deal with specific hostile virtual characters, etc. There are specific virtual props and consumables, collectibles, gems / coins in the game level; in addition, the game level also contains puzzle props, specific vehicles, specific abilities (magic, skills, etc.); and environmental interaction items, such as virtual explosives, virtual obstacles, virtual buildings, etc. Among them, each specific level in the game level has unique virtual elements, which usually have unique functions or effects, and can affect the player's game experience.

[0523] In some embodiments, based on the virtual elements respectively possessed in the plurality of virtual scenes, a hotword list corresponding to each of the plurality of virtual scenes is prepared; or, a hotword list corresponding to the plurality of virtual scenes is prepared, and the hotword list includes scene hotwords corresponding to the plurality of virtual scenes respectively.

[0524] Optionally, in order to obtain more rich scene hotwords, a hotword management system can be designed to achieve this goal, which has preprocessed all game assets and can calculate the position and direction of the main virtual character and / or NPC in real time interaction of the player. A hotword library (or hotword list) is generated according to asset preprocessing, which contains scene hotwords of various virtual scenes.

[0525] In some embodiments, based on the object perception range, virtual elements existing in the object perception range from the first virtual scene where the main virtual character and / or NPC is located are determined. That is, virtual elements in the object perception range are determined.

[0526] Illustratively, when the virtual scene changes, the hotword management system can automatically adjust the required scene hotwords, and according to the current scene, first select candidate scene hotwords from the predefined hotword library, and then select them again according to the player's field of view (object perception range).

[0527] Optionally, taking the field of view range as an example of the object awareness range, when determining the field of view range of the host virtual character and / or NPC, the field of view range is determined by the frustum of the camera. The frustum is a spatial region defined by the camera position and the field of view (FOV). In order to improve efficiency, spatial partitioning techniques such as quadtree, octree or grid can be used to organize virtual elements in the virtual scene. These data structures can help quickly determine which virtual elements are within the field of view range. When the spatial range is determined, frustum culling is performed to determine which virtual elements need to be rendered by calculating which virtual elements are within the frustum.

[0528] Illustratively, six planes (top, bottom, left, right, front, back) are extracted from the frustum of the camera to extract the frustum planes. For each virtual element, its bounding box is tested with the frustum planes. If the bounding box is completely outside the frustum, the element name, action name, etc. of the virtual element do not need to be loaded into the first scene hotword.

[0529] Illustratively, according to the results of frustum culling, it is determined which virtual elements corresponding to the element name, action name, etc. need to be loaded into the first scene hotword that needs to be used, and which virtual elements corresponding to the hotword can be unloaded from the memory. Optionally, a threshold for hotword loading and hotword unloading can also be set to avoid frequent hotword loading and hotword unloading operations.

[0530] Hierarchical Frustum Culling: In combination with spatial partitioning techniques, the element name and action name of virtual elements in large blocks are unloaded first, and then the element name and action name of virtual elements in small blocks are unloaded, etc.

[0531] Among them, the scene hotword is a frequently used vocabulary in a specific application scenario. By setting the scene hotword, the system can more accurately identify these words, thereby improving the overall recognition accuracy. These hotwords either only appear in a specific game level of a certain game, or are frequently used by players in a specific game level. These words are difficult to perform speech-to-text translation based on pronunciation alone or the frequency of use of these words is very high, thereby affecting the user experience.

[0532] Optionally, an initial hotword list is collected. This hotword list includes the following content.

[0533] (1) Scene-specific item words: such as "A area", "virtual stable", "virtual crash site", "virtual oil tank car", etc.

[0534] (2) Scene-specific character words: such as "Virtual Hero 1", "Virtual Soldier 2", "Virtual Monster", etc.

[0535] (3) Common instructions: such as "attack", "defense", "jump", "run", etc.

[0536] (4) Specific task words: such as "open door", "unlock treasure chest", "dialogue", etc.

[0537] (5) Environment interaction words: such as "view map", "use props", "save game", etc.

[0538] Step 1530, match the hot word list.

[0539] Illustratively, after determining the virtual elements within the object perception range, the hot word list is matched to determine the scene hot word corresponding to the virtual element as the first scene hot word.

[0540] Optionally, taking the example of multiple virtual scenes corresponding to one hot word list, the first virtual scene corresponding first hot word list is determined based on the first virtual scene, and at least one first scene hot word is determined from the first hot word list based on the virtual elements within the object perception range. For example: taking the element name describing the virtual element as the first scene hot word, or taking the action name of the interactive action interacting with the virtual element as the first scene hot word, etc.

[0541] Optionally, taking the example of multiple virtual scenes corresponding to one hot word list, at least one first scene hot word is determined from the hot word list based on the virtual elements within the object perception range. For example: taking the element name describing the virtual element as the first scene hot word, or taking the action name of the interactive action interacting with the virtual element as the first scene hot word, etc.

[0542] Step 1540, ASR system hot word weight.

[0543] Illustratively, the ASR system is an automatic speech recognition (Automatic Speech Recognition) system commonly used in speech recognition (Speech Recognition) technology, which aims to convert human voice form instructions into text form instructions. It involves the process of processing, analyzing and understanding sound signals, so that the computer can "understand" and perform corresponding operations.

[0544] The ASR system is the above-mentioned pre-trained natural language analysis model, and the ASR system includes the following several key components.

[0545] (1) Speech recognition acoustic model (acoustic network in the above-mentioned natural language analysis model);

[0546] An acoustic model is used to map the feature vectors of an audio signal to basic speech units (e.g., phonemes). It is usually trained by machine learning techniques (e.g., deep neural networks) using a large amount of annotated speech data to learn the relationship between speech signals and phonemes.

[0547] (2) Speech recognition language model (language network in the above natural language analysis model);

[0548] A language model is used to predict the probability of a sequence of words, helping the system to choose the most likely one among multiple possible sequences of words. Some language model information is usually included in the end-to-end network, while statistical method-based (n-gram model) decoding constraints are used to improve accuracy.

[0549] (3) Speech recognition hotword (first scene hotword in the above preset dictionary in the natural language analysis model, which can also be selected vocabulary including general vocabulary);

[0550] A hotword refers to a word or phrase that needs special attention and priority recognition in a specific virtual scene. For example, in a game scene, "virtual stable" or "virtual crash site" are hotwords. Hotwords usually require higher recognition accuracy, so there may be special models or algorithms to handle them. These hotwords often have a scene correlation with the virtual scene in the game, and if only the pronunciation is used, their corresponding text cannot be written.

[0551] (4) Speech recognition decoder (decoding network mentioned above);

[0552] A decoder is one of the core components of a speech recognition system, which combines the outputs of an acoustic model and a language model to find the most likely sequence of words using a search algorithm (e.g., the Viterbi algorithm). The task of the decoder is to convert a sequence of feature vectors into the final text output.

[0553] In an optional embodiment, a weighted finite-state transducer (WFST) is used to handle the relationship between states and transitions. In the speech recognition process, a state can represent different meanings based on different inputs, and a transition represents the transformation between states. For example, when processing multiple speech units, each speech unit can be considered a state, and the transition between speech units (e.g., from speech unit 1 to speech unit 2) is a transition between states; in the analysis of candidate word sequences, each word in the candidate word sequence can be considered a state, and the transition between words (e.g., from word 1 to word 2) is a transition between states, etc.

[0554] Illustratively, a WFST can assign a weight to each transition. The weight typically represents the cost or probability of the transition, e.g., the weight of the transition from phonetic unit 1 to phonetic unit 2; or the weight of the transition from word 1 to word 2, etc.

[0555] In some embodiments, WFSTs are used to represent various models and transition relationships, such as acoustic models, language models, lexicons, etc. In a speech recognition system, WFSTs are used to combine different models (e.g., acoustic models, language models, and lexicons) into a unified decoding network that can efficiently search for the most likely sequence of words.

[0556] Illustratively, an acoustic model represents the probability of a phoneme or other basic phonetic unit as a WFST; a language model represents the probability of a sequence of candidate words as a WFST; and a lexicon represents the mapping of a word to its corresponding phoneme sequence (phonetic unit sequence) as a WFST. The relationship between the acoustic model, the language model, and the lexicon is that in a speech recognition system, the acoustic model, the language model, and the lexicon are typically represented as independent WFSTs, which are then combined into a comprehensive WFST using specific combination and optimization algorithms, and the comprehensive WFST is used for decoding and recognition.

[0557] Optionally, an acoustic model (H) maps a plurality of acoustic features corresponding to a natural language command in speech form to a phoneme or other basic phonetic unit, resulting in a plurality of phonetic unit sequences. This mapping can be represented as a WFST, referred to as H. A context-dependent model (C) is typically used within the acoustic model to obtain phonetic unit sequences that are more closely tied to context.

[0558] A lexicon (L) maps words to phonetic unit sequences, resulting in a plurality of candidate word sequences, and this mapping can also be represented as a WFST, referred to as L.

[0559] A language model (G) represents the probability of a sequence of candidate words as a WFST, referred to as G.

[0560] By combining H, C, L, and G, a speech recognition system corresponding to HCLG is obtained, i.e., a comprehensive WFST is generated, referred to as HCLG. This comprehensive WFST represents the complete mapping from feature vectors to word sequences. In the decoding process, the ASR system uses the comprehensive WFST HCLG to search for the most likely sequence of candidate words as the command analysis result.

[0561] Illustratively, the natural language command in the form of speech is converted into a plurality of acoustic features, the plurality of acoustic features is scored using an acoustic model (H) and a context-dependent model (C) in the HCLG to generate phoneme probabilities; the phoneme probabilities are then mapped to a sequence of words using a pre-defined lexicon (L) in the HCLG; and finally, the command analysis result is obtained by analyzing the sequence of words through a language model (G), such as finding the most likely sequence of words as the command analysis result through a decoder using a search algorithm (such as the Viterbi algorithm).

[0562] In an optional embodiment, voice data and text data related to the game level need to be prepared. These data will be used to train the acoustic model and the language model. The voice data includes recorded and collected audio data in the form of speech that the player may use in the game; the text data includes previously inputted text by the player and text related to the game level inputted by the large language model, such as the name of the level, action instructions, etc.

[0563] Optionally, the Mel Frequency Cepstral Coefficients (MFCC) and fbank filter bank extraction features are extracted from the natural language command in the form of speech to train the acoustic model. After the acoustic model is trained using the extracted features in the acoustic model training, a model probability at the Connectionist Temporal Classification (CTC) level is obtained. The language model (LM) is trained using the data described above. Commonly used tools include the SRI Language Modeling Toolkit, the Ken Language Model Toolkit (KenLM), etc. A pronunciation lexicon is constructed, containing all possible words and their corresponding phoneme sequences. The pronunciation lexicon format is that each line contains a word and its phoneme representation. Then, the speech recognition decoding word graph is constructed using the construction tools arpa2fst, fstcompose, fstreduced, etc. will be obtained respectively.

[0564] Step 1550, input correction according to the large language model.

[0565] In an optional embodiment, the large language model is used to generate sample data for training the ASR system, so that the ASR system is trained by the sample data before the ASR system is applied; or, the ASR system is periodically trained by the sample data during the application of the ASR system; or, the ASR system is trained by the sample data in real time during the application of the ASR system, etc.

[0566] Illustratively, a large language model (LLM) refers to a natural language processing (NLP) model that is trained on a large amount of text data, has a high degree of complexity, and has high performance. Large language models use deep learning techniques, especially the Transformer architecture, to understand and generate natural language commands. Large language models perform well in various NLP tasks, such as text generation, translation, question answering, and summarization. Large-scale data training: Large language models are usually trained on a large amount of text data, which may include tens of billions or even hundreds of billions of words. Large-scale data enables the model to learn more rich language features and patterns. Complex model architecture: These models usually have a complex architecture, including multiple layers of Transformers, attention mechanisms, etc., to capture subtle differences in language. High performance: Due to the use of large amounts of data and complex model architecture, large language models usually have high understanding and generation capabilities, and can perform well in various language tasks.

[0567] Large language models have prompts, which are initial input texts provided to the model when using large language models for natural language processing tasks. The design and selection of prompts have an important impact on the quality and accuracy of the model's output. By carefully designing prompts, you can guide the model to generate more expected text. That is, prompts provide context information to the model, helping the model understand the user's intent and generate relevant text. For example, given a prompt "Write an article about climate change", the model will generate an article related to climate change. Prompts can also control the output style and format, such as prompts that contain specific instructions or format requirements to help control the output style and format of the model. For example, the prompt "Explain quantum mechanics in simple language" will guide the model to generate a simple and understandable explanation.

[0568] In addition, prompts can also include previous dialogue content to help the model understand the context of the current dialogue and generate more coherent replies. Effective prompt design needs to consider the following aspects: clear instructions; prompts should contain clear instructions to tell the model what task needs to be completed. For example, the prompt "Please summarize the main points of the following article" is more specific than the prompt "Summarize". Prompts can also provide enough context information to help the model understand the task. For example, in a dialogue system, prompts can include previous dialogue content. In addition, if there are specific format requirements for the output, they can be clearly stated in the prompt. For example, the prompt "Please list the answers to the following questions in a list format".

[0569] Provide context: Give the model enough context information so that it can generate reasonable speech input. Use examples: Guide the model to generate similar outputs by providing some examples. Provide explicit format: Make sure to specify explicitly in the prompt words.

[0570] In some embodiments, the process of generating sample data based on large language models is described as follows.

[0571] Illustratively, assume there is an adventure game, and the player can command NPCs through natural language commands in the form of speech, so it is hoped that the large language model can generate some possible sample data to train the natural language analysis model more targeted.

[0572] Optionally, the instruction template can be preset to splice the text corpus. For example: assume that the template is a string containing placeholders for inserting specific player instruction parameters. The placeholder can be represented by {}. For example: "{player_name} go {direction} {location} {action}", which represents the player's name (such as the name of the NPC) going in a certain direction, to a certain location, and taking a certain action; assume that the following player instruction parameters are: player_name = "No. 1 teammate"; direction = "north"; location = "forest"; action = "throw a virtual hand grenade"; splice the instruction to get template = "No. 1 teammate {player_name} in {location} {action}", that is: No. 1 teammate goes north to the forest and throws a virtual hand grenade.

[0573] Optionally, if the player is a player in an adventure game, he can command virtual characters through natural language commands. For example, the example instructions are: 1. Walk north; 2. Attack the enemy in front; 3. Use healing potion; 4. Open backpack; 5. View map. The instructions can command virtual characters.

[0574] Based on this example instruction, some possible sample data in the form of speech can be generated based on the large language model as data to train the natural language analysis model.

[0575] Illustratively, if the generated voice input is desired to be more specific, more context information can be provided. For example: the prompt word provided for the large language model is "You are a player in an adventure game, your character is a warrior, and you are currently in a forest. You can give instructions to your character through voice input. Here are some example instructions: 1. Walk north; 2. Attack the enemy in front; 3. Use healing potion; 4. Open backpack; 5. Check map; Current game status: Character: Warrior; Location: Forest; Task: Find hidden treasure; Now, please generate some possible player voice inputs."

[0576] Illustratively, if the generated voice input is desired to be more specific, more context information can be provided. For example: the prompt word provided for the large language model is "You are a player in an adventure game, your character is a warrior, and you are currently in a forest. You can give instructions to your character through voice input. Here are some example instructions: 1. Walk north; 2. Attack the enemy in front; 3. Use healing potion; 4. Open backpack; 5. Check map; Current game status: Character: Warrior; Location: Forest; Task: Find hidden treasure; Now, please generate some possible player voice inputs."

[0577] Optionally, based on the above prompt word, the large language model can generate the following text: 1. "Walk east; 2. Use fireball attack enemy; 3. Check task log; 4. Drink a bottle of magic potion; 5. Search the area nearby; 6. "Trade with merchant", 7. "Enter the castle", 8. "Equip a new sword", etc.

[0578] Through the above process, a large amount of text simulating player speech in various scenarios can be generated by the large language model based on human speech habits built-in, and the text can be used as sample data; and / or, the text can be converted into voice form to obtain voice input as sample data, so that the natural language analysis model can be trained using more abundant sample data.

[0579] In an optional embodiment, the text generated by the large language model is used as sample data to update the language model in the ASR system before the ASR system is applied.

[0580] Step 1560, activate the language model of the game level where the player is located.

[0581] In an optional embodiment, the language model (G) in the HCLG includes a plurality of language models, and the plurality of language models correspond to a plurality of virtual scenarios respectively. Based on the first virtual scenario where the host virtual character and / or NPC is located, a first language model corresponding to the first virtual scenario is determined from the plurality of language models (i.e., the first language sub-network corresponding to the first virtual scenario is determined from the language network).

[0582] Optionally, the plurality of language models are models obtained by training corresponding virtual scenes. In the process of training the language model with the text generated by the large language model as sample data, the text corresponding to each virtual scene is obtained, and the language model corresponding to the virtual scene is trained through the text.

[0583] Step 1570, the ASR system selects a language model during decoding.

[0584] Illustratively, when the ASR system performs decoding processing, the first language model corresponding to the first virtual scene is selected to analyze the content input by the acoustic model through the first language model.

[0585] In some embodiments, a plurality of phonetic units corresponding to the natural language command are obtained through an acoustic model in the ASR system, the phonetic unit being a basic constituent unit of a word pronunciation; a sequence matching relationship between the plurality of phonetic units and selected vocabularies in a preset vocabulary dictionary is analyzed through a language model in the ASR system to obtain a command analysis result in text form as a decoding result, the selected vocabularies including at least one first scene hotword and a general vocabulary.

[0586] In an optional embodiment, when the player-controlled master virtual character and / or NPC moves between different buildings, the spatial query logic provides complete location information to help the player quickly locate the position of the enemy virtual character. For example, when the master virtual character and / or NPC is outside the motel, the virtual gunshots are heard from the second floor of the motel, the system will prompt "the virtual gunshots are on the second floor of the motel". The player moves within the same building: when the player moves within the same building, the spatial query logic will simplify the feedback information and only provide information of the sub-level area. For example, when the master virtual character and / or NPC is on the first floor of the motel and the enemy virtual character is on the second floor, the system will prompt "the enemy is on the second floor". The player is in the same room: when the master virtual character and / or NPC and the enemy virtual character are in the same room, the spatial query logic provides more accurate position information to help the player react quickly. For example, when the player and the enemy are both on the second floor of the motel, the system will prompt "the enemy is in the right front position". Through the hierarchical feedback information, the player can more intuitively understand the position of the enemy virtual character, and the most relevant position information is provided according to the position of the player.

[0587] In an optional embodiment, in game speech recognition, different language models and scene hotwords are loaded according to different virtual scenes, which can significantly improve the accuracy and response speed of speech recognition. The following are the steps and methods to achieve this goal.

[0588] First, the ability to recognize the virtual scene in the current game is needed. The loading of the language model decoding word graph will be achieved in the following ways: (1) game state detection: by detecting the state variables of the game (such as player position, task progress, current activity, etc.) to determine the first virtual scene where the main virtual character and / or NPC is located; (2) event triggering: according to specific events (such as entering a certain area, starting a certain task) to switch scenes, thereby determining the first virtual scene where the main virtual character and / or NPC is located; (3) player input: the player can actively inform the current first virtual scene through voice or other input methods, etc.

[0589] Among them, different language models are trained or selected for different virtual scenes. For example, the language model in the battle scene can pay more attention to virtual battle-related vocabulary and phrases, while the language model in the exploration scene can pay more attention to navigation and interaction-related vocabulary.

[0590] Among them, a set of scene hot words are generated according to the virtual elements (such as virtual items, virtual characters, etc.) contained in the line of sight, which are the most frequently triggered or most important vocabulary in the current player scene. For example: in the battle scene, hot words may include "enemy", "attack", "retreat", etc., and in the trading scene, hot words may include "buy", "sell", "item", etc.

[0591] Illustratively, according to the identified virtual scene, the corresponding language model and scene hot words are dynamically loaded.

[0592] Among them, when a scene change is detected, the language model corresponding to the virtual scene is switched, which can be achieved by calling the Application Programming Interface (API) of the speech recognition engine; in addition, when a scene change is detected, the hot word list currently used is updated, so that the speech recognition engine can more accurately recognize these words.

[0593] In an optional embodiment, according to the virtual scene where the main virtual character and / or NPC is located in the game, the decoding strategy of the speech recognition system can also be dynamically adjusted, thereby improving the accuracy of speech recognition and user experience. The following are some methods and steps.

[0594] Illustratively, first, the game state, player position, current task, etc. Information is used to identify and classify the scene where the player is currently located.

[0595] Optionally, different decoding strategies are defined according to different virtual scenes. These strategies include: (1) vocabulary adjustment: adjusting the vocabulary of speech recognition according to the virtual scene, i.e. the selected vocabulary in the above-mentioned preset dictionary; (2) language model adjustment: using different language models to process different virtual scenes; (3) confidence threshold adjustment: adjusting the confidence threshold of speech recognition according to the virtual scene, such as manually adjusting the confidence threshold, and only when the predicted probability is higher than the confidence threshold, the candidate vocabulary sequence corresponding to the maximum preset probability is output as the instruction result.

[0596] Among them, by dynamically adjusting the decoder during the game running process, the decoding strategy of the speech recognizer can be dynamically adjusted according to the virtual scene where the host virtual character and / or NPC is located, thereby improving the flexibility of decoding.

[0597] In an optional embodiment, after the ASR system is online, it can be continuously optimized and adjusted according to the actual use of the player, so as to make it more suitable for the needs of different virtual scenes. The optimization of the language model includes the following operations: (1) hot word adjustment: according to the player feedback and use data, collect the text recognized by the ASR system, and automatically correct it with the correction engine. Through the dynamically adjusted hot word list of the player text after correction, it is ensured that it covers the most commonly used and most important vocabulary, so as to more accurately obtain the first scene hot word in the subsequent; (2) model update: retrain the language model and / or sound model and adjust the parameters of the model; (3) test and deployment: the updated model is subjected to sufficient contrast test to ensure its performance and stability. Once the test is passed, the trained model can be deployed in the game; (4) continuous iteration: speech recognition is a continuous improvement process, feedback is collected and analyzed regularly, and then the model is updated to adapt to the changes of the player's behavior and the game. In the actual game environment, test, collect data and feedback, and continuously iterate and improve the speech recognition system, thereby improving the accuracy of the command analysis result.

[0598] In an optional embodiment, by adjusting the language model of voice recognition according to the virtual scene, it is helpful to better understand and parse natural language commands in the form of voice. A suitable language model is selected from a predefined language model library, and when the scene changes, the language model analyzed can be automatically adjusted, and the process can be realized through a preset language model management system. In addition, the system can also establish an effective feedback mechanism, and players can provide feedback and correct recognition errors or add new scene hot words through a cloud large language model. These feedbacks will be updated online to improve the natural language analysis model, and the hot words and language models will be updated in time. In addition, in order to ensure the reliability and stability of the system, the existing large language model can also be used for offline testing and evaluation. Combined with the automatic testing of the system in different scenes and the feedback data of the players, the hot words and language models are further optimized according to the test feedback data to meet the needs of different scenes. For example: when the player enters an indoor room in a city room, "turn on the light", "turn off the light", "turn up the temperature" and the like may be scene hot words of the virtual indoor scene.

[0599] It is worth noting that the above is only an illustrative example, and the embodiments of the present application are not limited thereto.

[0600] Step 1580, based on the command analysis result output by the language model, the behavior intention recognition is performed, and the NPC is controlled to execute the virtual activity through the behavior intention.

[0601] Illustratively, the command analysis result in the form of text is analyzed by the hierarchical prediction network to determine the behavior intention corresponding to the command analysis result, and the NPC is controlled to execute the virtual activity conforming to the behavior intention through the behavior intention.

[0602] In an optional embodiment, taking a virtual soldier including a master virtual character and a non-player virtual character in a virtual environment as an example, the process of obtaining a command analysis result in the form of text based on a natural language command in the form of voice is described as follows.

[0603] Illustratively, as shown in FIG. 16, the virtual environment includes a master virtual character 1610 manually controlled by a player, a virtual soldier 1611 (referred to as No. 1) and a virtual soldier 1612 (referred to as No. 2) controlled by voice.

[0604] Optionally, after receiving the natural language command, it is determined that the first virtual scene in which the master virtual character 1610 is located is a "virtual outdoor scene", and at least one first scene hot word corresponding to the first virtual scene is obtained, including "truck", "virtual grass", "virtual house", "enter", "open truck", "defend", "desert", "attack", "oak tree", "stable", etc.

[0605] If there is no constraint of the first scene hotword, the natural language command is easily recognized as "No. 1, go to the front and make a sound"; but based on the at least one first scene hotword, the natural language command is converted into the command analysis result in the text form, the command analysis result is "No. 1, go to the front and defend"; the behavior intention is "front defense" obtained by performing the intention recognition on the command analysis result, and then the virtual soldier 1611 can be commanded to move to the front to make a defensive posture, so as to avoid that the enemy virtual character attacks the host virtual character first.

[0606] Or, if there is no constraint of the first scene hotword, the natural language command is easily recognized as "No. 2, go to the mountains and explore"; but based on the at least one first scene hotword, the natural language command is converted into the command analysis result in the text form, the command analysis result is "No. 2, go to the desert and explore"; the behavior intention is "desert exploration" obtained by performing the intention recognition on the command analysis result, and then the virtual soldier 1612 can be commanded to move to the desert to explore whether there is an enemy virtual character or a virtual treasure chest.

[0607] Or, if there is no constraint of the first scene hotword, the natural language command is easily recognized as "No. 1 and No. 2, supply me"; but based on the at least one first scene hotword, the natural language command is converted into the command analysis result in the text form, the command analysis result is "No. 1 and No. 2, attack me"; the behavior intention is "attack" obtained by performing the intention recognition on the command analysis result, and then the virtual soldier 1611 and the virtual soldier 1612 can be commanded to attack the virtual character that can be attacked nearby.

[0608] Or, if there is no constraint of the first scene hotword, the natural language command is easily recognized as "No. 1, go to find the nearby project book"; but based on the at least one first scene hotword, the natural language command is converted into the command analysis result in the text form, the command analysis result is "No. 1, go to find the nearby oak tree"; the behavior intention is "find the oak tree" obtained by performing the intention recognition on the command analysis result, and then the virtual soldier 1611 can be commanded to find the oak tree in the nearby area, so as to complete the game task or find the oak tree that can be avoided to be attacked.

[0609] Or, if there is no constraint of the first scene hotword, the natural language command is easily recognized as "No. 1 and No. 2, go there, it is just fine"; but based on the at least one first scene hotword, the natural language command is converted into the command analysis result in the text form, the command analysis result is "No. 1 and No. 2, go to the stable there"; the behavior intention is "search the stable" obtained by performing the intention recognition on the command analysis result, and then the virtual soldier 1611 and the virtual soldier 1612 can be commanded to search the nearby stable and move to the position where the stable is.

[0610] It is worth noting that the above is only an illustrative example, and the embodiments of the present application are not limited thereto.

[0611] An example embodiment of the present application provides a control method of a virtual role. The method comprises:

[0612] displaying at least one of a master virtual role and an NPC in a virtual environment;

[0613] Illustratively, the master virtual role is a virtual role directly controlled by a user in the virtual environment. There is also one or more non-player characters (NPCs) in the virtual environment; the NPCs and the master virtual role belong to the same virtual camp and are teammates of the master virtual role, and the NPCs follow the master virtual role to move in the virtual environment. In addition, the NPCs can also be follower roles, pet roles, etc. controlled by the master virtual role. In addition, the NPCs can also be neutral roles and only perform cooperative actions with the master virtual role under certain conditions.

[0614] obtaining a natural language command;

[0615] Illustratively, the natural language command includes a natural semantic of commanding the NPC; the natural language command can be directly input by the user or extracted from voice information input by the user, and the present application does not limit the obtaining method of the natural language command.

[0616] The natural language command includes at least one of a behavior intention and a scene entity; the behavior intention is used to indicate a type of virtual activity performed by the NPC, such as indicating a type of virtual activity performed in the virtual environment; the scene entity includes a first scene object, such as indicating a scene object in the virtual environment to which the virtual activity is performed. The behavior intention and the scene entity included in the natural language command are part of the words in the natural language command.

[0617] In an example, the natural language command controls the NPC from two aspects of performing a type of virtual activity and performing the virtual activity on a scene object in the virtual environment; it is a complex instruction for controlling the NPC.

[0618] controlling the NPC to perform a virtual activity in response to the natural language command;

[0619] For example, the control of the NPC to perform the virtual activity is determined based on the environment perception information of the host virtual character and / or the NPC. In some examples, the host virtual character and / or the NPC obtains the environment perception information within the perception range of the host virtual character and / or the NPC in the visual, auditory, action trajectory, etc. For example, the virtual activity is close to the current perception of the host virtual character and / or the NPC to the virtual environment, and the NPC is controlled based on the current perception of the host virtual character and / or the NPC to the virtual environment. For example, the NPC has autonomous behavior, and the user only instructs the NPC, and the user gives a command, and the NPC understands and executes the command based on its autonomous behavior.

[0620] In summary, the method provided in the embodiment controls the complex instructions of the NPC from two aspects of performing what type of virtual activity and performing the virtual activity on what scene object in the virtual environment, determines the virtual activity to be performed based on the environment perception information of the host virtual character and / or the NPC, and ensures the cooperation of the NPC and the host virtual character in performing the virtual activity based on the environment perception information of the host virtual character and / or the NPC.

[0621] FIG. 17 shows a flowchart of a virtual character control method according to an example embodiment of the present application. The method includes:

[0622] Step 1701: obtaining a spatial data set;

[0623] The spatial data set of the virtual environment includes a visual label of a scene object, and the visual label is used to describe the visual features of the scene object in at least one dimension, such as the material, transparency, color, shape, etc.

[0624] Step 1702: displaying at least one of the host virtual character and the NPC in the virtual environment;

[0625] For example, the host virtual character is a virtual character directly controlled by the user in the virtual environment. There is one or more NPC in the virtual environment; the NPC and the host virtual character belong to the same virtual camp and are teammates of the host virtual character.

[0626] Step 1703: obtaining a natural language command in the form of voice;

[0627] The natural language command includes a natural semantic used to instruct the NPC, and the command information includes a behavior intention and a scene entity; the behavior intention is used to indicate the type of virtual activity performed by the NPC, and the scene entity is used to indicate the target entity of the virtual activity;

[0628] Step 1704: converting the natural language command in the form of voice into a natural language command in the form of text;

[0629] Exemplarily, the Automatic Speech Recognition (ASR) processing is performed on the natural language command in the form of voice to determine the natural language command in the form of text. The ASR processing usually includes invoking components such as an acoustic model and a language model to realize recognition of pronunciation, vocabulary and syntax structure and the like of the natural language command in the form of voice, and conversion into the natural language command in the form of text.

[0630] Step 1705: performing intent recognition on the natural language command to obtain a first classification label;

[0631] Exemplarily, the first classification label corresponds to a first intent indicating a command intent for a non-player role in one intent dimension; for example, at least one of the following: indicating which non-player role is commanded by the command text among a plurality of non-player roles, indicating whether a virtual activity performed by the non-player role is related to initiating a virtual attack, and indicating what virtual activity is performed by the non-player role.

[0632] Step 1706: performing entity recognition on the natural language command to obtain a target entity;

[0633] The target entity is determined from the virtual environment in combination with environment perception information, and the environment perception information includes information perceived by at least one of the non-player role and the virtual role from the virtual environment. The target entity can be an entity currently seen by the player, or an entity historically seen by the player.

[0634] Step 1707: in response to the natural language command, controlling the NPC to perform a virtual activity according to the first intent corresponding to the first classification label, or controlling the NPC to perform a virtual activity associated with the target entity, or controlling the NPC to perform a virtual activity associated with the target entity according to the first intent corresponding to the first classification label;

[0635] Exemplarily, the NPC is controlled to perform a virtual activity according to the indication of the first classification label and / or the target entity.

[0636] Step 1708: broadcasting feedback information of the NPC;

[0637] Exemplarily, the game application program can generate corresponding feedback information in real time according to the environment perception information of the non-player role, and broadcast the feedback information. For example, the game application program can generate corresponding feedback information in real time according to the environment perception information triggering the broadcast condition when the environment perception information of the non-player role triggers the broadcast condition; or the game application program can generate corresponding feedback information in real time according to the environment perception information of the non-player role when the non-player role or the virtual role triggering the broadcast condition triggers the broadcast condition.

[0638] The preprocessing stage of step 1701 can be implemented as:

[0639] Substep 1: Obtain attribute text of a scene object in the virtual scene;

[0640] Exemplarily, the attribute text is used to introduce the inherent attributes of the scene object in the virtual scene. On the one hand, the attribute text realizes the description of the scene object in the text mode, and provides semantic information for the prediction of the visual label of the scene object while describing the scene object. On the other hand, the attribute text is the description of the scene object in the virtual scene, and in the case of a large number of scene objects in the virtual scene and the reuse of object models, the inherent attributes of the scene object in the virtual scene can be more accurately described.

[0641] In an optional implementation manner, the attribute text includes at least one of a name of the scene object in the virtual scene and a size of the scene object in the virtual scene.

[0642] Substep 2: Obtain appearance images of the scene object in the virtual scene;

[0643] Exemplarily, the appearance images are used to describe the style of the scene object. The appearance images carry the appearance styles such as color, texture and shape of the scene object, and the mutual positional relationship between the subparts in the picture mode, and can comprehensively describe the scene object from the picture mode (or visual mode). In an optional implementation manner, the appearance images of the scene object include images obtained by observing the scene object from at least two perspectives.

[0644] Substep 3: calling a multi-modal model to perform prediction on the attribute text and the appearance images of the scene object to obtain a visual label of the scene object;

[0645] Exemplarily, the multi-modal model has the ability to perform model prediction on text information and picture information of different modes. In the embodiment, the input parameters of the multi-modal model are the attribute text and the appearance images of the scene object; the multi-modal model predicts the visual label of the scene object from the two modes of the text mode and the picture mode, and the visual label is used to describe the visual features of the scene object in at least one dimension.

[0646] Optionally, the multi-modal model includes a visual question and answer model; the question sentence carries the attribute text of the scene object; the question sentence is used to guide the visual question and answer model to convert the appearance images into the visual label of the scene object, and at the same time, the question sentence provides supplementary information of the scene object in the text mode.

[0647] Obtaining expected information of the scene object;

[0648] Exemplarily, the expectation information is used to indicate a description dimension of the scene object expected in the visual label, and / or a desired format of the visual label; in one example, the expectation information used to indicate the description dimension of the scene object expected in the visual label includes but is not limited to at least one of a type description, a material, a transparency, a color, a surface feature, a shape of the scene object. In another example, the expectation information used to indicate the desired format of the visual label is at least one of a Comma Separated Values (CSV), a JavaScript Object Notation (JSON), an eXtensible Markup Language (XML).

[0649] constructing a question sentence of the appearance image according to the expectation information and the attribute text;

[0650] Exemplarily, the first subpart in the question sentence is supplementary introduction information of the scene object, carrying the attribute text of the scene object; the second subpart in the question sentence is an answer guiding sentence for the visual question answering model, carrying the expectation information.

[0651] Optionally, the spatial data set includes a matching label of the scene object in the virtual scene; correspondingly: performing a quasi-spoken rewriting on the visual label of the scene object to obtain the matching label conforming to the spoken expression of the natural language.

[0652] Exemplarily, as introduced above, the visual label is used to describe the visual features of the scene object in at least one dimension, and the appearance image of the scene object presents rich visual features of the scene object. However, the description of the scene object in the spoken expression of the natural language cannot cover the visual features of the scene object in each dimension. The purpose of performing the quasi-spoken rewriting on the visual label of the scene object is to obtain the matching label more similar to the spoken expression; it can be understood that performing the quasi-spoken rewriting can be deleting part of the content in the visual label, or changing the visual label to a label with the same semantics but different text expression.

[0653] In an optional implementation manner, the quasi-spoken rewriting is performed by calling a natural language model; the visual label of the scene object is input into the natural language model to predict the matching label conforming to the spoken expression of the natural language. The quasi-spoken rewriting on the visual label is realized based on calling the natural language model. Exemplarily, the natural language model carries prior knowledge of the spoken expression of the natural language. For example, the natural language model is implemented as a Large Language Model (LLM).

[0654] Optionally, the spatial data set further comprises spatial information of the scene object, and correspondingly further comprises:

[0655] obtaining the spatial position of the scene object in the virtual scene, and determining the spatial position of the scene object as the auxiliary information of the visual label of the scene object.

[0656] Illustratively, the spatial position of the scene object in the virtual scene is used to indicate the deployment of the scene object in the virtual scene, and the spatial information is used to indicate the position, size, etc. of the scene object in the virtual scene after the scene object is deployed in the virtual scene.

[0657] Further, at least one of the coordinate position, orientation information, bounding box information, and cover point information of the scene object in the virtual scene is obtained; wherein the coordinate position (Location) is used to indicate the position of the scene object in the virtual scene, such as the coordinate information of the center point or the preset point of the scene object in the virtual scene; the orientation (Rotation) is used to indicate the direction faced by the scene object in the virtual scene, such as the direction faced by the front of the scene object in the virtual scene; the bounding box (Bounding Box) is used to indicate the size of the scene object in the virtual scene; the cover point (Cover points) is used to indicate the recommended position of the virtual character when the virtual character approaches the scene object, so as to realize that the scene object can mask the virtual character.

[0658] The intent recognition stage of step 1705 can be implemented as:

[0659] Sub-step 4: calling at least one hierarchical prediction network to perform intent recognition on the natural language command to obtain a first classification label of the natural language command commanding the NPC;

[0660] Illustratively, the hierarchical prediction network has the ability to predict the classification label corresponding to the command text. Illustratively, the hierarchical prediction network includes at least two sub-networks, an upper network and a lower network, which are cascaded, and the lower network further performs classification label prediction according to the prediction result output by the upper network. The at least two sub-networks are constructed based on the tree structure of multiple classification labels, and the hierarchical prediction network constructs at least two sub-networks of corresponding hierarchical structure for the tree structure of classification labels, and divides the prediction task of a variety of classification labels into prediction sub-tasks performed by at least two sub-networks, so as to reduce the classification prediction complexity of each sub-network.

[0661] In an optional implementation, each of the at least one hierarchical prediction network is configured to predict a command text for a non-player character in a first intent dimension; the first intent dimension comprises at least one of a subject dimension, a semantic dimension, and a behavior dimension. For example, each of the at least one hierarchical prediction network comprises an i-th layer sub-network configured to predict a first-level behavior label of the command text in the first intent dimension, and an (i+1)-th layer sub-network configured to predict a second-level behavior label of the command text in the first-level behavior label, where i is a positive integer; a high-level behavior label (e.g., the first-level behavior label) predicted by a high-level sub-network (e.g., the i-th layer sub-network) of the hierarchical prediction network comprises a plurality of sub-labels or subordinate low-level labels, and a corresponding low-level sub-network (e.g., the (i+1)-th layer sub-network) needs to be invoked to further perform classification prediction (e.g., to predict the second-level behavior label). The prediction task of a variety of classification labels is divided into prediction sub-tasks performed by at least two sub-networks, so as to reduce the complexity of classification prediction of each sub-network.

[0662] Further, the first intent dimension comprises the subject dimension, and the hierarchical prediction network comprises a hierarchical subject prediction network; the subject prediction network is configured to predict a subject type in the natural language command; the subject type is used to represent an identity of the NPC commanded by the natural language command; for example, the subject prediction network is configured to predict which of a plurality of NPCs is commanded by the command text; the number of the NPCs can be one or more.

[0663] Further, the first intent dimension comprises the semantic dimension, and the hierarchical prediction network comprises a hierarchical semantic prediction network; the semantic prediction network is configured to predict a semantic type in the natural language command; the semantic type is used to represent a control manner of the NPC for initiating a virtual attack; for example, the subject prediction network is configured to predict whether a virtual activity performed by the NPC commanded by the command text is related to initiating a virtual attack.

[0664] Further, the first intent dimension comprises the behavior dimension, and the hierarchical prediction network comprises a hierarchical behavior prediction network; the behavior prediction network is configured to predict a behavior type in the natural language command; the behavior type is used to represent a behavior manner of the NPC for performing a virtual activity; for example, the behavior prediction network is configured to predict a virtual activity performed by the NPC.

[0665] Correspondingly, in step 1707, in response to the natural language command, the NPC performs a virtual activity according to a first intent corresponding to the first classification label.

[0666] For example, the first classification label corresponds to a first intention indicating a command intention for a non-player role in one intention dimension; such as at least one of: indicating which non-player role is commanded by the command text, indicating whether the virtual activity performed by the non-player role is related to initiating a virtual attack, and indicating what virtual activity is performed by the non-player role. According to the indication of the first classification label, the NPC is controlled to perform the virtual activity.

[0667] The target entity identification process for step 1706 can be implemented as:

[0668] Sub-step 5: querying the target entity from the entity set of the virtual environment according to the target entity information of the target entity expressed by the natural language command and the environment perception information.

[0669] The target entity is determined from the virtual environment in combination with the environment perception information, and the environment perception information includes information perceived by the master virtual role and / or the non-player role from the virtual environment.

[0670] Since the target entity information in the natural language command is usually ambiguous, for example, the natural language command can be "move to the red truck", and the target entity information in it is "red truck", and there can be many red trucks in the virtual environment, and the target entity cannot be accurately determined from the virtual environment according to the target entity information in the natural language command.

[0671] Therefore, the embodiments of the present application provide a method for determining the target entity in combination with the target entity information and the environment perception information. For the above example, the "red truck" expressed by the player in the natural language command should be the red truck that the player can see, and therefore, in combination with the visual range of the master virtual role, the red truck located in the visual range of the master virtual role can be screened from the multiple red trucks in the virtual environment, and the red truck is the target entity indicated by the player in the natural language command.

[0672] Since the natural language command is issued by the player based on the perception of the virtual environment, in order to accurately identify the target entity indicated in the natural language command, the game application program will identify the target entity in the natural language command in combination with the environment perception information when the player issues the natural language command. Based on the player's perception of the virtual environment, the entity closest to the target entity information that the player can perceive in the virtual environment is inferred as the target entity.

[0673] For example, the virtual environment includes a plurality of candidate entities matching the natural language command, and the target entity is an entity selected from the plurality of candidate entities based on the environment perception information of the host virtual character or the non-player character. For example, the game application first selects a plurality of candidate entities matching the target entity information in the natural language command from the entity set, and then selects an entity that can be perceived by the host virtual character or the non-player character as the target entity according to the environment perception information.

[0674] For example, the environment perception information of the host virtual character can include: the picture or entity that the host virtual character can see when observing the virtual environment, the environmental sound that the host virtual character can hear, the source direction of the environmental sound, the type and sound size of the environmental sound, the perception information obtained by the host virtual character using the perception skill (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and the perception information obtained by the host virtual character using the sensing device (for example, sensor sensing signal, positioning signal of positioning props...

Claims

A control method of a virtual role, executed by a computer device, the method comprising: displaying at least one of a master virtual role and an NPC in a virtual environment; acquiring a natural language command in a speech form, the natural language command comprising a natural semantic for controlling the NPC; controlling the NPC to perform a virtual activity according to a behavior intention expressed by the natural language command; wherein the behavior intention is identified from the natural language command according to environment perception information, the environment perception information comprising information perceived by at least one of the master virtual role and the NPC in the virtual environment. The method of claim 1, wherein, The controlling the NPC to perform a virtual activity according to a behavior intention expressed by the natural language command comprises: converting the natural language command into a command analysis result in a text form, the command analysis result being a result of speech-to-text performed on the natural language command based on the environment perception information; wherein the command analysis result is used to obtain the behavior intention by intention recognition, the behavior intention being intention information expressed in the command analysis result; controlling the NPC to perform the virtual activity based on the behavior intention. The method according to claim 1 or 2, wherein The converting the natural language command into a command analysis result in a text form comprises: acquiring at least one first scene hotword based on a first virtual scene in which the master virtual role and / or the NPC is located, the at least one first scene hotword being a scene-related vocabulary of the first virtual scene; translating the natural language command into the command analysis result in a text form based on the at least one first scene hotword. The method according to any one of claims 1 to 3, wherein The acquiring at least one first scene hotword based on a first virtual scene in which the master virtual role and / or the NPC is located comprises: acquiring a first scene type corresponding to the first virtual scene in which the master virtual role and / or the NPC is located; acquiring a hotword set corresponding to the first scene type; and taking at least one vocabulary in the hotword set as the at least one first scene hotword, the hotword set being a set of scene-related vocabularies collected based on the first scene type; or acquiring an object perception range of the master virtual role and / or the NPC in the first virtual scene, the object perception range being a three-dimensional spatial range in which the master virtual role and / or the NPC has perception of other scene elements; and acquiring the at least one first scene hotword based on environment perception information in the object perception range; or acquiring a first scene type corresponding to the first virtual scene; acquiring a hotword set corresponding to the first scene type; acquiring an object perception range of the master virtual role and / or the NPC in the first virtual scene; and acquiring at least one vocabulary from the hotword set as the at least one first scene hotword based on environment perception information in the object perception range. The method according to any one of claims 1 to 4, wherein The object perception range comprises at least one of: a visual perception range of the master virtual role and / or the NPC; an auditory perception range of the master virtual role and / or the NPC. An olfactory perception range of the master virtual role and / or the NPC; A perception skill or a perception range of the master virtual role and / or the NPC. The method according to any one of claims 1 to 5, wherein The method further includes: In a case where the object perception range includes a visual perception range, generating a preset visual cone as the visual perception range, the preset visual cone being a three-dimensional region for measuring the visual perception range, the preset visual cone having a starting point at a position of the master virtual role and / or the NPC and a cone central line at an orientation of the master virtual role and / or the NPC. The method according to any one of claims 1 to 6, wherein The environmental perception information includes virtual elements; The method further includes: Highlighting the virtual elements in the first virtual scene within the object perception range; The element name of the virtual element is used as the first scene hotword, or the action name of an interactive action corresponding to the virtual element is used as the first scene hotword. The method according to any one of claims 1 to 7, wherein The first virtual scene is one of a plurality of virtual scenes in a virtual environment; The method further includes: Determining the first virtual scene in which the master virtual role and / or the NPC is located from the plurality of virtual scenes based on a position of the master virtual role and / or the NPC in the virtual environment. The method according to any one of claims 1 to 8, wherein The method further includes: In a case where the command analysis result is not obtained, in response to the master virtual role and / or the NPC moving from the first virtual scene to a second virtual scene, obtaining at least one second scene hotword corresponding to the second virtual scene; Converting the natural language command into the command analysis result in text form based on the at least one second scene hotword. The method according to any one of claims 1 to 9, wherein The method further includes: Displaying the command analysis result in text form; The method further includes: Displaying a reply statement in which the NPC replies to the command analysis result, the reply statement being a text statement generated based on performing semantic analysis on the command analysis result. The method according to any one of claims 1 to 10, wherein The method further includes: Obtaining a plurality of scene hotwords, the plurality of scene hotwords corresponding to at least two virtual scenes; Obtaining the at least one first scene hotword corresponding to the first virtual scene from the plurality of scene hotwords based on the first virtual scene in which the master virtual role and / or the NPC is located. The method according to any one of claims 1 to 11, wherein The plurality of scene hotwords correspond to scene identifiers, the scene identifiers being used to represent virtual scenes when the scene hotwords are collected. The method further includes: Obtaining the at least one first scene hotword corresponding to the first virtual scene from the plurality of scene hotwords based on the first virtual scene in which the master virtual role and / or the NPC is located. obtaining, based on the first virtual scene in which the host virtual character and / or the NPC is located, a plurality of candidate scene hotwords with a first scene identifier from the plurality of scene hotwords, the first scene identifier being a scene identifier corresponding to the first virtual scene; selecting, based on an object perception range of the host virtual character and / or the NPC in the first virtual scene, at least one candidate scene hotword from the plurality of candidate scene hotwords as the first scene hotword. The method according to any one of claims 1 to 12, wherein The conversion of the natural language command into the command analysis result in text form based on the at least one first scene hotword comprises: obtaining a pre-trained natural language analysis model, the pre-trained natural language analysis model comprising an acoustic network, a language network, and a preset dictionary, the preset dictionary comprising the at least one first scene hotword, other scene hotwords, and general vocabulary; obtaining, by the acoustic network, a plurality of phonetic units corresponding to the natural language command, the phonetic unit being a basic unit of vocabulary pronunciation; analyzing, by the language network, a sequence matching relationship between the plurality of phonetic units and selected vocabulary in the preset dictionary, to obtain the command analysis result in text form, the selected vocabulary comprising at least the at least one first scene hotword and the general vocabulary. The method according to any one of claims 1 to 13, wherein The selected vocabulary further comprises the other scene hotwords. The vocabulary in the preset dictionary comprises an analysis weight, a first analysis weight of the at least one first scene hotword being higher than a second analysis weight of the other scene hotwords, the analysis weight referring to a degree of attention of vocabulary in sequence matching. The method according to any one of claims 1 to 14, wherein The analysis, by the language network, of the sequence semantic of the plurality of candidate vocabulary sequences to obtain at least one candidate vocabulary sequence as the command analysis result in text form comprises: analyzing a matching relationship between the plurality of phonetic units and the selected vocabulary in the preset dictionary to obtain a plurality of candidate vocabulary sequences, the plurality of candidate vocabulary sequences comprising at least one vocabulary from the plurality of selected vocabulary, the candidate vocabulary sequence being a vocabulary sequence obtained based on a change relationship between the plurality of phonetic units; analyzing, by the language network, a sequence semantic of the plurality of candidate vocabulary sequences to obtain at least one candidate vocabulary sequence as the command analysis result in text form. The method according to any one of claims 1 to 15, wherein The language network comprises at least two language sub-networks corresponding to at least two virtual scenes respectively. The analysis, by the language network, of the sequence semantic of the plurality of candidate vocabulary sequences to obtain at least one candidate vocabulary sequence as the command analysis result in text form comprises: determining, based on the first virtual scene in which the host virtual character and / or the NPC is located, a first language sub-network corresponding to the first virtual scene from the at least two language sub-networks; analyzing, by the first language sub-network, a sequence semantic of the plurality of candidate vocabulary sequences to obtain at least one candidate vocabulary sequence as the command analysis result in text form. The method according to any one of claims 1 to 16, wherein The NPC has a plurality of behavior capabilities in the virtual scene, and the plurality of behavior capabilities correspond to a plurality of classification labels; The method further comprises: calling at least one hierarchical prediction network to perform behavior intention recognition on the command analysis result, obtaining a first classification label of the command analysis result intention as the behavior intention, which instructs the NPC; the hierarchical prediction network comprises at least two sub-networks, and the at least two sub-networks are constructed based on a tree structure of the plurality of classification labels; The method further comprises: controlling the NPC to perform the virtual activity according to the first intention corresponding to the first classification label. The method according to any one of claims 1 to 17, wherein The behavior intention comprises at least one intention dimension; each hierarchical prediction network in the at least one hierarchical prediction network is used to predict the command analysis result in one intention dimension for the instruction intention of the NPC; The i-th layer sub-network in each hierarchical prediction network is used to predict a first-level behavior label of the command analysis result in the one intention dimension, and the i+1-th layer sub-network in the hierarchical prediction network is used to predict a second-level behavior label corresponding to the first-level behavior label in the command analysis result, i being a positive integer; The intention dimension comprises at least one of a subject dimension, a semantic dimension, and a behavior dimension. According to any one of claims 1 to 18, wherein The intention dimension comprises the subject dimension, the hierarchical prediction network comprises a hierarchical structure subject prediction network, the subject prediction network has the ability to predict the subject type in the command analysis result; the subject type is used to represent the identity of the NPC instructed by the command analysis result; And / or, the intention dimension comprises the semantic dimension, the hierarchical prediction network comprises a hierarchical structure semantic prediction network, the semantic prediction network has the ability to predict the semantic type in the command analysis result; the semantic type is used to represent the control mode of the NPC for initiating a virtual attack; And / or, the intention dimension comprises the behavior dimension, the hierarchical prediction network comprises a hierarchical structure behavior prediction network, the behavior prediction network has the ability to predict the behavior type in the command analysis result; the behavior type is used to represent the behavior mode of the NPC for performing a virtual activity. A control method of a virtual character, wherein, The method is executed by a server, and the method comprises: obtaining a natural language command in the form of voice, the natural language command comprising a natural semantic for controlling an NPC; converting the natural language command into a command analysis result in the form of text, the command analysis result being obtained by performing voice-to-text on the natural language command based on environment perception information, the environment perception information comprising information perceived by at least one of a master virtual character and the NPC in a virtual environment; performing behavior intention recognition on the command analysis result to obtain a behavior intention, the behavior intention being intention information expressed in the command analysis result, and the behavior intention being used to control the NPC to perform a virtual activity. The method of claim 20, wherein, The method further comprises: obtaining at least one first scene hotword based on the first virtual scene in which the host virtual character and / or the NPC is located, the at least one first scene hotword being a scene-related vocabulary of the first virtual scene; converting the natural language command into the command analysis result in text form based on the at least one first scene hotword. The method according to claim 20 or 21, wherein The obtaining at least one first scene hotword based on the first virtual scene in which the host virtual character and / or the NPC is located includes: obtaining a first scene type corresponding to the first virtual scene in which the host virtual character and / or the NPC is located; obtaining a hotword set corresponding to the first scene type; taking at least one vocabulary in the hotword set as the at least one first scene hotword, the hotword set being a set of scene-related vocabularies collected based on the first scene type; or obtaining an object perception range of the host virtual character and / or the NPC in the first virtual scene, the object perception range being a three-dimensional space range in which the host virtual character and / or the NPC has perception of other scene elements; obtaining the at least one first scene hotword based on environmental perception information in the object perception range; or obtaining a first scene type corresponding to the first virtual scene; obtaining a hotword set corresponding to the first scene type; obtaining an object perception range of the host virtual character and / or the NPC in the first virtual scene; and obtaining at least one vocabulary from the hotword set as the at least one first scene hotword based on environmental perception information in the object perception range. The method according to any one of claims 20 to 22, wherein The object perception range includes at least one of: a visual perception range of the host virtual character and / or the NPC; an auditory perception range of the host virtual character and / or the NPC; an olfactory perception range of the host virtual character and / or the NPC; a range perceived by a perception skill or a perception device possessed by the host virtual character and / or the NPC. The method according to any one of claims 20 to 23, wherein The obtaining the object perception range of the host virtual character and / or the NPC in the first virtual scene includes: in a case where the object perception range includes a visual perception range, generating a preset visual cone as the visual perception range, the preset visual cone being a three-dimensional region for measuring the visual perception range, with a position of the host virtual character and / or the NPC as a starting point and an orientation of the host virtual character and / or the NPC as a cone center line. The method according to any one of claims 20 to 24, wherein The environmental perception information includes a virtual element; The obtaining at least one first scene hotword based on the first virtual scene in which the host virtual character and / or the NPC is located includes: determining the virtual element in the first virtual scene that is within the object perception range; taking an element name of the virtual element as the first scene hotword; or taking an action name of an interactive action corresponding to the virtual element as the first scene hotword. The method according to any one of claims 20 to 25, wherein The first virtual scene is one of a plurality of virtual scenes in a virtual environment; The method further includes: determine, based on the position of the host virtual role and / or the NPC in the virtual environment, the first virtual scene in which the host virtual role and / or the NPC is located from the plurality of virtual scenes. The method according to any one of claims 20 to 26, wherein The method further comprises: In the absence of the command analysis result, in response to the host virtual role and / or the NPC moving from the first virtual scene to a second virtual scene, obtaining at least one second scene hotword corresponding to the second virtual scene; based on the at least one second scene hotword, converting the natural language command into the command analysis result in text form. The method according to any one of claims 20 to 27, wherein The at least one first scene hotword is obtained based on the first virtual scene in which the host virtual role and / or the NPC is located, comprising: obtain a plurality of scene hotwords, the plurality of scene hotwords corresponding to at least two virtual scenes; based on the first virtual scene in which the host virtual role and / or the NPC is located, obtaining the at least one first scene hotword corresponding to the first virtual scene from the plurality of scene hotwords. The method according to any one of claims 20 to 28, wherein The plurality of scene hotwords correspond to scene identifiers, and the scene identifiers are used to represent the virtual scene when the scene hotword is collected. The at least one first scene hotword is obtained based on the first virtual scene in which the host virtual role and / or the NPC is located, comprising: based on the first virtual scene in which the host virtual role and / or the NPC is located, obtaining a plurality of candidate scene hotwords with a first scene identifier from the plurality of scene hotwords, the first scene identifier being the scene identifier corresponding to the first virtual scene; based on the object perception range of the host virtual role and / or the NPC in the first virtual scene, selecting at least one candidate scene hotword from the plurality of candidate scene hotwords as the first scene hotword. The method according to any one of claims 20 to 29, wherein The at least one first scene hotword is obtained based on the first virtual scene in which the host virtual role and / or the NPC is located, comprising: obtain a pre-trained natural language analysis model, the pre-trained natural language analysis model comprising an acoustic network, a language network and a preset dictionary, the preset dictionary comprising the at least one first scene hotword, other scene hotwords and general vocabulary; obtain a plurality of phonetic units corresponding to the natural language command through the acoustic network, the phonetic unit being a basic unit of vocabulary pronunciation; through the language network, analyze the sequence matching relationship between the plurality of phonetic units and selected vocabulary in the preset dictionary to obtain the command analysis result in text form, the selected vocabulary at least including the at least one first scene hotword and the general vocabulary. The method according to any one of claims 20 to 30, wherein The selected vocabulary further includes the other scene hotwords; wherein the vocabulary in the preset dictionary comprises an analysis weight, the first analysis weight of the at least one first scene hotword being higher than the second analysis weight of the other scene hotwords, and the analysis weight refers to the degree of attention of vocabulary participating in sequence matching. The method according to any one of claims 20 to 31, wherein The sequence matching relationship between the plurality of speech units and the selected vocabulary in the preset vocabulary dictionary is analyzed through the language network to obtain the command analysis result in text form, including: The matching relationship between the plurality of speech units and the selected vocabulary in the preset vocabulary dictionary is analyzed to obtain a plurality of candidate vocabulary sequences, at least one of the plurality of selected vocabulary is included in the plurality of candidate vocabulary sequences, and the candidate vocabulary sequence is a vocabulary sequence obtained based on the change relationship between the plurality of speech units; The sequence semantics of the plurality of candidate vocabulary sequences is analyzed through the language network to obtain at least one candidate vocabulary sequence from the plurality of candidate vocabulary sequences as the command analysis result in text form. The method according to any one of claims 20 to 32, wherein The language network includes a language sub-network corresponding to at least two virtual scenes respectively; The sequence semantics of the plurality of candidate vocabulary sequences is analyzed through the language network to obtain at least one candidate vocabulary sequence from the plurality of candidate vocabulary sequences as the command analysis result in text form, including: Based on the first virtual scene in which the host virtual character and / or the NPC is located, a first language sub-network corresponding to the first virtual scene is determined from at least two language sub-networks; The sequence semantics of the plurality of candidate vocabulary sequences is analyzed through the first language sub-network to obtain at least one candidate vocabulary sequence from the plurality of candidate vocabulary sequences as the command analysis result in text form. The method according to any one of claims 20 to 33, wherein The NPC has a plurality of behavior capabilities in the virtual scene, and the plurality of behavior capabilities correspond to a plurality of classification labels; The behavior intention recognition is performed on the command analysis result to obtain a behavior intention, including: At least one hierarchical prediction network is called to perform behavior intention recognition on the command analysis result, and a first classification label of the command analysis result intention commanding the NPC is obtained as the behavior intention. The hierarchical prediction network includes at least two sub-networks, and the at least two sub-networks are constructed based on a tree structure of the plurality of classification labels. The method according to any one of claims 20 to 34, wherein The behavior intention includes at least one intention dimension; each hierarchical prediction network in the at least one hierarchical prediction network is used to predict the command analysis result in one intention dimension for the command intention of the NPC; The i-th layer sub-network in each hierarchical prediction network is used to predict a first-level behavior label of the command analysis result in the one intention dimension, and the i+1-th layer sub-network in the hierarchical prediction network is used to predict a second-level behavior label corresponding to the first-level behavior label in the command analysis result, i is a positive integer; The intention dimension includes at least one of a subject dimension, a semantic dimension, and a behavior dimension. The method according to any one of claims 20 to 35, wherein The intention dimension includes the subject dimension, the hierarchical prediction network includes a hierarchical structure of a subject prediction network, the subject prediction network has the ability to predict the subject type in the command analysis result; and the subject type is used to represent the identity of the NPC commanded by the command analysis result. And / or, the intention dimension includes the semantic dimension, the hierarchical prediction network includes a hierarchical semantic prediction network, the semantic prediction network has the ability to predict semantic types in the command analysis result; the semantic types are used to characterize the control mode of the NPC for initiating virtual attacks; And / or, the intention dimension includes the behavior dimension, the hierarchical prediction network includes a hierarchical behavior prediction network, the behavior prediction network has the ability to predict behavior types in the command analysis result; the behavior types are used to characterize the behavior mode of the NPC for performing virtual activities. A virtual role control device, the device comprising: a display module configured to display at least one of a master virtual role and an NPC in a virtual environment; an acquisition module configured to acquire a natural language command in the form of speech, the natural language command including a natural semantic for controlling the NPC; a control module configured to control the NPC to perform a virtual activity through a behavior intention expressed by the natural language command, wherein the behavior intention is identified from the natural language command according to environment perception information, and the environment perception information includes information perceived by at least one of the master virtual role and the NPC in the virtual environment. A virtual role control device, the device comprising: an acquisition module configured to acquire a natural language command in the form of speech, the natural language command including a natural semantic for controlling the NPC; a conversion module configured to convert the natural language command into a command analysis result in the form of text, the command analysis result being a result of performing speech-to-text on the natural language command based on environment perception information, and the environment perception information including information perceived by at least one of a master virtual role and the NPC in a virtual environment; an identification module configured to perform behavior intention identification on the command analysis result to obtain a behavior intention, the behavior intention being intention information expressed in the command analysis result, and the behavior intention being used to control the NPC to perform the virtual activity. A computer device, comprising a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the virtual role control method according to any one of claims 1 to 36. A computer readable storage medium, the storage medium storing at least one program, the at least one program being loaded and executed by a processor to implement the virtual role control method according to any one of claims 1 to 36. A computer program product, comprising computer instructions, the computer instructions being executed by a processor to implement the virtual role control method according to any one of claims 1 to 36.

Citation Information

Patent Citations

  • Speech recognition method and related product

    CN115547337A

  • Virtual object control method and device, equipment and storage medium

    CN116983639A

  • Non-player character interaction method and system

    CN118079400A

  • Game operation method with natural language

    TW200521727A

  • Dynamic interaction menus from natural language representations

    US20070288404A1