Virtual character control method and apparatus, processing method and apparatus, device, and storage medium

By acquiring command information and combining it with environmental awareness information, and using multiple semantic understanding methods to control NPCs, the problem of the single control method for virtual characters is solved, and flexible and diversified control of virtual characters in the virtual environment is realized.

WO2026031857A1PCT designated stage Publication Date: 2026-02-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/104778
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-08-09
Filing Date
2025-06-27
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Current technologies offer limited control methods for virtual characters, making it difficult to flexibly control them using natural language.

Method used

By acquiring command information and utilizing the environmental awareness information of the main virtual character and NPCs, multiple semantic understanding methods, including natural language and environmental awareness information, are employed to control the virtual activities of NPCs.

Benefits of technology

It enables flexible control of virtual characters in the virtual environment, enhancing the diversity of NPC behavior and responsiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025104778_12022026_PF_FP_ABST
    Figure CN2025104778_12022026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of human-computer interaction, and discloses a virtual character control method and apparatus, a processing method and apparatus, a device, and a storage medium. The control method comprises: displaying a user-controlled virtual character or an NPC located in a virtual environment; acquiring command information, wherein the command information comprises natural semantics for controlling the NPC, and the command information corresponds to at least two semantic understanding modes; and in response to the command information, controlling the NPC to execute a virtual event according to a first semantic understanding mode among the at least two semantic understanding modes, wherein the first semantic understanding mode is determined from among the at least two semantic understanding modes on the basis of environment sensing information of the user-controlled virtual character and / or the NPC.
Need to check novelty before this filing date? Find Prior Art

Description

Virtual role control method, processing method, device, equipment and storage medium

[0001] Priority information

[0002] The present application claims priority to the Chinese patent application No. 2024110983968, filed on August 9, 2024, and entitled "Virtual role control method, processing method, device, equipment and storage medium", the whole content of which is incorporated herein by reference. TECHNICAL FIELD

[0003] The present application relates to the field of human-computer interaction, and in particular to a virtual role control method, a processing method, a device, an equipment and a storage medium. BACKGROUND

[0004] A large number of scene objects are deployed in an application program including a virtual environment, which provides a virtual space for virtual activities of a virtual role.

[0005] In the related art, natural semantics and behaviors of a virtual role are usually determined in a mapping manner, for example, a fixed password "forward" is triggered to control the virtual role to move forward.

[0006] However, the above-mentioned fixed password is single in type, and how to flexibly control the virtual role in a natural language manner is a problem to be solved. SUMMARY

[0007] The present application provides a virtual role control method, a processing method, a device, an equipment and a storage medium, and the technical solution is as follows:

[0008] According to an aspect of the present application, a virtual role control method is provided, which comprises:

[0009] Displaying a main control virtual role or an NPC located in a virtual environment;

[0010] Obtaining command information, the command information including natural semantics for controlling the NPC, and the command information corresponding to at least two semantic understanding manners;

[0011] In response to the command information, controlling the NPC to perform a virtual activity according to a first semantic understanding manner in the at least two semantic understanding manners;

[0012] The first semantic understanding manner is determined in the at least two semantic understanding manners based on environment perception information of the main control virtual role and / or the NPC.

[0013] According to an aspect of the present application, a virtual role control method is provided, which comprises:

[0014] display a master virtual character or NPC located in a virtual environment;

[0015] obtain command information, the command information comprising natural language semantics for controlling the NPC, the command information comprising a behavior intention and a scene entity; the behavior intention being used to indicate a type of virtual activity performed by the NPC, and the scene entity being used to indicate a target entity to which the virtual activity is directed;

[0016] in response to the command information, control the NPC to perform a virtual activity;

[0017] wherein the virtual activity is determined according to environmental perception information of the master virtual character and / or the NPC.

[0018] According to an aspect of the present application, a method for processing command information is provided, the method comprising:

[0019] obtain command information, the command information comprising natural language semantics for controlling the NPC, the command information corresponding to at least two semantic understanding manners;

[0020] determine a first semantic understanding manner from the at least two semantic understanding manners of the command information based on environmental perception information of a master virtual character and / or the NPC;

[0021] based on the command information, control the NPC to perform a virtual activity according to the first semantic understanding manner from the at least two semantic understanding manners.

[0022] According to an aspect of the present application, a method for processing command information is provided, the method comprising:

[0023] obtain command information, the command information comprising natural language semantics for controlling the NPC, the command information comprising a behavior intention and a scene entity; the behavior intention being used to indicate a type of virtual activity performed by the NPC, and the scene entity being used to indicate a target entity to which the virtual activity is directed;

[0024] based on the command information, control the NPC to perform a virtual activity;

[0025] wherein the virtual activity is determined according to environmental perception information of the master virtual character and / or the NPC.

[0026] According to an aspect of the present application, a method for controlling a virtual character is provided, the method comprising:

[0027] display at least one of a master virtual character and an NPC located in a virtual environment;

[0028] acquire a natural language command, the natural language command comprising natural semantics for instructing the NPC, the command information comprising a behavior intention and a scene entity; the behavior intention is used to indicate a type of virtual activity performed by the NPC, and the scene entity is used to indicate a target entity to which the virtual activity is directed;

[0029] in response to the natural language command, control the NPC to perform the virtual activity;

[0030] wherein the virtual activity is determined according to environmental perception information of the host virtual character and / or the NPC.

[0031] In an optional implementation of the application, the method further comprises:

[0032] querying the target entity from an entity set of the virtual environment according to target entity information of the target entity indicated by the natural language command and the environmental perception information;

[0033] the response to the natural language command, the control of the NPC to perform the virtual activity, comprises:

[0034] in response to the natural language command, control the NPC to perform the virtual activity associated with the target entity;

[0035] wherein the environmental perception information comprises information perceived by at least one of the host virtual character and the NPC from the virtual environment.

[0036] In an optional implementation of the application, the querying of the target entity from the entity set of the virtual environment according to the target entity information of the target entity indicated by the natural language command and the environmental perception information comprises:

[0037] parsing the natural language command to obtain the target entity information; the target entity information comprises at least one of the following: entity type, entity name, entity position, and entity feature;

[0038] calculate the similarity of the target entity information and each entity information in the entity set;

[0039] determine the target entity from the entity set according to the similarity and the environmental perception information.

[0040] In an optional implementation of the application, the determining of the target entity from the entity set according to the similarity and the environmental perception information comprises:

[0041] determine a target perception range according to the target entity information and the environmental perception information;

[0042] filtering the target entity from entities perceived within the target perception range according to the similarity.

[0043] In an optional implementation of the present application, the similarity includes a text similarity; and the entity set includes entity information of a first entity.

[0044] The calculating of the similarity between the target entity information and each entity information in the entity set includes:

[0045] tokenizing the target entity information to obtain at least one target entity label;

[0046] converting the at least one target entity label into at least one target embedding vector;

[0047] obtaining entity information of the first entity, the entity information including a text embedding vector, the text embedding vector being an embedding vector converted based on a text label of the first entity;

[0048] respectively calculating a text parent similarity between the at least one target embedding vector and the text embedding vector to obtain at least one text parent similarity corresponding to the at least one target embedding vector respectively;

[0049] determining a sum of the at least one text parent similarity as a text similarity between the target entity information and the entity information of the first entity.

[0050] In an optional implementation of the present application, the similarity includes an image similarity; and the entity set includes entity information of a first entity.

[0051] The calculating of the similarity between the target entity information and each entity information in the entity set includes:

[0052] tokenizing the target entity information to obtain at least one target entity label;

[0053] converting the at least one target entity label into at least one target embedding vector;

[0054] obtaining entity information of the first entity, the entity information including an image embedding vector, the image embedding vector being an embedding vector extracted based on an image of the first entity;

[0055] respectively calculating an image parent similarity between the at least one target embedding vector and the text embedding vector to obtain at least one image parent similarity corresponding to the at least one target embedding vector respectively;

[0056] The sum of the at least one image parent similarity is determined as an image similarity between the target entity information and entity information of the first entity.

[0057] In an optional implementation of the present application, the method further includes:

[0058] The at least one hierarchical prediction network is called to perform intent recognition on the natural language command, to obtain a first classification label of an intent of the natural language command commanding the NPC; the hierarchical prediction network includes at least two sub-networks, and the at least two sub-networks are constructed based on a tree structure of the multiple classification labels;

[0059] The NPC is controlled to perform the virtual activity in response to the natural language command, including:

[0060] The NPC is controlled to perform the virtual activity according to a first intent corresponding to the first classification label in response to the natural language command.

[0061] In an optional implementation of the present application, each hierarchical prediction network in the at least one hierarchical prediction network is used to predict a commanding intent of the natural language command on the NPC in one intent dimension;

[0062] An i-th sub-network in each hierarchical prediction network is used to predict a first-level behavior label of the natural language command in the one intent dimension, and an i+1-th sub-network in the hierarchical prediction network is used to predict a second-level behavior label corresponding to the first-level behavior label of the natural language command, i being a positive integer;

[0063] The intent dimension includes at least one of a subject dimension, a semantic dimension, and a behavior dimension.

[0064] In an optional implementation of the present application, the intent dimension includes the subject dimension, the hierarchical prediction network includes a hierarchical-structure subject prediction network, the subject prediction network has a capability of predicting a subject type in a natural language command, and the subject type is used to indicate an identity of the NPC commanded by the natural language command.

[0065] The intent dimension includes the semantic dimension, the hierarchical prediction network includes a hierarchical-structure semantic prediction network, the semantic prediction network has a capability of predicting a semantic type in a natural language command, and the semantic type is used to indicate a control mode of the NPC for initiating a virtual attack.

[0066] And / or, the intention dimension includes the behavior dimension, the hierarchical prediction network includes a hierarchical behavior prediction network, and the behavior prediction network has the ability to predict a behavior intention type in the natural language command; the behavior intention type is used to indicate a behavior manner of the NPC for performing a virtual activity.

[0067] In an optional implementation of the present application, the method further includes:

[0068] The information structure prediction network is called to perform a classification prediction on the natural language command, and it is predicted that the natural language command corresponds to a first control label; the information structure prediction network has the ability to predict whether the natural language command carries natural semantics for commanding the NPC, and the first control label is used to indicate that the natural language command carries natural semantics for commanding the NPC.

[0069] In an optional implementation of the present application, the method further includes:

[0070] The attribute text of the scene object in the virtual scene is obtained, and the attribute text is used to introduce inherent attributes of the scene object in the virtual scene.

[0071] The appearance image of the scene object in the virtual scene is obtained, and the appearance image is used to describe a style of the scene object.

[0072] A multimodal model is called to perform a prediction on the attribute text and the appearance image of the scene object, to obtain a visual label of the scene object, and the visual label is used to describe visual features of the scene object in at least one dimension.

[0073] In an optional implementation of the present application, the multimodal model includes a visual question and answer model, and the calling of the multimodal model to perform the prediction on the attribute text and the appearance image of the scene object to obtain the visual label of the scene object includes:

[0074] A question sentence for the appearance image is constructed, and the question sentence carries the attribute text of the scene object.

[0075] The appearance image and the question sentence of the scene object are input into the visual question and answer model to obtain an answer sentence, and the answer sentence is taken as the visual label of the scene object.

[0076] In an optional implementation of the present application, the construction of the question sentence for the appearance image includes:

[0077] obtaining expected information of the scene object, the expected information being used to indicate a description dimension of the scene object expected in the visual label, and / or a desired format of the visual label;

[0078] constructing the question sentence of the appearance image according to the expected information and the attribute text;

[0079] wherein a first subpart in the question sentence is supplementary introduction information of the scene object, carrying the attribute text of the scene object; and a second subpart in the question sentence is an answer guiding sentence for the visual question answer model, carrying the expected information.

[0080] In an optional implementation of the present application, the attribute text of the scene object includes at least one of a name of the scene object in the virtual scene and a size of the scene object in the virtual scene.

[0081] And / or, the appearance image of the scene object includes images obtained by observing the scene object from at least two perspectives.

[0082] In an optional implementation of the present application, the method further includes:

[0083] performing pseudo-spontaneous rewriting on the visual label of the scene object to obtain a matching label conforming to a spontaneous language expression.

[0084] In an optional implementation of the present application, the environment perception information further includes spatial information of the scene object, and the method further includes:

[0085] obtaining a spatial position of the scene object in the virtual scene, and determining the spatial position of the scene object as auxiliary information of the visual label of the scene object.

[0086] In an optional implementation of the present application, the obtaining of the spatial position of the scene object in the virtual scene includes:

[0087] obtaining at least one of a coordinate position, orientation information, bounding box information and occlusion point information of the scene object in the virtual scene;

[0088] wherein the coordinate position is used to indicate a position of the scene object in the virtual scene, the orientation information is used to indicate a direction faced by the scene object in the virtual scene, the bounding box information is used to indicate a size of the scene object in the virtual scene, and the occlusion point information indicates a recommended virtual character standing point when a virtual character approaches the scene object.

[0089] In an optional implementation of the present application, the method further comprises:

[0090] The feedback information of the NPC is broadcasted, and the feedback text corresponding to the feedback information is non-fixed text generated based on the environment perception information of the host virtual character and / or the NPC.

[0091] In an optional implementation of the present application, the feedback information comprises a reply broadcast; and the broadcasting of the feedback information of the NPC comprises:

[0092] In a case where the behavior intention of the natural language command is a query, the reply broadcast of the NPC is broadcasted; and the reply broadcast comprises reply content for the natural language command generated based on the environment perception information of the NPC.

[0093] The natural language command is used to query information from the non-player character.

[0094] In an optional implementation of the present application, the feedback information comprises a feedback broadcast; and the broadcasting of the feedback information of the NPC comprises:

[0095] In a case where the behavior intention of the natural language command is to execute a task, the feedback broadcast of the NPC is broadcasted; and the feedback broadcast comprises an execution status of the NPC executing the task according to the environment perception information of the NPC.

[0096] The natural language command is used to instruct the NPC to execute a task.

[0097] In an optional implementation of the present application, the feedback information comprises a dynamic broadcast; and the broadcasting of the feedback information of the NPC comprises:

[0098] In a case where the environment perception information of the NPC meets a dynamic broadcast condition, a dynamic broadcast is broadcasted; and the dynamic broadcast comprises an abnormal situation identified according to the environment perception information of the NPC.

[0099] In an optional implementation of the present application, the feedback information is a spatial voice audio of the non-player character.

[0100] The broadcasting of the feedback information of the NPC comprises:

[0101] The feedback text corresponding voice audio is obtained;

[0102] Based on the relative positional relationship between the non-player character and the host virtual character in the three-dimensional virtual environment, an audio effect parameter corresponding to the non-player character is obtained;

[0103] adjust the human voice audio based on the sound effect parameter, generate, and broadcast the spatial human voice audio corresponding to the non-player character, the spatial human voice audio being used to represent the perception effect of the audio generated by the master virtual character on the non-player character in the relative positional relationship.

[0104] In an optional implementation of the present application, the obtaining of the sound effect parameter corresponding to the non-player character based on the relative positional relationship between the non-player character and the master virtual character in the three-dimensional virtual environment comprises:

[0105] obtaining first positional information of the non-player character in the three-dimensional virtual environment; and obtaining second positional information of the master virtual character in the three-dimensional virtual environment;

[0106] determining the relative positional relationship between the non-player character and the master virtual character based on the first positional information and the second positional information;

[0107] generating the sound effect parameter corresponding to the non-player character based on the relative positional relationship.

[0108] In an optional implementation of the present application, the natural language command is a voice form natural language command, and the method further comprises:

[0109] obtaining a plurality of first scene hotwords corresponding to a first virtual scene in which a target virtual character is located, the target virtual character comprising at least one of the NPC and the master virtual character, the plurality of first scene hotwords being scene-related vocabularies of the first virtual scene;

[0110] converting the voice form natural language command into a text form based on the plurality of first scene hotwords.

[0111] In an optional implementation of the present application, the obtaining of the plurality of first scene hotwords corresponding to the first virtual scene in which the target virtual character is located comprises:

[0112] obtaining an object perception range of the target virtual character in the first virtual scene, the object perception range being a three-dimensional spatial range in which the target virtual character has perception of other scene elements, the object perception range comprising at least one of: a visual perception range of the target virtual character; an auditory perception range of the target virtual character; an olfactory perception range of the target virtual character; a range perceived by a perception skill or a perception device possessed by the target virtual character;

[0113] obtaining the plurality of first scene hotwords based on environmental perception information in the object perception range.

[0114] In an optional implementation of the present application, the converting the natural language command in the voice form into the text form based on the plurality of first scene hotwords comprises:

[0115] In the case that the natural language command is analyzed by a natural language analysis model, a plurality of speech units corresponding to the natural language command are obtained by an acoustic network, the speech unit is a basic constituting unit of a word pronunciation, the natural language analysis model comprises the acoustic network, a preset dictionary, and language sub-networks corresponding to at least two virtual scenes respectively, and the virtual environment comprises the at least two virtual scenes;

[0116] Based on the first virtual scene in which the target virtual role is located, a first language sub-network corresponding to the first virtual scene is determined from at least two language sub-networks;

[0117] The sequence matching relationship between the plurality of speech units and selected words in the preset dictionary is analyzed by the first language sub-network, the selected words at least comprising the plurality of first scene hotwords and the general words, and the natural language command in the voice form is converted into the text form.

[0118] According to another aspect of the present application, a virtual role control device is provided, the device comprising:

[0119] A display module configured to display a main control virtual role or NPC located in a virtual environment;

[0120] An acquisition module configured to acquire command information, the command information comprising natural semantics for controlling the NPC, and the command information corresponding to at least two semantic understanding manners;

[0121] A control module configured to control the NPC to perform a virtual activity according to a first semantic understanding manner among the at least two semantic understanding manners in response to the command information.

[0122] The first semantic understanding manner is determined based on environmental perception information of the main control virtual role and / or the NPC among the at least two semantic understanding manners.

[0123] In an optional design of the present application, the command information corresponds to at least two semantic understanding manners of scene entities, and the control module is further configured to:

[0124] In response to the command information, the NPC is controlled to perform the virtual activity on a first scene object located within a perception range of the main control virtual role and / or the NPC.

[0125] The virtual environment has a plurality of scene objects belonging to the same type, and the plurality of scene objects include the first scene object.

[0126] In an optional design of the present application, the control module is further configured to:

[0127] In response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in the visual field range.

[0128] In an optional design of the present application, the visual field range is a visual field range observed by the host virtual character and / or the NPC; and / or, the visual field range is an extended visual field range obtained by using a perception skill or a perception device.

[0129] In an optional design of the present application, in response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in the perception range of the host virtual character and / or the NPC, including:

[0130] In response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in the first virtual terrain, and the first virtual terrain is a virtual terrain to which a location of the host virtual character and / or the NPC belongs.

[0131] In an optional design of the present application, the control module is further configured to at least one of:

[0132] In a case where the first virtual terrain is an indoor terrain, in response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in a first area, and the first area is an indoor connected area including a first location where the host virtual character and / or the NPC is located.

[0133] In a case where the first virtual terrain is a street terrain adjacent to a building, in response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in a second area, and the second area includes a building surrounding a second location where the host virtual character and / or the NPC is located.

[0134] In a case where the first virtual terrain is an outdoor terrain away from a building, in response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in a third area, and the third area is an area with the same virtual vegetation distribution as a third location where the host virtual character and / or the NPC is located.

[0135] In an optional design of the present application, the command information corresponds to semantic understanding manners of at least two word pronunciations, and the control module is further configured to:

[0136] In response to the command information, the command information is converted into command text according to natural language pronunciations of scene objects in a perception range, the perception range being a virtual area in which the host virtual character and / or the NPC obtain the environmental perception information;

[0137] According to the command text, the NPC is controlled to perform the virtual activity according to the first semantic understanding manner.

[0138] In an optional design of the present application, the control module is further configured to:

[0139] When a similarity between a pronunciation of a first phoneme in the command information and the natural language pronunciation of the scene object exceeds a similarity threshold, and candidate text of the first phoneme exists, the text information of the first phoneme is determined as the name text and / or the introduction text of the scene object, and the command text is constructed according to the text information;

[0140] The candidate text of the first phoneme does not include the name text and / or the introduction text of the scene object.

[0141] In an optional design of the present application, the command information corresponds to semantic understanding manners of at least two behavior intentions, and the control module is further configured to:

[0142] In response to the command information, the NPC is controlled to perform the virtual activity according to a first behavior intention; the first behavior intention corresponds to a type of virtual activity permitted to be performed in a perception range, the perception range being a virtual area in which the host virtual character and / or the NPC obtain the environmental perception information.

[0143] In an optional design of the present application, the control module is further configured to:

[0144] In response to the command information, a first direction in the perception range is taken as a reference direction, and the NPC is controlled to perform the virtual activity based on a second direction according to a direction word in the command information;

[0145] The direction word is used to describe an angle between the second direction and the first direction, and / or describe the second direction in a reference system constructed based on the first direction.

[0146] In an optional design of the present application, the control module is further configured to:

[0147] According to the command information, command text in a natural language form is determined.

[0148] invoke a conditional random field to perform entity prediction on the command text, to obtain an entity label carried by the command text;

[0149] determine a first scene object in the perception range according to the entity label;

[0150] invoke a prediction network to perform intent recognition on the command text, to obtain a classification label of the command text; the prediction network is configured to predict a control manner of the command text on at least one behavior dimension for the NPC;

[0151] determine a control command of the NPC according to the classification label, and control the NPC to perform the virtual activity on the first scene object in the first semantic understanding manner based on the control command.

[0152] In an optional design of the present application, the control module is further configured to:

[0153] invoke a hierarchical prediction network to perform intent recognition on the command text, to obtain the classification label of the command text;

[0154] The hierarchical prediction network includes at least two sub-networks, and the at least two sub-networks are constructed based on a tree structure of multiple classification labels.

[0155] In an optional design of the present application, the acquisition module is further configured to:

[0156] acquire visual labels of multiple candidate scene objects in the perception range;

[0157] The control module is further configured to:

[0158] respectively calculate similarities between the visual labels of the multiple candidate scene objects and the entity label, to determine the first scene object with the highest similarity among the multiple candidate scene objects.

[0159] In an optional design of the present application, the acquisition module is further configured to:

[0160] acquire spatial positions of each scene object in the virtual environment in the virtual environment;

[0161] The control module is further configured to:

[0162] screen the multiple candidate scene objects in the perception range from the each scene object in the virtual environment according to the spatial positions.

[0163] In an optional design of the present application, the acquisition module is further configured to:

[0164] obtain at least one of coordinate position, orientation information, bounding box information, and cover point information of each scene object in the virtual environment;

[0165] The coordinate position is used to indicate a position of the scene object in the virtual environment, the orientation information is used to indicate a direction faced by the scene object in the virtual environment, the bounding box information is used to indicate a size of the scene object in the virtual environment, and the cover point information indicates a recommended standing point of a virtual character when the virtual character approaches the scene object.

[0166] In an optional design of the present application, the control module is further configured to:

[0167] obtain attribute text of the scene object in the virtual environment, the attribute text being used to introduce inherent attributes of the scene object in the virtual environment;

[0168] obtain appearance images of the scene object in the virtual environment, the appearance images being used to describe styles of the scene object;

[0169] invoke a multi-modal model to perform prediction on the attribute text and the appearance images of the scene object, to obtain a visual label of the scene object, the visual label being used to describe the scene object in at least one dimension;

[0170] construct a spatial database of the virtual environment based on the visual labels of the scene objects in the virtual environment;

[0171] filter, based on the perception range, the visual labels of the candidate scene objects in the perception range in the spatial database.

[0172] In an optional design of the present application, the device further includes:

[0173] a presentation module configured to present feedback information of the NPC to the command information, the feedback information being used to indicate that the NPC receives the command information and / or that the NPC has executed the virtual activity indicated by the command information.

[0174] According to another aspect of the present application, a processing device of command information is provided, and the device includes:

[0175] an obtaining module configured to obtain command information, the command information including natural semantics for controlling an NPC, and the command information corresponding to at least two semantic understanding manners;

[0176] The processing module is configured to determine a first semantic understanding manner from the at least two semantic understanding manners of the command information based on environment perception information of the host virtual character and / or the NPC.

[0177] The control module is configured to control the NPC to perform a virtual activity according to the first semantic understanding manner from the at least two semantic understanding manners of the command information.

[0178] In an optional design of the present application, the command information corresponds to at least two semantic understanding manners of scene entities.

[0179] The processing module is further configured to:

[0180] determine the command information as being for a first scene object located within a perception range of the host virtual character and / or the NPC based on the environment perception information; and

[0181] wherein a plurality of scene objects belonging to the same type exist in a virtual environment in which the host virtual character and / or the NPC is located, the plurality of scene objects include the first scene object, and the perception range is a virtual area in which the host virtual character and / or the NPC acquires the environment perception information.

[0182] In an optional design of the present application, the processing module is further configured to:

[0183] determine the command information as being for a first scene object located in a visual field range based on the environment perception information; and

[0184] or, determine the command information as being for a first scene object located in a first virtual terrain based on the environment perception information; the first virtual terrain is a virtual terrain to which a location of the host virtual character and / or the NPC belongs.

[0185] In an optional design of the present application, the command information corresponds to at least two semantic understanding manners of word pronunciations.

[0186] The processing module is further configured to:

[0187] determine the command information as carrying a natural language pronunciation of a scene object in a perception range based on the environment perception information, and convert the command information into a command text according to the first semantic understanding manner; and

[0188] wherein the perception range is a virtual area in which the host virtual character and / or the NPC acquires the environment perception information, and the NPC performs the virtual activity based on an indication of the command text.

[0189] In an optional design of the present application, the command information corresponds to semantic understanding manners of at least two behavior intentions.

[0190] The processing module is further configured to:

[0191] Based on the environment perception information, the command information is determined to be executed according to a first behavior intention, the first behavior intention corresponding to a virtual activity type permitted to be executed in a perception range, the perception range being a virtual area in which the host virtual character and / or the NPC obtain the environment perception information.

[0192] In an optional design of the present application, the processing module is further configured to:

[0193] According to the command information, a command text in a natural language form is determined;

[0194] Conditional random fields are invoked to perform entity prediction on the command text, to obtain entity labels carried by the command text;

[0195] According to the entity labels, a first scene object is determined in the perception range;

[0196] A prediction network is invoked to perform intention recognition on the command text, to obtain a classification label of the command text, the prediction network being configured to predict a control manner of the command text on at least one behavior dimension for the NPC;

[0197] The control of the NPC according to the first semantic understanding manner of the at least two semantic understanding manners based on the command information includes:

[0198] According to the classification label corresponding to the command information and the first scene object, a control command of the NPC is determined, and the control command is sent, the control command being configured to control the NPC to execute the virtual activity according to the first semantic understanding manner for the first scene object.

[0199] In an optional design of the present application, the processing module is further configured to:

[0200] A hierarchical prediction network is invoked to perform intention recognition on the command text, to obtain the classification label of the command text;

[0201] The hierarchical prediction network includes at least two sub-networks, and the at least two sub-networks are constructed based on a tree structure of multiple classification labels.

[0202] In an optional design of the present application, the acquisition module is further configured to acquire visual labels of multiple candidate scene objects in the perception range;

[0203] The processing module is further configured to calculate similarity between the visual label of each of the candidate scene objects and the entity label respectively, and determine the first scene object with the highest similarity among the candidate scene objects.

[0204] In an optional design of the present application, the obtaining module is further configured to obtain spatial positions of each scene object in the virtual environment.

[0205] The processing module is further configured to filter the candidate scene objects in the perception range from the scene objects in the virtual environment according to the spatial positions.

[0206] In an optional design of the present application, the obtaining module is further configured to:

[0207] obtain at least one of coordinate positions, orientation information, bounding box information, and cover point information of the scene object in the virtual environment.

[0208] The coordinate positions are used to indicate the position of the scene object in the virtual environment, the orientation information is used to indicate the direction faced by the scene object in the virtual environment, the bounding box information is used to indicate the size of the scene object in the virtual environment, and the cover point information is used to indicate the recommended virtual character standing point when the virtual character approaches the scene object.

[0209] In an optional design of the present application, the obtaining module is further configured to:

[0210] obtain attribute text of the scene object in the virtual environment, the attribute text being used to introduce inherent attributes of the scene object in the virtual environment.

[0211] obtain appearance images of the scene object in the virtual environment, the appearance images being used to describe the style of the scene object.

[0212] invoke a multi-modal model to perform prediction on the attribute text and the appearance images of the scene object, to obtain a visual label of the scene object, the visual label being used to describe the scene object in at least one dimension.

[0213] construct a spatial database of the virtual environment based on the visual labels of each scene object in the virtual environment.

[0214] filter the visual labels of the candidate scene objects in the perception range based on the perception range in the spatial database.

[0215] According to another aspect of the present application, a virtual character control device is provided, the device comprising:

[0216] a display module configured to display at least one of the master virtual character and the NPC in the virtual environment;

[0217] an acquisition module configured to acquire a natural language command, the natural language command comprising natural semantics for instructing the NPC, the command information comprising a behavior intention and a scene entity, the behavior intention being used to indicate a type of virtual activity performed by the NPC, and the scene entity being used to indicate a target entity to which the virtual activity is directed;

[0218] a control module configured to control the NPC to perform the virtual activity in response to the natural language command.

[0219] The virtual activity is determined according to environmental perception information of the master virtual character and / or the NPC.

[0220] According to another aspect of the present application, a computer device is provided, the computer device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the virtual character control method and / or the command information processing method according to the above aspects.

[0221] According to another aspect of the present application, a computer readable storage medium is provided, the readable storage medium storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by a processor to implement the virtual character control method and / or the command information processing method according to the above aspects.

[0222] According to another aspect of the present application, a computer program product is provided, the computer program product comprising computer instructions stored in a computer readable storage medium, the computer instructions being read and executed by a processor from the computer readable storage medium to implement the virtual character control method and / or the command information processing method according to the above aspects.

[0223] The technical solutions provided by the present application have at least the following beneficial effects:

[0224] According to the embodiment of the present application, the virtual activity is executed according to the first semantic understanding mode by referring to the environment perception information of the host virtual character and / or the NPC, and in the case that the command information has at least two semantic understanding modes, on the one hand, the user can control the NPC to execute the virtual activity according to the command information without supplementing more information, and on the other hand, the environment perception information of the host virtual character and / or the NPC is referred to, so that the NPC and the host virtual character can execute the virtual activity cooperatively. BRIEF DESCRIPTION OF DRAWINGS

[0225] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.

[0226] Fig. 1 is a schematic diagram of controlling a non-player character in a virtual environment according to an example embodiment of the present application;

[0227] Fig. 2 is a structural block diagram of a computer system according to an example embodiment of the present application;

[0228] Fig. 3 is a flowchart of a virtual character control method according to an example embodiment of the present application;

[0229] Fig. 4 is a flowchart of a virtual character control method according to an example embodiment of the present application;

[0230] Fig. 5 is a flowchart of a virtual character control method according to an example embodiment of the present application;

[0231] Fig. 6 is a flowchart of a virtual character control method according to an example embodiment of the present application;

[0232] Fig. 7 is a flowchart of a virtual character control method according to an example embodiment of the present application;

[0233] Fig. 8 is a schematic diagram of command information according to an example embodiment of the present application;

[0234] Fig. 9 is a schematic diagram of a visual tag of a scene object according to an example embodiment of the present application;

[0235] Fig. 10 is a flowchart of a virtual character control method according to an example embodiment of the present application;

[0236] Fig. 11 is a flowchart of a command information processing method according to an example embodiment of the present application;

[0237] FIG. 12 is a schematic diagram of a method for processing command information according to an example embodiment of the present application;

[0238] FIG. 13 is a schematic diagram of a behavior tree of an NPC according to an example embodiment of the present application;

[0239] FIG. 14 is a flowchart of a method for controlling a virtual character according to an example embodiment of the present application;

[0240] FIG. 15 is a block diagram of a control device for a virtual character according to an example embodiment of the present application;

[0241] FIG. 16 is a block diagram of a processing device for command information according to an example embodiment of the present application;

[0242] FIG. 17 is a block diagram of a terminal according to an example embodiment of the present application.

[0243] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the present application. DETAILED DESCRIPTION

[0244] To make the objects, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0245] The example embodiments will be described in detail herein with reference to the accompanying drawings. The following description is presented for purposes of illustration and description, and is not intended to limit the application, as understood by persons of ordinary skill in the art. The implementation described in the following example embodiments is merely representative of the devices and methods consistent with aspects of the present application, as detailed in the following claims.

[0246] The terminology used in the present disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used in the present disclosure and the appended claims, the singular forms "a," "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0247] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of countries and regions. For example, the command information and other information involved in the present application are obtained under sufficient authorization.

[0248] It should be understood that although the terms first, second, etc. can be used in this disclosure to describe various information, these information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, a first parameter can also be referred to as a second parameter without departing from the scope of the present disclosure, and similarly, a second parameter can also be referred to as a first parameter. Depending on the context, the word "if" as used herein can be interpreted as "when" or "upon" or "in response to determining".

[0249] The embodiments of the present application provide a scheme for a user to give natural language commands in the form of voice to instruct a non-player character (NPC) in a virtual environment. In the scheme, the user usually has one or more NPCs as teammates in addition to a master virtual character controlled by the user. The user can use human conversation in a relatively casual and non-mechanical manner to instruct the NPC to make feedback on scene objects in the virtual environment as desired by the user, so as to instruct the NPC to complete a task in cooperation with the master virtual character. For example, the NPC has autonomous behavior capability, and the user only instructs the NPC, the user gives a command, and the NPC understands and executes the command based on its autonomous behavior capability. For example, referring to FIG. 1, a user controls a master virtual character 10 to play a game, and the master virtual character 10 has an NPC teammate 20. There is a car 30 in the visual range of the master virtual character 10, and the user says a natural language command in the form of voice: "No. 2, go and hide behind the car 30", and then the NPC teammate 20 will move to hide behind the car 30 by itself. It should be noted that there can be multiple cars in the scene shown in FIG. 1, and the NPC teammate 20 will accurately understand the car said by the user as the car 30 in the visual range of the user. That is, the NPC teammate 20 has relatively intelligent natural language understanding capability.

[0250] Compared with the traditional technology of instructing the NPC by using mechanical fixed instructions, the embodiments of the present application instruct the NPC by using relatively casual natural language commands, which can provide the user with more natural, more complex and more flexible language instruction capability.

[0251] In another aspect, the natural language understanding capability of the embodiments of the present application is embodied in that the NPC teammate not only considers the literal information of the natural language command when understanding the natural language command. The NPC teammate also combines the environmental perception information of the virtual environment of the host virtual character and / or the NPC itself to assist in understanding the semantics in the natural language command. That is, the NPC not only considers the information of the natural language command in one modality, but also considers the perception information of other modalities such as visual field, hearing, radar, etc. to assist in understanding and executing the natural language command. Since the natural language command in the form of voice is a command in the form of spoken language, not a command in the form of written language, there may be multiple candidate understanding ways or unclear places when only considering the literal information of the natural language command. The NPC teammate combines the environmental perception information of the virtual environment of the host virtual character and / or the NPC itself to determine a reasonable understanding way among the multiple candidate understanding ways or eliminate doubts about unclear places, thereby realizing a more intelligent natural language understanding capability. The environmental perception information mentioned above includes at least one of the following:

[0252] · visual information perceived within the visual field range of the host virtual character;

[0253] · hearing information perceived within the hearing range of the host virtual character;

[0254] · information perceived by the perception skill or perception device possessed by the host virtual character;

[0255] · visual information perceived within the visual field range of the NPC;

[0256] · hearing information perceived within the hearing range of the NPC;

[0257] · information perceived by the perception skill or perception device possessed by the NPC.

[0258] In combination with reference to FIG. 1, the scheme includes at least one of the following five stages:

[0259] Stage one: preprocessing of spatial data;

[0260] The server will preprocess the scene objects in the virtual scene and construct the spatial data of the virtual scene.

[0261] The spatial data of the virtual scene includes the visual label of the scene object and the spatial position of the scene object, and in some examples, the spatial data also includes the appearance image of the scene object. The scene object is any object appearing in the virtual scene, such as the car, wall, box, etc. in FIG. 1.

[0262] Firstly, the pre-acquisition process of the visual label of the scene object is introduced. The attribute information of the scene object in the virtual scene is acquired, such as at least one of the name, size and the like of the scene object; and the appearance image of the scene object is acquired, the appearance image including images obtained by observing the scene object from at least two perspectives, and observing the scene object from multiple perspectives can ensure that the appearance image carries comprehensive appearance information of the scene object.

[0263] The server calls the multi-modal model to perform prediction on the attribute information and the appearance image of the scene object, extracts hidden layer features from the attribute information in the text dimension and the appearance image in the image dimension based on the ability of the multi-modal model to process information in the text dimension and the image dimension, and performs decoding on the hidden layer features to predict the visual label of the scene object, the visual label being used to describe the scene object in at least one dimension in a text manner; such as describing the scene object in the dimensions of material, transparency, color, shape, etc. The natural language model is called to perform label optimization on the visual label to obtain the visual label of the scene object; the natural language model has a text generation capability, and performs rewriting on the visual label input into the natural language model in a manner conforming to the spoken expression of natural language, so as to realize label optimization of the visual label and obtain the optimized visual label of the scene object. The optimized visual label conforms to the spoken expression of natural language, and describes the scene object in at least one dimension in a manner close to the spoken expression.

[0264] Taking a virtual bed in a virtual scene as an example, usually the color, material, and placement position of the virtual bed are concerned, while usually the surface features of the head and side plate of the bed, such as carving and embossing, are ignored. The purpose of performing pseudo-spoken expression rewriting on the visual label of the scene object is to obtain a matching label closer to the spoken expression; such as the label of the color, material, and placement position dimensions that are concerned in the spoken expression, and the label of the surface feature dimension of the head and side plate of the bed that is ignored in the spoken expression.

[0265] Next, the spatial position of the scene object is introduced. The spatial position includes at least one of the following of the scene object: coordinate position (Location), such as coordinate information of a center point or a preset point of the scene object in the virtual scene; orientation (Rotation), such as a direction in which a front surface of the scene object faces in the virtual scene; bounding box (Bounding Box), used to indicate the size of the scene object in the virtual scene; cover point (Cover points), used to indicate a recommended virtual character position point when a virtual character approaches the scene object, so as to realize that the scene object can mask the virtual character.

[0266] Next, the pre-acquisition process of the appearance image of the scene object is introduced. The appearance image is acquired in the process of predicting the visual label of the scene object, and the appearance image includes images obtained by observing the scene object from at least two perspectives.

[0267] Optionally, the spatial data of each scene object in the entire virtual environment is pre-acquired and arranged into a data set for use in the subsequent query process.

[0268] Stage two: voice recognition and intent recognition;

[0269] In the process of controlling the master virtual role by the user, the terminal receives the natural language command input by the user in the voice; the server performs voice recognition and intent recognition on the natural language command to analyze the input instruction of the user;

[0270] Voice recognition is a text conversion on the natural language instruction input by the user in the voice to obtain the instruction text corresponding to the natural language instruction, and converts the voice information into text information. The encoding network is called to perform feature encoding on the instruction text to obtain the feature representation of the natural language instruction in the hidden layer space; the text segmentation network is called to perform segmentation on the feature representation to obtain multiple clauses in the instruction text, and then the intent recognition is performed on each clause in the multiple clauses. In the case that the natural language instruction input by the user in the voice is a long sentence, the segmentation of the instruction text corresponding to the natural language instruction can be realized, and the demand of analyzing the input instruction of the user in the long sentence scenario can be met.

[0271] The intent recognition of each clause is introduced as follows: the input instruction of the user is obtained by intent recognition, and the input instruction of the user includes at least one entity, semantic type, subject type and intent type in the clause.

[0272] On the one hand, the Conditional Random Field (CRF) method is used to recognize the entity in the clause, such as recognizing the entity in the clause as a scene object in the virtual scene, such as a virtual building, a virtual article, virtual vegetation, etc. The entity in the clause is used to indicate the scene object to which the virtual activity is directed, such as the entity in the clause is a virtual building, and the clause carries a semantic indicating that the NPC initiates a virtual attack, then the virtual attack is performed on the virtual building, that is, the virtual attack is initiated on the virtual building. On the other hand, the prediction network is called to predict the semantic type, subject type and intent type in the clause respectively. Among them, the semantic type includes whether to indicate that the NPC initiates a virtual attack; the subject type is used to indicate that the subject of the clause is one or more NPCs or a user-controlled master virtual role; and the intent type is a desired virtual activity indicated by the clause, such as moving, initiating a virtual attack, etc.

[0273] Stage three: query of spatial data;

[0274] For the entity in the input instruction, it is necessary to determine the spatial position of the entity in the virtual environment. In the case where there are multiple candidate entities in the virtual environment that match the entity in the input instruction, the target entity that matches the input instruction can be uniquely determined among the multiple candidate entities in combination with the environment perception information of the host virtual character and / or NPC.

[0275] For example, when the environment perception information is the field of view information of the host virtual character, the spatial data in the virtual environment is the query range, and the real-time position and orientation of the host virtual character are the query reference conditions, the spatial position of the target entity corresponding to the input instruction in the virtual environment is queried. In order to subsequently control the NPC to move to the vicinity of the target entity, or perform the behavior indicated by the input instruction on the target entity.

[0276] However, there can be multiple environment perception information of the host virtual character and / or NPC, and multiple environment perception information can be fused to assist the determination process of the target entity, or multiple environment perception information can be used according to priority to assist the determination process of the target entity, which is not limited.

[0277] Stage four: voice feedback of NPC;

[0278] The NPC will make voice feedback to the user. The types of voice feedback include at least one of immediate feedback, instruction execution feedback, and dynamic feedback. Among them, the immediate feedback refers to that the NPC makes voice feedback to the natural language instruction immediately after the user issues the natural language instruction to the NPC, such as "received", "start execution", "good" and the like; the execution instruction feedback refers to that the NPC makes voice feedback to the execution result of the control instruction after executing the control instruction indicated by the intention of the natural language instruction, such as "successfully executed", "failed to successfully execute" and the like; the dynamic feedback refers to that the NPC makes voice feedback according to real-time environment perception when the user does not issue the natural language instruction to the NPC.

[0279] Among them, the voice feedback is obtained by text-to-speech of the text content. The text content is the content inferred by the large language model in combination with the spatial data and dynamic environment perception information. Using the large language model to dynamically generate the text content of the voice feedback can break through the monotony, mechanicalness and easy repetition of the fixed template voice feedback on the one hand; on the other hand, it does not need to pre-save too many pre-produced template voices, which can avoid the problem of too large data volume of the client.

[0280] Stage five: behavior control of NPC;

[0281] The input instruction of the user is input, the spatial data of the virtual scene and the runtime information corresponding to the master virtual role are referenced, and an AI model is used to determine the target entity in the virtual space and the control instruction of the NPC.

[0282] The control instruction is an instruction that can be executed by the behavior tree of the NPC. The subsequent action indicated by the control instruction that the NPC needs to perform can be a single action or a sequence of actions. In the case of a sequence of multiple actions, the sequence of multiple actions is used to indicate that the NPC performs a sequence of actions, and the data structure for storing the sequence of multiple actions is in the form of a list. The instructions to be executed in the sequence of instructions are cached. When the behavior tree of the NPC executes the current instruction, the cached instructions to be executed are called back and executed.

[0283] Hotword update function:

[0284] For phase two, the hotword system is also introduced to improve the accuracy of speech recognition.

[0285] The hotword system is used to provide scene hotwords related to the scene involved in the natural language instruction in the speech recognition process, so that when a natural language instruction of a voice input is received, the hotword library can be searched for words as an instruction analysis result based on the currently running virtual scene, so that the matching degree between the instruction analysis result and the virtual scene is higher.

[0286] The virtual scene includes at least one of a virtual battle scene, a virtual transaction scene, a virtual office scene, a virtual kitchen scene, and a plurality of virtual scenes. The frequently used words in the plurality of virtual scenes, the specific virtual elements existing in the virtual scene (such as virtual buildings and virtual characters existing only in a specific scene type) are analyzed in advance as scene hotwords, and a scene hotword library composed of a plurality of scene hotwords is obtained. The plurality of scene hotwords correspond to the virtual scene. In addition, different scene language models are trained in advance for different virtual scenes, such as a scene language model in a battle scene that pays more attention to words related to battles and a scene language model in a transaction scene that pays more attention to words related to transactions. The scene language model can more specifically analyze the natural language command received in the virtual scene.

[0287] For example, in the process of user controlling a non-player character, a natural language command is received, based on the game state, object position and current task completion, it is determined that a first virtual scene in which a target virtual character (at least one of a master virtual character and a non-player character) is currently located is a virtual battle scene, a battle scene language model corresponding to the virtual battle scene is obtained, and based on a field of view range of the target virtual character, first scene hot words in the virtual battle scene are obtained, including "truck", "virtual grass", "virtual house", "enter", "open truck", "defend", "desert", "attack", "oak tree", "stable", "firearm", etc., the natural language command is decoded and processed by a pre-trained natural language analysis model (including an acoustic network, a language network, a preset dictionary, etc.) under the constraint of the first scene hot words, and an instruction analysis result is output to accurately control the non-player character. The instruction analysis result is illustrated as follows.

[0288] (1) The instruction analysis result is "No. 1, go to the front and defend", which can be used to control No. 1 to move to the front and make a defensive posture, avoiding the enemy virtual character attacking the master virtual character first; without the constraint of the first scene hot words, the natural language command is easily identified as "No. 1, go to the front and make a sound", which affects the virtual battle plan of the player.

[0289] (2) The instruction analysis result is "No. 2, go to the desert and explore", which can be used to control No. 2 to move to the desert position to explore whether there is an enemy virtual character or a virtual treasure chest; without the constraint of the first scene hot words, the natural language command is easily identified as "No. 2, go to the mountain and explore", which makes No. 2 move to the wrong position, and the human-computer interaction efficiency is low.

[0290] (3) The instruction analysis result is "No. 1 and No. 2, attack me", which can be used to control No. 1 and No. 2 to virtually attack the nearby virtual characters that can be attacked; without the constraint of the first scene hot words, the natural language command is easily identified as "No. 1 and No. 2, supply me", which makes No. 1 and No. 2 make an error behavior that does not meet the player's expectation, which cannot protect the master virtual character, and easily makes the enemy virtual character win the virtual game.

[0291] (4) The instruction analysis result is "No. 1, find the nearby oak tree", which can be used to control No. 1 to find the oak tree in the nearby area, so as to complete the game task or find the oak tree that can be used to avoid attack; without the constraint of the first scene hot words, the natural language command is easily identified as "No. 1, find the nearby project book", which makes No. 1 search for a virtual element that does not meet the player's expectation, and cannot meet the virtual battle demand of the player.

[0292] (5) The instruction analysis result is "1 and 2, go to the stable over there", which can be used to control No. 1 and No. 2 to search for a stable nearby and move to the location of the stable. Without the constraint of the first scene hotword, the natural language command is easily identified as "1 and 2, go there, it will be done soon", so that No. 1 and No. 2 mistakenly think that the host virtual character currently wants to complete the virtual game by himself, and thus the host virtual character cannot be provided with better game assistance.

[0293] (6) The instruction analysis result is "No. 2, pick up the weapon in front", which can be used to control No. 2 to search for a weapon in front and perform a pick-up action on the weapon. Without the constraint of the first scene hotword, the natural language command is easily identified as "No. 2, pick up the crab in front", so that No. 2 cannot accurately find the virtual element that the host virtual character wants to pick up, which may either prompt the player that "the crab cannot be found" or pick up the wrong virtual element "crab" for the player, thereby interfering with the virtual combat process of the player.

[0294] Environmental sound effect function:

[0295] In addition, for the environmental sound effect of the entire virtual scene, the scheme also provides a spatial audio enhancement scheme.

[0296] The audio played by the terminal includes environmental audio and NPC audio. The environmental audio is audio generated based on the scene characteristics of the virtual scene currently occupied by the host virtual object, so as to give the user a sense of being there. The NPC audio is audio generated based on the characteristics of the NPC, so that the user can intuitively feel the emotion and physical state of the character through the sound heard.

[0297] Environmental audio: identify the scene elements in the virtual scene where the host virtual object is located, generate / choose appropriate element sound effects from a sound effect library in real time according to the scene elements, and synthesize the element sound effects to obtain the environmental audio. For example, the virtual scene is a forest at night, and the scene elements include trees, owls, insects, etc. The environmental audio includes rustling of leaves, owl calls, insect calls, etc.

[0298] NPC audio: determine the NPC contained in the virtual scene, generate the corresponding character voice based on the character type, current emotion and current behavior of the NPC. For example, when the NPC is a middle-aged man who is running, the generated character voice is a deep male voice with a breathing sound effect when running.

[0299] When the terminal plays the environmental audio and / or NPC audio to the user, the terminal performs audio enhancement processing on the environmental audio and / or NPC audio to improve the realism of the audio.

[0300] FIG. 2 shows a structural block diagram of a computer system according to an example embodiment of the present application. The computer system 100 includes a first terminal 110, a server 120, and a second terminal 130.

[0301] The first terminal 110 is installed and runs a client 111 supporting a virtual environment, which can be a multiplayer online battle program. When the first terminal runs the client 111, a user interface of the client 111 is displayed on a screen of the first terminal 110. The client 111 can be any one of a battle royale shooting game, a virtual reality (VR) application, an augmented reality (AR) program, a three-dimensional map program, a virtual reality game, an augmented reality game, a casual game, a party game, a first-person shooting game (FPS), a third-person shooting game (TPS), a multiplayer online battle arena game (MOBA), and a simulation game (SLG). In this embodiment, the client 111 is taken as an example of an FPS game. The first terminal 110 is a terminal used by a first user 112, and the first user 112 uses the first terminal 110 to control a first virtual object in a virtual environment, which can be referred to as a virtual object of the first user 112. The activities of the first virtual object include, but are not limited to, at least one of moving, jumping, teleporting, releasing a skill, using a prop, adjusting a body posture, crawling, walking, running, riding, flying, jumping, driving, picking up, shooting, attacking, and throwing. Illustratively, the first virtual object is a first virtual object, such as a simulated character role or an animation character role.

[0302] The second terminal 130 installs and runs a client 131 supporting a virtual environment, which can be a multiplayer online battle program. When the second terminal 130 runs the client 131, a user interface of the client 131 is displayed on a screen of the second terminal 130. The client can be any one of a battle royale shooting game, a VR application, an AR program, a three-dimensional map program, a virtual reality game, an augmented reality game, an FPS, a TPS, a MOBA, and a SLG, and in this embodiment, the client is taken as an example of a MOBA game. The second terminal 130 is a terminal used by a second user 132, and the second user 132 uses the second terminal 130 to control a second virtual object located in a virtual environment for activities, which can be referred to as a virtual object of the second user 132. Illustratively, the second virtual object is a virtual object, such as a simulated character role or an animation character role.

[0303] Optionally, the first virtual object and the second virtual object are in the same virtual environment. Optionally, the first virtual object and the second virtual object can belong to the same camp, the same team, the same organization, have a friendship relationship, or have temporary communication authority. Optionally, the first virtual object and the second virtual object can belong to different camps, different teams, different organizations, or have an enemy relationship.

[0304] Optionally, the clients installed on the first terminal 110 and the second terminal 130 are the same, or the clients installed on the two terminals are the same type of clients on different operating system platforms (Android or IOS). The first terminal 110 can be taken as one of a plurality of terminals, and the second terminal 130 can be taken as another of the plurality of terminals, and in this embodiment, the first terminal 110 and the second terminal 130 are taken as examples. The device types of the first terminal 110 and the second terminal 130 are the same or different, and the device types include at least one of a smart phone, a tablet computer, an e-book reader, an MP3 player, an MP4 player, a laptop computer, and a desktop computer.

[0305] Only two terminals are shown in FIG. 2, but in different embodiments, a plurality of other terminals 140 can access the server 120. Optionally, one or more terminals 140 are developer corresponding terminals, and a development and editing platform of the client supporting the virtual environment is installed on the terminal 140. A developer can edit and update the client on the terminal 140, and transmit an updated client installation package to the server 120 through a wired or wireless network. The first terminal 110 and the second terminal 130 can download the client installation package from the server 120 to update the client.

[0306] The first terminal 110, the second terminal 130, and the other terminals 140 are connected to the server 120 through a wireless network or a wired network.

[0307] The server 120 comprises at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. The server 120 is configured to provide background services for clients supporting a three-dimensional virtual environment. Optionally, the server 120 undertakes major computing work, and the terminal undertakes secondary computing work; or the server 120 undertakes secondary computing work, and the terminal undertakes major computing work; or the server 120 and the terminal adopt a distributed computing architecture for collaborative computing.

[0308] In an illustrative example, the server 120 comprises a processor 122, a user account database 123, a battle service module 124, and a user-oriented input / output interface (I / O interface) 125. The processor 122 is configured to load instructions stored in the server 120 and process data in the user account database 123 and the battle service module 124. The user account database 123 is configured to store data of user accounts used by the first terminal 110, the second terminal 130, and other terminals 140, such as avatars of the user accounts, nicknames of the user accounts, battle power indexes of the user accounts, and service areas where the user accounts are located. The battle service module 124 is configured to provide multiple battle rooms for users to battle, such as 1V1 battles, 3V3 battles, 5V5 battles, and the like. The user-oriented I / O interface 125 is configured to establish communication with the first terminal 110 and / or the second terminal 130 through a wireless network or a wired network and exchange data.

[0309] The method provided in the present application can be applied to at least one of the following scenarios, but is not limited to: a virtual reality application program, a three-dimensional map program, a casual game, a party gathering game, a first-person shooting game (FPS), a third-person shooting game (TPS), a multiplayer online battle arena game (MOBA), a multiplayer gun battle survival game, and the like. The following embodiments are exemplarily described in the application of a game. The method provided in the present application can be completely executed by a client, such as a single-player game; can be completely executed by a server, such as a cloud game in which the server performs all calculations and the client is only responsible for displaying and collecting user operation behaviors; or can be cooperated by a client and a server, in which the client is responsible for a part and the server is responsible for a part.

[0310] FIG. 3 shows a flowchart of a virtual role control method according to an example embodiment of the present application. The method can be executed by a computer device. The method comprises:

[0311] Step 510: display a master virtual character or NPC in the virtual environment;

[0312] Illustratively, the master virtual character is a virtual character directly controlled by the user in the virtual environment. There is one or more non-player characters (NPC) in the virtual environment; the NPC and the master virtual character belong to the same virtual camp and are teammates of the master virtual character.

[0313] Step 520: obtain command information;

[0314] Illustratively, the command information includes natural language semantics for controlling the NPC; the natural language semantics for controlling the NPC is obtained by inputting a command text in the virtual scene to command the non-player character, and the command text can be directly input by the user or extracted from voice information input by the user, and the application does not limit the control mode of the command text. Illustratively, the command information includes natural language semantics for controlling the NPC, also known as natural language commands or voice form natural language commands.

[0315] Illustratively, in addition to the master virtual character controlled by the user in the virtual environment, the user also has one or more non-player characters (NPC) as teammates. The user has control over the above-mentioned non-player characters, and through natural language form command text, instructs the non-player characters to make their own desired feedback to the scene objects in the virtual environment, so as to instruct the non-player characters to complete the task in cooperation with the master virtual character.

[0316] The command information corresponds to at least two semantic understanding modes; the at least two semantic understanding modes correspond to different control modes of the NPC in the virtual environment. Illustratively, at least one of the scene entities, word pronunciation, and behavior intention indicated in the at least two semantic understanding modes is different.

[0317] Illustratively, in the case where there are at least two semantic understanding modes for the word pronunciation of the command information, there are at least two semantic understanding modes for the command information itself from the dimension of natural language. In the case where there are at least two semantic understanding modes for the scene entities / behavior intention of the command information, the command information is combined with the virtual environment, and there are at least two semantic understanding modes for controlling the NPC in the virtual environment.

[0318] Step 530: in response to the command information, controlling the NPC to perform a virtual activity according to a first semantic understanding mode of the at least two semantic understanding modes;

[0319] Exemplarily, the first semantic understanding manner is one of the at least two semantic understanding manners corresponding to the command information. The first semantic understanding manner is determined based on environment perception information of the host virtual character and / or the NPC, from the at least two semantic understanding manners. In some examples, the host virtual character and / or the NPC obtains the environment perception information within a perception range of a visual, an auditory, a movement trajectory, or the like. Based on the environment perception information, the first semantic understanding manner determined from the at least two semantic understanding manners is close to a current perception situation of the host virtual character and / or the NPC to the virtual environment, and the NPC is controlled with reference to the current perception situation of the host virtual character and / or the NPC to the virtual environment.

[0320] Exemplarily, since the first semantic understanding manner is determined based on the environment perception information of the host virtual character and / or the NPC, from the at least two semantic understanding manners, and there are at least two situations in the virtual environment (at least one of positions, virtual life values, virtual prop holding amounts, virtual skill cooling times, virtual levels, or the like of the host virtual character and / or the NPC in the above two situations are different), the first semantic understanding manner determined is different in the face of the same command information. That is to say, as the situation in the virtual environment changes, the NPC performs different virtual activities in response to the same command information.

[0321] In summary, the method provided by the embodiment determines the virtual activity performed according to the first semantic understanding manner by taking the environment perception information of the host virtual character and / or the NPC as a reference from the at least two semantic understanding manners of the command information. In the case that there are at least two semantic understanding manners of the command information, on the one hand, the user does not need to supplement more basis for the command information, and the NPC can be controlled to perform the virtual activity according to the command information, improving the human-computer interaction efficiency; on the other hand, taking the environment perception information of the host virtual character and / or the NPC as a reference ensures that the NPC and the host virtual character cooperatively perform the virtual activity.

[0322] Next, the at least two semantic understanding manners corresponding to the command information are introduced. In one implementation manner, the command information corresponds to at least two semantic understanding manners of scene entities; in another implementation manner, the command information corresponds to at least two semantic understanding manners of word pronunciations; and in yet another implementation manner, the command information corresponds to at least two semantic understanding manners of behavior intentions.

[0323] The above three implementation manners will be introduced in the following with separate embodiments.

[0324] The command information corresponds to at least two semantic understanding manners of scene entities.

[0325] FIG. 4 shows a flowchart of a method for controlling a virtual role according to an example embodiment of the present application. The method can be performed by a computer device. That is, in the embodiment shown in FIG. 3, step 530 can be implemented as step 532:

[0326] Step 532: in response to the command information, controlling the NPC to perform a virtual activity on a first scene object located within a perception range of the host virtual role and / or the NPC;

[0327] By way of example, the perception range is a virtual area in which the host virtual role and / or the NPC can obtain environmental perception information. The perception range can be a connected area in the virtual environment, or a plurality of independent areas that are not connected to each other, and the present application does not limit the shape of the perception range. In the perception range, the host virtual role and / or the NPC can obtain information in the virtual environment through visual, auditory, action trajectory, and other perception methods.

[0328] In the present embodiment, there are a plurality of scene objects of the same type in the virtual environment, and the at least two semantic understanding manners corresponding to the command information are used to indicate that the virtual activity is performed on the plurality of scene objects of the same type; that is, there are at least two semantic understanding manners of the scene entity. The plurality of scene objects include the first scene object. By way of example, since the command text does not explicitly indicate which scene object of the plurality of scene objects the virtual activity is performed on, the command text corresponds to at least two semantic understanding manners.

[0329] By way of example, taking a plurality of scene objects as virtual vehicles in the virtual environment, there are a plurality of virtual vehicles in the virtual environment, which are distributed at various positions in the virtual environment. The command information is used to control the NPC to move to a position where a virtual vehicle is located. Since the command information does not explicitly indicate which position of a virtual vehicle to move to, there are a plurality of semantic understanding manners, that is, moving to any position of a virtual vehicle satisfies the natural language semantics of the command information. In an implementation manner provided by the present application, among the plurality of semantic understanding manners, a first semantic understanding manner is determined according to the perception range, and the NPC is controlled to move to a position where a virtual vehicle in the perception range is located according to the first semantic understanding manner, so as to determine one semantic understanding manner among the plurality of semantic understanding manners with the perception range as a reference, and to control the NPC to perform a virtual activity on the first scene object; on the one hand, the user does not need to supplement more basis for the command information, thereby improving the human-computer interaction efficiency; on the other hand, the NPC will not perform a virtual activity on a scene object outside the perception range of the NPC or the host virtual role, thereby causing the NPC to deviate from the cooperation with the host virtual role or to enter a range (i.e., a strange range / dangerous range) for which no perception information is obtained; and the NPC and the host virtual role can cooperatively perform a virtual activity.

[0330] In an optional implementation, the NPC performs the virtual activity on the first scene object in the field of view range; correspondingly, the step can be implemented as:

[0331] · in response to the command information, controlling the NPC to perform the virtual activity on the first scene object in the field of view range;

[0332] For example, the field of view range is the observation range of the virtual environment by the master virtual character and / or the NPC. For example, the field of view range can be the visual perception range of the master virtual character and / or the NPC itself, or the visual perception range of the perception skill or the perception device used.

[0333] The first semantic understanding mode is used to instruct the NPC to perform the virtual activity on the first scene object in the field of view range, fully considering that visual perception is one of the perception ways to obtain information in the virtual environment, and is specially designed for controlling the NPC in a natural semantic way, fully considering that the user directly controls the master virtual character in the virtual environment; fully considering the will of controlling the NPC to perform the virtual activity, which is generated after the master virtual character observes the scene object in the virtual environment, such as observing a virtual vehicle in the virtual environment, generating a will to control the NPC to move to realize the virtual vehicle as a cover to mask the NPC and help the NPC avoid being discovered by the enemy virtual character. Only when the master virtual character observes the virtual vehicle in the virtual environment, the will to use the virtual vehicle as a cover is generated. The first semantic understanding mode determined according to the field of view range is used to instruct the NPC to perform the virtual activity on the first scene object in the field of view range, which limits the scene object of the virtual activity to the field of view range, and can combine the command text in the natural semantic way and the current visual observation of the virtual environment to flexibly control the virtual character.

[0334] Optionally, the field of view range is the field of view range observed by the master virtual character and / or the NPC, and is the range of the observation picture obtained by the virtual camera model following the movement of the master virtual character and / or the NPC; and / or, the field of view range is an extended field of view range obtained by using a perception skill or a perception device, and the user of the perception skill (such as an extended field of view skill) or the perception device (such as a virtual telescope, a virtual drone, etc.) includes the master virtual character and / or the NPC.

[0335] In an optional implementation, the NPC performs the virtual activity on the first scene object in the first virtual terrain; correspondingly, the step can be implemented as:

[0336] · in response to the command information, controlling the NPC to perform the virtual activity on the first scene object in the first virtual terrain;

[0337] Exemplarily, the first virtual terrain is a virtual terrain to which a location where the host virtual character and / or the NPC is located belongs; and exemplarily, the first virtual terrain is a connected region including the location where the host virtual character and / or the NPC is located.

[0338] Next, the case where the first virtual terrain is an indoor terrain, a street terrain, and an outdoor terrain is introduced respectively.

[0339] In the case where the first virtual terrain is an indoor terrain, the NPC is controlled to perform the virtual activity on the first scene object located in the first region in response to the command information.

[0340] In the case where the first virtual terrain is an indoor terrain, the first location where the host virtual character and / or the NPC is located belongs to the indoor terrain; and the first region is a connected region of the indoor terrain including the first location where the host virtual character and / or the NPC is located.

[0341] After the host virtual character and / or the NPC enters the indoor region, the virtual activity is performed on the first scene object in the indoor region (i.e., the first region). It is fully considered that the indoor region in the virtual environment is a region rich in virtual resources and a region providing shelter for the virtual character; even if the indoor region is blocked by walls, there are many indoor objects, etc., the NPC can perform the virtual activity in the indoor region, and is not limited to one virtual room where the host virtual character and / or the NPC is located.

[0342] In the case where the first virtual terrain is a street terrain adjacent to a building, the NPC is controlled to perform the virtual activity on the first scene object located in the second region in response to the command information.

[0343] In the case where the first virtual terrain is a street terrain, the first location where the host virtual character and / or the NPC is located is close to the building; and the second region includes the building around the second location where the host virtual character and / or the NPC is located.

[0344] In the case where the location where the host virtual character and / or the NPC is located is adjacent to the building, the virtual activity performed by the NPC can be performed on the building. It is fully considered that the indoor region in the virtual environment is a region rich in virtual resources and a region providing shelter for the virtual character; in the case where the location where the host virtual character and / or the NPC is located is adjacent to the building, the NPC can approach or even enter the building to obtain virtual resources or get shelter of the building.

[0345] In the case where the first virtual terrain is an outdoor terrain away from the building, the NPC is controlled to perform the virtual activity on the first scene object located in the third region in response to the command information.

[0346] In a case where the first virtual terrain is an outdoor terrain, the first position where the host virtual character and / or the NPC is located belongs to the outdoor terrain; and the third region is a region having the same virtual vegetation distribution as the third position where the host virtual character and / or the NPC is located.

[0347] In a case where the host virtual character and / or the NPC is located in an outdoor terrain, the virtual activity performed by the NPC can be performed on a first scene object in a region having the same virtual vegetation distribution as the third position where the host virtual character and / or the NPC is located. The virtual terrain in which the virtual activity is performed by the NPC has the same virtual vegetation distribution as the third position, so that the NPC can be hidden by the virtual vegetation, and the hiding effect of the virtual vegetation obtained by the NPC in the process of performing the virtual activity is the same as the hiding effect of the host virtual character and / or the hiding effect of the NPC before performing the virtual activity.

[0348] In summary, the method provided in the embodiment can determine, in at least two semantic understanding manners of the command information, a virtual activity performed on a first scene object in a perception range of the host virtual character and / or the NPC, with reference to the perception range. In a case where there are at least two semantic understanding manners of the command information, on the one hand, the user does not need to supplement more information for the command information, and the NPC can perform the virtual activity according to the command information, improving the human-computer interaction efficiency; on the other hand, the NPC will not perform the virtual activity on scene objects outside the perception range of the host virtual character or the NPC, ensuring that the NPC and the host virtual character perform the virtual activity cooperatively.

[0349] The command information corresponds to at least two semantic understanding manners of word pronunciation.

[0350] FIG. 5 shows a flowchart of a virtual character control method provided in an example embodiment of the present application. The method can be executed by a computer device. That is, in the embodiment shown in FIG. 3, step 530 can be implemented as steps 534 and 535:

[0351] Step 534: in response to the command information, converting the command information into command text according to a natural language pronunciation of a scene object in the perception range;

[0352] For example, the perception range is a virtual region in which the host virtual character and / or the NPC obtains environmental perception information; the perception range can be one connected region in the virtual environment, or multiple independent regions that are not connected to each other, and the shape of the perception range is not limited in the present application. In the perception range, the host virtual character and / or the NPC can obtain information in the virtual environment through visual, auditory, and action trajectory perception modes.

[0353] In this embodiment, the command information is voice information. Due to the richness of natural language, the command information corresponds to at least two semantic understanding manners of word pronunciation. For example, in the voice expression of Chinese, the pronunciation "qiangxie" can be the pronunciation of a virtual shooting prop "gun", or the pronunciation of a food "choking crab" (a kind of sashimi food of crab). It can be understood that due to the above-mentioned situation of the same or similar pronunciation but different meanings in natural language, the command information corresponds to at least two semantic understanding manners of word pronunciation.

[0354] In an implementation provided in the present application, in the plurality of semantic understanding manners, a first semantic understanding manner is determined according to the perception range, and according to the first semantic understanding manner, the command information is converted into command text according to the natural language pronunciation of the scene object in the perception range. It is realized that in the plurality of semantic understanding manners of word pronunciation, the command text is converted according to the natural language pronunciation of the scene object in the perception range; accordingly, the command text includes the name text and / or introduction text of the scene object.

[0355] On the one hand, based on the natural language pronunciation of the scene object, the complex artificial neural network for predicting the semantics in the voice information is avoided, and the calculation complexity of converting the command information into the command text is reduced. On the other hand, it is fully considered that the command information is often highly accompanied by information in the virtual environment when controlling the NPC, that is, the voice control of the NPC needs to indicate the purpose of the virtual activity, such as picking up a virtual gun, and the virtual gun is a scene object appearing in the virtual environment. The scene object in the perception range is taken as a reference to flexibly control the NPC.

[0356] In an optional implementation, the present step can be implemented as:

[0357] · in the case that the similarity between the pronunciation of the first phoneme in the command information and the natural language pronunciation of the scene object exceeds a similarity threshold, and the first phoneme has a candidate text, the text information of the first phoneme is determined as the name text and / or introduction text of the scene object, and the command text is constructed according to the text information;

[0358] For example, in the case that the similarity between the pronunciation of the first phoneme in the command information and the natural language pronunciation of the scene object exceeds a similarity threshold, the first phoneme in the command information is similar to the pronunciation of the name text and / or introduction text of the scene object. The first phoneme has a candidate text, which is used to indicate that another candidate text has a pronunciation similar to the first phoneme. The candidate text of the first phoneme does not include the name text and / or introduction text of the scene object.

[0359] In one example, the virtual environment is a shooting game, the similarity between the first phoneme "qiangxie" in the command information and the natural language pronunciation of the scene object virtual gun exceeds a similarity threshold, and the first phoneme has a candidate text "chacai". The text information of the first phoneme is determined as the name text "gun" of the scene object.

[0360] Step 535: According to the command text, the NPC performs a virtual activity according to the first semantic understanding mode.

[0361] For example, the command text is converted according to the first semantic understanding mode, and accordingly, the NPC performs a virtual activity according to the first semantic understanding mode.

[0362] In summary, the method provided in the embodiment determines the command text of the command information according to the natural language pronunciation of the scene object in the perception range in at least two semantic understanding modes of the command information. In the case where the command information has at least two semantic understanding modes, on the one hand, the user does not need to supplement more basis for the command information, and can control the NPC to perform a virtual activity according to the command information, improving the human-computer interaction efficiency; on the other hand, the method fully considers the case that the command information is often highly accompanied by information in the virtual environment when controlling the NPC, and realizes flexible control of the NPC.

[0363] The command information corresponds to at least two semantic understanding modes of the behavior intention.

[0364] FIG. 6 shows a flowchart of a virtual character control method according to one example embodiment of the present application. The method can be executed by a computer device. That is, in the embodiment shown in FIG. 3, step 530 can be implemented as step 536:

[0365] Step 536: In response to the command information, the NPC performs a virtual activity according to the first behavior intention.

[0366] For example, the perception range is a virtual area in which the host virtual character and / or the NPC obtain environmental perception information. The perception range can be one connected area in the virtual environment, or a plurality of independent areas that are not connected to each other, and the shape of the perception range is not limited in the present application. In the perception range, the host virtual character and / or the NPC can obtain information in the virtual environment through visual, auditory, and action trajectory perception modes.

[0367] In the embodiment, due to the richness of natural language, the command information corresponds to at least two semantic understanding manners of behavior intention. For example, the command information includes walking forward. Since the command information does not explicitly indicate the reference system under which the forward movement is performed, the understanding manner of the behavior intention can be walking forward in the direction in which the NPC faces, or walking forward to the marked position point as the destination. Illustratively, the first behavior intention determined in the semantic understanding of the multiple behavior intentions according to the perception range corresponds to the virtual activity type permitted to be performed in the perception range, so as to ensure that the virtual activity can be performed in the virtual environment.

[0368] In an optional implementation, the step can be implemented as:

[0369] · in response to the command information, taking a first direction in the perception range as a reference direction, and according to a direction word in the command information, controlling the NPC to perform the virtual activity based on a second direction;

[0370] Illustratively, the first direction includes, but is not limited to, the direction in which the virtual character or the NPC faces. The command information includes a direction word, which is used to describe the angle between the second direction and the first direction, and / or describe the second direction in the reference system constructed based on the first direction.

[0371] Illustratively, taking the first direction as the reference direction is to determine the first direction as the front direction; and the reference system constructed based on the first direction is a plane coordinate system and / or a three-dimensional coordinate system taking the first direction as the positive half-axis direction of x.

[0372] In summary, the method provided in the embodiment performs the virtual activity according to the virtual activity type permitted to be performed in the perception range in at least two semantic understanding manners of the command information, so as to ensure that the virtual activity can be successfully performed in the virtual environment. The first direction in the perception range and the direction word included in the command information are used to determine the virtual activity performed based on the second direction; without the user supplementing more basis for the command information, the NPC can perform the virtual activity according to the command information, the human-computer interaction efficiency is improved, and the problem that the direction of the virtual activity cannot be determined according to the command information when the command information corresponds to at least two semantic understanding manners of behavior intention is avoided.

[0373] FIG. 7 shows a flowchart of a virtual character control method provided in an example embodiment of the present application. The method can be performed by a computer device. That is, in the embodiment shown in FIG. 3, the step 530 can be implemented as steps 542 to 550:

[0374] Step 542: determining the command text in natural language form according to the command information;

[0375] Exemplarily, the command information is subjected to automatic speech recognition (ASR) processing to determine a command text in natural language form. The automatic speech recognition processing generally includes invoking acoustic models, language models and the like components to recognize the pronunciation, vocabulary and syntax structure and the like information of the command information and convert the information into the command text in text format.

[0376] Exemplarily, the command information is obtained in response to a triggering operation of the button-type control. The command information is recorded in response to the triggering operation of the button-type control. Exemplarily, the command information is obtained when the user controls the host virtual character to perform the virtual activity in the virtual environment. The microphone recording function is in an open state during the process in which the user controls the host virtual character to perform the virtual activity in the virtual environment.

[0377] Step 544: invoking a conditional random field to perform entity prediction on the command text to obtain an entity label carried by the command text.

[0378] Exemplarily, the conditional random field (CRF) is a statistical model that processes the complex dependency relationship between the input and the output by defining a feature function and a corresponding parameter weight. The conditional random field can extract the corresponding entity label from a plurality of expressions in natural semantics. The conditional random field can obtain the context information of the input command text sequence, calculate a conditional probability distribution, and obtain the entity label carried by the command text according to the conditional probability distribution.

[0379] Step 546: determining a first scene object in the perception range according to the entity label.

[0380] Exemplarily, the perception range is a virtual area in which the host virtual character and / or the NPC obtains environmental perception information. The perception range can be one connected area in the virtual environment, or a plurality of independent areas that are not connected to each other. The shape of the perception range is not limited in the present application. In the perception range, the host virtual character and / or the NPC can obtain information in the virtual environment through visual, auditory, action trajectory and the like perception manners.

[0381] The first scene object is determined from the virtual environment in combination with the environmental perception information. The environmental perception information includes information perceived by at least one of the host virtual character and the non-player character from the perception range in the virtual environment.

[0382] The virtual environment includes a plurality of candidate entities matching the entity label carried by the command text, and the first scene object is an entity selected from the plurality of candidate entities based on the environment perception information of the host virtual character or the non-player character. For example, the game application first selects a plurality of candidate entities matching the entity label carried by the command text from the entity set, and then selects an entity that can be perceived by the host virtual character or the non-player character as the first scene object according to the environment perception information. For example, the environment perception information includes at least one of the following: visual information perceived within the visual range of the host virtual character; auditory information perceived within the auditory range of the host virtual character; information perceived by the perception skill or the perception device possessed by the host virtual character; visual information perceived within the visual range of the non-player character; auditory information perceived within the auditory range of the non-player character; information perceived by the perception skill or the perception device possessed by the non-player character.

[0383] Step 548: calling the prediction network to perform intent recognition on the command text to obtain a classification label of the command text;

[0384] The prediction network has the ability to predict the classification label corresponding to the command text. The prediction network is used to predict the control mode of the command text on the NPC in at least one behavior dimension;

[0385] Optionally, a hierarchical prediction network is called to perform intent recognition on the command text to obtain a classification label of the command text;

[0386] For example, the hierarchical prediction network includes at least two sub-networks, the upper network and the lower network are cascaded, and the lower network further performs classification label prediction according to the prediction result output by the upper network. The at least two sub-networks are constructed based on the tree structure of a plurality of classification labels, and the hierarchical prediction network constructs at least two sub-networks with corresponding hierarchical structure for the tree structure of the classification labels, and divides the prediction task of a variety of classification labels into prediction sub-tasks performed by at least two sub-networks, so as to reduce the classification prediction complexity of each sub-network.

[0387] Step 550: determining a control command of the NPC according to the classification label, and controlling the NPC to perform a virtual activity on the first scene object according to the first semantic understanding mode based on the control command;

[0388] Taking the classification label of the command text as input, the target entity in the virtual space and the control command of the NPC are determined; for example, each classification label in the command text has a corresponding control command, and the control command of the NPC is determined according to the corresponding relationship between the classification label and the control command.

[0389] The control command indicates a subsequent action to be performed by the NPC, which can be a single action or a sequence of actions. In the case of a multi-action sequence instruction, the multi-action sequence instruction is used to instruct the NPC to perform a sequence of actions, and the data structure for storing the multi-action sequence instruction is in the form of a list. The instructions to be executed in the sequence instruction are cached. When the behavior tree of the NPC executes the current instruction, the cache is called back, and the instructions to be executed in the cache are continued to be issued and executed.

[0390] To sum up, the method provided in the embodiment, by taking the environment perception information of the host virtual character and / or the NPC as a reference, determines to perform the virtual activity according to the first semantic understanding manner; in the case that there are at least two semantic understanding manners for the command information, on the one hand, the user does not need to supplement more basis for the command information, and the NPC can perform the virtual activity according to the command information, improving the human-computer interaction efficiency; on the other hand, taking the environment perception information of the host virtual character and / or the NPC as a reference, ensures that the NPC and the host virtual character perform the virtual activity cooperatively.

[0391] In an optional implementation, before step 546, the method further includes:

[0392] obtaining spatial positions of each scene object in the virtual environment;

[0393] For example, the spatial position of the scene object in the virtual scene indicates the deployment of the scene object in the virtual scene, and the spatial information indicates the position, size, etc. of the scene object in the virtual scene after being deployed in the virtual scene.

[0394] Optionally, the spatial position is introduced as follows.

[0395] obtaining at least one of the coordinate position, the orientation information, the bounding box information, and the cover point information of each scene object in the virtual environment;

[0396] The coordinate position (Location) indicates the position of the scene object in the virtual scene, such as the coordinate information of the center point or the preset point of the scene object in the virtual scene; the orientation (Rotation) indicates the direction faced by the scene object in the virtual scene, such as the direction faced by the front of the scene object in the virtual scene; the bounding box (Bounding Box) indicates the size of the scene object in the virtual scene; and the cover point (Cover points) indicates the recommended position of the virtual character when the virtual character approaches the scene object, so that the scene object can mask the virtual character.

[0397] • according to the spatial position, screen out a plurality of candidate scene objects in the perception range from the various scene objects in the virtual environment;

[0398] For example, the spatial position of the scene object in the virtual scene is used to indicate the deployment of the scene object in the virtual scene, and the spatial information is used to indicate the position, size, etc. of the scene object in the virtual scene after the scene object is deployed in the virtual scene.

[0399] For example, the perception range is a virtual area in which the host virtual character and / or NPC obtain environmental perception information. The perception range can be a connected area in the virtual environment, or a plurality of independent areas that are not connected to each other, and the application does not limit the shape of the perception range. In the perception range, the host virtual character and / or NPC can obtain information in the virtual environment through visual, auditory, and action trajectory perception methods.

[0400] According to the spatial position, a plurality of scene objects in the virtual environment located in the perception range are determined, and the plurality of scene objects located in the perception range are determined as a plurality of candidate scene objects.

[0401] • obtain the visual label of the plurality of candidate scene objects in the perception range;

[0402] The visual label is used to describe the scene object in at least one dimension. The description dimension of the visual label on the scene object includes but is not limited to at least one of the type description, material, transparency, color, surface feature, and shape of the scene object.

[0403] Correspondingly, step 546 can be implemented as:

[0404] • respectively calculate the similarity between the visual label of the plurality of candidate scene objects and the entity label, and determine a first scene object with the highest similarity from the plurality of candidate scene objects.

[0405] For example, the entity label is used to describe the command information indicating which scene object to execute the virtual activity. For example, part of the words in the command information correspond to the entity label.

[0406] Based on the similarity between the entity label and the visual label of the plurality of candidate scene objects, the first scene object corresponding to the command information is screened out from the plurality of candidate scene objects in the perception range. The command information is used to command the non-player character in the virtual scene; the command information not only has language to introduce that the virtual activity is an activity for which virtual scene object (for example, the command information indicates to move to the virtual building, that is, the virtual activity of moving is executed for the virtual building), but also has words indicating the identity of the non-player character executing the virtual activity.

[0407] Fig. 8 shows a schematic diagram of command information according to an example embodiment of the present application. The input words included in the command information 600 include “blue” and “car”; the virtual scene includes two scene objects; it can be understood that a larger number of scene objects can be included in the virtual scene in different embodiments. The visual label of the scene object A 601 includes metal, old, car, truck, blue, and broken; the visual label of the scene object B 602 includes metal, scratch, car, broken, cash truck, and black. The similarity scores between the input words included in the command information 600 and the visual labels of the scene objects are calculated respectively; the plurality of similarity scores correspond to the plurality of input words one by one; for the input word “blue” in the command information 600, the similarity score with the scene object A 601 is 1.0; for the input word “car” in the command information 600, the similarity score with the scene object A 601 is 1.0; the similarity between the scene object A 601 and the command information 600 is 1.0+1.0=2.0. Similarly, the similarity between the scene object B 602 and the command information 600 is calculated to be 1.0+0.91=1.91. The similarity between the scene object A 601 and the command information 600 is greater than the similarity between the scene object B 602 and the command information 600, and the scene object A 601 is determined as the first scene object.

[0408] Next, the process of obtaining the visual label is introduced. Taking a scene object in a virtual scene as an example:

[0409] · Obtain the attribute text of the scene object in the virtual scene;

[0410] For example, the attribute text is used to introduce the inherent attributes of the scene object in the virtual scene; on the one hand, the attribute text realizes the description of the scene object in the text mode, and at the same time provides semantic information for the prediction of the visual label of the scene object. On the other hand, the attribute text is a description of the scene object in the virtual scene, in the face of a large number of scene objects in the virtual scene, in the case of reuse of object models, it can more accurately describe the inherent attributes of the scene object in the virtual scene. For example, in order to avoid redesigning the object model of the scene object in the virtual environment, at least one of the transformation operations such as scaling, stretching, and rotating the object model of the virtual table is performed, and the transformed object model is deployed in the virtual environment as a virtual bench; in the face of the above model reuse, the attribute text can describe the inherent characteristics of the scene object in the virtual scene.

[0411] In an optional implementation, the attribute text includes at least one of a name of the scene object in the virtual scene and a size of the scene object in the virtual scene; for example, the size of the scene object in the virtual scene is used to indicate virtual space occupied by the scene object in the virtual scene, and to avoid the case of reusing the object model, the size of the object model is changed, and the influence on the description of the scene object. The name in the virtual scene can reflect the use of the scene object in the virtual scene, and to avoid the case of reusing the object model, when deploying two different scene objects based on the same object model, the influence on the description of the scene object.

[0412] • Obtain an appearance image of the scene object in the virtual scene;

[0413] For example, the appearance image is used to describe the style of the scene object; based on the richness of natural language, the player can describe the scene object in the virtual scene from multiple semantic angles; when constructing the visual label of the scene object, comprehensive description information of the scene object needs to be obtained, the appearance image carries the appearance style of the scene object such as color, texture, shape, and the mutual position relationship between each part on the picture modal, and can comprehensively describe the scene object from the picture modal (or visual modal).

[0414] In an optional implementation, the appearance image of the scene object includes images obtained by observing the scene object from at least two perspectives. The observation of the scene object from at least two perspectives in the appearance image can avoid the problem that the spatial structure of the scene object is blocked and cannot present all appearance styles in the scene object at a single perspective.

[0415] It should be noted that the appearance image can be obtained in the virtual scene, such as intercepting the appearance image of the scene object in the scene object; the appearance image can also be obtained outside the virtual scene, such as intercepting the appearance image of the scene object in the development interface during the development and design process of the scene object.

[0416] • Calling a multi-modal model to perform prediction on the attribute text and the appearance image of the scene object to obtain a visual label of the scene object;

[0417] For example, the multi-modal model has the ability to perform model prediction on text information and picture information of different modalities. In this embodiment, the input parameters of the multi-modal model are the attribute text and the appearance image of the scene object; the multi-modal model predicts the visual label of the scene object from two modalities of the text modality and the picture modality, and the visual label is used to describe the scene object in at least one dimension.

[0418] Exemplarily, the multi-modal model comprises an artificial neural network (ANN), and the network of the artificial neural network comprises but is not limited to at least one of a convolutional neural network (CNN), a recurrent neural network (RNN), a temporal convolutional network (TCN), a long short-term memory (LSTM), a multilayer perceptron (MLP), and a support vector machine (SVM). The network structure of the multi-modal network is not limited in the present application. Exemplarily, the multi-modal model has the ability to extract visual features of a scene object from input information of a text modality and a picture modality, and obtain visual labels of the scene object. Exemplarily, the multi-modal model can be a classification model that assigns a scene object to a preset type label, or a prediction model that predicts a label capable of expressing visual features of a scene object for the scene object.

[0419] Optionally, the visual label of the scene object is subjected to personification rewriting to obtain a rewritten visual label conforming to a spoken expression of natural language. The visual label is used to describe the scene object in at least one dimension, and the appearance image of the scene object presents rich visual features of the scene object. However, the description of the scene object in the spoken expression of natural language cannot cover the visual features of the scene object in each dimension. For example, a virtual bed in a virtual scene is described in a spoken expression, which usually focuses on the color, material, and placement position of the virtual bed, and usually neglects the process, such as carving and embossing, used for the head of the bed and the side plate of the bed. The purpose of performing personification rewriting on the visual label of the scene object is to obtain a matching label more similar to the spoken expression. It can be understood that performing personification rewriting can be deleting part of the content in the visual label, or changing the visual label to a label with the same semantics but different text expression. Exemplarily, performing personification rewriting on the visual label is based on calling a natural language model. Exemplarily, the natural language model carries prior knowledge of the spoken expression of natural language. For example, the natural language model is implemented as a large language model (LLM).

[0420] Further, performing personification rewriting by calling the natural language model comprises:

[0421] obtaining a first sample label pair;

[0422] According to the first sample label pair and the visual label, a rewriting guide sentence is constructed;

[0423] The rewriting guide sentence is input into the natural language model, and a matching label conforming to the natural language spoken expression is predicted;

[0424] In this example, the first sample label pair includes a first label before personification rewriting and a second label obtained through personification rewriting; taking a virtual bed in a virtual scene as an example, the first label in the first sample label pair is: “Type description: wooden double bed; Material: [wood]; Transparency: opaque; Color: [brown]; Surface feature: [carving]; Shape: [cuboid]”. The second label is: “bed, brown, wooden, rectangular”.

[0425] The rewriting guide sentence has a natural semantic of rewriting the visual label with reference to the first sample label pair; in one example, the visual label that needs to be rewritten is: “Type description: green plant potted; Material: [plant]; Transparency: opaque; Color: [brown, green]; Surface feature: [smooth]; Shape: [irregular shape]”. Accordingly, according to the first sample label pair and the visual label, the guide sentence obtained is: example: rewrite “Type description: wooden double bed; Material: [wood]; Transparency: opaque; Color: [brown]; Surface feature: [carving]; Shape: [cuboid]” to “bed, brown, wooden, rectangular”. According to the above example, rewrite “Type description: green plant potted; Material: [plant]; Transparency: opaque; Color: [brown, green]; Surface feature: [smooth]; Shape: [irregular shape]”.

[0426] FIG. 9 shows a schematic diagram of a visual label of a scene object provided by one example embodiment of the present application. In one example, the question sentence 402 is: “Below is a rendering diagram of a three-dimensional model named A with a size of B; please understand what the three-dimensional model is through the model name, image, and size information, and summarize the overall characteristics, and finally output the visual label in the Json format.”. Among them, taking the semicolon “;” in the question sentence 402 as a division, the first half is the first subpart, which is the supplementary introduction information of the scene object. The supplementary introduction information of the scene object includes the attribute text of the scene object obtained from the data set 400, such as the name of the scene object in the virtual scene is A, and the size of the scene object in the virtual scene is B.

[0427] The second subpart is a second half part, which is a response guiding sentence for the visual question answering model. Further, the question sentence includes a preset sentence template, and the attribute text of the scene object is filled into the preset sentence template to obtain the question sentence. The filling position of the attribute text of the scene object is the position of the word "A" and the word "B" in the question sentence. The name of the scene object in the virtual scene in the attribute text is marked as an asset name, and the size of the scene object in the virtual scene in the attribute text is marked as an asset size. The attribute text is filled into the preset sentence template to obtain the question sentence.

[0428] The appearance image 404 of the scene object includes images obtained by observing the scene object from nine perspectives. The question sentence 402 and the appearance image 404 of the scene object are input into the visual question answering model 410 to obtain the visual label 415 of the scene object. Exemplarily, the visual label 415 of the scene object includes: "type description: wooden double bed; material: [wood]; transparency: opaque; color: [brown]; surface feature: [carving]; shape: [cuboid]".

[0429] Further, the visual label 415 of the scene object is used to supplement the data set 400. After being supplemented by the visual label 415, the data set 400 includes the appearance image 404, the visual label 415 and the spatial position 420 of the scene object in the virtual scene, and the spatial position 420 includes a coordinate position, orientation information, a bounding box information and a mask point information, which will be introduced in a separate embodiment below. In various embodiments of the present application, the data set 400 is also referred to as a spatial information library of the virtual scene.

[0430] · constructing a spatial database of the virtual environment based on the visual labels of the scene objects in the virtual environment;

[0431] The spatial database includes the visual labels of the scene objects in the virtual environment.

[0432] · filtering the visual labels of the multiple candidate scene objects in the perception range based on the perception range in the spatial database.

[0433] Exemplarily, the spatial database includes the spatial position introduced above. According to the spatial position, the multiple scene objects in the virtual environment located in the perception range are determined, and the multiple scene objects located in the perception range are determined as the multiple candidate scene objects. The visual labels of the multiple candidate scene objects are obtained in the spatial database.

[0434] To sum up, the method provided in this embodiment realizes the description information of the scene object in the text dimension and the image dimension, and predicts the visual label of the scene object by calling the multi-modal model to perform prediction on the attribute text and the appearance image. In the process of predicting the visual label of the scene object, the description information of the scene object in multiple dimensions is used, so that the characteristics of the scene object can be comprehensively described from the natural language semantics in the text dimension and the visual information in the image dimension, the accuracy of the visual label is ensured, and the efficiency of obtaining the visual label of the scene object is improved.

[0435] FIG. 10 shows a flowchart of a virtual role control method provided in an example embodiment of the present application. The method can be executed by a computer device. That is, on the basis of the embodiment shown in FIG. 3, the method further includes step 560:

[0436] Step 560: presenting feedback information of the NPC to the command information.

[0437] For example, the feedback information is used to indicate that the NPC receives the command information and / or the NPC has executed the virtual activity indicated by the command information.

[0438] For example, the feedback information is presented when the NPC starts to execute the virtual activity instructed by the command information, so as to prompt the user that the command information has been issued, and avoid the user repeatedly issuing the same command information in an unclear situation. For example, the feedback information is presented when the NPC executes the virtual activity instructed by the command information and / or completes the virtual activity, so as to inform the user of the execution situation of the command information.

[0439] For example, the feedback information is usually voice information, and the feedback information corresponds to the classification label obtained by recognizing the execution intention of the command information one by one. The feedback information has different feedback contents according to whether the virtual activity instructed by the command information is successfully completed.

[0440] For example, the feedback information can be saved in a local audio library of the terminal, and the feedback information is obtained and played in the local audio library. The feedback information can also be generated by a Text-to-Speech (TTS) technology according to at least one of the classification label obtained by recognizing the execution intention of the command information, whether the virtual activity instructed by the command information is successfully completed, and the like, so as to obtain feedback information of various types.

[0441] To sum up, the method provided in this embodiment determines that the virtual activity is executed in the first semantic understanding manner by taking the environment perception information of the NPC and / or the host virtual role as a reference among at least two semantic understanding manners of the command information; and presents the feedback information of the NPC to the command information, which can improve the human-computer interaction efficiency of the NPC and the host virtual role in information transmission, and increase the immersion of the user.

[0442] An example embodiment of the present application provides a control method of a virtual role. The method comprises:

[0443] displaying a master virtual role or NPC in a virtual environment;

[0444] An example master virtual role is a virtual role directly controlled by a user in a virtual environment. There is one or more non-player characters (NPCs) in the virtual environment; the NPCs and the master virtual role belong to the same virtual camp and are teammates of the master virtual role.

[0445] obtaining command information;

[0446] An example command information comprises natural language semantics for controlling the NPCs; the NPCs are commanded in the virtual scene through command texts, which can be directly input by the user or extracted from voice information input by the user, and the control manner of the command texts is not limited in the present application.

[0447] The command information comprises a behavior intention and a scene entity; the behavior intention is used to indicate a type of virtual activity performed by the NPCs, such as indicating a type of virtual activity performed in the virtual environment; the scene entity comprises a first scene object, such as indicating a scene object in the virtual environment to which the virtual activity is performed. The behavior intention and the scene entity comprised in the command information are part of the words in the command information.

[0448] An example command information controls the NPCs from two aspects of a type of virtual activity performed and a scene object in the virtual environment to which the virtual activity is performed; it is a complex instruction for controlling the NPCs.

[0449] controlling the NPCs to perform a virtual activity in response to the command information;

[0450] An example controlling the NPCs to perform a virtual activity is determined based on environmental perception information of the master virtual role and / or the NPCs. In some examples, the master virtual role and / or the NPCs obtain the environmental perception information within a perception range of visual, auditory, action trajectory, etc.

[0451] An example virtual activity is close to a current perception situation of the master virtual role and / or the NPCs to the virtual environment, and the NPCs are controlled with reference to the current perception situation of the virtual role and / or the NPCs to the virtual environment.

[0452] To sum up, the method provided in the embodiment controls the complex instructions of the NPC from two aspects of executing what type of virtual activity and executing the virtual activity for what scene object in the virtual environment, determines the virtual activity to be executed by taking the environment perception information of the host virtual role and / or the NPC as reference, and ensures the NPC to execute the virtual activity in cooperation with the host virtual role by taking the environment perception information of the host virtual role and / or the NPC as reference.

[0453] In an optional implementation, the scene entity includes a first scene object; and the above-mentioned controlling the NPC to execute the virtual activity in response to the command information can be implemented as:

[0454] controlling the NPC to execute the virtual activity for the first scene object located in the perception range of the host virtual role and / or the NPC in response to the command information;

[0455] wherein, there are multiple scene objects belonging to the same type in the virtual environment, the multiple scene objects include the first scene object, and the perception range is a virtual area in which the host virtual role and / or the NPC acquire the environment perception information.

[0456] Furthermore, the above-mentioned controlling the NPC to execute the virtual activity in response to the command information can be implemented as:

[0457] controlling the NPC to execute the virtual activity for the first scene object located in the first virtual terrain in response to the command information, the first virtual terrain being a virtual terrain to which a location of the host virtual role and / or the NPC belongs.

[0458] In another optional implementation, the behavior intention corresponds to a type of virtual activity permitted to be executed in the perception range; and the above-mentioned controlling the NPC to execute the virtual activity in response to the command information can be implemented as:

[0459] controlling the NPC to execute the virtual activity according to the first behavior intention in response to the command information; and the perception range is a virtual area in which the host virtual role and / or the NPC acquire the environment perception information.

[0460] In yet another optional implementation, the command information further includes a natural language pronunciation of a scene object in the perception range; and the above-mentioned controlling the NPC to execute the virtual activity in response to the command information can be implemented as:

[0461] converting the command information into a command text according to the natural language pronunciation of the scene object in the perception range in response to the command information, the perception range being a virtual area in which the host virtual role and / or the NPC acquire the environment perception information;

[0462] controlling the NPC to execute the virtual activity according to the command text.

[0463] It should be noted that the above unfinished introduction of the scene entity, the behavior intention, and the natural language pronunciation can refer to the corresponding embodiments of FIG. 4 to FIG. 6 in the above. Here, it will not be repeated one by one.

[0464] In an optional implementation, the above-mentioned response to the command information, controlling the NPC to perform the virtual activity, can be implemented as:

[0465] According to the command information, determining the command text in the natural language form;

[0466] Calling the conditional random field to perform entity prediction on the command text, to obtain the entity label carried by the command text;

[0467] In the perception range, determining the first scene object according to the entity label;

[0468] Calling the prediction network to perform intention recognition on the command text, to obtain the classification label of the command text; the prediction network is used to predict the control mode of the command text on at least one behavior dimension for the NPC;

[0469] According to the classification label, determining the control command of the NPC, and based on the control command, controlling the NPC to perform the virtual activity on the first scene object.

[0470] For the above-mentioned content, please refer to the corresponding embodiments of FIG. 7 to FIG. 9 in the above. Here, it will not be repeated one by one.

[0471] FIG. 11 shows a flowchart of a method for processing command information according to an example embodiment of the present application. The method can be executed by a computer device. The method comprises:

[0472] Step 710: obtaining command information;

[0473] For example, the command information includes natural semantics for controlling the NPC; the command text can be directly input by the user, or can be extracted from the voice information input by the user, which can be used to command the non-player character in the virtual scene. The control mode of the command text is not limited in the present application. For example, the command information includes natural semantics for controlling the NPC, which is also called natural language command or voice form natural language command.

[0474] For example, the main virtual character is a virtual character directly controlled by the user in the virtual environment. There is one or more non-player characters (NPC) in the virtual environment; the NPC and the main virtual character belong to the same virtual camp and are teammates of the main virtual character.

[0475] In one example, the method is performed by a server, and the command information acquired by the server is reported by a terminal. The server implements the processing of the command information.

[0476] Step 720: determining a first semantic understanding manner from at least two semantic understanding manners of the command information based on the environment perception information of the master virtual character and / or the NPC.

[0477] The command information corresponds to at least two semantic understanding manners. The at least two semantic understanding manners correspond to different control manners of the NPC in the virtual environment. For example, at least one of the scene entity, the word pronunciation, and the behavior intention indicated in the at least two semantic understanding manners is different.

[0478] For example, the first semantic understanding manner is one of the at least two semantic understanding manners corresponding to the command information. The first semantic understanding manner is determined from the at least two semantic understanding manners based on the environment perception information of the master virtual character and / or the NPC. In some examples, the master virtual character and / or the NPC acquires the environment perception information within the perception range of the perception manner such as vision, hearing, and action trajectory. Based on the environment perception information, the first semantic understanding manner determined from the at least two semantic understanding manners is close to the current perception situation of the master virtual character and / or the NPC to the virtual environment, and the NPC is controlled based on the current perception situation of the master virtual character and / or the NPC to the virtual environment.

[0479] Step 730: controlling the NPC to perform a virtual activity according to the first semantic understanding manner from the at least two semantic understanding manners based on the command information.

[0480] For example, since the first semantic understanding manner is determined from the at least two semantic understanding manners based on the environment perception information of the master virtual character and / or the NPC, the first semantic understanding manner is different when facing the same command information in at least two situations in the virtual environment (at least one of the position, the virtual life value, the virtual prop holding amount, the virtual skill cooling time, and the virtual level of the master virtual character and / or the NPC in the above two situations is different). That is to say, different virtual activities are performed by the NPC in response to the same command information as the situation in the virtual environment changes.

[0481] In one example, the server needs to send the command information to the terminal, so that the terminal controls the NPC to perform a virtual activity according to the first semantic understanding manner from the at least two semantic understanding manners based on the indication of the command information.

[0482] To sum up, the method provided in the embodiment determines the virtual activity to be executed according to the first semantic understanding manner by taking the environment perception information of the host virtual character and / or the NPC as a reference in at least two semantic understanding manners of the command information. In the case that there are at least two semantic understanding manners of the command information, on the one hand, the user does not need to supplement more basis for the command information, and can control the NPC to execute the virtual activity according to the command information, thereby improving the human-computer interaction efficiency; on the other hand, taking the environment perception information of the host virtual character and / or the NPC as a reference, the NPC and the host virtual character are ensured to cooperatively execute the virtual activity.

[0483] In an optional implementation, similar to the description in FIGS. 4-6 above, in an implementation, the command information corresponds to at least two semantic understanding manners of scene entities; in another implementation, the command information corresponds to at least two semantic understanding manners of word pronunciations; in yet another implementation, the command information corresponds to at least two semantic understanding manners of behavior intentions.

[0484] In an implementation, step 720 in FIG. 11 can be implemented as:

[0485] determining, based on the environment perception information, the command information to be executed on a first scene object located in a perception range of the host virtual character and / or the NPC, the command information corresponding to at least two semantic understanding manners of scene entities;

[0486] wherein there are multiple scene objects belonging to the same type in a virtual environment where the host virtual character and / or the NPC are located, the multiple scene objects include the first scene object, and the perception range is a virtual area where the host virtual character and / or the NPC obtain the environment perception information.

[0487] Further optionally, determining, based on the environment perception information, the command information to be executed on the first scene object located in the perception range of the host virtual character and / or the NPC can be implemented as:

[0488] determining, based on the environment perception information, the command information to be executed on the first scene object located in the field of view range;

[0489] or, determining, based on the environment perception information, the command information to be executed on the first scene object located in a first virtual terrain; the first virtual terrain is a virtual terrain to which a location of the host virtual character and / or the NPC belongs.

[0490] In another implementation, step 720 in FIG. 11 can be implemented as:

[0491] • based on the environmental perception information, determine the command information as natural language utterances of scene objects in a perception range, and convert the command information into command text according to a first semantic understanding manner; the command information corresponds to a semantic understanding manner of at least two word utterances;

[0492] The perception range is a virtual area in which the host virtual character and / or the NPC obtain the environmental perception information; the NPC performs a virtual activity based on an instruction of the command text.

[0493] In another implementation, step 720 in FIG. 11 can be implemented as:

[0494] • based on the environmental perception information, determine the command information as performing a virtual activity according to a first behavior intention; the first behavior intention corresponds to a type of virtual activity permitted to be performed in a perception range, and the perception range is a virtual area in which the host virtual character and / or the NPC obtain the environmental perception information. The command information corresponds to a semantic understanding manner of at least two behavior intentions;

[0495] It should be noted that the semantic understanding manner of the scene entity, the semantic understanding manner of the word utterance, and the semantic understanding manner of the behavior intention are described in the above corresponding embodiments of FIGS. 4 to 6, and will not be repeated here.

[0496] In an optional implementation, step 720 in FIG. 11 can be implemented as:

[0497] • determine the command text in a natural language form according to the command information;

[0498] • perform entity prediction on the command text by calling a conditional random field to obtain entity labels carried by the command text;

[0499] • determine a first scene object in the perception range according to the entity labels;

[0500] • perform intention recognition on the command text by calling a prediction network to obtain a classification label of the command text; the prediction network is used to predict a control manner of the command text on the NPC in at least one behavior dimension;

[0501] Correspondingly, step 730 in FIG. 11 can be implemented as:

[0502] • determine a control command of the NPC according to the classification label corresponding to the command information and the first scene object, and send the control command; the control command is used to control the NPC to perform a virtual activity on the first scene object according to the first semantic understanding manner.

[0503] In one example, the method is performed by a server, and the command information obtained by the server is reported by a terminal. The server implements processing on the command information. In one example, the server needs to send the command information to the terminal, so that the terminal controls the NPC to perform the virtual activity according to the first semantic understanding manner in the at least two semantic understanding manners according to the indication of the command information.

[0504] Optionally, the above calling the prediction network to perform the intention recognition on the command text to obtain the classification label of the command text can be implemented as:

[0505] calling the hierarchical prediction network to perform the intention recognition on the command text to obtain the classification label of the command text.

[0506] The hierarchical prediction network includes at least two sub-networks, and the at least two sub-networks are constructed based on a tree structure of multiple classification labels.

[0507] Optionally, the method further comprises:

[0508] obtaining visual labels of a plurality of candidate scene objects in the perception range;

[0509] determining the first scene object in the perception range according to the entity label, comprising:

[0510] respectively calculating similarities between the visual labels of the plurality of candidate scene objects and the entity label to determine the first scene object with the highest similarity in the plurality of candidate scene objects.

[0511] Optionally, the method further comprises:

[0512] obtaining spatial positions of each scene object in the virtual environment in the virtual environment;

[0513] screening the plurality of candidate scene objects in the perception range from the each scene object in the virtual environment according to the spatial positions.

[0514] Optionally, the above obtaining the spatial positions of each scene object in the virtual environment in the virtual environment can be implemented as:

[0515] obtaining at least one of coordinate positions, orientation information, bounding box information, and cover point information of each scene object in the virtual environment;

[0516] The coordinate positions are used to indicate positions of the scene object in the virtual environment, the orientation information is used to indicate a direction faced by the scene object in the virtual environment, the bounding box information is used to indicate sizes of the scene object in the virtual environment, and the cover point information indicates a recommended standing point of the virtual character when the virtual character approaches the scene object.

[0517] Optionally, the method further comprises:

[0518] obtain attribute text of a scene object in the virtual environment, the attribute text being used to introduce inherent attributes of the scene object in the virtual environment;

[0519] obtain appearance images of the scene object in the virtual environment, the appearance images being used to describe styles of the scene object;

[0520] invoke a multi-modal model to perform prediction on the attribute text and the appearance images of the scene object, to obtain a visual label of the scene object, the visual label being used to describe the scene object in at least one dimension;

[0521] construct a spatial database of the virtual environment based on the visual labels of the scene objects in the virtual environment;

[0522] obtain visual labels of a plurality of candidate scene objects in a perception range, comprising:

[0523] filter, based on the perception range, the visual labels of the plurality of candidate scene objects in the perception range in the spatial database.

[0524] It should be noted that the command text, entity prediction, entity label, prediction network, intent recognition, control command and other information in this optional implementation manner are not fully described, please refer to the embodiments corresponding to FIG. 7 in the foregoing, and the embodiments after the embodiments corresponding to FIG. 7; hereinafter will not be repeated one by one.

[0525] An example embodiment of the present application provides a processing method of command information. The method comprises:

[0526] obtaining command information;

[0527] Illustratively, the main control virtual role is a virtual role directly controlled by the user in the virtual environment. There is one or more non-player characters (NPC) in the virtual environment; the NPC and the main control virtual role belong to the same virtual camp and are teammates of the main control virtual role.

[0528] Illustratively, the command information includes natural semantics for controlling the NPC; the natural semantics for controlling the NPC is used to instruct the non-player character in the virtual scene through the command text, and the command text can be directly input by the user or extracted from the voice information input by the user, and the control mode of the command text is not limited in the present application.

[0529] The command information includes a behavior intent and a scene entity; the behavior intent is used to indicate a type of virtual activity performed by the NPC, such as indicating a type of virtual activity performed in the virtual environment; the scene entity includes a first scene object, such as indicating a type of virtual activity performed on a scene object in the virtual environment. The behavior intent and the scene entity included in the command information are part of the words in the command information.

[0530] The command information is used to control the NPC to perform the virtual activity in an example, and the virtual activity is performed on a scene object in a virtual environment.

[0531] The NPC is controlled to perform the virtual activity based on the command information.

[0532] In an example, the NPC is controlled to perform the virtual activity based on environment perception information of the master virtual character and / or the NPC. In some examples, the environment perception information is obtained by the master virtual character and / or the NPC within a perception range in a perception mode such as vision, hearing, action trajectory, etc.

[0533] In an example, the virtual activity is close to the current perception of the virtual environment by the master virtual character and / or the NPC, and the NPC is controlled based on the current perception of the virtual environment by the master virtual character and / or the NPC. In an example, the method is executed by a server, and the command information obtained by the server is reported by a terminal. The processing of the command information is implemented by the server.

[0534] Optionally, the scene entity includes a first scene object; and / or, the behavior intention corresponds to a type of virtual activity permitted to be performed in the perception range; and / or, the command information further includes a natural language pronunciation of the scene object in the perception range; for details of the scene entity, the behavior intention, and the natural language pronunciation, please refer to the corresponding embodiments of FIGS. 4 to 6 in the foregoing description. Here, the repeated description is omitted.

[0535] In an example, the command text in the natural language form is determined according to the command information.

[0536] The conditional random field is called to perform entity prediction on the command text, and an entity label carried by the command text is obtained.

[0537] The first scene object is determined in the perception range according to the entity label.

[0538] The prediction network is called to perform intention recognition on the command text, and a classification label of the command text is obtained. The prediction network is used to predict a control mode of the command text on the NPC in at least one behavior dimension.

[0539] The control command of the NPC is determined according to the classification label, and the NPC is controlled to perform the virtual activity on the first scene object based on the control command.

[0540] For details of the above description, please refer to the corresponding embodiments of FIGS. 7 to 9 in the foregoing description. Here, the repeated description is omitted.

[0541] To sum up, the method provided in the embodiment controls the complex instructions of the NPC from two aspects of executing what type of virtual activity and executing the virtual activity on what scene object in the virtual environment, determines the virtual activity to be executed by taking the environment perception information of the host virtual character and / or the NPC as reference, and ensures the NPC to execute the virtual activity in cooperation with the host virtual character by taking the environment perception information of the host virtual character and / or the NPC as reference.

[0542] FIG. 12 shows a schematic diagram of a processing method of command information provided in an example embodiment of the present application. The command information 1002 is obtained, which is a control voice of the user to the NPC. An Automatic Speech Recognition (ASR) service 1102 is requested to perform conversion on the command information 1002, and command text 1004 corresponding to the command information 1002 is obtained.

[0543] The auxiliary data 1006 is obtained, which includes static data of the pre-constructed virtual environment and dynamic data of the virtual environment at the current time.

[0544] The static data includes a visual label of the scene object, which is used to describe the scene object in at least one dimension, such as material, transparency, color, shape, etc. The static data includes a spatial position of the scene object, which includes at least one of the following of the scene object: a coordinate position (Location), such as coordinate information of a center point or a preset point of the scene object in the virtual scene; a rotation (Rotation), such as a direction that a front surface of the scene object faces in the virtual scene; a bounding box (Bounding Box), which is used to indicate the size of the scene object in the virtual scene; cover points (Cover points), which are used to indicate a recommended position point of the virtual character when the virtual character approaches the scene object, so as to realize that the scene object can mask the virtual character.

[0545] The dynamic data includes the basic information (such as coordinate position, orientation, occupied space size, etc.) of the host virtual character and / or the ally character (NPC) and the enemy character of the user at the current time when the command information 1002 is issued, which is obtained in real time; the instructions, body posture, etc. that the host virtual character and / or the ally character is currently executing.

[0546] The request model inference service 1104 performs inference prediction on the command text 1004 and the auxiliary data 1006 to obtain instruction data 1010 and feedback text 1012; wherein the instruction data 1010 is an instruction that can be used for the behavior tree execution of the NPC, also known as a control instruction. The feedback text 1012 is text presented in a voice manner for indicating that the NPC receives the command information, and / or the NPC has executed the virtual activity indicated by the command information.

[0547] Exemplarily, the classification label of the command text 1004 is input to determine the target entity in the virtual space and the control command of the NPC; exemplarily, there is a corresponding control command for each classification label in the command text 1004, and the control command of the NPC is determined according to the corresponding relationship between the classification label and the control command.

[0548] Exemplarily, the model inference service 1104 has the ability to predict the classification label and the entity label of the command text 1004, such as calling a conditional random field to perform entity prediction on the command text to obtain the entity label carried by the command text; calling a prediction network to perform intent recognition on the command text to obtain the classification label of the command text; the entity label is used to determine the first scene object in the perception range, and the control command is used to control the NPC to execute the virtual activity according to the first semantic understanding manner for the first scene object.

[0549] The instruction data 1010 and the feedback text 1012 are cached to the first component 1020 of the terminal, and the terminal acquires the command information 1002. The first component 1020 of the terminal calls the second component 1120 on the server side based on Unreal Engine Remote Procedure Call (RPC); the first component 1020 and the second component 1120 are both used to cache the instruction data 1010 and the feedback text 1012.

[0550] The control command indicates a subsequent action that the NPC needs to perform, which can be a single action or a sequence of actions. In the case of a multi-action sequence instruction as the control command, the multi-action sequence instruction is used to indicate that the NPC executes a sequence of actions, and the data structure for storing the multi-action sequence instruction is in the form of a list, and the to-be-executed instruction in the sequence instruction is cached. When the behavior tree of the NPC executes the current instruction, the cached instruction is called back for execution, and the to-be-executed instruction in the cache is continued to be issued for execution. The second component 1120 issues the instruction data 1010 and the feedback text 1012 to the first voice feedback logic 1202, the activity execution logic 1204, the activity judgment logic 1206, and the second voice feedback logic 1208 in sequence according to the order of the sequence instruction.

[0551] The first voice feedback logic 1202 is configured to instruct the voice feedback presented in the case that the NPC receives the command information; the activity execution logic 1204 is configured to control the NPC to execute the virtual activity according to the instruction data 1010; the activity judgment logic 1206 is configured to determine whether the virtual activity executed by the NPC is ended; the second voice feedback logic 1208 is configured to instruct the voice feedback presented in the case that the NPC has executed the virtual activity indicated by the command information; for example, the voice-formatted feedback information obtained according to the feedback text 1012 can be converted by requesting the text-to-voice service 1132, or can be an existing audio obtained by querying the audio library 1134.

[0552] For example, the NPC has a behavior tree 1300 to control the NPC to act in the virtual environment according to the preset logic (such as automatically avoiding virtual attacks, etc.); the behavior tree 1300 is connected with the third voice feedback logic 1210, and the third voice feedback logic 1210 is configured to provide voice feedback when executing the virtual activity according to the preset logic, and broadcast the activity automatically executed by the NPC in the virtual environment.

[0553] It should be noted that the text-to-voice service 1102, the model inference service 1104 and the text-to-voice service 1132 in the embodiment can be implemented as a separate server, or can be implemented as a server cluster together with a background server supporting the virtual environment.

[0554] Next, the control command of the NPC is introduced.

[0555] As shown in the introduction in FIG. 12, the request model inference service performs inference prediction on the command text and auxiliary data to obtain instruction data and feedback text. The request packet input into the request model inference service includes at least one of the following fields: a Content field used to indicate the command text; a CurDecision field used to indicate the NPC currently executing the instruction; a CurReceiver field used to indicate the identity of the NPC of the last command; a CurTargetLocation field used to indicate the target position of the NPC of the last command; a MoveType field used to indicate the current moving manner of the NPC (running, slow walking, normal walking, etc.); a PostureType field used to indicate the current posture of the NPC (standing, squatting, lying, etc.); a Location field used to indicate the position of the host virtual character and the NPC; a Rotation field used to indicate the orientation of the host virtual character and the NPC; a TraceInfo field used to indicate the ray information of the host virtual character, such as the front, back, left, right, up and down six directions, used to indicate whether there is an obstruction, whether it is indoors; an EnemyLocation field used to indicate the position of the current enemy character; an EnemyRotation field used to indicate the orientation of the current enemy character; a GunFireLocation field used to indicate the position of the virtual attack sound (such as gunshots) in the current virtual environment; and a NormalSoundLocation field used to indicate the position of the environment sound in the current virtual environment.

[0556] The model inference service has the ability to predict the type of virtual activity that needs to be performed, the target entity to which the virtual activity is performed, the identity of the controlled NPC performing the virtual activity, and the like. The content inferred by the model inference service includes at least one of the following fields (also referred to as a response content): a Base field for indicating a reference frame for performing the virtual activity (such as walking in front of an A object, and the A object being the reference frame); a Fire field for indicating whether to allow initiating a virtual attack; an AbsLoc field for indicating a relative position (such as front, back, left, and right) for performing the virtual activity; an AbsLocDis field for indicating a relative distance (such as 5 meters in front) for performing the virtual activity; an Item field for indicating the identity of an item used (such as indicating a hand grenade, a smoke bomb, or a medical kit); a Receiver field for indicating a receiver of command information, that is, the identity of an NPC performing the virtual activity; a VolumeId field for indicating a range of an area for performing the virtual activity; a MoveType field for indicating a moving manner for performing the virtual activity; a PostureType field for indicating a body posture for performing the virtual activity; a TargetLocation field for indicating a moving destination coordinate for performing the virtual activity; a TargetYaw field for indicating a desired orientation for performing the virtual activity; a TargetType field for indicating the type of the destination (whether it is an item, an evacuation point, or an indoor site); a FeedbackText field for indicating a feedback text; a Content_with_breakpoint field for indicating a clause segmentation result of a command text (No. 2 goes in front, No. 3 comes back, which is segmented into No. 2 goes in front and No. 3 comes back); and a Decision field for indicating the type of the virtual activity.

[0557] For example, when the value of the Decision field is -2, the type of the virtual activity is to stay in place; when the value of the Decision field is -1, the type of the virtual activity is to keep the virtual activity type of the historical command; when the value of the Decision field is 0, the type of the virtual activity is to go to a specified location, return after reaching the location or finding an enemy; when the value of the Decision field is 1, the type of the virtual activity is to go to a specified location, not to be interrupted by finding an enemy, and to stay in place after reaching the location; when the value of the Decision field is 2, the type of the virtual activity is to circle around the host virtual character, and to face the possible direction of encountering an enemy; when the value of the Decision field is 3, the type of the virtual activity is to stop any current behavior; when the value of the Decision field is 4, the type of the virtual activity is to adjust the body posture / movement speed of the NPC only; when the value of the Decision field is 5, the type of the virtual activity is to throw / use an item; when the value of the Decision field is 6, the type of the virtual activity is to inquire about a location; when the value of the Decision field is 7, the type of the virtual activity is to inquire about intelligence; when the value of the Decision field is 8, the type of the virtual activity is to open a door; when the value of the Decision field is 9, the type of the virtual activity is to close a door; when the value of the Decision field is 10, the type of the virtual activity is to launch a virtual attack; when the value of the Decision field is 11, the type of the virtual activity is to dodge a virtual attack; and when the value of the Decision field is 12, the type of the virtual activity is to retreat from an enemy character.

[0558] FIG. 13 shows a schematic diagram of a behavior tree of an NPC according to an example embodiment of the present application. The behavior tree of the NPC includes a general service 1500, which controls the NPC to observe, pick up items, etc. in a virtual environment. The child nodes of the general service 1500 include a preset service 1502, an interruption service 1504, and a control service 1506.

[0559] The preset service 1502 is a service automatically executed by the NPC in the virtual environment, such as the NPC automatically avoiding a virtual grenade thrown. The interrupt service 1504 is used to indicate the NPC to stop the current virtual activity, such as when a teammate character is attacked, the NPC stops the current virtual activity and assists the teammate character. The control service 1506 is a virtual activity executed in response to the command information of the user. The control service includes the following types: a standby activity 1510, an inquiry activity 1520, a movement activity 1530, a protection activity 1540, a reconnaissance activity 1550, a door interaction activity 1560, and a virtual prop usage activity 1570. The seven types of virtual activities have corresponding Decision fields, and the virtual activity on the corresponding behavior tree branch is executed according to the Decision field in the response packet of the model inference service. It can be understood that some or all of the seven types of virtual activities also include subtypes, such as the standby activity 1510 including a change direction activity 1511 and a wait preset time activity 1512. The inquiry activity 1520 includes a position inquiry 1521 and an intelligence inquiry 1522. That is, the behavior tree of the NPC can include a more loaded tree structure, which is not limited by the present application.

[0560] An example embodiment of the present application provides a control method of a virtual character. The method includes:

[0561] · displaying at least one of a main control virtual character and an NPC in a virtual environment;

[0562] For example, the main control virtual character is a virtual character directly controlled by the user in the virtual environment. There is one or more non-player characters (NPC) in the virtual environment; the NPC and the main control virtual character belong to the same virtual camp and are teammates of the main control virtual character, and the NPC follows the main control virtual character to move in the virtual environment. In addition, the NPC can also be a follower character, a pet character, etc. controlled by the main control virtual character. In addition, the NPC can also be a neutral character and only cooperates with the main control virtual character under certain conditions.

[0563] · obtaining a natural language command;

[0564] For example, the natural language command includes a natural language semantic for commanding the NPC; the natural language command for commanding the NPC in the virtual scene can be directly input by the user or extracted from the voice information input by the user, and the present application does not limit the obtaining method of the natural language command.

[0565] The natural language command comprises at least one of a behavior intention and a scene entity; the behavior intention is used to indicate a type of virtual activity performed by the NPC, such as being used to indicate what type of virtual activity is performed in the virtual environment; the scene entity comprises a first scene object, such as being used to indicate what scene object in the virtual environment the virtual activity is performed for. The behavior intention and the scene entity comprised in the natural language command are part of words in the natural language command.

[0566] In one example, the natural language command controls the NPC from two aspects of what type of virtual activity is performed and what scene object in the virtual environment the virtual activity is performed for; it is a complex instruction for controlling the NPC.

[0567] • controlling the NPC to perform the virtual activity in response to the natural language command;

[0568] In examples, the controlling the NPC to perform the virtual activity is determined based on environment perception information of the master virtual character and / or the NPC. In some examples, the master virtual character and / or the NPC obtains the environment perception information within a perception range in a perception manner such as vision, hearing, action trajectory, etc. In examples, the virtual activity is close to a current perception situation of the virtual environment by the master virtual character and / or the NPC, and the NPC is controlled with reference to the current perception situation of the virtual environment by the master virtual character and / or the NPC. In examples, the NPC has autonomous behavior capability, and the user only commands the NPC, and the user gives a command, and the NPC understands and executes the command based on its own autonomous behavior capability.

[0569] In summary, the method provided by the embodiment controls the NPC from two aspects of what type of virtual activity is performed and what scene object in the virtual environment the virtual activity is performed for, determines the virtual activity to be performed with reference to the environment perception information of the master virtual character and / or the NPC, and ensures that the NPC cooperates with the master virtual character to perform the virtual activity with reference to the environment perception information of the master virtual character and / or the NPC.

[0570] FIG. 14 shows a flowchart of a virtual character control method provided by one example embodiment of the present application. The method comprises:

[0571] Step 201: obtaining a spatial data set;

[0572] The spatial data set of the virtual environment comprises a visual label of a scene object, and the visual label is used to describe visual features of the scene object in at least one dimension, such as describing the scene object in dimensions of material, transparency, color, shape, etc.

[0573] Step 202: displaying at least one of a master virtual character and an NPC located in a virtual environment;

[0574] Exemplarily, the master virtual role is a virtual role directly controlled by the user in the virtual environment. There is one or more NPCs in the virtual environment; the NPCs and the master virtual role belong to the same virtual camp and are teammates of the master virtual role.

[0575] Step 203: obtaining a natural language command in a voice form;

[0576] The natural language command includes a natural semantic for commanding the NPC, and the command information includes a behavior intention and a scene entity; the behavior intention is used to indicate a type of virtual activity performed by the NPC, and the scene entity is used to indicate a target entity to which the virtual activity is directed;

[0577] Step 204: converting the natural language command in the voice form into a natural language command in a text form;

[0578] Exemplarily, the natural language command in the voice form is subjected to an automatic speech recognition (ASR) process to determine the natural language command in the text form. The automatic speech recognition process usually includes calling components such as an acoustic model and a language model to realize recognition of pronunciation, vocabulary and syntax structure and the like of the natural language command in the voice form, and conversion into the natural language command in the text form.

[0579] Step 205: performing intention recognition on the natural language command to obtain a first classification label;

[0580] Exemplarily, the first classification label corresponds to a first intention indicating a commanding intention for a non-player role in one intention dimension; for example, at least one of the following: indicating which non-player role of a plurality of non-player roles is commanded by the command text, indicating whether a virtual activity performed by the non-player role is related to a virtual attack, and indicating a type of virtual activity performed by the non-player role.

[0581] Step 206: performing entity recognition on the natural language command to obtain a target entity;

[0582] The target entity is determined from the virtual environment in combination with environment perception information, and the environment perception information includes information perceived by at least one of the master virtual role and the non-player role from the virtual environment. The target entity can be an entity currently seen by the player, or an entity historically seen by the player.

[0583] Step 207: in response to the natural language command, controlling the NPC to perform a virtual activity according to the first intention corresponding to the first classification label, or controlling the NPC to perform a virtual activity associated with the target entity, or controlling the NPC to perform a virtual activity associated with the target entity according to the first intention corresponding to the first classification label;

[0584] Exemplarily, the NPC is controlled to perform a virtual activity according to the first classification label and / or the indication of the target entity.

[0585] Step 208: broadcasting feedback information of the NPC;

[0586] Exemplarily, the game application program can generate corresponding feedback information in real time according to the environmental perception information of the non-player character, and broadcast the feedback information. For example, the game application program can generate corresponding feedback information in real time according to the environmental perception information triggering the broadcasting condition when the environmental perception information of the non-player character triggers the broadcasting condition; or the game application program can generate corresponding feedback information in real time according to the environmental perception information of the non-player character when the non-player character or the master virtual character triggers the broadcasting condition.

[0587] The preprocessing stage of step 201 can be implemented as follows:

[0588] Substep 1: obtaining attribute text of a scene object in the virtual scene;

[0589] Exemplarily, the attribute text is used to introduce the inherent attribute of the scene object in the virtual scene; on the one hand, the attribute text realizes the description of the scene object in the text mode, and provides semantic information for the prediction of the visual label of the scene object while describing the scene object. On the other hand, the attribute text is the description of the scene object in the virtual scene, and in the case of a large number of scene objects in the virtual scene and the reuse of object models, the inherent attribute of the scene object in the virtual scene can be described more accurately.

[0590] In an optional implementation, the attribute text includes at least one of a name of the scene object in the virtual scene and a size of the scene object in the virtual scene.

[0591] Substep 2: obtaining appearance images of the scene object in the virtual scene;

[0592] Exemplarily, the appearance images are used to describe the style of the scene object; the appearance images carry the appearance style such as color, texture, shape and mutual position relationship between subparts of the scene object in the picture mode, and can comprehensively describe the scene object from the picture mode (or visual mode). In an optional implementation, the appearance images of the scene object include images obtained by observing the scene object from at least two perspectives.

[0593] Substep 3: calling a multi-modal model to perform prediction on the attribute text and the appearance images of the scene object, to obtain a visual label of the scene object;

[0594] Exemplarily, the multi-modal model has the capability of performing model prediction on text information and picture information of different modalities. In the embodiment, the input parameters of the multi-modal model are the attribute text and the appearance image of the scene object; the multi-modal model predicts the visual label of the scene object from two modalities of the text modality and the picture modality, and the visual label is used to describe the visual features of the scene object in at least one dimension.

[0595] Optionally, the multi-modal model comprises a visual question and answer model; the question sentence carries the attribute text of the scene object; the question sentence is used to guide the visual question and answer model to convert the appearance image into the visual label of the scene object, and at the same time, the question sentence provides the supplementary information of the scene object in the form of text.

[0596] The expected information of the scene object is obtained;

[0597] Exemplarily, the expected information is used to indicate the description dimension of the scene object expected in the visual label, and / or the expected format of the visual label; in one example, the expected information used to indicate the description dimension of the scene object expected in the visual label includes but is not limited to at least one of the type description, the material, the transparency, the color, the surface feature, and the shape of the scene object. In another example, the expected information used to indicate the expected format of the visual label is at least one of the following: Comma Separated Values (CSV), JavaScript Object Notation (JSON), and eXtensible Markup Language (XML).

[0598] The question sentence of the appearance image is constructed according to the expected information and the attribute text;

[0599] Exemplarily, the first subpart in the question sentence is the supplementary introduction information of the scene object, which carries the attribute text of the scene object; the second subpart in the question sentence is the answer guiding sentence for the visual question and answer model, which carries the expected information.

[0600] Optionally, the spatial data set comprises a matching label of the scene object in the virtual scene; accordingly, the visual label of the scene object is subjected to paralinguistic rewriting to obtain a matching label conforming to the spoken language.

[0601] Exemplarily, as introduced above, the visual label is used to describe the visual features of the scene object in at least one dimension, and the appearance image of the scene object presents rich visual features of the scene object. However, the description of the scene object in the natural language spoken expression cannot cover the visual features of the scene object in each dimension. The purpose of performing the quasi-spoken rewriting on the visual label of the scene object is to obtain a matching label that is more similar to the spoken expression. It can be understood that performing the quasi-spoken rewriting can be deleting part of the content in the visual label, or changing the visual label to a label with the same semantics but different text expression.

[0602] In an optional implementation, the quasi-spoken rewriting is performed by calling a natural language model; the visual label of the scene object is input into the natural language model to predict a matching label conforming to the spoken expression of the natural language. The quasi-spoken rewriting on the visual label is implemented based on calling the natural language model. Exemplarily, the natural language model carries prior knowledge of the spoken expression of the natural language. For example, the natural language model is implemented as a large language model (LLM).

[0603] Optionally, the spatial data set further includes spatial information of the scene object, and correspondingly further includes:

[0604] The spatial position of the scene object in the virtual scene is obtained, and the spatial position of the scene object is determined as auxiliary information of the visual label of the scene object.

[0605] Exemplarily, the spatial position of the scene object in the virtual scene is used to indicate the deployment situation of the scene object in the virtual scene, and the spatial information is used to indicate the position, size, and the like of the scene object in the virtual scene after the scene object is deployed in the virtual scene.

[0606] Further, at least one of the coordinate position, the orientation information, the bounding box information, and the cover point information of the scene object in the virtual scene is obtained;

[0607] The coordinate position (Location) is used to indicate the position of the scene object in the virtual scene, such as the coordinate information of the center point or the preset point of the scene object in the virtual scene; the orientation (Rotation) is used to indicate the direction faced by the scene object in the virtual scene, such as the direction faced by the front of the scene object in the virtual scene; the bounding box (Bounding Box) is used to indicate the size of the scene object in the virtual scene; and the cover point (Cover points) is used to indicate the recommended position point of the virtual character when the virtual character approaches the scene object, so that the scene object can mask the virtual character.

[0608] The intent recognition stage of step 205 can be implemented as:

[0609] Sub-step 4: calling at least one hierarchical prediction network to perform intent recognition on the natural language command to obtain a first classification label of the natural language command instructing the NPC;

[0610] For example, the hierarchical prediction network has the ability to predict the classification label corresponding to the command text. For example, the hierarchical prediction network includes at least two sub-networks, an upper network and a lower network, which are cascaded, and the lower network further performs prediction of the classification label according to the prediction result output by the upper network. The at least two sub-networks are constructed based on a tree structure of multiple classification labels, and the hierarchical prediction network constructs at least two sub-networks of corresponding hierarchical structure for the tree structure of classification labels, and divides the prediction task of a variety of classification labels into prediction sub-tasks performed by the at least two sub-networks, so as to reduce the complexity of classification prediction of each sub-network.

[0611] In an optional implementation, each hierarchical prediction network in the at least one hierarchical prediction network is configured to predict a command text instructing a non-player character in an intent dimension; the intent dimension includes at least one of a subject dimension, a semantic dimension, and a behavior dimension. For example, in the at least one hierarchical prediction network, an i-th sub-network in each hierarchical prediction network is configured to predict a first-level behavior label of the command text in an intent dimension, and an i+1-th sub-network in the hierarchical prediction network is configured to predict a second-level behavior label corresponding to the first-level behavior label in the command text, where i is a positive integer; a high-level behavior label (such as a first-level behavior label) predicted by a high-level sub-network (such as an i-th sub-network) in the hierarchical prediction network includes multiple sub-labels, or subordinate lower-level labels; and a corresponding lower-level sub-network (such as an i+1-th sub-network) needs to be called to further perform classification prediction (such as to predict a second-level behavior label). The prediction task of a variety of classification labels is divided into prediction sub-tasks performed by the at least two sub-networks, so as to reduce the complexity of classification prediction of each sub-network.

[0612] Further, the intent dimension includes a subject dimension, and the hierarchical prediction network includes a hierarchical structure of a subject prediction network, and the subject prediction network has the ability to predict a subject type in the natural language command; the subject type is used to indicate the identity of the NPC instructed by the natural language command; for example, the subject prediction network is configured to predict which of a plurality of non-player characters is instructed by the command text, and the number of non-player characters instructed by the command text can be one or more.

[0613] And / or, the intent dimension comprises a semantic dimension, and the hierarchical prediction network comprises a hierarchical semantic prediction network, the semantic prediction network having a capability of predicting a semantic type in the natural language command; the semantic type being indicative of a control manner of the NPC for initiating a virtual attack; illustratively, the subject prediction network is configured to predict whether a virtual activity performed by the non-player character under the command text is related to initiating a virtual attack.

[0614] And / or, the intent dimension comprises a behavior dimension, and the hierarchical prediction network comprises a hierarchical behavior prediction network, the behavior prediction network having a capability of predicting a behavior intent type in the natural language command; the behavior intent type being indicative of a behavior manner of the NPC for performing a virtual activity. Illustratively, the behavior prediction network is configured to predict which virtual activity is performed by the non-player character.

[0615] Correspondingly, in step 207, in response to the natural language command, the NPC is controlled to perform a virtual activity according to a first intent corresponding to the first classification label;

[0616] Illustratively, the first intent corresponding to the first classification label is indicative of a command intent for the non-player character in one intent dimension; such as at least one of the following: indicative of which non-player character is the non-player character under the command text, indicative of whether a virtual activity performed by the non-player character is related to initiating a virtual attack, and indicative of which virtual activity is performed by the non-player character. The NPC is controlled to perform a virtual activity according to the indication of the first classification label.

[0617] The target entity recognition process for step 206 can be implemented as follows:

[0618] Sub-step 5: querying the target entity from an entity set of the virtual environment according to target entity information of the target entity indicated by the natural language command and the environment perception information;

[0619] The target entity is determined from the virtual environment in combination with the environment perception information, and the environment perception information comprises information perceived by the host virtual character and / or the non-player character from the virtual environment.

[0620] Since the target entity information in the natural language command is usually ambiguous, for example, the natural language command can be "move to the red truck", and the target entity information therein is "red truck", and there can be many red trucks in the virtual environment, and the target entity cannot be accurately determined from the virtual environment according to the target entity information in the natural language command.

[0621] Therefore, the embodiment of the present application provides a method for determining a target entity by combining target entity information and environment perception information. In the above example, the "red truck" expressed by the player in the natural language command should be the red truck that the player can see. Therefore, by combining the visual field range of the virtual character controlled by the player, the red truck located in the visual field range of the virtual character controlled by the player can be screened from the multiple red trucks in the virtual environment, and the red truck is the target entity indicated by the player in the natural language command.

[0622] Since the natural language command is issued by the player based on the perception of the virtual environment, in order to accurately identify the target entity indicated in the natural language command, the game application program identifies the target entity in the natural language command by combining the environment perception information of the player when issuing the natural language command. Based on the perception of the virtual environment by the player, the entity in the virtual environment that the player can perceive and that is closest to the target entity information is inferred as the target entity.

[0623] For example, the virtual environment includes multiple candidate entities matching the natural language command, and the target entity is an entity screened from the multiple candidate entities based on the environment perception information of the virtual character controlled by the player or the non-player character. For example, the game application program first selects multiple candidate entities matching the target entity information in the natural language command from the entity set, and then selects an entity that can be perceived by the virtual character controlled by the player or the non-player character from the multiple candidate entities as the target entity according to the environment perception information.

[0624] For example, the environment perception information of the virtual character controlled by the player can include: a picture or entity that can be seen by the virtual character controlled by the player when observing the virtual environment, an environmental sound that can be heard by the virtual character controlled by the player, a source direction of the environmental sound, a type and sound size of the environmental sound, perception information obtained by the virtual character controlled by the player by using a perception skill (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and perception information obtained by the virtual character controlled by the player by using a sensing device (for example, a sensing signal of a sensor, a positioning signal of a positioning prop, etc.).

[0625] The environment perception information of the non-player character can include: a picture or entity that can be seen by the non-player character when observing the virtual environment, an environmental sound that can be heard by the non-player character, a source direction of the environmental sound, a type and sound size of the environmental sound, perception information obtained by the non-player character by using a perception skill (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and perception information obtained by the non-player character by using a sensing device (for example, a sensing signal of a sensor, a positioning signal of a positioning prop, etc.).

[0626] It should be noted that the environment perception information can include information that the host virtual character and the non-player character perceive in real time when the natural language command is received, or can include information that the host virtual character and the non-player character perceive historically before the natural language command is received. That is, the target entity can be an entity that the player currently sees, or can be an entity that the player has historically seen. Therefore, the game application needs to combine the real-time environment perception information and the historical environment perception information of the host virtual character and / or the non-player character to identify the target entity.

[0627] For example, when the natural language command is "Let's go back to the hotel we just passed by", the game application needs to query the hotel that the host virtual character has passed by according to the historical environment perception information of the host virtual character.

[0628] In an optional embodiment, the client queries the target entity from the entity set of the virtual environment according to the target entity information indicated by the natural language command and the environment perception information. In another optional embodiment, the client reports the natural language command to the server, and the server queries the target entity from the entity set of the virtual environment according to the target entity information indicated by the natural language command and the environment perception information.

[0629] Optionally, the game application infers the target entity according to the target entity information, the environment perception information, and the preprocessed virtual environment data. The preprocessed virtual environment data includes the entity set of the virtual environment.

[0630] The entity set includes entity information of each entity in the virtual environment. The entity information includes at least one of the following: name, type, location, feature, text label, embedding vector of the text label, image, embedding vector of the image.

[0631] The text label can be at least one word obtained by tokenizing the feature (for example, a description text of the appearance of the entity), and the embedding model is called to obtain the embedding vector corresponding to each text label.

[0632] The image can include an image obtained by observing a three-dimensional model of the entity from at least one direction, for example, the image can include three views of the entity. The embedding vector of the image is an embedding vector of an image feature label. The image feature label is obtained by calling a multi-modal model to perform feature recognition based on the image and the text label of the entity.

[0633] Exemplarily, the entity set includes first-class entities and second-class entities, the first-class entities include at least one of a region and a building, and the second-class entities include physical objects. The text label of the first-class entity is manually marked, for example, the text label of a certain room of a certain building is manually marked as: a certain region, a certain building, a certain floor, and a room. The text label of the second-class entity is automatically generated by calling a large language model.

[0634] The method for generating the text label of the second-class entity can include: obtaining at least one image of the entity, inputting the at least one image into the large language model to obtain a description text of the entity; performing word segmentation on the description text to obtain at least one text label of the entity, each text label including one word obtained after word segmentation. Performing a vectorization operation on the text label to obtain an embedding vector corresponding to the text label.

[0635] The embedding vector is a high-dimensional vector data of the text label, and the embedding vector of the text label is generated in advance, so that it is not necessary to repeatedly generate the vector of the text label of each entity when matching the target entity based on the target entity information from the entity set, and the matching efficiency of the target entity information and the entity information can be improved.

[0636] It should be noted that, instead of directly taking the description text as the text label, the word segmentation result of the description text is taken as the text label, because the description text usually has different lengths, and for a longer description text, the embedding vector thereof has a poor effect in actual search and comparison of similarity. For example, when searching for a “truck” using “a blue small truck with rusted body”, the “a blue small truck with rusted body” is probably ranked after “a car” in the recall result, but actually, for the search of “truck”, no matter how complex the additional description is, the “truck” should be searched first, and therefore, when performing the similarity search, the comparison should be performed on individual feature words instead of the description text.

[0637] In addition, in order to further extract the image features of the entity, the method provided in the embodiments of the present application can further extract hidden visual features in the image of the entity based on the text label and the image of the entity. Taking the first entity as an example, the method includes: obtaining at least one perspective image of the first entity; obtaining at least one text label of the first entity; calling a multi-modal model to extract image features of the first entity based on the at least one perspective image and the at least one text label of the first entity, to obtain an image embedding vector of the first entity.

[0638] Since the text label is manually annotated or generated by a large language model based on human language characteristics, the text label extracted based on human language habits may ignore some features of the entity. For example, when manually annotating the text label for an oil drum, the more noticeable text labels such as "metal" and "rust" may be annotated, and the detailed features such as "rust", "blue", "yellow", and "peeling paint in the lower right corner" may be ignored. Therefore, the method provided in the embodiments of the present application also uses a multi-modal model to extract more image features from the text label and the image of the entity based on the multi-modal model, and performs image similarity matching with the image of the target entity based on the image features to improve the recognition accuracy of the target entity.

[0639] The training method of the multi-modal model can be: inputting the sample entity image and the sample label into the pre-trained multi-modal model, fine-tuning the pre-trained multi-modal model according to the loss of the predicted label output by the pre-trained multi-modal model and the sample label, to obtain a multi-modal model capable of outputting image feature labels based on input text labels and images. The image feature label includes more detailed description text extracted from the image. For example, the text description of "a metal oil drum" can only be associated with the two text labels "metal" and "oil drum", but the Clip (multi-modal) model can also identify hidden visual information in the image of the oil drum, such as the image feature labels "rust", "blue", and "yellow". These image features are not present in the text label, so combining the Clip model for image feature search can further improve the accuracy of entity search.

[0640] In an optional embodiment, the game application program can query the target entity by the following method.

[0641] (1) Analyzing the natural language command to obtain target entity information; the target entity information includes at least one of the following: entity type, entity name, entity position, and entity feature.

[0642] Optionally, the game application program calls a large language model to analyze the natural language command to obtain the target entity information of the target entity. The embedding model is called to perform vectorization processing on the target entity information to obtain a target embedding vector of the target entity information.

[0643] For example, the natural language command is "come here in front of the blue truck", and the target entity information analyzed by the large language model includes the position information "in front of" and the entity name "blue truck".

[0644] Exemplarily, a named entity recognition technique in the field of natural language processing can also be called to extract target entity information from the natural language command. The named entity recognition technique is used to identify entities with specific meanings in text, for example, to identify names, place names, adjectives, and the like.

[0645] For example, the natural language command is "find a paper box behind the red sofa on the first floor of the motel". Using the named entity recognition technique, the building name "motel first floor", the article name "sofa" and "paper box", the adverb "behind", and the adjective "red" can be identified. After logical construction, a hierarchical scene query call form with search type, content, adverb, and constraint information is formed. Among them, the search type is used to narrow the data retrieval range based on similarity matching, the search content is the specific entity description, and the adjective is combined with the description text for query. The floor constraint is determined by the player's location, which needs to narrow the range up and down in the indoor scene, thereby avoiding finding entities that are not visible across floors.

[0646] (2) Calculate the similarity of the target entity information and the entity information of each entity in the entity set.

[0647] Exemplarily, the entity set includes entity information of at least one entity, and the entity information of each entity can include at least one of a text embedding vector of a text label and an embedding vector of an image feature. The target entity information can be calculated with the text embedding vector and the image embedding vector, respectively, and the entity with higher similarity is determined as the target entity.

[0648] The methods of similarity matching with the text embedding vector and the image embedding vector are given below, respectively.

[0649] 1) The similarity includes the text similarity of the target entity information and the text embedding vector.

[0650] Taking the first entity in the entity set as an example, the game application program (client or server) tokenizes the target entity information to obtain at least one target entity label; converts the at least one target entity label into at least one target embedding vector; obtains the entity information of the first entity, which includes a text embedding vector converted based on the text label of the first entity; calculates the text parent similarity of the at least one target embedding vector and the text embedding vector, respectively, to obtain at least one text parent similarity corresponding to the at least one target embedding vector, respectively; and determines the sum of the at least one text parent similarity as the text similarity of the target entity information and the entity information of the first entity.

[0651] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to text embedding vector 1, and entity 2 corresponds to text embedding vector 2. Then, similarity 1 of target entity label 1 and text embedding vector 1 is calculated, similarity 2 of target entity label 2 and text embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the text similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and text embedding vector 2 is calculated, similarity 4 of target entity label 2 and text embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the text similarity of target entity information and entity 2.

[0652] For example, the entity information of the first entity includes at least one text embedding vector, and the at least one target embedding vector includes a first target embedding vector. The game application calculates text sub-similarities of the first target embedding vector and the at least one text embedding vector respectively, and obtains at least one text sub-similarity. The highest value in the at least one text sub-similarity is determined as a text parent-similarity corresponding to the first target embedding vector.

[0653] For example, the number of target entity labels is one: target entity label 1. The number of entities in the entity set is also one: entity 1. Entity 1 corresponds to text embedding vector 1 and text embedding vector 3. Then, similarity 1 of target entity label 1 and text embedding vector 1 is calculated, similarity 5 of target entity label 1 and text embedding vector 3 is calculated, and the greater value of similarity 1 and similarity 5 is determined as the similarity of target entity information (target entity label 1) and entity 1.

[0654] For example, for the input target entity information "blue car", first perform word segmentation processing on the target entity information to decompose into target entity labels "blue" and "car", representing two features of the query target entity. Then, the similarity of each feature with each entity in the entity set is calculated, the maximum value of the similarity of each feature is taken, and finally the sum of the similarity scores corresponding to all query features is taken, representing the text similarity of the target entity information and the queried entity.

[0655] For example, the target entity information includes two target entity labels "blue" and "car", and the two target entity labels are converted into two target embedding vectors. Then, the entity information in the entity set is obtained, for example, the entity set includes entity 1 and entity 2, the text labels of entity 1 include "metal", "old", "car", "truck", "blue", and "damaged", and the text labels of entity 2 include "metal", "scratch", "car", "damaged", "money truck", and "black".

[0656] The similarity between the target entity information and the text of entity 1 is calculated: first, the similarity between the target entity label "blue" and each text label of entity 1 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "blue" and the text label "blue" of entity 1 is 1 (1 is the maximum value); then, the similarity between the target entity label "car" and each text label of entity 1 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "car" and the text label "car" of entity 1 is 1, and then the sum of the two similarities of the target entity labels "car" and "blue" is taken to obtain the final similarity between the target entity information and the entity 1 as 2.

[0657] Similarly, the similarity between the target entity information and the text of entity 2 is calculated: first, the similarity between the target entity label "blue" and each text label of entity 2 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "blue" and the text label "black" of entity 1 is 0.91; then, the similarity between the target entity label "car" and each text label of entity 1 is calculated, and the maximum value is taken, for example, the similarity between the target entity label "car" and the text label "car" of entity 2 is 1 (1 is the maximum value), and then the sum of the two similarities of the target entity labels "car" and "blue" is taken to obtain the final similarity between the target entity information and the entity 2 as 1.91.

[0658] It can be seen that the similarity between the target entity information and entity 1 is 2, which is higher than the similarity between the target entity information and entity 2, which is 1.91.

[0659] 2) The similarity includes an image similarity between the target entity information and the image embedding vector.

[0660] Taking the first entity in the entity set as an example, the game application (client or server) tokenizes the target entity information to obtain at least one target entity label; converts the at least one target entity label into at least one target embedding vector; obtains the entity information of the first entity, and the entity information includes an image embedding vector, which is an embedding vector extracted based on the image of the first entity; calculates the image parent similarity between the at least one target embedding vector and the text embedding vector respectively to obtain at least one image parent similarity corresponding to the at least one target embedding vector respectively; and determines the sum of the at least one image parent similarity as the image similarity between the target entity information and the entity information of the first entity.

[0661] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to image embedding vector 1, and entity 2 corresponds to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the image similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the image similarity of target entity information and entity 2.

[0662] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to image embedding vector 1, and entity 2 corresponds to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the image similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the image similarity of target entity information and entity 2.

[0663] For example, the number of target entity labels is two, target entity label 1 and target entity label 2, and the number of entities in the entity set is also two, entity 1 and entity 2, entity 1 corresponds to image embedding vector 1, and entity 2 corresponds to image embedding vector 2. Then, similarity 1 of target entity label 1 and image embedding vector 1 is calculated, similarity 2 of target entity label 2 and image embedding vector 1 is calculated, and the sum of similarity 1 and similarity 2 is determined as the image similarity of target entity information and entity 1. Similarity 3 of target entity label 1 and image embedding vector 2 is calculated, similarity 4 of target entity label 2 and image embedding vector 2 is calculated, and the sum of similarity 3 and similarity 4 is determined as the image similarity of target entity information and entity 2.

[0664] In an optional embodiment, the calculation of the image similarity can use a target image embedding vector, that is, the target image embedding vector is used instead of the target embedding vector in the above method. The target image embedding vector can be obtained by the following method: tokenizing the target entity information to obtain at least one target entity label; inputting the at least one target entity label into a multi-modal model to obtain the target image embedding vector.

[0665] 3) The similarity includes a text similarity and an image similarity.

[0666] In the case where the first entity and the target entity correspond to a text similarity and an image similarity, the game application determines the average of the text similarity and the image similarity as the similarity between the first entity and the target entity.

[0667] Alternatively, in the case where the first entity and the target entity correspond to a text similarity and an image similarity, the greater value of the text similarity and the image similarity is determined as the similarity between the first entity and the target entity.

[0668] Alternatively, in a case where the first entity and the target entity correspond to a text similarity and an image similarity, a sum of the text similarity and the image similarity is determined as the similarity of the first entity and the target entity.

[0669] (3) determining the target entity from the entity set according to the similarity and the environmental perception information.

[0670] In an optional embodiment, the game application first performs similarity searching according to the target entity information to obtain at least one candidate entity with a higher similarity, and sorts the candidate entities according to the similarity; then performs filtering and sorting of the perception based on the position, orientation, direction, etc. of the master virtual character and / or the non-player character, and finally screens the target entity.

[0671] For example, the game application determines a target perception range according to the target entity information and the environmental perception information; and screens the target entity from the entities perceived within the target perception range according to the similarity.

[0672] Alternatively, the game application determines x entities with the highest similarity from the entity set according to the similarity to form a candidate entity list; determines the target entity from the entities within the target perception range according to the target entity information and the environmental perception information. x is a positive integer.

[0673] For example, the game application determines the entity with the highest similarity within the perception range as the target entity.

[0674] It should be noted that when there are multiple entities with equal similarity within the target perception range, the multiple entities can be sorted according to the distance between the entities and the master virtual character. For example, in a case where the number of entities with the highest similarity within the perception range is at least two, the entity with the highest similarity within the perception range and closest to the master virtual character is determined as the target entity.

[0675] In an optional embodiment, the game application determines the entity with the highest similarity within the perception range and higher than a threshold value as the target entity. For example, the number of perception ranges is at least one. For example, the perception range includes at least one of the following: the visual range of the master virtual character, the auditory range of the master virtual character, the perception range of the master virtual character for perceiving virtual props, the visual range of the non-player character, the auditory range of the non-player character, the perception range of the non-player character for perceiving virtual props.

[0676] For example, the game application can sequentially traverse the at least one perception range described above, or the game application can sequentially traverse the at least one perception range determined according to the natural language command. For example, the natural language command is "Can you see the truck in front of you? Move to the truck", the game application can traverse the visual field range of the master virtual role first, and then traverse the visual field range of the non-player character, and match the entity with the highest similarity and higher than the threshold.

[0677] For example, when there are at least two perception ranges, the traversal order of the at least two perception ranges can be preset, or the traversal order of the at least two perception ranges can be determined according to the intention indicated in the natural language command. For example, when the intention is to perform a search for an item, the visual field range is traversed first. When the intention is to find a combat scene, the auditory range is traversed first.

[0678] If there are multiple levels of nested instructions in the natural language command, the first target entity is searched first, and then the next target entity is searched according to the position of the previous target entity. In the case where the natural language command includes a second target entity determined based on the position of the first target entity, the first target entity is queried from the entity set of the virtual environment according to the first target entity information of the first target entity indicated by the natural language command and the environmental perception information; and then the second target entity is queried from the entity set of the virtual environment according to the second target entity information of the second target entity, the first target entity and the environmental perception information.

[0679] In response to the natural language command, the NPC performs a virtual activity associated with the target entity.

[0680] For example, the behavior intention of the natural language command is recognized by calling a large language model, a control instruction or a control instruction sequence is generated according to the behavior intention, and the non-player character is controlled to complete an activity based on the target entity according to the control instruction or the control instruction sequence.

[0681] The NPC feedback information broadcast in step 208 can be implemented as:

[0682] Sub-step 6: Broadcast the feedback information of the non-player character, and the feedback text corresponding to the feedback information is non-fixed text generated based on the environmental perception information of the non-player character;

[0683] Exemplarily, the game application program can generate corresponding feedback information in real time according to the environmental perception information of the non-player character, and play the feedback information. For example, the game application program can generate corresponding feedback information in real time according to the environmental perception information triggering the playback condition when the environmental perception information of the non-player character triggers the playback condition; or the game application program can generate corresponding feedback information in real time according to the environmental perception information of the non-player character when the non-player character or the master virtual character triggers the playback condition.

[0684] Exemplarily, the feedback text is obtained by reasoning based on the static entity data of the three-dimensional virtual environment and the dynamic environmental perception information of the non-player character by calling a large language model. The large language model generates different feedback texts for different perception situations of the non-player character.

[0685] Exemplarily, the feedback text can also be obtained by reasoning based on the static entity data of the three-dimensional virtual environment and the dynamic environmental perception information of the non-player character by the large language model, and the feedback text conforms to the personality characteristics of the non-player character. For example, when the non-player character is a robust uncle image, the feedback text can adopt a bold tone; when the non-player character is a reporter, the feedback text can adopt a news report tone.

[0686] It should be noted that due to the randomness of the generation result of the large language model, in the same scene, when the environmental perception information of the non-player character is the same, the generated feedback text can also be different.

[0687] The environmental perception information can include information perceived by the non-player character in real time, and can also include information perceived by the non-player character in history.

[0688] Exemplarily, the feedback information is obtained by reasoning by the game application program based on the environmental perception information of the non-player character and the static entity data of the three-dimensional virtual environment. The feedback information can be a voice form of playback, or a text form of playback. The static entity data includes data of entities that are relatively constant in the three-dimensional virtual environment, for example, includes model information of various buildings, terrains, and vehicles in the three-dimensional virtual environment.

[0689] Exemplarily, the feedback information can include at least one of the following: immediate feedback, execution feedback, and dynamic feedback. The immediate feedback and the execution feedback are feedback information generated for the natural language command, and the dynamic feedback is feedback information generated spontaneously by the non-player character.

[0690] 1. The immediate feedback is feedback information generated immediately when the natural language command is received. The immediate feedback can include reply playback and immediate playback. The reply playback is used to restore the query in the natural language command; the immediate playback is used to feed back the receiving situation of the natural language command.

[0691] For example, the client can instruct the non-player character to act in the three-dimensional virtual environment according to the intention of the natural language command. Then, the client can generate instant feedback corresponding to the natural language command when receiving the natural language command, and the instant feedback is used to indicate that the natural language command has been received, or the instant feedback is used to reply to the natural language command immediately.

[0692] In an optional embodiment, the feedback information includes a reply broadcast. The reply broadcast includes a reply to the inquiry natural language command generated for the inquiry natural language command proposed by the player. For example, the inquiry natural language command includes an inquiry about the location of the target entity, and the corresponding reply broadcast should include the query result of the location of the target entity.

[0693] The terminal broadcasts the reply broadcast of the non-player character in the case that the behavior intention of the received natural language command is inquiry; the reply broadcast includes the reply content to the natural language command generated according to the environmental perception information of the non-player character to the three-dimensional virtual environment.

[0694] The natural language command is used to inquire information from the non-player character. This natural language command can be referred to as an inquiry natural language command. The inquiry natural language command is a natural language command with an inquiry intention. The inquiry natural language command contains at least one question.

[0695] The natural language command is a command indication conveyed using human natural language. The terminal receives the natural language command issued by the player, and instructs the non-player character according to the intention corresponding to the natural language command. For example, the terminal receives the voice audio of the player, converts the voice audio into text, and obtains the natural language command. Or, the terminal receives the text of the natural language command input by the player.

[0696] The player can use spoken language or written language to issue natural language commands. The game application program calls a large language model to understand the natural language command, extracts the intention expressed in the natural language command, and generates a control instruction according to the intention to instruct the activity of the non-player character.

[0697] For example, the natural language command can be “come to me”, and the game application program can parse the intention of the natural language command as “control the non-player character to move to the location of the master virtual character”. Then, the game application program obtains the location of the master virtual character, generates a navigation route for the non-player character to move to the location of the master virtual character, and controls the non-player character to move to the location of the master virtual character according to the navigation route.

[0698] For example, the natural language command can be "go pick up the treasure chest", and the game application can parse the intention of the natural language command as "control the non-player character to move to the location of the treasure chest and pick up the treasure chest". The game application can obtain the location of the treasure chest, control the non-player character to move to the location of the treasure chest, and perform the operation of picking up the treasure chest after arriving at the location of the treasure chest.

[0699] For example, in the case that the inquiry natural language command includes an intelligence inquiry to the target entity, based on the perception of the target entity by the non-player character, a first reply broadcast is broadcasted; the first reply broadcast includes the intelligence perception result of the target entity by the non-player character.

[0700] Alternatively, in the case that the inquiry natural language command includes a location inquiry to the target entity, based on the perception of the target entity by the non-player character, a second reply broadcast is broadcasted, and the second reply broadcast includes a description of the location of the target entity.

[0701] For example, the large language model is called to parse the received natural language command to obtain the intention of the natural language command. When the intention is an inquiry type, it can be determined that the natural language command is an inquiry natural language command. When the natural language command is an inquiry natural language command, the large language model can infer according to the static entity data and the environmental perception information according to the intention of the natural language command to obtain the reply text. That is, the large language model is called to parse the received natural language command, and the intention of the natural language command is an inquiry, and the reply text of the natural language command. The game application calls the text-to-speech service according to the processing logic of the inquiry intention, converts the reply text into a reply broadcast, and the reply broadcast is an audio form of broadcast. The reply audio is sent to the client for broadcast.

[0702] For example, when the large language model identifies the intention of the natural language command as an inquiry, since there can be thousands of questions from users, it is impossible to store the reply to each question locally on the client side. Therefore, the game application calls the text-to-speech service to generate a reply broadcast in real time according to the reply text returned by the large language model, so that the game application can generate a corresponding reply broadcast in response to the player's question in real time.

[0703] For example, the game application (client or server) calls the large language model to parse the inquiry natural language command, infers the reply text based on the static environment data and the environmental perception information of the three-dimensional virtual environment, generates human voice based on the reply text, and obtains the reply broadcast.

[0704] For example, the client reports the received inquiry natural language command to the server, the server invokes the large language model to analyze the inquiry natural language command, infers the reply text based on the static environment data and the environment perception information of the three-dimensional virtual environment, generates the human voice audio based on the reply text to obtain the reply broadcast, and sends the reply broadcast to the client.

[0705] 2. The execution feedback is generated when the natural language command is executed according to the intention indicated by the natural language command. The execution feedback can also be referred to as execution broadcast.

[0706] For example, when the natural language command is used to instruct the non-player character to execute the target task, the client can generate the execution feedback during the execution of the target task by the non-player character, and the execution feedback is used to indicate the execution state of the target task; or the client can generate the execution feedback after the non-player character completes the target task, and the execution feedback is used to indicate the execution result of the target task.

[0707] In an optional embodiment, the feedback information includes feedback broadcast (which can also be referred to as execution feedback). The feedback broadcast includes the broadcast of the task execution state and the task execution result generated for the task natural language command proposed by the player. For example, the task natural language command includes moving to the target entity, and the corresponding feedback broadcast can include moving to the target entity, or having arrived at the target entity.

[0708] The terminal broadcasts the feedback broadcast in the case that the behavior intention of the received natural language command is to execute the task; the feedback broadcast includes the execution state after the non-player character executes the natural language command according to the environment perception information of the three-dimensional virtual environment.

[0709] The natural language command is used to instruct the non-player character to execute the task. The natural language command can be referred to as a task natural language command. The task natural language command includes the description text of the task, for example, the description text of the task can include at least one of the following: task name, action required to be executed by the non-player character, method for executing the task, task target, and location of the task target.

[0710] For example, in the case that the intention of the task natural language command includes searching for the target entity, the first feedback broadcast is broadcasted, and the first feedback broadcast includes the search result of the target entity by the non-player character.

[0711] Or, in the case that the intention of the task natural language command includes using the virtual prop, the second feedback broadcast is broadcasted, and the second feedback broadcast is used to indicate that the virtual prop is ready.

[0712] Or, in the case where the intention of the task natural language command includes using a virtual prop, a third feedback broadcast is broadcasted, the third feedback broadcast being used to indicate a use result of the virtual prop.

[0713] Or, in the case where the intention of the task natural language command includes controlling the non-player character to move, a fourth feedback broadcast is broadcasted, the fourth feedback broadcast being used to indicate a movement result of the movement.

[0714] Or, in the case where the intention of the task natural language command includes performing an interaction with a target entity, a fifth feedback broadcast is broadcasted, the fifth feedback broadcast being used to indicate at least one of a search result of the non-player character on the target entity, an interaction result of the non-player character with the target entity.

[0715] Or, in the case where the intention of the task natural language command includes an item interaction, a sixth feedback broadcast is broadcasted, the sixth feedback broadcast being used to indicate an item interaction result of the non-player character performing the item interaction.

[0716] For example, the large language model is called to parse the received natural language command to obtain the intention of the natural language command. When the intention is a task, it can be determined that the natural language command is a task natural language command. When the natural language command is a task natural language command, the large language model can infer according to the intention of the natural language command, the static entity data and the environmental perception information to obtain a task instruction sequence, the task instruction sequence being used to control the non-player character to execute the task. The game application program receives the task instruction sequence, and controls the non-player character to execute the task according to the task instructions in the task instruction sequence. During the control of the non-player character according to the task instructions, the corresponding execution situation can be fed back according to the feedback broadcast logic corresponding to different task instructions.

[0717] For example, since the instructions executable by the non-player character in the three-dimensional virtual environment are limited and traversable, the client can locally store the feedback broadcasts corresponding to each character instruction. When the non-player character executes the corresponding instruction, the client can read the local feedback broadcast to perform voice broadcast.

[0718] Or, when there is variable content in the feedback broadcast corresponding to a certain task instruction, for example, the target entity in the feedback broadcast is variable, or the position information in the feedback broadcast is variable; the client can also generate a broadcast voice of the variable content according to the target entity or the position information returned by the large language model, and splice the broadcast voice of the variable content with the broadcast voice of the non-variable content stored locally to obtain the final feedback broadcast.

[0719] For example, when the task natural language command is "find the nearby treasure chest", calling the large language model to parse the task natural language command can obtain that its intention is a task, the task target is to find a target entity, and the target entity is a treasure chest; based on the static entity data and the environment perception information, it can be inferred that there is a treasure chest in the left front of the host virtual role. In the case where the target entity can be found, the feedback report corresponding to the task target is "find aaa at xxx", where xxx is the location of the target entity, and aaa is the name of the target entity. The game application can call the online text-to-speech technology to generate the location voice according to the location information "left front" in the inference result; call the online text-to-speech technology to generate the name voice according to the name "treasure chest" of the target entity, and then splice the location voice and the name voice into the feedback report to obtain the final feedback report.

[0720] For example, the game application (client or server) calls the large language model to parse the task natural language command, infers the task execution instruction based on the static environment data and the environment perception information of the three-dimensional virtual environment, controls the non-player character to execute the task according to the task execution instruction, generates the feedback text according to the execution state of the non-player character executing the task, and generates the human voice audio based on the feedback text to obtain the feedback report.

[0721] For example, the client reports the task natural language command to the server; the server parses the task natural language command, infers the task execution instruction based on the static environment data and the environment perception information of the three-dimensional virtual environment, controls the non-player character to execute the task according to the task execution instruction, generates the feedback text according to the execution state of the non-player character executing the task, generates the human voice audio based on the feedback text to obtain the feedback report, and sends the feedback report to the client; the client receives and reports the feedback report.

[0722] 3. The dynamic feedback is the feedback information generated spontaneously according to the environment perception information of the non-player character. The dynamic feedback can also be referred to as dynamic report.

[0723] For example, the client can spontaneously generate the dynamic feedback based on the perception of the three-dimensional virtual environment without receiving the natural language command, and the dynamic feedback is used to indicate the abnormal situation found by the non-player character in the three-dimensional virtual environment. For example, the abnormal situation can include finding an enemy virtual character, finding that the state of a friendly virtual character has changed, finding a dangerous situation, finding a battle trace or a looting trace, etc.

[0724] In an alternative embodiment, the feedback information comprises dynamic announcement (may also be referred to as "dynamic feedback"). The dynamic announcement comprises announcement of a perceived abnormal situation of the non-player character. For example, when the non-player character perceives that it is under attack, the announcement is that it is under attack; when the non-player character perceives that there is a dangerous situation ahead, the announcement is that there is danger ahead.

[0725] The terminal announces the dynamic announcement when the environmental perception information of the non-player character in the three-dimensional virtual environment satisfies a dynamic announcement condition.

[0726] The dynamic announcement is an announcement generated according to an abnormal situation when the game application determines that there is an abnormal situation that needs to be announced based on the environmental perception information of the non-player character. The dynamic announcement includes reminder information of the abnormal situation. For example, the abnormal situation can include at least one of the following: discovery of an enemy virtual character, discovery of a change in the state of a friendly virtual character, discovery of a new trace.

[0727] For example, in a case where the non-player character perceives an enemy virtual character, a first dynamic announcement is announced; the first dynamic announcement includes the position of the enemy virtual character perceived by the non-player character.

[0728] Alternatively, in a case where the non-player character perceives a dangerous situation, a second dynamic announcement is announced; the second dynamic announcement is used to prompt the dangerous situation.

[0729] Alternatively, in a case where the non-player character perceives a change in the state of a friendly virtual character, a third dynamic announcement is announced; the third dynamic announcement is used to prompt the change in the state of the friendly virtual character.

[0730] Alternatively, in a case where the non-player character discovers a new trace in the three-dimensional virtual environment, a fourth dynamic announcement is announced; the fourth dynamic announcement is used to prompt the new trace. The new trace can be a battle trace and / or a looting trace.

[0731] For example, since the number of dynamic announcements that can be triggered by the non-player character in the three-dimensional virtual environment is limited and traversable, the client can locally store dynamic announcement speech, and the client can read the dynamic announcement speech locally when triggering the corresponding dynamic announcement speech to announce. Similarly, some dynamic announcement speech can have variable content, and the client can generate variable content in real time through text-to-speech technology and splice it with a dynamic announcement template to obtain the final dynamic announcement.

[0732] In an alternative embodiment, when generating the feedback information, it can be necessary to query the position of a target entity in the three-dimensional virtual environment, or to determine the target entity indicated in the natural language command from the three-dimensional virtual environment. At this time, an entity query method needs to be used to query the target entity.

[0733] Before performing the target entity query, static entity data of the three-dimensional virtual environment needs to be constructed in advance. The static entity data construction can serve both the inference of the large language model and the generation of real-time feedback text on the game application side after the inference result is returned.

[0734] In an alternative embodiment, entities in the three-dimensional virtual environment are traversed, and entity information of each entity is exported from the game engine to obtain the static entity data.

[0735] For example, for in-game scenes, an editor tool is developed to traverse the entire scene StaticMesh, special Actor traversal (such as interactive doors, which do not belong to the StaticMesh category), vegetation traversal, and export the position, orientation, and bounding box size of the entity as static entity data for subsequent tagging by the large language model and inference by the large language model.

[0736] Among them, StaticMesh is a type of static geometry resource in Unreal Engine 4 (UE4) game engine, used to represent unchangeable three-dimensional models such as buildings, props, etc. Actor is a basic object class in Unreal Engine 4 (UE4) game engine, representing an entity in the game world, such as characters, objects, light sources, etc. An actor can contain multiple components to achieve different functions.

[0737] For example, the static entity data can also include spatial data, which is obtained by manual labeling. Spatial data is more complex than static entities, as there are multiple floors and multiple layers of nesting within a floor in the space (such as a motel area containing the second floor, and the second floor area containing room 201). For such data, manual calibration is used to construct. Under the editor, add Volume to the scene, divide the required area of the entire graph and label it accordingly. Volume is a special Actor class in Unreal Engine 4 (UE4) game engine, representing a three-dimensional area with specific functions, such as trigger areas, audio areas, etc.

[0738] After obtaining the static entity data, entity or spatial queries can be performed based on the static entity data during game execution.

[0739] For example, when the player issues a natural language command, the natural language command will be accompanied by the constructed static entity data and the real-time captured runtime data (which includes non-player characters and / or environment perception information of the host virtual character) to request inference from the large language model. In addition, when the NPC needs to generate dynamic feedback based on runtime data, entity queries are also performed to generate feedback text based on the current state.

[0740] When the target entity is included in the natural language command, the game application program queries the target entity according to the current position and orientation of the host virtual character, the position and orientation of the teammates or enemies, and the target entity information of the target entity described in the natural language command, and the environmental perception information of the host virtual character and / or the non-player character, and filters the target entity matching the target entity information from at least one candidate entity within the perception range of the host virtual character and / or the non-player character.

[0741] Optionally, the natural language command is used to instruct the non-player character to perform an activity related to the target entity. For example, the target entity can be the moving destination of the non-player character, the target entity can be an object that needs to be observed by the non-player character, the target entity can be a target that needs to be attacked by the non-player character, or the target entity can be an object that needs to be interacted with by the non-player character.

[0742] The target entity is a virtual object existing in the three-dimensional virtual environment. The target entity is an entity described in the natural language command. The entity can refer to a virtual object that is fixed in the three-dimensional virtual environment, for example, the target entity can be a virtual building, a virtual terrain, a virtual vehicle, a virtual plant, a virtual prop, a virtual item, etc. The entity can also refer to a virtual object that can interact with a virtual character in the three-dimensional virtual environment, for example, the target entity can be an interaction point (e.g., a door, a window, a cabinet, a cellar, etc.), a virtual prop, a virtual character, a virtual light source, etc. in the three-dimensional virtual environment.

[0743] For example, the natural language command can be "please help me open the door", and the target entity can be "the door", and the non-player character needs to perform an activity related to the door "open the door". Or, the natural language command can be "is the kitchen safe", and the target entity can be "the kitchen", and the non-player character needs to perform an activity related to the kitchen "check whether there are enemy virtual characters or other dangerous situations in the kitchen".

[0744] Optionally, the natural language command includes target entity information of the target entity. The target entity information can be a description text of the target entity. For example, the target entity information can include at least one of the following: the name, type, location, and characteristics of the target entity. The game application program can find the target entity from a plurality of entities in the three-dimensional virtual environment based on the target entity information in the natural language.

[0745] The game application program analyzes the natural language command, extracts the target entity information therein, determines the target entity from a set of entities in the three-dimensional virtual environment based on the target entity information, and then controls the non-player character to perform an activity related to the target entity according to the intention of the natural language command.

[0746] The target entity is determined from the three-dimensional virtual environment in combination with environmental perception information, and the environmental perception information includes information perceived by at least one of the master virtual character or the non-player character from the three-dimensional virtual environment.

[0747] Since the target entity information in the natural language command is usually ambiguous, for example, the natural language command can be "move to the red truck", and the target entity information therein is "red truck", and there can be many red trucks in the three-dimensional virtual environment, and the target entity cannot be accurately determined from the three-dimensional virtual environment according to the target entity information in the natural language command.

[0748] Therefore, the embodiments of the present application provide a method of determining a target entity in combination with target entity information and environmental perception information. For the above example, the "red truck" expressed by the player in the natural language command should be a red truck that the player can see, and therefore, in combination with the field of view of the master virtual character, the red truck located in the field of view of the master virtual character can be filtered from the multiple red trucks in the three-dimensional virtual environment, and the red truck is the target entity indicated by the player in the natural language command.

[0749] Since the natural language command is issued by the player based on the perception of the three-dimensional virtual environment, in order to accurately identify the target entity indicated in the natural language command, the game application program will combine the environmental perception information when the player issues the natural language command to identify the target entity in the natural language command. Based on the perception of the three-dimensional virtual environment by the player, the entity closest to the target entity information that the player can perceive in the three-dimensional virtual environment is inferred, which is the target entity.

[0750] For example, the three-dimensional virtual environment includes multiple candidate entities matching the natural language command, and the target entity is an entity selected from the multiple candidate entities based on the environmental perception information of the master virtual character or the non-player character. For example, the game application program first selects multiple candidate entities matching the target entity information in the natural language command from the entity set, and then selects an entity that can be perceived by the master virtual character or the non-player character from the multiple candidate entities as the target entity according to the environmental perception information.

[0751] For example, the perception range of the master virtual character is determined according to the position and orientation of the master virtual character. For example, the perception range is a sector region in front of the master virtual character with a certain field of view size; entities not in the sector range are filtered out. Then, the target entity is selected from the entities with higher similarity scores according to the similarity between the target entity information and the entity information of the candidate entities. When there are two entities with the highest similarity scores in the perception range, for example, entity 2 and entity 3 have a similarity score of 2, and the entity closer to the master virtual character, entity 2, is selected as the target entity.

[0752] When it is necessary to generate feedback information according to the position description of the target entity, the game application can perform spatial query to generate the position description according to the positional relationship between the master virtual character and the target entity.

[0753] The main application scenario of spatial query is to obtain the positions of the master virtual character, non-player characters, enemy virtual characters, gunshots, etc. For example, a broadcast containing spatial information such as "enemy found on the second floor of the motel" can be made.

[0754] The spatial query logic defines the concepts of parent area and child area. The parent area refers to a large area in a three-dimensional virtual environment that covers multiple buildings or multiple isolated spaces, for example, a motel with a front yard and a back yard and containing multiple rooms and a basement. The child area refers to an independent space in the parent area, which can be a closed space or a special functional space, for example, a guest room, a kitchen, etc. in the motel.

[0755] In order to make the position description returned by the spatial query closer to human expression, when the master virtual character and the queried position are in different parent areas, the query result returns complete text information (for example, the master virtual character is outside the motel and the queried position is on the second floor of the motel, and the query result returns "enemy on the second floor of the motel"). When the master virtual character and the queried position are in the same parent area but different child areas, the query result only returns the child area information (for example, the master virtual character is on the first floor of the motel and the queried position is on the second floor, and the query result returns "enemy on the second floor"). When the master virtual character and the queried position are in the same parent area and the same child area, the query result only returns the direction relative to the master player (for example, the master virtual character and the queried position are both on the second floor of the motel, and the query result returns "enemy in the right front position").

[0756] Exemplarily, in a case that the master virtual role and the target entity are located in different parent spaces, the position description includes a parent space and a child space in which the target entity is located. In a case that the master virtual role and the target entity are located in different child spaces of the same parent space, the position description includes a child space in which the target entity is located. In a case that the master virtual role and the target entity are located in the same child space of the same parent space, the position description includes a relative position of the target entity and the master virtual role; wherein, one parent space includes at least one child space.

[0757] For the broadcasting of the NPC feedback information in step 208, the feedback information is implemented as spatial human voice audio processed by spatial sound effect, and a process of the spatial sound effect processing can be implemented as:

[0758] Sub-step 7: obtaining human voice audio corresponding to the feedback text;

[0759] In the embodiment of the present application, the generation process of the feedback text is as shown in the foregoing sub-step 6, and will not be described here.

[0760] In some embodiments, the process of obtaining human voice audio based on the feedback text can be implemented as: obtaining feedback text and role characteristic information corresponding to the non-player role; generating human voice audio matching the role characteristic information according to the feedback text.

[0761] Exemplarily, the role characteristic information is used to describe the attribute of the non-player role, that is, the role characteristic information is information used to describe the current state and characteristics of the non-player role. In some embodiments, the attribute of the non-player role includes at least one of a basic attribute and a scenario performance attribute of the non-player role. The basic attribute is an attribute pre-configured for the non-player role, that is, an attribute that does not change due to the three-dimensional virtual environment or the current scenario, for example, age information of the non-player role, male / female information of the non-player role, faction information of the non-player role, etc. The scenario performance attribute is an attribute of the non-player role determined in real time in the scenario of the three-dimensional virtual environment, that is, an attribute associated with the current scenario, for example, emotion information and behavior information of the non-player role in the current scenario.

[0762] Optionally, the role characteristic information includes at least one of role basic information, role emotion information, and role behavior information. The role basic information is used to indicate a basic situation of the non-player character, and the role basic information is information of a basic attribute of the non-player character, for example, age information of the non-player character, male / female information of the non-player character, faction information of the non-player character, and the like. The role emotion information is used to indicate an emotional state of the non-player character in a dialogue scene corresponding to the feedback text, and the role emotion information is a scene performance attribute of the non-player character, for example, the role emotion information indicates that the non-player character is in an excited state, an angry state, a sad state, or the like. The role behavior information is used to indicate an action performed by the non-player character in the dialogue scene, and the role behavior information is a scene performance attribute of the non-player character, for example, the non-player character is in a running state, an attacking state, a wounded treatment state, or the like.

[0763] In some embodiments, the generation of the human voice audio is implemented by using a pre-trained voice generation model, and the voice generation model is used to generate a voice matching the role characteristic information of the non-player character. Illustratively, the feedback text and the role characteristic information are input into the pre-trained voice generation model to obtain the human voice audio as the human voice audio.

[0764] Optionally, the voice generation model can be implemented by using a neural network model such as a convolutional neural network, a feedforward neural network, a residual network, a Transformer, a multimodal large language model (MLLM), or the like, which is not specifically limited herein.

[0765] In some embodiments, the voice generation model includes a text encoder, a role encoder, and a decoder. The text encoder is used to perform text encoding on the input feedback text, that is, the feedback text is input into the text encoder to obtain a text encoding representation. The role encoder is used to perform feature encoding on the input role characteristic information, that is, the role characteristic information is input into the role encoder to obtain a role encoding representation.

[0766] Illustratively, after obtaining the text encoding representation and the role encoding representation, the text encoding representation and the role encoding representation are fused to obtain an encoding representation input into the decoder, that is, the text encoding representation and the role encoding representation are fused to obtain a joint encoding representation, and the joint encoding representation is input into the decoder to generate the human voice audio.

[0767] In some embodiments, the role encoder comprises at least one of a first sub-encoder, a second sub-encoder, and a third sub-encoder. The first sub-encoder is configured to perform feature encoding based on the base attribute of the non-player character, i.e., in a case where the role feature information comprises role base information, the role base information is input into the first sub-encoder to obtain a first role encoding representation; the second sub-encoder is configured to perform feature encoding based on the emotional state of the non-player character, i.e., in a case where the role feature information comprises role emotional information, the role emotional information is input into the second sub-encoder to obtain a second role encoding representation; and the third sub-encoder is configured to perform feature encoding based on the behavior state of the non-player character, i.e., in a case where the role feature information comprises role behavior information, the role behavior information is input into the third sub-encoder to obtain a third role encoding representation.

[0768] Sub-step 8: based on the relative positional relationship between the non-player character and the master virtual character in the three-dimensional virtual environment, obtaining an audio effect parameter corresponding to the non-player character;

[0769] Illustratively, the relative positional relationship between the non-player character and the master virtual character in the three-dimensional virtual environment is used to indicate the relative relationship between a first position corresponding to the non-player character and a second position corresponding to the master virtual character in the three-dimensional virtual environment.

[0770] Optionally, the relative positional relationship comprises at least one of a positional distance, a positional direction, and a spatial surrounding condition between the non-player character and the master virtual character, wherein the positional distance is used to indicate the proximity relationship between the non-player character and the master virtual character, the positional direction is used to indicate the angle size relationship between the first position where the non-player character is located and the orientation direction of the master virtual character, and the spatial surrounding condition is used to indicate the surrounding condition of the master virtual character by the space element forming a virtual space in the three-dimensional virtual environment, and the propagation effect of the audio emitted by the non-player character on the audio in the virtual space.

[0771] In some embodiments, the determination of the relative positional relationship between the non-player character and the master virtual character can be implemented by: obtaining first position information of the non-player character in the three-dimensional virtual environment, obtaining second position information of the master virtual character in the three-dimensional virtual environment, and determining the relative positional relationship between the non-player character and the master virtual character based on the first position information and the second position information.

[0772] In some embodiments, the first position information of the non-player character and the second position information of the master virtual character are position information determined based on the same preset coordinate system. Optionally, the preset coordinate system can be a world coordinate system corresponding to the three-dimensional virtual environment, or the preset coordinate system can be a coordinate system established with the master virtual character as the origin.

[0773] In some embodiments, the sound effect parameter corresponding to the non-player character is obtained from a database when the sound effect parameter corresponding to the non-player character is pre-generated and stored in the database. In other embodiments, the sound effect parameter corresponding to the non-player character can also be generated in real time, i.e., based on the relative positional relationship, the sound effect parameter corresponding to the non-player character is generated.

[0774] Optionally, the sound effect parameter corresponding to the non-player character can include at least one of the following types:

[0775] Firstly, a distance type;

[0776] Illustratively, the sound effect parameter of the distance type is used to indicate that the sound effect of the human voice audio is adjusted based on the distance between the non-player character and the master virtual character. In some embodiments, the sound effect parameter of the distance type can adjust the audio volume of the human voice audio to simulate the effect of distance on the size of the sound; and / or, the sound effect parameter of the distance type can adjust the audio time delay of the human voice audio to simulate the effect of distance on the propagation time of the sound.

[0777] Secondly, a direction type;

[0778] Illustratively, the sound effect parameter of the direction type is used to indicate that the sound effect of the human voice audio is adjusted based on the direction of the non-player character relative to the master virtual character. In some embodiments, the sound effect parameter of the direction type can adjust the audio volume of the human voice audio to simulate whether the master virtual character is facing the scene audio element; and / or, the sound effect parameter of the direction type can adjust the volume parameters of the human voice audio in different channels to simulate the direction of the scene sound source element emitting the audio.

[0779] Thirdly, a spatial effect type;

[0780] Illustratively, the sound effect parameter of the spatial effect type is used to indicate that the sound effect in the virtual space formed by the spatial element is adjusted, i.e., the sound effect performance of the non-player character emitting the audio in the virtual space is indicated by the sound effect parameter of the spatial effect type. In some embodiments, the sound effect parameter of the spatial effect type can increase the echo effect and the reverberation effect to simulate the effect of the sound in the space.

[0781] Optionally, the sound effect parameter of the space effect type includes at least one of a reverberation time parameter, a pre-delay parameter, a wet / dry mix parameter, a room width parameter, and a distance effect parameter. The reverberation time parameter is used to indicate the speed of audio decay in the virtual space; the pre-delay parameter is used to indicate the time difference between the direct arrival of audio to the main virtual character and the first reflection arrival to the main virtual character; the wet / dry mix parameter is used to indicate the proportion between directly propagated audio and reflected propagated audio; the room width parameter is used to indicate the degree of diffusion of audio on the horizontal plane in the virtual space; and the distance effect parameter is used to indicate the attenuation of audio with the propagation distance in the propagation process.

[0782] Fourth, a custom type;

[0783] Illustratively, the sound effect parameter of the custom type is a parameter for sound effect adjustment customized by the user. Optionally, the user can customize the overall volume of the scene audio, whether to add background music, and the volume of different types of audio.

[0784] In some embodiments, the client provides a custom interface for the user to configure the custom sound effect parameter.

[0785] In some embodiments, the server is configured with a parameter prediction model for personalized learning of the custom sound effect parameter for the user. Illustratively, the custom sound effect parameters of a plurality of candidate accounts are obtained, the plurality of custom sound effect parameters are input into the parameter prediction model to be trained, the parameter prediction model to be trained is iteratively trained to obtain the parameter prediction model, and the sound effect parameters generated by the system (e.g., the sound effect parameters of the distance type, the direction type, and the space effect type) are optimized through the parameter prediction model.

[0786] Substep 9: adjusting the human voice audio based on the sound effect parameter to generate and play the spatial human voice audio corresponding to the non-player character;

[0787] The spatial human voice audio is used to represent the perception effect of the audio generated by the main virtual character to the non-player character in the relative positional relationship.

[0788] In the embodiments of the present application, the spatial human voice audio corresponding to the non-player character is generated by adjusting the human voice audio according to the sound effect parameter corresponding to the non-player character. That is, the spatial human voice audio is the audio data corresponding to the non-player character obtained by adjusting the human voice audio through the sound effect parameter. In some embodiments, the spatial human voice audio includes at least one of the audio adjusted by the sound effect parameter, and audio duration information, a playback start timestamp of the audio, a playback end timestamp of the audio, playback condition information of the audio, and element identification information corresponding to the audio.

[0789] Optionally, when the sound effect parameter indicates to adjust the volume size of the human voice audio of the non-player character, the volume size of the human voice audio of the non-player character on at least one sound channel is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates to adjust the playing time delay of the human voice audio of the non-player character, the audio playing start time corresponding to the human voice audio of the non-player character is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates to adjust the sound speed size of the human voice audio of the non-player character, the playing speed of the human voice audio of the non-player character is adjusted based on the sound effect parameter. Optionally, when the sound effect parameter indicates the tone of the human voice audio of the non-player character, the frequency in the frequency spectrum of the human voice audio of the non-player character is adjusted based on the sound effect parameter. When the sound effect parameter indicates the timbre of the human voice audio of the non-player character, the filter corresponding to the non-player character is determined based on the sound effect parameter, and the timbre corresponding to the human voice audio is adjusted through the filter.

[0790] Optionally, the client plays the spatial human voice audio to realize the broadcast of the feedback information of the non-player character.

[0791] The speech-to-text stage of step 204 can be implemented as:

[0792] Sub-step 10: Obtain a plurality of first scene hotwords corresponding to a first virtual scene in which the target virtual character is located.

[0793] The target virtual character includes at least one of an NPC and a master virtual character.

[0794] Optionally, the master virtual character is a virtual character commanded by the player, and the NPC is a virtual character assisting the master virtual character in a virtual game. In general, the virtual game includes one master virtual character and at least one NPC. There is a certain difference between the master virtual character and the NPC.

[0795] Optionally, the master virtual character is a virtual character mainly controlled by the player, and the NPC is a virtual character selected and commanded by the player. For example, the master virtual character is a virtual character manually controlled by the player during the game process, and the manual operation of the terminal interface is used to control the master virtual character. The NPC is a virtual character commanded by the player in the form of voice, and the player occasionally issues a natural language command in the form of voice to command the NPC.

[0796] In some embodiments, the first virtual scene is determined based on the target virtual character, and the target virtual character is at least one of an NPC and a master virtual character, that is, the process of determining the first virtual scene is implemented as at least one of the following.

[0797] (1) If the master virtual character is selected as the target virtual character, the virtual scene in which the master virtual character is located is taken as the first virtual scene.

[0798] (2) If the NPC is selected as the target virtual character, the virtual scene in which the NPC is located is taken as the first virtual scene; if there are multiple NPCs, the virtual scene including the most NPCs can be taken as the first virtual scene, the NPC closest to the host virtual character can also be taken as the target virtual character, and the virtual scene in which the NPC is located is taken as the first virtual scene, or the NPC with the largest character attribute value (such as at least one of the virtual blood volume, the virtual magic value, the virtual defense value, etc.) can also be taken as the target virtual character, and the virtual scene in which the NPC is located is taken as the first virtual scene.

[0799] (3) If the host virtual character and the NPC are selected as the target virtual character, the virtual scene in which the host virtual character and the NPC are located can be taken as the first virtual scene, etc.

[0800] In some embodiments, the target virtual character is determined based on the selection of the player; or the target virtual character is set by default by the system.

[0801] In some embodiments, a plurality of first scene hot words are determined based on the first virtual scene.

[0802] The plurality of first scene hot words are scene-related words of the first virtual scene. That is, there is an association between the first scene hot word and the first virtual scene, and the first scene hot word is a word describing the first virtual scene.

[0803] Optionally, the first scene hot word includes a scene state of the first virtual scene, such as the first virtual scene being a virtual restaurant, the first scene hot word including an operating state, a resting state, a closed state, etc.

[0804] Optionally, the first scene hot word includes an element name of a virtual element in the first virtual scene, such as the first virtual scene being a virtual restaurant, the first scene hot word including a virtual table, a virtual chair, a virtual kitchen utensil, a virtual dish, etc.

[0805] Optionally, the first scene hot word includes an interactive word for interacting with a virtual element in the first virtual scene, such as the first virtual scene being a virtual battlefield, the first scene hot word including attack, attack virtual character, defense, build barrier, etc.

[0806] In some embodiments, after receiving the natural language command, a plurality of first scene hot words are collected based on the first virtual scene; or after receiving the natural language command, a plurality of first scene hot words corresponding to the first virtual scene are filtered from a plurality of scene hot words obtained in advance based on the first virtual scene, etc.

[0807] In an optional embodiment, an object perception range of the target virtual character in the first virtual scene is obtained.

[0808] wherein the object perception range is a three-dimensional spatial range in which the target virtual character has a perception of other scene elements.

[0809] Optionally, when the target virtual character is a master virtual character, the object perception range of the master virtual character in the first virtual scene is obtained; when the target virtual character is an NPC, the object perception range of the NPC in the first virtual scene is obtained; or when the target virtual character is a master virtual character and an NPC, the object perception range of the master virtual character in the first virtual scene is obtained and the object perception range of the NPC in the first virtual scene is obtained.

[0810] In some embodiments, the object perception range comprises at least one of the following.

[0811] (1) a visual perception range of the target virtual character.

[0812] Illustratively, the visual perception range refers to a spatial range that an observer can perceive or understand visually, and is usually used to describe the size of an area that a biological body or a technical device can cover or perceive visually. The visual perception range of the target virtual character refers to a three-dimensional spatial range that the target virtual character can perceive through observation as an observer.

[0813] Optionally, the visual perception range of the target virtual character is obtained based on the position of the target virtual character in the first virtual scene and the visual observation capability.

[0814] (2) an auditory perception range of the target virtual character.

[0815] Illustratively, the auditory perception range refers to a spatial range that a listener can perceive or understand audibly, and is usually used to describe the size of a sound range that a biological body or a technical device can cover or perceive audibly. The auditory perception range of the target virtual character refers to a three-dimensional spatial range that the target virtual character can perceive through auditory perception as a listener.

[0816] Optionally, the auditory perception range of the target virtual character is obtained based on the position of the target virtual character in the first virtual scene and the auditory perception capability.

[0817] (3) an olfactory perception range of the target virtual character.

[0818] Illustratively, the olfactory perception range refers to a spatial range in which an odor is perceived or identified in the sense of smell. The olfactory perception range of the target virtual character refers to a three-dimensional spatial range that the target virtual character can perceive in the sense of smell.

[0819] Optionally, based on the position of the target virtual role in the first virtual scene and the olfactory perception capability, an olfactory perception range of the target virtual role is obtained.

[0820] (4) a perception skill possessed by the target virtual role or a range perceived by a perception device.

[0821] Illustratively, the perception skill is a capability of the target virtual role to perform scene perception on the first virtual scene, and the perception skill includes at least one of a perception acquisition form and a perception enhancement form.

[0822] Optionally, the perception acquisition form is used to represent a process of acquiring a non-possessed perception skill. After the perception skill is acquired through the perception acquisition form, the target virtual role can perform targeted scene perception on the first virtual scene through the perception skill. Optionally, the perception enhancement form is used to represent a process of enhancing a certain perception capability by possessing the perception skill in the case of possessing the certain perception capability.

[0823] Illustratively, the perception device is a virtual prop used to perceive the first virtual scene, and the perception device includes at least one of a perception acquisition form and a perception enhancement form.

[0824] Optionally, different perception skills or perception devices each correspond to a preset effective range, representing a range interval that can be perceived when the perception skill or the perception device is applied, and the preset effective range represents the object perception range.

[0825] In some embodiments, the object perception range is a spherical range, a ring-shaped range, an irregular three-dimensional space range, etc., and the shape of the object perception range is not limited here.

[0826] In an optional embodiment, a plurality of first scene hotwords are obtained based on environmental perception information within the object perception range.

[0827] Illustratively, the environmental perception information is used to represent environmental information perceived and obtained by the target virtual role within the object perception range. Optionally, the environmental perception information includes virtual elements determined by a ray-shooting method, and can also include region structure information determined by analyzing geometric structure and layout information of the virtual environment, etc.

[0828] Illustratively, after the object perception range corresponding to the target virtual role is determined, based on the range of the object perception range in the first virtual scene and the environmental perception information determined by the target virtual role within the object perception range based on perception, a plurality of first scene hotwords representing the state of the target virtual role are obtained.

[0829] In an optional embodiment, the environmental perception information includes virtual elements; and virtual elements within the object perception range in the first virtual scene are determined.

[0830] The object perception range is illustratively a part of the three-dimensional space range in the first virtual scene; the first virtual scene includes a large number of virtual elements, and the virtual elements are elements constituting the virtual scene, such as virtual ground, virtual buildings, virtual trees, virtual characters, virtual stones, etc.

[0831] In some embodiments, in the case of determining the object perception range, from the plurality of virtual elements corresponding to the first virtual scene, the virtual elements in the object perception range are determined.

[0832] In some embodiments, the element name of the virtual element is taken as the first scene hotword; or the action name of the interactive action corresponding to the virtual element is taken as the first scene hotword.

[0833] Sub-step 11: converting the natural language command in the form of voice into text form based on the plurality of first scene hotwords;

[0834] Illustratively, the first scene hotword is based on the vocabulary highly associated with the first virtual scene, and therefore in the case of wishing to better coordinate the NPC with the main control virtual character activity through the natural language command, the first scene hotword can be taken as a constraint condition on the basis of the first virtual scene, so that in the process of converting the natural language command into text form, the text content of the command analysis result is more consistent with the first virtual scene, avoiding obtaining a command analysis result that is not adapted to the first virtual scene.

[0835] In an optional embodiment, in the case of analyzing the natural language command through the natural language analysis model, the plurality of speech units corresponding to the natural language command are obtained through the acoustic network.

[0836] The speech unit is illustratively a basic constituent unit of vocabulary pronunciation.

[0837] In some embodiments, the command feature representation corresponding to the natural language command is extracted; and the command feature representation is analyzed through the acoustic network.

[0838] Illustratively, taking the command feature representation representing a plurality of acoustic features as an example, the target of the acoustic network is to map the continuous acoustic feature sequence to the speech unit sequence, such as phonemes or phonetic segments. It learns to predict the possible speech unit (phoneme or phonetic segment) sequence under the given acoustic feature.

[0839] Optionally, the acoustic network outputs one or more possible speech unit sequences, and the speech unit sequence includes a plurality of speech units. Different speech unit sequences may have different speech units, or only the order of the speech units may be different. The plurality of speech unit sequences represent the arrangement of the speech units considered by the language network to be the most possible when understanding the input natural language command.

[0840] The natural language analysis model includes an acoustic network, a preset dictionary, and language sub-networks corresponding to the at least two virtual scenes respectively.

[0841] Illustratively, the virtual environment includes at least two virtual scenes. For example, the virtual environment is a large virtual world, which includes virtual kitchen 1, virtual kitchen 2, virtual office, etc., and each of virtual kitchen 1, virtual kitchen 2, virtual office, etc. can be regarded as a virtual scene.

[0842] In an optional embodiment, based on the first virtual scene in which the target virtual role is located, a first language sub-network corresponding to the first virtual scene is determined from the at least two language sub-networks.

[0843] Illustratively, the plurality of virtual scenes correspond to a language sub-network respectively, and different language sub-networks are used for analyzing the corresponding virtual scenes. For example, virtual...

Claims

1. A method for controlling a virtual character, characterized in that, The method includes: Displays the main virtual character or non-player NPC character located in the virtual environment; Obtain command information, which includes natural semantics for controlling the NPC, and the command information corresponds to at least two semantic understanding methods; In response to the command information, the NPC is controlled to perform virtual activities according to the first semantic understanding method among the at least two semantic understanding methods; The first semantic understanding method is determined from the at least two semantic understanding methods based on the environmental perception information of the main virtual character and / or the NPC.

2. The method according to claim 1, characterized in that, The command information corresponds to at least two semantic understanding methods for scene entities. Responding to the command information by controlling the NPC to perform virtual activities according to the first semantic understanding method among the at least two semantic understanding methods includes: In response to the command information, the NPC is controlled to perform the virtual activity on a first scene object located within the perception range of the main virtual character and / or the NPC; The virtual environment contains multiple scene objects of the same type, including the first scene object, and the perception range is the virtual area where the main virtual character and / or the NPC obtains the environmental perception information.

3. The method according to claim 2, characterized in that, The step of responding to the command information by controlling the NPC to perform the virtual activity on a first scene object located within the perception range of the main virtual character and / or the NPC includes: In response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located within the field of view.

4. The method according to claim 3, characterized in that, The field of view is the field of view observed by the main virtual character and / or the NPC; and / or, the field of view is an extended field of view obtained by using perception skills or sensory tools.

5. The method according to claim 2, characterized in that, The step of responding to the command information by controlling the NPC to perform the virtual activity on a first scene object located within the perception range of the main virtual character and / or the NPC includes: In response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in the first virtual terrain, where the first virtual terrain is the virtual terrain to which the location of the main virtual character and / or the NPC belongs.

6. The method according to claim 5, characterized in that, The step of controlling the NPC to perform the virtual activity in response to the command information includes at least one of the following: When the first virtual terrain is an indoor terrain, in response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in the first area, the first area being an indoor connected area including the first location of the main virtual character and / or the NPC. In the case that the first virtual terrain is a block terrain adjacent to buildings, in response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in the second area, the second area including the main virtual character and / or the buildings surrounding the second location where the NPC is located; In the case that the first virtual terrain is outdoor terrain far from buildings, in response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in the third region, the third region being an area with the same virtual vegetation distribution as the third location of the main virtual character and / or the NPC.

7. The method according to claim 1, characterized in that, The command information corresponds to at least two semantic understanding methods for word pronunciation. Responding to the command information by controlling the NPC to perform virtual activities according to the first semantic understanding method among the at least two semantic understanding methods includes: In response to the command information, the command information is converted into command text according to the natural language pronunciation of the scene objects in the perception range. The perception range is the virtual area where the main virtual character and / or the NPC has acquired the environmental perception information. Based on the command text, control the NPC to execute the virtual activity in accordance with the first semantic understanding method.

8. The method according to claim 7, characterized in that, The step of responding to the command information by converting the command information into command text according to the natural language pronunciation of objects in the scene within the perception range includes: If the similarity between the pronunciation of the first phoneme in the command information and the natural language pronunciation of the scene object exceeds a similarity threshold, and if there are candidate texts for the first phoneme, the text information of the first phoneme is determined as the name text and / or description text of the scene object, and the command text is constructed based on the text information. The candidate text for the first phoneme does not include the name text and / or description text of the scene object.

9. The method according to claim 1, characterized in that, The command information corresponds to at least two semantic understanding methods for behavioral intent. Responding to the command information, controlling the NPC to perform virtual activities according to the first semantic understanding method among the at least two semantic understanding methods includes: In response to the command information, the NPC is controlled to perform the virtual activity according to the first behavioral intent; the first behavioral intent corresponds to the type of virtual activity that is permitted to be performed within the perception range, the perception range being the virtual area where the main virtual character and / or the NPC has acquired the environmental perception information.

10. The method according to claim 9, characterized in that, The step of responding to the command information and controlling the NPC to perform the virtual activity according to the first intention includes: In response to the command information, taking the first direction within the perception range as the reference direction, and based on the directional words in the command information, the NPC is controlled to perform the virtual activity based on the second direction; Wherein, the directional term is used to describe the angle between the second direction and the first direction, and / or to describe the second direction in a reference frame constructed based on the first direction.

11. The method according to claim 1, characterized in that, The step of responding to the command information and controlling the NPC to perform virtual activities according to a first semantic understanding method among the at least two semantic understanding methods includes: Based on the command information, determine the command text in natural language form; The conditional random field is invoked to perform entity prediction on the command text to obtain the entity labels carried by the command text; Within the perception range, determine the first scene object based on the entity label; The prediction network is invoked to perform intent recognition on the command text to obtain the classification label of the command text; the prediction network is used to predict the control method of the command text for the NPC in at least one behavioral dimension; Based on the classification tags, the control commands for the NPC are determined, and based on the control commands, the NPC is controlled to perform the virtual activities on the first scene objects in accordance with the first semantic understanding method.

12. The method according to claim 11, characterized in that, The invocation prediction network performs intent recognition on the command text to obtain classification labels for the command text, including: A hierarchical prediction network is invoked to perform intent recognition on the command text, thereby obtaining the classification label of the command text; The hierarchical prediction network includes at least two sub-networks, which are constructed based on a tree structure with multiple classification labels.

13. The method according to claim 11, characterized in that, The method further includes: Obtain visual labels for multiple candidate scene objects within the perception range; Determining the first scene object within the perception range based on the entity label includes: Calculate the similarity between the visual labels and entity labels of the multiple candidate scene objects, and determine the first scene object with the highest similarity among the multiple candidate scene objects.

14. The method according to claim 13, characterized in that, The method further includes: Obtain the spatial position of each scene object in the virtual environment; Based on the spatial location, the plurality of candidate scene objects within the perception range are selected from the various scene objects in the virtual environment.

15. The method according to claim 14, characterized in that, The step of obtaining the spatial position of each scene object in the virtual environment includes: Obtain at least one of the following: coordinate position, orientation information, bounding box information, and cover point information of each scene object in the virtual environment; The coordinate position is used to indicate the position of the scene object in the virtual environment, the orientation information is used to indicate the direction the scene object faces in the virtual environment, the bounding box information is used to indicate the size of the scene object in the virtual environment, and the cover point information indicates the recommended virtual character position when the virtual character approaches the scene object.

16. The method according to claim 13, characterized in that, The method further includes: Obtain the attribute text of the scene object in the virtual environment, the attribute text being used to describe the inherent attributes of the scene object in the virtual environment; Obtain the appearance image of the scene object in the virtual environment, the appearance image being used to describe the style of the scene object; A multimodal model is invoked to predict the attribute text and appearance image of the scene object to obtain a visual label for the scene object, the visual label being used to describe the scene object in at least one dimension; Based on the visual labels of objects in each scene within the virtual environment, a spatial database of the virtual environment is constructed. The step of obtaining visual labels for multiple candidate scene objects within the perception range includes: Based on the perception range, visual labels for multiple candidate scene objects within the perception range are obtained from the spatial database.

17. The method according to any one of claims 1 to 16, characterized in that, The method further includes: The system presents feedback information from the NPC to the command information, which indicates that the NPC has received the command information and / or that the NPC has executed the virtual activity indicated by the command information.

18. A method for controlling a virtual character, characterized in that, The method includes: Displays the main virtual character or non-player NPC character located in the virtual environment; Obtain command information, which includes natural semantics for controlling the NPC, and includes behavioral intent and scene entities; the behavioral intent is used to indicate the type of virtual activity performed by the NPC, and the scene entities are used to indicate the target entity targeted by the virtual activity; In response to the command information, control the NPC to perform virtual activities; The virtual activity is determined based on the environmental awareness information of the main virtual character and / or the NPC.

19. The method according to claim 18, characterized in that, The scene entity includes a first scene object; the step of controlling the NPC to perform virtual activities in response to the command information includes: In response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located within the perception range of the main virtual character and / or the NPC; The virtual environment contains multiple scene objects of the same type, including the first scene object, and the perception range is the virtual area where the main virtual character and / or the NPC obtains the environmental perception information.

20. The method according to claim 19, characterized in that, The step of responding to the command information by controlling the NPC to perform the virtual activity on the first scene object located within the perception range of the main virtual character and / or the NPC includes: In response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located within the field of view; Alternatively, in response to the command information, the NPC is controlled to perform the virtual activity on the first scene object located in the first virtual terrain, where the first virtual terrain is the virtual terrain to which the location of the main virtual character and / or the NPC belongs.

21. The method according to claim 18, characterized in that, The behavioral intent includes the first behavioral intent; The step of controlling the NPC to perform virtual activities in response to the command information includes: In response to the command information, the NPC is controlled to perform the virtual activity according to the first behavioral intention; the perception range is the virtual area where the main virtual character and / or the NPC has acquired the environmental perception information.

22. The method according to claim 18, characterized in that, The command information also includes natural language pronunciations of objects in the scene within the perception range; the step of controlling the NPC to perform virtual activities in response to the command information includes: In response to the command information, the command information is converted into command text according to the natural language pronunciation of the scene objects in the perception range. The perception range is the virtual area where the main virtual character and / or the NPC has acquired the environmental perception information. Based on the command text, control the NPC to perform the virtual activity.

23. The method according to claim 18, characterized in that, The step of controlling the NPC to perform virtual activities in response to the command information includes: Based on the command information, determine the command text in natural language form; The conditional random field is invoked to perform entity prediction on the command text to obtain the entity labels carried by the command text; Within the perception range, determine the first scene object based on the entity label; The prediction network is invoked to perform intent recognition on the command text to obtain the classification label of the command text; the prediction network is used to predict the control method of the command text for the NPC in at least one behavioral dimension; Based on the classification tags, the control commands for the NPC are determined, and based on the control commands, the NPC is controlled to perform the virtual activities on the first scene objects.

24. A method for processing command information, characterized in that, The method includes: Obtain command information, which includes natural semantics for controlling non-player NPC characters, and the command information corresponds to at least two semantic understanding methods; Based on the environmental awareness information of the main virtual character and / or the NPC, a first semantic understanding method is determined among the at least two semantic understanding methods of the command information; Based on the command information, the NPC is controlled to perform virtual activities according to the first semantic understanding method among the at least two semantic understanding methods.

25. A method for processing command information, characterized in that, The method includes: Obtain command information, which includes natural semantics for controlling non-player character NPCs, and includes behavioral intent and scene entities; the behavioral intent is used to indicate the type of virtual activity performed by the NPC, and the scene entities are used to indicate the target entity targeted by the virtual activity; Based on the command information, control the NPC to perform virtual activities; The virtual activities are determined based on the environmental awareness information of the main virtual character and / or the NPC.

26. A control device for a virtual character, characterized in that, The device includes: The display module is used to display the main virtual character or non-player NPC character located in the virtual environment; The acquisition module is used to acquire command information, which includes natural semantics for controlling the NPC, and the command information corresponds to at least two semantic understanding methods; The control module is configured to respond to the command information and control the NPC to perform virtual activities according to the first semantic understanding method among the at least two semantic understanding methods; The first semantic understanding method is determined from the at least two semantic understanding methods based on the environmental perception information of the main virtual character and / or the NPC.

27. A command information processing device, characterized in that, The device includes: The acquisition module is used to acquire command information, which includes natural semantics for controlling non-player character NPCs, and the command information corresponds to at least two semantic understanding methods; The processing module is used to determine a first semantic understanding method among the at least two semantic understanding methods of the command information based on the environmental perception information of the main virtual character and / or the NPC; The control module is used to control the NPC to perform virtual activities according to the first semantic understanding method among the at least two semantic understanding methods, based on the command information.

28. A control device for a virtual character, characterized in that, The device includes: The display module is used to display the main virtual character or non-player NPC character located in the virtual environment; The acquisition module is used to acquire command information, which includes natural semantics for controlling the NPC, and includes behavioral intent and scene entities; the behavioral intent is used to indicate the type of virtual activity performed by the NPC, and the scene entities are used to indicate the target entity targeted by the virtual activity; The control module is used to control the NPC to perform virtual activities in response to the command information; The virtual activity is determined based on the environmental awareness information of the main virtual character and / or the NPC.

29. A command information processing device, characterized in that, The device includes: The acquisition module is used to acquire command information, which includes natural semantics for controlling non-player character NPCs. The command information includes behavioral intentions and scene entities. The behavioral intentions are used to indicate the type of virtual activity performed by the NPC, and the scene entities are used to indicate the target entity targeted by the virtual activity. The control module is used to control the NPC to perform virtual activities based on the command information; The virtual activities are determined based on the environmental awareness information of the main virtual character and / or the NPC.

30. A computer device, characterized in that, The computer device includes: a processor and a memory, wherein the memory stores at least one program; the processor is configured to execute the at least one program in the memory to implement the virtual character control method as described in any one of claims 1 to 23, and / or the command information processing method as described in claim 24 or 25.

31. A computer-readable storage medium, characterized in that, The readable storage medium stores executable instructions, which are loaded and executed by a processor to implement the virtual character control method as described in any one of claims 1 to 23, and / or the command information processing method as described in claim 24 or 25.

32. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, wherein a processor reads from and executes the computer instructions to implement the virtual character control method as described in any one of claims 1 to 23, and / or the command information processing method as described in claim 24 or 25.

Citation Information

Patent Citations

  • Task guiding method and device in virtual scene, equipment, medium and program product

    CN114247141A

  • Human-computer interaction method and device, electronic equipment and storage medium

    CN114461775A

  • Method and device for controlling virtual object in game, electronic equipment and storage medium

    CN118161861A

  • Virtual character control method and device, virtual character processing method and device, equipment and storage medium

    CN119003699A

  • Object processing method and apparatus in virtual scene, device, and storage medium

    US20230338854A1