Voice conversion method and device based on virtual scene, equipment, medium and product
By acquiring scene hot words from virtual scenes and converting them into natural language commands, the problem of low command accuracy in virtual scenes by speech recognition systems is solved, enabling more efficient and accurate virtual character control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-09
- Publication Date
- 2026-03-31
AI Technical Summary
In existing technologies, speech recognition systems suffer from low accuracy in processing natural language commands due to issues such as unclear pronunciation, excessively fast or slow speech speed. This results in an inability to accurately recognize natural language commands and affects the accuracy of controlling virtual characters.
By acquiring multiple scene hot words in the virtual scene where the target virtual character is located, and using these hot words to convert natural language commands into text-based command analysis results, the accuracy of command analysis is improved by combining them with a pre-trained natural language analysis model.
By analyzing natural language commands under the constraints of the first virtual scene, the ambiguity of the analysis is reduced, the accuracy of the command analysis results is improved, and non-player characters can accurately act according to natural language commands, thereby improving the efficiency of human-computer interaction and the comfort of operation.
Smart Images

Figure CN119049463B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of machine learning, and in particular to a speech conversion method, apparatus, device, medium and product based on virtual scenes. Background Technology
[0002] As people's cultural and entertainment standards improve, their expectations and demands for the virtual world also increase. Virtual scenes are provided in various forms such as terminal games, virtual reality (VR), and augmented reality (AR), allowing players to control virtual characters to move around in the virtual scene manually or by voice.
[0003] In related technologies, if a player controls a virtual character via voice, after receiving a natural language command in audio form, the player will perform text analysis on the natural language command through a pre-trained speech recognition system, and control the virtual character based on the analyzed text content. The model in the speech recognition system is trained on a large number of text words, and the text content is a combination of some text words from the large number of text words.
[0004] However, natural language commands may have quality issues, such as unclear pronunciation, speaking too fast or too slow, which may prevent the speech recognition system from accurately recognizing the natural language commands, resulting in low accuracy of the text content. Summary of the Invention
[0005] This application provides a speech conversion method, apparatus, device, medium, and product based on a virtual scene, which enables a better match between command analysis results and a first virtual scene, thereby improving the accuracy of the command analysis results. The technical solution is as follows.
[0006] On the one hand, a speech conversion method based on a virtual scene is provided, the method comprising:
[0007] Acquire natural language commands in the form of speech, the natural language commands being used to direct a non-player character, the target virtual character being in a first virtual scene, the target virtual character including at least one of the non-player character and the master virtual character;
[0008] Obtain multiple first scene hot words corresponding to the first virtual scene where the target virtual character is located, wherein the multiple first scene hot words are scene-related words of the first virtual scene;
[0009] The natural language command is converted into a text form based on the multiple first-scene hot words.
[0010] On the other hand, a voice conversion device based on a virtual scene is provided, the device comprising:
[0011] The acquisition module is used to acquire natural language commands in the form of voice, the natural language commands being used to instruct a non-player character, the target virtual character being in a first virtual scene, the target virtual character including at least one of the non-player character and the main control virtual character;
[0012] The acquisition module is also used to acquire multiple first scene hot words corresponding to the first virtual scene where the target virtual character is located, wherein the multiple first scene hot words are scene-related words of the first virtual scene;
[0013] The analysis module is used to convert the natural language command into a text form command analysis result based on the multiple first scenario hot words.
[0014] In an optional embodiment, the acquisition module is further configured to acquire a first scene type corresponding to the first virtual scene where the target virtual character is located; acquire a set of hot words corresponding to the first scene type, and use the words in the hot word set as the plurality of first scene hot words, wherein the hot word set is a set of scene-related words collected based on the first scene type; or, acquire the object perception range of the target virtual character in the first virtual scene, wherein the object perception range is the three-dimensional spatial range in which the target virtual character perceives other scene elements; acquire the plurality of first scene hot words based on environmental perception information within the object perception range; or, acquire a first scene type corresponding to the first virtual scene; acquire a set of hot words corresponding to the first scene type; acquire the object perception range of the target virtual character in the first virtual scene; and acquire at least two words from the hot word set as the plurality of first scene hot words based on environmental perception information within the object perception range.
[0015] In an optional embodiment, the object-aware range includes at least one of the following:
[0016] The visual perception range of the target virtual character;
[0017] The auditory perception range of the target virtual character;
[0018] The olfactory perception range of the target virtual character;
[0019] The range of perception skills or sensory tools possessed by the target virtual character.
[0020] In an optional embodiment, the acquisition module is further configured to, when the object perception range includes the visual perception range, generate a preset visual cone as the visual perception range, with the position of the target virtual character as the starting point and the orientation of the target virtual character as the cone midline, wherein the preset visual cone is a three-dimensional region that measures the visual perception range.
[0021] In an optional embodiment, the environment-aware information includes virtual elements;
[0022] The acquisition module is further configured to determine the virtual elements in the first virtual scene that are within the object's perception range; use the element name of the virtual element as the first scene hot word; or, use the action name of the interactive action corresponding to the virtual element as the first scene hot word.
[0023] In an optional embodiment, the first virtual scene is one of a plurality of virtual scenes in a virtual environment;
[0024] In an optional embodiment, the acquisition module is further configured to determine the first virtual scene in which the target virtual character is located from the plurality of virtual scenes based on the position of the target virtual character in the virtual environment.
[0025] In an optional embodiment, the acquisition module is further configured to, in response to the target virtual character moving from the first virtual scene to the second virtual scene without generating the command analysis result, acquire a plurality of second scene hot words corresponding to the second virtual scene; and convert the natural language command into the command analysis result in text form based on the plurality of second scene hot words.
[0026] In an optional embodiment, the acquisition module is further configured to determine the first virtual scene corresponding to the first virtual game based on the first virtual game in which the target virtual character participates; wherein, multiple virtual games each correspond to a virtual scene, and the multiple virtual games are a simulated battle environment provided for the target virtual character.
[0027] In an optional embodiment, the acquisition module is further configured to acquire multiple scene hot words, the multiple scene hot words corresponding to at least two virtual scenes; based on the first virtual scene in which the target virtual character is located, acquire the multiple first scene hot words corresponding to the first virtual scene from the multiple scene hot words.
[0028] In an optional embodiment, the plurality of scene hot words correspond to scene identifiers, and the scene identifiers are used to represent the virtual scene when collecting scene hot words;
[0029] The acquisition module is further configured to acquire multiple candidate scene hot words with a first scene identifier from the multiple scene hot words based on the first virtual scene where the target virtual character is located, wherein the first scene identifier is the scene identifier corresponding to the first virtual scene; and select at least two candidate scene hot words as the first scene hot words based on the object perception range of the target virtual character in the first virtual scene.
[0030] In an optional embodiment, the analysis module is further configured to acquire a pre-trained natural language analysis model, the pre-trained natural language analysis model including an acoustic network, a language network, and a preset dictionary, the preset dictionary including the plurality of first scenario hot words, other scenario hot words, and general vocabulary; acquire multiple speech units corresponding to the natural language command through the acoustic network, the speech units being the basic building blocks of word pronunciation; analyze the sequence matching relationship between the plurality of speech units and selected words in the preset dictionary through the language network to obtain the command analysis result in text form, the selected words including at least the plurality of first scenario hot words and the general vocabulary.
[0031] In an optional embodiment, the selected vocabulary also includes the other scenario hot words;
[0032] The vocabulary in the preset dictionary includes analysis weights. The first analysis weight of the multiple first scenario hot words is higher than the second analysis weight of the other scenario hot words. The analysis weight refers to the degree of attention a word receives when participating in sequence matching.
[0033] In an optional embodiment, the analysis module is further configured to analyze the matching relationship between the plurality of speech units and the selected words in the preset dictionary, obtain a plurality of candidate word sequences, wherein the plurality of candidate word sequences include at least one word from the plurality of selected words, and the candidate word sequences are word sequences obtained based on the change relationship between the plurality of speech units; analyze the sequence semantics of the plurality of candidate word sequences through the language network, and obtain at least one candidate word sequence from the plurality of candidate word sequences as the command analysis result in text form.
[0034] In an optional embodiment, the analysis module is further configured to determine a first language subnetwork corresponding to the first virtual scene from at least two language subnetworks based on the first virtual scene in which the target virtual character is located; analyze the sequence semantics of the plurality of candidate word sequences through the first language subnetwork, and obtain at least one candidate word sequence from the plurality of candidate word sequences as the command analysis result in text form.
[0035] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the virtual scene-based speech conversion method as described in any of the embodiments of this application above.
[0036] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the virtual scene-based speech conversion method as described in any of the embodiments of this application above.
[0037] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the virtual scene-based speech conversion method described in any of the above embodiments.
[0038] The beneficial effects of the technical solutions provided in this application include at least the following:
[0039] When directing non-player characters via natural language commands, multiple hot words from the first virtual scene where the target virtual character is located are obtained. Natural language commands are then analyzed under the constraints of these hot words, allowing for more easily understood command analysis results that guide the non-player character to act according to the natural language commands. Using the first virtual scene where the target virtual character is located as a condition for analyzing natural language commands provides more targeted text information through the hot words, reducing ambiguity caused by poor command quality. This not only facilitates more accurate understanding of natural language commands but also ensures a better match between command analysis results and the first virtual scene, improving the accuracy of command analysis results. This significantly avoids ineffective or erroneous commands to non-player characters, providing players with a more efficient and comfortable operating environment and improving human-computer interaction efficiency. Using a smaller number of hot words from the first scene also reduces the amount of data processing during command analysis, improving command analysis efficiency. Attached Figure Description
[0040] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0041] Figure 1 This is a flowchart illustrating how a voice command (natural language command) in the form of natural language is used to instruct a non-player character in a virtual environment, as provided in an exemplary embodiment of this application.
[0042] Figure 2 This is a structural block diagram of a conversion system provided in an exemplary embodiment of this application;
[0043] Figure 3 This is a flowchart of a speech conversion method based on a virtual scene provided in an exemplary embodiment of this application;
[0044] Figure 4 This is a flowchart of a speech conversion method based on a virtual scene provided in another exemplary embodiment of this application;
[0045] Figure 5 This is a schematic diagram of an interface for determining the perceived range of an object, provided in an exemplary embodiment of this application.
[0046] Figure 6 This is a flowchart of a speech conversion method based on a virtual scene provided in another exemplary embodiment of this application;
[0047] Figure 7 This is a flowchart illustrating the analysis of natural language commands using a natural language analysis model, provided in an exemplary embodiment of this application.
[0048] Figure 8 This is a technical flowchart of a speech conversion method based on a virtual scene provided in an exemplary embodiment of this application;
[0049] Figure 9 This is a schematic diagram of the interface of a speech conversion method based on a virtual scene provided in an exemplary embodiment of this application;
[0050] Figure 10 This is an overall flowchart of a speech conversion method based on a virtual scene provided in an exemplary embodiment of this application;
[0051] Figure 11 This is a flowchart of a virtual character control method provided in an exemplary embodiment of this application;
[0052] Figure 12 This is a structural block diagram of a speech conversion device based on a virtual scene provided in an exemplary embodiment of this application;
[0053] Figure 13 This is a schematic diagram of the structure of a server provided in an exemplary embodiment of this application. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0055] First, a brief introduction to the terms used in the embodiments of this application will be given.
[0056] Virtual scene: A virtual scene is a scene displayed (or provided) by an application when it runs on a terminal. This virtual scene can be a simulation of a real scene, a semi-simulated / semi-fictional scene, or a purely fictional scene. A virtual scene can be any of a two-dimensional virtual scene, a 2.5-dimensional virtual scene, or a three-dimensional virtual scene; this application does not limit it to any particular type. The following embodiments use a three-dimensional virtual scene as an example.
[0057] Virtual models are models used in virtual scenes to mimic real-world scenes. For example, a virtual model occupies a certain volume within a virtual scene. Examples of virtual models include: terrain models, building models, plant and animal models, virtual prop models, virtual vehicle models, and virtual object models. For instance, terrain models include: ground, mountains, rivers, rocks, steps, etc.; building models include: houses, walls, containers, and fixed facilities inside buildings: tables, chairs, cabinets, beds, etc.; plant and animal models include: trees, flowers, birds, etc.; virtual prop models include: virtual attack tools, first-aid kits, airdrops, etc.; virtual vehicle models include: cars, ships, helicopters, etc.; and virtual object models include: people, animals, anime characters, etc.
[0058] Virtual characters / objects: These refer to movable objects in a virtual scene. These movable objects can be virtual objects, virtual animals, anime characters, etc., such as people, animals, plants, oil drums, walls, and stones displayed in a 3D virtual scene. Optionally, virtual objects are 3D models created based on animation skeletal technology. Each virtual object has its own shape and volume in the 3D virtual scene, occupying a portion of the space within the 3D virtual scene.
[0059] In related technologies, if a player controls a virtual character via voice, after receiving a natural language command in audio format, a pre-trained speech recognition system analyzes the command to determine the text content, which then guides the player's control of the virtual character. The speech recognition system's model is trained on a large vocabulary of text, and the text content is a combination of selected words from that vocabulary. However, natural language commands may suffer from quality issues, such as unclear pronunciation, excessively fast or slow speech, which can prevent the speech recognition system from accurately recognizing the commands, resulting in low accuracy of the text content.
[0060] This application provides a virtual scene-based speech conversion method. This method uses the first virtual scene in which the target virtual character is located as a condition for analyzing natural language commands. By using hot words from the first scene, it provides more targeted text information for the analysis process, making the text-based command analysis results more closely match the first virtual scene, improving the accuracy of the command analysis results, and thus improving the accuracy of commands given to non-player characters. The virtual scene-based speech conversion method provided in this application can be applied to virtual scenes requiring speech conversion, such as terminal games, virtual reality, augmented reality, virtual tourism, virtual meetings, virtual shopping, and virtual fitness. This application does not limit its application to these applications.
[0061] In one optional embodiment, the application of a virtual scene-based speech conversion method to a terminal game is used as an example for illustration.
[0062] To illustrate, a game application is installed on the terminal. During gameplay, the application displays a virtual scene where the player is playing a virtual match. Players can use voice control to command non-player characters (NPCs) within this virtual scene. Taking a target virtual character in a first virtual scene as an example, the target virtual character includes at least one of a non-player character and a main virtual character. That is, when commanding a non-player character, the first virtual scene can be determined based on the non-player character, the main virtual character, or a combination of both. Optionally, multiple "first scene hot words" corresponding to the first virtual scene where the target virtual character is located are obtained. These hot words are scene-related terms for the first virtual scene. For example, if the player manually controls the main virtual character on the terminal screen, the first virtual scene is determined based on the main virtual character. Multiple hot words include the names of virtual elements within the first virtual scene; for example, if the first virtual scene is a virtual kitchen, multiple hot words include virtual knives, virtual plates, etc. Optionally, the natural language commands are converted into text-based command analysis results based on multiple first-scene hot words. That is, the speech-to-text process is constrained by multiple first-scene hot words, which makes it easier to convert the natural language commands in speech form into text-based instruction analysis text that is more in line with the first virtual scene. For example, the instruction analysis text obtained based on the constraints of multiple first-scene hot words is "defend with a kitchen knife", avoiding recognition as "defend with a cutting knife", improving the accuracy of instruction analysis text, and also making it easier to better command non-player characters to use virtual kitchen knives to defend against attacks from other virtual characters, improving the efficiency and accuracy of character command.
[0063] In one optional embodiment, the application of a virtual scene-based speech conversion method to a virtual reality scene is used as an example for illustration.
[0064] In a schematic representation, a VR headset can transmit and present a virtual scene received from a computer device. Players observe this virtual scene by wearing the VR headset and can move to other locations within the virtual scene to explore it by moving and adjusting body parts. Taking the player wearing the VR headset as the primary virtual character, the virtual scene presented by the VR headset can also contain other virtual characters controlled and / or directed by other players, as well as non-player characters. Optionally, players can command non-player characters using natural language commands. The VR headset can receive these natural language commands via a deployed microphone array, autonomously analyze the commands, or send them to a connected computer device for analysis. The natural language commands are used to command non-player characters, and the target virtual character is located in a first virtual scene, including at least one of a non-player character and a primary virtual character. For example, the first virtual scene is a virtual scene shared by the primary virtual character and the non-player character to be commanded. Optionally, multiple hot keywords corresponding to the first virtual scene where the target virtual character is located are obtained. For example, if the target virtual character is the main control virtual character, and the player, wearing a VR headset, explores to a virtual battlefield, they need to engage in virtual combat with enemy virtual character A. The obtained hot keywords for the first virtual scene include virtual swords, attack, and attack enemy virtual character A. Optionally, based on these hot keywords, natural language commands are converted into text-based command analysis results. The resulting command analysis result is "Attack me," avoiding the identification as "Supply me," thus improving the accuracy of command analysis text. This also facilitates players in better commanding non-player characters to effectively attack enemy virtual characters, improving the accuracy of character commands, enriching the fun of virtual reality scenes, and enabling players to have a more realistic and immersive virtual experience.
[0065] It should be noted that this application may display prompt interfaces, pop-ups, or output voice prompts before and during the collection of user data. These prompt interfaces, pop-ups, or voice prompts are used to inform the user that their data is being collected. This ensures that the application only begins the steps for collecting user data after receiving confirmation from the user regarding the prompt interface or pop-up; otherwise (i.e., without user confirmation), the steps for collecting user data end, meaning no user data is collected. In other words, all user data collected in this application is collected with the user's consent and authorization, and the collection, use, and processing of related user data must comply with the relevant laws, regulations, and standards of the relevant regions.
[0066] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant regions. For example, the natural language commands and scene hot words involved in this application were obtained with full authorization.
[0067] This application provides a solution for users to use natural language voice commands to control non-player characters (NPCs) in a virtual environment. In this solution, in addition to the user's own main virtual character, the user typically has one or more NPCs as teammates. Users can use relatively casual, rather than mechanical, human-like conversation to instruct the NPCs to provide desired responses to objects in the virtual environment, thereby directing the NPCs to cooperate with the main virtual character to complete tasks. Exemplary References Figure 1 The user controls a main virtual character 10 to play the game. The main virtual character 10 has an NPC teammate 20. Within the main virtual character 10's field of vision, there is a car 30. The user speaks a natural language command: "Number 2, go and ambush behind car 30." The NPC teammate 20 will then automatically move to the back of car 30 to ambush. It should be noted that... Figure 1 In the scene shown, there may be multiple cars. The NPC teammate 20 will accurately understand that the car mentioned by the user is the car 30 located within the user's field of vision. That is, the NPC teammate 20 has a relatively intelligent natural language understanding ability.
[0068] On the one hand, compared with the mechanical fixed instructions used to command NPCs in traditional technologies, the embodiments of this application use relatively arbitrary natural language commands to command NPCs, which can provide users with more natural, more complex and more flexible language command capabilities.
[0069] On the other hand, the intelligence of the natural language understanding capability provided in this application embodiment is reflected in the fact that when the NPC teammate understands natural language commands, it does not only consider the literal information of the natural language commands. The NPC teammate also combines the environmental perception information of the main virtual character and / or the NPC itself in the virtual environment to assist in understanding the semantics of the natural language commands. That is, the NPC not only considers the information of the natural language command as a single modality, but also considers the perception information of other modalities such as vision, hearing, and radar to assist in the understanding and execution of the natural language commands. Since the natural language commands in speech form are spoken commands, rather than written commands, there may be multiple candidate understanding methods or unclear points when only considering the literal information of the natural language commands. However, the NPC teammate, by combining the environmental perception information of the main virtual character and / or the NPC itself in the virtual environment, can determine the reasonable understanding method from multiple candidate understanding methods, or eliminate doubts about unclear points, thereby achieving a more intelligent natural language understanding capability. The environmental perception information mentioned above includes, but is not limited to, at least one of the following:
[0070] • Visual information perceived within the field of vision of the main virtual character;
[0071] • Auditory information perceived within the hearing range of the main virtual character;
[0072] • The perceptual skills or information perceived by the main virtual character's sensory devices;
[0073] • Visual information perceived within the NPC's field of vision;
[0074] • Auditory information perceived within the NPC's hearing range;
[0075] • Information perceived by NPCs through their perception skills or sensing tools.
[0076] Reference Figure 1 The scheme includes at least one of the following five phases:
[0077] Phase 1: Spatial data preprocessing;
[0078] The server preprocesses the objects in the virtual scene to construct the spatial data of the virtual scene.
[0079] Spatial data for a virtual scene includes matching visual labels for scene objects and their spatial locations. In some examples, the spatial data also includes images of the scene objects' appearance. Scene objects are any objects that appear in the virtual scene, such as... Figure 1 Cars, walls, boxes, etc.
[0080] First, we introduce the pre-acquisition process of visual tags for matching objects in the scene. This involves acquiring attribute information of the objects in the virtual scene, such as at least one of their names and dimensions; and acquiring appearance images of the objects, which include images obtained from viewing the objects from at least two perspectives. Viewing the objects from multiple perspectives ensures that the appearance images carry comprehensive appearance information about the objects.
[0081] The server invokes a multimodal model to predict the attribute information and appearance images of scene objects. Leveraging the multimodal model's ability to process both text and image dimensions, it extracts hidden features from the text-level attribute information and the image-level appearance images, and decodes these hidden features to predict visual object labels for the scene objects. These visual object labels describe the scene objects in at least one dimension using textual language, such as material, transparency, color, or shape. A natural language model is then invoked to optimize these visual object labels, resulting in matching visual labels for the scene objects. The natural language model, with its text generation capabilities, rewrites the input visual object labels in a conversational style, optimizing the visual object labels to obtain optimized matching visual labels for the scene objects. These optimized visual labels conform to conversational language, describing the scene objects in at least one dimension in a way that closely resembles spoken language.
[0082] Taking a virtual bed in a virtual scene as an example, when describing it in spoken language, people usually focus on its color, material, and placement, but often neglect the craftsmanship used on the headboard and side panels, such as carving or embossing. The purpose of rewriting the visual labels of scene objects in a colloquial style is to obtain matching labels that are closer to spoken language. For example, this involves enriching the labels for dimensions like color, material, and placement that are considered in spoken language with a variety of words with the same semantic meaning, and removing labels for surface features of the headboard and side panels that are often overlooked in spoken language.
[0083] Next, we will introduce the spatial position of scene objects. Spatial position includes at least one of the following: coordinates (Location), such as the coordinates of the center point or preset point of the scene object in the virtual scene; orientation (Rotation), such as the direction the front of the scene object faces in the virtual scene; bounding box, used to indicate the size of the scene object in the virtual scene; and cover points, used to indicate the recommended position points for a virtual character when approaching a scene object, so that the scene object can cover the virtual character.
[0084] Next, we will introduce the pre-acquisition process for appearance images of scene objects. Appearance images are acquired during the process of predicting matching visual labels for scene objects, and include images of scene objects observed from at least two viewpoints.
[0085] Optionally, spatial data of various scene objects in the entire virtual environment will be pre-acquired and organized into a dataset for use in subsequent query stages.
[0086] Phase Two: Speech Recognition and Intent Recognition;
[0087] During the process of user control of the main virtual character, the terminal receives natural language commands input by the user via voice; the server performs speech recognition and intent recognition on the natural language commands to analyze the user's input instructions;
[0088] Speech recognition converts user-inputted natural language commands into text, obtaining the corresponding instruction text. It transforms speech information into text information. An encoding network encodes the instruction text to obtain its feature representation in the hidden space. A text segmentation network then segments this feature representation into multiple clauses within the instruction text. Intent recognition is then performed on each clause separately. This approach enables the segmentation of long sentences in the user's speech input, meeting the need for analyzing user input commands in long sentence scenarios.
[0089] The following describes the intent recognition for each clause; the user's input instructions are obtained through intent recognition, and the user's input instructions include at least one entity, semantic type, subject type and intent type in the clause.
[0090] On one hand, a Conditional Random Field (CRF) approach is used to identify entities in the clauses, such as buildings, virtual items, and virtual vegetation in the virtual scene. Entities in the clauses indicate the scene objects targeted by the virtual activity. For example, if the entity in the clause is a virtual building, and the clause carries the semantics of instructing an NPC to launch a virtual attack, then the virtual attack is executed against the virtual building. On the other hand, a prediction network is invoked to predict the semantic type, subject type, and intent type in the clauses. The semantic type includes whether it instructs an NPC to launch a virtual attack; the subject type indicates that the subject of the clause is one or more NPCs or a user-controlled virtual character; and the intent type is the desired virtual activity indicated by the clause, such as moving or launching a virtual attack.
[0091] Phase 3: Spatial data query;
[0092] Using user input commands as input, and taking the spatial data of the virtual scene and the runtime information of the main virtual character as references, an AI model is used to determine the target entities in the virtual space and the control commands for the NPCs.
[0093] For entities in input commands, it is necessary to determine the spatial location of the entity in the virtual environment. For example, the text input command is first parsed to extract the entity within it. If multiple candidate entities in the virtual environment match the entity in the input command, the target entity matching the input command can be uniquely identified from among the candidate entities by combining the environmental awareness information of the main virtual character and / or NPC.
[0094] For example, taking the environmental perception information as the field of vision information of the main virtual character, the spatial data of the scene in the virtual environment is used as the query scope, and the real-time position and orientation of the main virtual character are used as query reference conditions to query the spatial location of the target entity in the virtual environment corresponding to the input command. This allows for subsequent control of the NPC to move to the vicinity of the target entity's location, or to execute the behavior indicated by the input command at the target entity's location.
[0095] However, since there may be multiple types of environmental perception information for the main virtual character and / or NPC, multiple types of environmental perception information can be integrated to assist in the determination process of the target entity, or multiple types of environmental perception information can be used according to priority to assist in the determination process of the target entity, without limitation.
[0096] Phase Four: NPC Voice Feedback;
[0097] NPCs provide voice feedback to users. The types of voice feedback include at least one of the following: instant feedback, command execution feedback, and dynamic feedback. Instant feedback refers to the NPC immediately responding to a natural language command issued by the user, such as "received," "start execution," or "okay." Command execution feedback refers to the NPC providing voice feedback on the execution result of a control command, such as "executed successfully" or "failed to execute successfully." Dynamic feedback refers to the NPC providing voice feedback based on real-time environmental awareness even when the user has not issued a natural language command.
[0098] The aforementioned voice feedback is obtained by converting text content into speech. This text content is inferred through a large language model combined with the aforementioned spatial data and dynamic runtime environment-aware information. Using a large language model to dynamically generate the text content for voice feedback overcomes the monotony, mechanical nature, and repetitiveness of fixed, template-based voice feedback. Furthermore, it eliminates the need to pre-save too many pre-made template voice messages, thus avoiding the problem of excessively large data volume on the client side.
[0099] Phase 5: NPC Behavior Control;
[0100] Using user input commands as input, and taking the spatial data of the virtual scene and the runtime information of the main virtual character as references, an AI model is used to determine the target entities in the virtual space and the control commands for the NPCs.
[0101] The aforementioned natural language commands typically include control instructions, which are instructions that can be executed by the NPC's behavior tree. Control instructions instruct the NPC to perform a subsequent action, which can be a single action or a sequence of actions. When the control instruction is a multi-action sequence instruction, it instructs the NPC to execute a sequence of actions. The data structure storing these multi-action sequence instructions is in list form, caching the instructions to be executed within the sequence. After the NPC's behavior tree completes the execution of the current instruction, the cache is called back, and the cached instructions to be executed are then reissued for further execution.
[0102] Hot word update function:
[0103] For Phase Two, this solution also introduces a hot word system to help improve the accuracy of sound recognition.
[0104] The hot word system is used to provide scene hot words related to the scenario involved in the natural language command during the speech recognition process. When a natural language command is received from voice input, the system can prioritize searching for words from the hot word library based on the currently running virtual scenario as the speech recognition result and command analysis result, so that the matching degree between the command analysis result, the speech recognition result and the virtual scenario is higher.
[0105] Virtual scenes encompass at least one of multiple scene types, such as virtual battle scenes, virtual trading scenes, virtual office scenes, and virtual kitchen scenes. Frequently used vocabulary and specific virtual elements (e.g., virtual buildings and virtual character objects that only exist in specific scene types) across multiple virtual scene types are pre-analyzed as scene hot words, resulting in a scene hot word library. Each scene hot word corresponds to a specific virtual scene type. Furthermore, different scene language models are pre-trained for different virtual scene types. For example, the scene language model for battle scenes focuses more on battle-related vocabulary, while the scene language model for trading scenes focuses more on trading-related vocabulary. These scene language models can perform more targeted analysis of natural language commands received within virtual scene types.
[0106] For example, when a user receives a natural language command while controlling a virtual character (not a player character), based on the game state, object location, and current task completion status, the first virtual scene where the target virtual character (at least one of the main virtual character and the non-player character) is currently located is determined to be a virtual combat scene, and the target scene type is determined to be a combat scene type. The combat scene language model corresponding to the virtual combat scene type is obtained. Furthermore, based on the field of vision of the main virtual character, the target first scene hot words under the virtual combat scene type are obtained, including: "truck," "virtual bushes," "virtual house," "enter," "open truck," "defend," "desert," "attack," "oak tree," "stable," etc., which are within the field of vision and contain "enemy A," "attack," etc. Under the constraint of the first target scene hot words, the natural language command is decoded by a pre-trained natural language analysis model (including an acoustic model network, a target scene language network model, a preset dictionary, etc.) and decoding network. The output text content is "Give me an attack," instead of the usual result obtained from speech recognition—"Give me supplies." This improves the accuracy of the user's input command and outputs the command analysis result to accurately control the non-player character. The command analysis result is illustrated below with an example.
[0107] (1) The command analysis result is “Number 1, go forward and defend”. This command analysis result can be used to control Number 1 to move forward and take a defensive stance to avoid the enemy virtual character attacking the main virtual character first. Without the constraint of the first scene hot words, the natural language command is easily identified as “Number 1, go forward and make a counterattack”, which will affect the player’s virtual combat plan.
[0108] (2) The command analysis result is “Number 2, go explore the desert”. This command analysis result can be used to control Number 2 to move to the desert location to explore whether there are enemy virtual characters or virtual treasure chests or other virtual elements. Without the constraint of the first scene hot words, the natural language command is easily identified as “Number 2, go explore the mountains”, which causes Number 2 to move to the wrong location, resulting in low human-computer interaction efficiency.
[0109] (3) The command analysis result is “1 and 2, attack me”. This command analysis result can be used to control 1 and 2 to launch virtual attacks on nearby attackable virtual characters. Without the constraint of the first scene hot words, the natural language command is easily identified as “1 and 2, supply me”, which causes 1 and 2 to make wrong behavior that does not meet the player’s expectations. This not only fails to protect the main virtual character, but also makes it easy for the enemy virtual character to win the virtual game.
[0110] (4) The command analysis result is "Number 1, go find the nearby oak tree". This command analysis result can be used to control Number 1 to search for oak trees in the nearby area so that he can complete the game task or find an oak tree to avoid attacks. Without the constraint of the first scene hot words, the natural language command is easily identified as "Number 1, go find the nearby project book", which makes Number 1 search for virtual elements that do not meet the player's expectations and cannot meet the player's virtual combat needs.
[0111] (5) The command analysis result is "1 and 2, go to the stable over there". This command analysis result can be used to control 1 and 2 to search for nearby stables and move to the location of the stables. Without the constraint of the first scene hot words, the natural language command is easily identified as "1 and 2, go there, it will be done soon". This makes 1 and 2 mistakenly believe that the main virtual character wants to complete the virtual game on its own, so it cannot provide good game assistance to the main virtual character.
[0112] Ambient sound effects function:
[0113] In addition, the solution also provides a spatial audio enhancement solution for the environmental sound effects of the entire virtual scene.
[0114] The audio played on the terminal includes ambient audio and NPC audio. Ambient audio is generated based on the characteristics of the virtual scene in which the main virtual object is currently located, so as to give users an immersive experience; NPC audio is generated based on the characteristics of the NPC, so that users can intuitively feel the character's emotions and physical state through the sounds they hear.
[0115] Ambient audio: Identifies scene elements within the virtual environment where the main virtual object is located, generates appropriate element sound effects in real time based on these scene elements, or selects them from a sound effects library, and synthesizes these sound effects to obtain ambient audio. For example, if the virtual scene is a forest at night, the scene elements include trees, owls, insects, etc., and the ambient audio includes the rustling of leaves, the call of an owl, the chirping of insects, etc.
[0116] NPC Audio: Identify the NPCs in the virtual scene and generate corresponding character voices based on the NPC's character type, current emotion, and current behavior. For example, if the NPC is a middle-aged man running, the generated character voice will be a deep male voice with breathing sounds while running.
[0117] When the terminal plays ambient audio and / or NPC audio to the user, it performs audio enhancement processing on the ambient audio and / or NPC audio to improve the realism of the audio.
[0118] In one optional embodiment, the conversion system involved in the embodiments of this application will be described. The speech conversion method based on virtual scenes provided in the embodiments of this application can be implemented by the terminal alone, by the server, or by the terminal and the server through data interaction. The embodiments of this application do not limit this. Optionally, the speech conversion method based on virtual scenes is described as being implemented by the interaction between the terminal and the server.
[0119] This is illustrative; please refer to it. Figure 2 The conversion system involves a terminal 210 and a server 220, which are connected via a communication network 230.
[0120] In some embodiments, a game application is installed on the terminal 210. During the operation of the game application, a virtual scene is displayed, and the player can control the virtual character in the virtual scene through manual control or voice control.
[0121] Optionally, taking the example of a player commanding a non-player character via voice, the terminal 210 is equipped with a microphone array that can acquire natural language commands in voice form. These natural language commands are used to command non-player characters, which are NPC characters used to assist players in virtual matches, such as virtual minions, virtual heroes, virtual monsters, and virtual mounts.
[0122] The target virtual character is located in the first virtual scene, and the target virtual character includes at least one of a non-player character and a main virtual character.
[0123] For illustrative purposes, the main virtual character is also a virtual character controlled by the player. Optionally, the main virtual character is the virtual character primarily controlled by the player, such as a virtual character manually controlled by the player and a non-player character controlled by the player via voice; or, the main virtual character is a virtual character chosen by the player and a non-player character is a virtual character configured by default by the system; or, both the main virtual character and the non-player character are virtual characters chosen by the player, where the main virtual character has the highest priority and the non-player character has a lower priority, etc.
[0124] Optionally, the virtual scene where the non-player character is located can be designated as the first virtual scene; or, the virtual scene where the main virtual character is located can be designated as the first virtual scene; or, the virtual scene where both the main virtual character and the non-player character are located can be designated as the first virtual scene, etc. For example, the virtual scene in which players participate in virtual matches can also be called a virtual environment. The virtual environment is large and includes multiple virtual scenes. The non-player character and the main virtual character may be in the same virtual scene or in different virtual scenes, which is not limited here.
[0125] In some embodiments, when the terminal 210 analyzes the natural language command itself, the terminal 210 obtains multiple first scene hot words corresponding to the first virtual scene where the target virtual character is located after receiving the natural language command; when the terminal 210 analyzes the natural language command with the help of the server 220, the terminal 210 sends the natural language command in voice form to the server 220 through the communication network 230, thereby obtaining multiple first scene hot words corresponding to the first virtual scene where the target virtual character is located through the server 220.
[0126] Among them, many of the hot words in the first scene are scene-related words in the first virtual scene.
[0127] For illustrative purposes, the first scene hot words are words describing the first virtual scene, or words describing virtual elements within the first virtual scene, or words describing interactions with virtual elements, etc. For example: if the first virtual scene is a virtual battle scene, the first scene hot words include attack, defense, virtual swords, virtual bows and arrows, etc.; or if the first virtual scene is a virtual kitchen, the first scene hot words include virtual kitchen knives, virtual scissors, virtual plates, etc.
[0128] In some embodiments, taking the analysis of natural language commands by server 220 as an example, server 220 converts natural language commands into command analysis results in text form based on multiple first scenario hot words.
[0129] As an illustration, server 220 uses multiple hot words from the first scenario as constraints during instruction analysis and converts natural language commands into text-based command analysis results using speech-to-text technology. Due to the constraints of the hot words from the first scenario, the instructions are prioritized as the vocabulary required for text conversion during the instruction analysis process. This establishes a strong correlation between the command analysis results and the first virtual scenario, making the command analysis results more closely match the first virtual scenario and improving the accuracy of the command analysis results.
[0130] Optionally, the server 220 can independently analyze the command results in text form to direct non-player characters to move more accurately, and generate activity animation data for directing character movement, which is then sent to the terminal 210 via the communication network 230 so that the activity animation corresponding to the activity animation data can be rendered and displayed on the screen interface of the terminal 210.
[0131] Optionally, the server 220 sends the command analysis results to the terminal 210 via the communication network 230. The terminal 210 can command the non-player character to act more accurately based on the command analysis results in text form, and render the non-player character's activity animation on the screen interface of the terminal 210, thereby improving the accuracy and efficiency of character command.
[0132] It is worth noting that the aforementioned terminals include, but are not limited to, mobile terminals such as mobile phones, tablets, portable laptops, smart voice interaction devices, smart home appliances, and in-vehicle terminals, and can also be desktop computers, etc.; the aforementioned servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers.
[0133] Based on the above introduction of terms and application scenarios, the speech conversion method based on virtual scenarios provided in this application will be explained, taking the application of this method to a server as an example. Figure 3 As shown, the method includes the following steps 310 to 330.
[0134] Step 310: Obtain natural language commands in speech form.
[0135] Speech form refers to the form of expression using speech or sound, that is, conveying information or intent through speech or sound. In the field of computer science, especially in human-computer interaction and natural language processing, speech form usually refers to the way users input information using spoken language.
[0136] In illustrative terms, natural language commands, as instructions in the form of speech, are typically implemented as audio data. For example, when a player speaks while operating a terminal, the terminal automatically captures the speech using a microphone array and obtains the audio data as a natural language command in speech form; or, when a player operates a terminal, they trigger an audio capture control and then speak, which the terminal receives and obtains as audio data as a natural language command in speech form, and so on.
[0137] Natural language commands are used to instruct non-player characters.
[0138] Indicatively, non-player characters possess autonomous action capabilities and can move independently within the virtual environment. Optionally, when a non-player character is not commanded by a player (e.g., not receiving natural language commands), the non-player character moves along a preset trajectory within the virtual environment; or, the non-player character attacks enemy virtual characters or assists the main virtual character in performing game tasks according to preset game logic. When a non-player character is commanded by a player (e.g., receiving natural language commands), the non-player character will perform targeted actions within the virtual environment according to the natural language commands to more purposefully assist the main virtual character.
[0139] For illustrative purposes, a non-player character is a virtual character displayed within the game, or a virtual character displayed in other virtual environments such as virtual reality scenes or augmented reality scenes.
[0140] Optionally, non-player characters are virtual characters commanded by players via voice. Non-player characters include at least one of several types of virtual characters, such as virtual minions, virtual heroes, virtual monsters, and virtual mounts. Based on natural language commands issued by players during gameplay, non-player characters in the virtual environment can be directed to act according to these commands.
[0141] In some embodiments, the process of commanding a non-player character is the process of unlocking the game account after it reaches a preset level. The game account is the account that the player logs into.
[0142] For example, if the preset level is 10, and the game account has not yet reached the preset level, the non-player character is an uncontrollable virtual character, meaning the player cannot command the non-player character via voice. Once the game account reaches the preset level, the non-player character becomes a controllable virtual character, meaning the player can now command the non-player character via voice.
[0143] In some embodiments, non-player characters are default virtual characters, such as defaulting virtual minions A and B as non-player characters that players can command; or, non-player characters are virtual characters selected by the player, such as providing multiple candidate control characters with different personalities and / or autonomous behavioral abilities. Personality reflects the attack strength of the candidate control character in virtual matches (e.g., a violent candidate control character has higher attack strength, a mild-mannered candidate control character has lower attack strength, etc.); autonomous behavioral abilities represent the candidate control character's skill set in virtual matches, such as candidate control character 1 being able to use ultimate skill A to attack enemy virtual characters within a circular area, and candidate control character 2 being able to use ultimate skill B to defend against attacks from nearby enemy virtual characters, etc. Players can select at least one candidate control character from among the multiple candidate control characters as a non-player character that can be commanded in the game.
[0144] The target virtual character is located in the first virtual scene, and the target virtual character includes at least one of a non-player character and a main virtual character.
[0145] For illustrative purposes, the main virtual character is the one commanded by the player, while the non-player character is a virtual character that assists the main virtual character in the virtual game. Typically, a virtual game includes one main virtual character and at least one non-player character. There are certain distinctions between the main virtual character and the non-player character.
[0146] Optionally, the main virtual character is the virtual character primarily controlled by the player, while the non-player character is the virtual character chosen and commanded by the player. For example, the main virtual character is the virtual character manually controlled by the player during the game, and the player's manual operations on the terminal interface are used to control the main virtual character; the non-player character is the virtual character commanded by the player via voice, and the player occasionally issues natural language commands in the form of voice to command the non-player character.
[0147] Optionally, the main virtual character is a virtual character selected by the player, such as selecting virtual hero A as the main virtual character through the hero configuration area before the start of the virtual game; non-player characters are virtual characters configured by default by the system, such as virtual minion 1 and virtual minion 2 configured by default by the game as non-player characters, etc.
[0148] Optionally, both the main virtual character and the non-player characters are virtual characters chosen by the player. The main virtual character has the highest priority, while the non-player characters have lower priority. For example, if a player selects virtual hero A, virtual hero B, and virtual hero C to participate in a virtual match, and virtual hero A is listed first or has the highest priority, then virtual hero A is considered the main virtual character, and virtual heroes B and C are considered non-player characters. Alternatively, if the player selects virtual hero B first or has the highest priority, then virtual hero B is considered the main virtual character, and virtual heroes A and C are considered non-player characters, and so on.
[0149] In some embodiments, a first virtual scene is determined based on a target virtual character, wherein the target virtual character is at least one of a non-player character and a master virtual character, that is, the process of determining the first virtual scene is implemented in at least one of the following ways.
[0150] Optionally, if the main virtual character is selected as the target virtual character, then the virtual scene in which the main virtual character is located is selected as the first virtual scene.
[0151] Alternatively, if a non-player character is chosen as the target virtual character, the virtual scene where the non-player character is located is designated as the first virtual scene. If there are multiple non-player characters, the virtual scene containing the most non-player characters can be designated as the first virtual scene. Alternatively, the non-player character closest to the main virtual character can be designated as the target virtual character, and its virtual scene can be designated as the first virtual scene. Alternatively, the non-player character with the highest character attribute value (such as at least one of virtual health, virtual mana, virtual defense, etc.) can be designated as the target virtual character, and its virtual scene can be designated as the first virtual scene.
[0152] Alternatively, if you choose to use both the main virtual character and the non-player character as the target virtual character, you can use the virtual scene where the main virtual character and the non-player character are both located as the first virtual scene, etc.
[0153] In some embodiments, the target virtual character is determined based on the player's choice; or, the target virtual character is a system default setting.
[0154] Optionally, during a virtual match, the terminal interface displays a character configuration area. This area is used to adjust and determine the target virtual character in the first virtual scene. The character configuration area includes a first character identifier corresponding to the main virtual character and a second character identifier corresponding to a non-player character. Selecting the first character identifier means the main virtual character is the target virtual character; selecting the second character identifier means a non-player character is the target virtual character, and so on. For example, if multiple non-player characters exist (e.g., character 1 and character 2), the corresponding non-player character is determined based on the selected second character identifier (e.g., the non-player character corresponding to the selected second character identifier is character 2), and this non-player character is used as the target virtual character (e.g., character 2 is used as the target virtual character), etc.
[0155] Optionally, the system defaults to setting the main virtual character as the target virtual character; or, the system defaults to setting a non-player character as the target virtual character; or, when multiple non-player characters exist, the system defaults to setting at least one of them as the target virtual character, etc.
[0156] Step 320: Obtain multiple hot words for the first virtual scene corresponding to the first virtual scene where the target virtual character is located.
[0157] In a schematic way, after determining the target virtual character and the first virtual scene in which the target virtual character is located, multiple hot words of the first scene are obtained based on the first virtual scene.
[0158] Among them, several of the hot words for the first scene are scene-related words for the first virtual scene. That is, there is a relationship between the hot words for the first scene and the first virtual scene, and the hot words for the first scene are words that describe the first virtual scene.
[0159] Optionally, the first scene hot words include the scene status of the first virtual scene, such as: if the first virtual scene is a virtual restaurant, the first scene hot words include the operating status, the resting status, the closed status, etc.
[0160] Optionally, the first scene hot words include the element names of virtual elements within the first virtual scene, such as: if the first virtual scene is a virtual restaurant, the first scene hot words include virtual tables, virtual chairs, virtual kitchen utensils, virtual plates, etc.
[0161] Optionally, the first scene hot words include interactive words that interact with virtual elements within the first virtual scene. For example, if the first virtual scene is a virtual battlefield, the first scene hot words include attack, attacking virtual characters, defense, building barriers, etc.
[0162] In some embodiments, after receiving a natural language command, multiple first scene hot words are collected based on the first virtual scene; or, after receiving a natural language command, multiple first scene hot words corresponding to the first virtual scene are selected from multiple pre-acquired scene hot words based on the first virtual scene.
[0163] Step 330: Convert natural language commands into text form based on multiple hot words from the first scenario and analyze the command results.
[0164] As an illustration, since the hot words of the first scene are words highly related to the first virtual scene, if we want to use natural language commands to guide non-player characters to better cooperate with the main virtual character's activities, we can use the hot words of the first scene as constraints based on the first virtual scene. In the process of converting natural language commands into text form command analysis results, the text content of the command analysis results will be more consistent with the first virtual scene, avoiding the acquisition of command analysis results that are not compatible with the first virtual scene, thereby reducing the error rate of command analysis results.
[0165] In some embodiments, a pre-trained natural language analysis model is used to analyze natural language commands. The natural language commands and multiple hot words of the first scenario are used as input to the model. The natural language analysis model extracts the instruction feature representation of the natural language commands and the hot word feature representation of the multiple hot words of the first scenario. The instruction feature representation and the hot word feature representation are fused to obtain a fused feature representation. The command analysis result in text form is output based on the fused feature representation.
[0166] In some embodiments, a pre-trained natural language analysis model is used to analyze natural language commands. The natural language commands and multiple first-scene hot words are used as input to the model. During the analysis of natural language commands by the natural language analysis model, an attention mechanism is used to determine the weight of the first-scene hot words in the command analysis process, thereby achieving the purpose of using multiple first-scene hot words as analysis constraints. Finally, the output of the natural language analysis model is used as the command analysis result.
[0167] Optionally, natural language analysis models include Gaussian Mixture Model-Hidden Markov Model (GMM-HMM), deep learning models such as Recurrent Neural Networks (RNN), deep learning models such as Transformer, and end-to-end deep learning speech recognition (ASR). Among them, end-to-end deep learning speech recognition includes connectionist temporal classification algorithms and sequence-to-sequence algorithms. The selection of natural language analysis models is not limited here.
[0168] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.
[0169] In summary, when directing non-player characters via natural language commands, multiple hot words from the first virtual scene where the target virtual character is located are obtained. Natural language commands are then analyzed under the constraints of these hot words, allowing for more easily understood command analysis results that guide the non-player character to act according to the natural language commands. Using the first virtual scene where the target virtual character is located as a condition for analyzing natural language commands provides more targeted text information through the hot words, reducing ambiguity caused by poor command quality. This not only facilitates a more accurate understanding of natural language commands but also ensures a better match between the command analysis results and the first virtual scene, improving the accuracy of the command analysis results. This significantly avoids ineffective or erroneous commands to non-player characters, providing players with a more efficient and comfortable operating environment and improving human-computer interaction efficiency. Furthermore, using a smaller number of hot words from the first scene reduces the amount of data processing during command analysis, improving command analysis efficiency.
[0170] In an optional embodiment, the first scene hot words can be obtained based on the first scene type of the first virtual scene, or based on the object perception range of the target virtual character in the first virtual scene, or a combination of the first scene type and the object perception range can be used to obtain the first scene hot words. (Illustrative example, such as...) Figure 4 As shown above, Figure 3 The step 320 shown can also be implemented as steps 410 to 411; or steps 420 to 421; or steps 430 or 432.
[0171] Step 410: Obtain the first scene type corresponding to the first virtual scene where the target virtual character is located.
[0172] For illustrative purposes, the first scene type is the scene type corresponding to the first virtual scene. The scene type is used to represent the scene state of the virtual scene. For example, scene types include various types that describe the scene state, such as office type, battle type, kitchen type, bedroom type, classroom type, shop type, factory type, etc.
[0173] Optionally, the first virtual scene corresponds to the first scene type, which is used to describe the scene state represented by the first virtual scene. For example, if the first virtual scene is a virtual office, then the first scene type is office type; or, if the first virtual scene is a virtual battle area, then the first scene type is battle type; or, if the first virtual scene is a virtual store, then the first scene type is store type, etc.
[0174] Step 411: Obtain the set of hot words corresponding to the first scenario type, and use the words in the hot word set as multiple hot words for the first scenario.
[0175] Indicatively, multiple scene types each correspond to a set of hot words. The set of hot words represents the set of scene-related words collected based on the scene type, and also represents the set of scene-related words collected based on at least one virtual scene under that scene type.
[0176] For example: the office type corresponds to hot word set 1. Hot word set 1 includes multiple scene-related words collected based on at least one virtual office scene under the office type. For example, hot word set 1 includes scene-related words collected based on virtual office 1, and scene-related words collected based on virtual office 2. Virtual office 1 and virtual office 2 are two virtual offices under the office type.
[0177] For example: the combat type corresponds to hot word set 2, which includes multiple scene-related words collected based on at least one virtual combat scene under the combat type. For example, hot word set 2 includes scene-related words collected based on virtual combat scene 1, and scene-related words collected based on virtual combat scene 2. Virtual combat scene 1 and virtual combat scene 2 are two virtual combat scenes under the combat type.
[0178] Optionally, after determining the first scene type corresponding to the first virtual scene, a set of hot words corresponding to the first scene type is obtained. This set of hot words is a collection of scene-related words collected based on the first scene type; that is, the set of hot words corresponding to the first scene type is a collection of scene-related words collected based on the first scene type. For example, if the first scene type of the first virtual scene is combat, then the set of hot words corresponding to combat (set 2) is used as the set of hot words corresponding to the first scene type.
[0179] In some embodiments, the words in the hot word set are referred to as scene-related words, or hot word words, or simply words. After obtaining the hot word set corresponding to the first scene type, the words in the hot word set are used as multiple first scene hot words. That is, when obtaining multiple first scene hot words based on the first virtual scene, multiple words in the hot word set corresponding to the first scene type to which the first virtual scene belongs are used as multiple first scene hot words corresponding to the first virtual scene.
[0180] In illustrative terms, the first virtual scene is virtual kitchen 1. The scene type corresponding to virtual kitchen 1 is determined to be kitchen type. The hot word set corresponding to kitchen type is obtained and the words in it are used as multiple first scene hot words. The hot word set corresponding to kitchen type includes not only the words obtained based on virtual kitchen 1, but also the words obtained based on virtual kitchen 2. Therefore, the multiple first scene hot words obtained in this way have a wider range of characteristics.
[0181] Step 420: Obtain the object perception range of the target virtual character in the first virtual scene.
[0182] Schematic illustration: The target virtual character includes at least one of a controlling virtual character and a non-player character. When determining the object's perceptual range based on the target virtual character, at least one of the following situations applies.
[0183] Optionally, when the target virtual character is the main virtual character, the object perception range of the main virtual character in the first virtual scene is obtained; when the target virtual character is a non-player character, the object perception range of the non-player character in the first virtual scene is obtained; or, when the target virtual character is both the main virtual character and a non-player character, the object perception range of the main virtual character in the first virtual scene and the object perception range of the non-player character in the first virtual scene are both obtained.
[0184] Among them, the object perception range is the three-dimensional spatial range within which the target virtual character perceives the existence of other scene elements.
[0185] In illustrative terms, the target virtual character is a virtual character in the first virtual scene, which can be regarded as a virtual element in the first virtual scene. The first virtual scene may also include other scene elements besides the target virtual character. The object perception of the target virtual character represents the target virtual character's perception of other scene elements. The object perception range is the range of object perception, which is implemented as a three-dimensional spatial range in the first virtual scene.
[0186] In one alternative embodiment, the object-aware range includes at least one of the following.
[0187] (1) The object perception range includes the visual perception range of the target virtual character.
[0188] Indicatively, the visual perception range refers to the spatial range that an observer can visually perceive or understand. It is typically used to describe the size of the area that a living organism or technological device can visually cover or perceive. The visual perception range of a target virtual character is the three-dimensional spatial range perceived through the observation of the target virtual character.
[0189] Optionally, the visual perception range of the target virtual character can be obtained based on the target virtual character's position in the first virtual scene and its visual observation capabilities.
[0190] Indicatively, "position" represents the location information of the target virtual character relative to the first virtual scene, such as using position coordinates; "visual observation capability" represents the target virtual character's ability to observe the first virtual scene, such as visual distance (i.e., the farthest distance that can be observed and perceived) and visual angle (i.e., the range of angles that can be observed and perceived). Combining position and visual observation capability, the visual perception range of the target virtual character can be roughly determined from the first virtual scene.
[0191] Optionally, the visual perception range can be a cone-shaped range, a sphere-shaped range, an irregular three-dimensional space range, etc., and the shape of the visual perception range is not limited here.
[0192] Optionally, during the game configuration phase, the virtual characters can be pre-configured with corresponding visual observation abilities, and different virtual characters can have the same visual perception abilities; or, different virtual characters can have different visual perception abilities, such as the main virtual character having stronger visual perception abilities and non-player characters having weaker visual perception abilities, etc., without any limitation here.
[0193] (2) The object perception range includes the auditory perception range of the target virtual character.
[0194] In a illustrative sense, auditory perception range refers to the spatial range that a listener can perceive or understand through hearing. It is typically used to describe the size of the range of sounds that a living organism or technological device can cover or perceive. The auditory perception range of a target virtual character is the three-dimensional spatial range perceived by the target virtual character as a listener, based on the target virtual character's auditory perception capabilities.
[0195] Optionally, the auditory perception range of the target virtual character can be obtained based on the target virtual character's position in the first virtual scene and its auditory perception ability.
[0196] Indicatively, "position" represents the location information of the target virtual character relative to the first virtual scene, such as using position coordinates. "Auditory perception ability" represents the target virtual character's ability to perceive changes in sound within the first virtual scene. This auditory perception ability includes auditory distance (the farthest distance at which sound can be heard), auditory angle (the angular range of sound coverage that can be determined by the direction of the sound source), and frequency range (the range of sound frequencies that can be perceived). By combining position and auditory perception ability, the auditory perception range of the target virtual character can be roughly determined from the first virtual scene.
[0197] Optionally, the auditory perception range can be a spherical range, a ring-shaped range, an irregular three-dimensional spatial range, etc. The shape of the auditory perception range is not limited here.
[0198] Optionally, during the game configuration phase, the virtual characters can be pre-configured with corresponding auditory perception abilities; different virtual characters can have the same auditory perception abilities; or, different virtual characters can have different auditory perception abilities, such as the main virtual character having stronger auditory perception abilities and non-player characters having weaker auditory perception abilities, etc., without any limitation here.
[0199] (3) The object perception range includes the olfactory perception range of the target virtual character.
[0200] In a schematic sense, the olfactory perception range refers to the spatial range within which an odor is perceived or identified by the sense of smell. The olfactory perception range of a target virtual character is the three-dimensional spatial range that the target virtual character can perceive by smell.
[0201] Optionally, the olfactory perception range of the target virtual character can be obtained based on the target virtual character's position in the first virtual scene and its olfactory perception ability.
[0202] Indicatively, "position" represents the location of the target virtual character relative to the first virtual scene, such as through coordinates. "Olive perception range" represents the target virtual character's ability to perceive changes in odor within the first virtual scene, including olfactory sensitivity (the ability to react to odor molecules to varying degrees) and olfactory distance (indirectly representing the source of the odor, for example, the target virtual character). By combining position and olfactory perception ability, the target virtual character's olfactory perception range can be roughly determined from the first virtual scene.
[0203] Optionally, the olfactory perception range can be a spherical range, a ring-shaped range, an irregular three-dimensional spatial range, etc. The shape of the olfactory perception ability is not limited here.
[0204] Optionally, during the game configuration phase, corresponding olfactory perception abilities can be pre-configured for virtual characters. Furthermore, olfactory propagation distances and intensity can be configured for multiple odor sources. Optionally, different virtual characters may have the same olfactory perception abilities; or, different virtual characters may have different olfactory perception abilities, such as the main virtual character having a stronger olfactory perception ability and non-player characters having a weaker olfactory perception ability. This is not limited here.
[0205] (4) The object perception range includes the range that the target virtual character can perceive with its perception skills or perception tools.
[0206] In illustrative terms, perceptual skills are the ability of a target virtual character to perceive a first virtual scene. Perceptual skills include at least one of perceptual acquisition forms and perceptual enhancement forms.
[0207] Optionally, the perception acquisition form is used to represent the process of acquiring a previously unpossessed perceptual skill. Once a perceptual skill is acquired through this process, the target virtual character can use it to perform targeted scene perception of the first virtual scene. For example, the target virtual character has visual perception capabilities by default, but lacks olfactory and auditory perception capabilities; if the target virtual character acquires olfactory perception skills based on completing a specific task, it can use these skills to perceive the first virtual scene by smell, and can also continue to use its existing visual perception capabilities to perceive the first virtual scene visually, etc.
[0208] Optionally, the perceptual enhancement form is used to represent the process of enhancing a perceptual ability by acquiring a perceptual skill, given the existing perceptual ability. For example, a target virtual character has visual and auditory perceptual abilities by default; if the target virtual character acquires visual perceptual skills based on completing a specific task, it can enhance its existing visual perceptual abilities through these skills, thereby enabling a higher level of visual perception of the first virtual scene, such as determining a larger visual perception range.
[0209] In illustrative terms, a sensory tool is a virtual prop used to perceive a first virtual scene. A sensory tool includes at least one of a perception acquisition form and a perception enhancement form; for example, a sensory tool with a perception acquisition form is a first sensory tool, and a sensory tool with a perception enhancement form is a second sensory tool.
[0210] Optionally, the target virtual character can acquire sensory tools by completing specific game tasks or reaching a certain game level. For example, if the target virtual character acquires a first sensory tool, it can gain additional sensory abilities it does not currently possess, such as additional olfactory perception; if the target virtual character acquires a second sensory tool, it can enhance the level of at least one existing sensory ability, such as improving its existing visual perception ability, and / or improving its existing auditory perception ability, etc.
[0211] Optionally, different sensing skills or sensing tools each correspond to a preset effective range, representing the range that can be perceived when applying that sensing skill or sensing tool. The preset effective range represents the object's perception range. For example, if sensing skill 1 is used to perceive environmental information within 5 meters, the preset effective range is a circular area with a radius of 5 meters centered on the target virtual character; or, if sensing tool 2 is used to perceive environmental information within 10 meters in front, the preset effective range is a semi-circular area with a radius of 5 meters centered on the target virtual character, and so on.
[0212] It is worth noting that the above object perception range is only an illustrative example, and the embodiments of this application do not limit it.
[0213] In an optional embodiment, when the object perception range includes the visual perception range, a preset visual cone is generated as the visual perception range, with the location of the target virtual character as the starting point and the orientation of the target virtual character as the cone's centerline.
[0214] Among them, the preset visual cone is a three-dimensional region that measures the range of visual perception.
[0215] Indicatively, when the object's perception range includes the visual perception range, the position and orientation of the target virtual character are determined. Taking the position as the starting point and the orientation as the center line of the cone, the preset visual cone is used as the visual perception range determined based on visual perception ability.
[0216] Optionally, the preset view cone is a cone shape, with the virtual eye at the location of the target virtual character as the starting point, the direction as the center line of the cone, and a preset value as the radius value of the bottom circular area, thus obtaining the cone as the object perception range.
[0217] Optionally, the preset view cone is a polygonal cone shape, with the virtual eye at the location of the target virtual character as the starting point, the orientation as the center line of the cone, and a preset value as the area value of the bottom polygonal region, thereby obtaining the polygonal cone as the object perception range, etc.
[0218] In an optional embodiment, taking the target virtual character as the main virtual character as an example, obtaining the object perception range of the target virtual character in the first virtual scene is equivalent to obtaining the object perception range of the main virtual character in the first virtual scene.
[0219] In some embodiments, the object perception range of the target virtual character is determined based on the first position of the master virtual character in the first virtual scene.
[0220] Indicatively, the visual perception range of a first preset visual cone, generated with the first position as the starting point and the first orientation as the cone's centerline, is used as the object perception range; and / or, the auditory perception range of a spherical region, obtained with the first position as the center and the first value as the auditory perception radius, is used as the object perception range; and / or, the olfactory perception range of an irregular three-dimensional spatial region, obtained with the first position as the center and the second value as the olfactory perception radius, is used as the object perception range, etc.
[0221] In an optional embodiment, taking a non-player character as an example, obtaining the object perception range of the target virtual character in the first virtual scene is equivalent to obtaining the object perception range of the non-player character in the first virtual scene.
[0222] In some embodiments, the object perception range of the target virtual character is determined based on the second position of the non-player character in the first virtual scene.
[0223] Indicatively, the visual perception range of a second preset visual cone generated with the second position as the starting point and the second orientation as the cone's centerline can be used as the object perception range; and / or, the auditory perception range of a spherical region obtained with the second position as the center and the third value as the auditory perception radius can be used as the object perception range; and / or, the olfactory perception range of an irregular three-dimensional spatial region obtained with the second position as the center and the fourth value as the olfactory perception radius can be used as the object perception range, etc.
[0224] In an optional embodiment, taking the target virtual character as the main control virtual character and the non-player character as an example, when the main control virtual character and the non-player character are in the same first virtual scene, the object perception range of the target virtual character in the first virtual scene is obtained, including: obtaining the first object perception range of the main control virtual character in the first virtual scene, and obtaining the second object perception range of the non-player character in the first virtual scene, and using the first object perception range and the second object perception range together as the object perception range.
[0225] To illustrate, if the main virtual character and the non-player character are both taken as target virtual characters, then the first object perception range can be obtained based on the main virtual character, and the second object perception range can be obtained based on the non-player character. The first object perception range and the second object perception range are used together as the object perception range for subsequent hot word acquisition.
[0226] Step 421: Obtain multiple hot words for the first scene based on environmental perception information within the object perception range.
[0227] Schematic, environmental perception information is used to represent the environmental information perceived and acquired by the target virtual character within the object's perception range. Optionally, environmental perception information includes virtual elements determined by ray emission, and may also include regional structure information determined by analyzing the geometric structure and layout information of the virtual environment.
[0228] In a schematic way, after determining the object perception range corresponding to the target virtual character, based on the range of the object perception range in the first virtual scene, and the environmental perception information determined by the target virtual character based on perception within the object perception range, multiple first scene hot words representing the state of the target virtual character are obtained.
[0229] In an optional embodiment, the environmental awareness information includes virtual elements; identifying virtual elements within the object awareness range in the first virtual scene.
[0230] Indicatively, the object perception range is a portion of the three-dimensional space within the first virtual scene; the first virtual scene includes a large number of virtual elements, which are the elements that constitute the virtual scene, such as virtual ground, virtual buildings, virtual trees, virtual characters, virtual stones, etc.
[0231] Optionally, given the object perception range is determined, the virtual elements within the object perception range are selected from among the multiple virtual elements corresponding to the first virtual scene.
[0232] Indicative, such as Figure 5 As shown, taking the visual perception range as the object perception range as an example, if the first virtual scene in which the main virtual character is located is displayed from the first-person perspective, then the position and orientation of the main virtual object can be determined as follows: Figure 5 The visual cone shown represents the visual perception range.
[0233] Optionally, the position and orientation of the main virtual object are characterized based on the position and orientation of the virtual camera 510; taking the position of the virtual camera 510 as the starting point and the orientation of the virtual camera as the center line of the cone, the following is obtained: Figure 5 The tetrahedral view frustum shown serves as the object's perception range. Alternatively, the two limiting planes in the preset camera rendering are... Figure 5 The near clipping plane 521 and the far clipping plane 522 are parallel to the XY plane of the virtual camera and are separated by a certain distance along their center lines. Any virtual element closer to the virtual camera than the near clipping plane 521 and any virtual element farther from the virtual camera than the far clipping plane 522 will not be rendered. Figure 5 The smaller tetrahedron 532 is cut off from the larger tetrahedron 531 shown, and the resulting solid frustum is used as the view frustum.
[0234] Schematic, the distance between the far clipping plane 522 and the virtual camera 510 is the farthest distance observed by the master virtual object. The view frustum is used as the object's perception range, which includes various virtual elements within the 3D scene, such as virtual grass, virtual buildings, and virtual cars. That is, virtual elements such as virtual cars, virtual grass, and virtual buildings are considered as virtual elements within the object's perception range. Figure 5 As shown, multiple virtual elements are projected onto the far clipping surface 522 as an example. In reality, multiple virtual elements may be distributed and displayed in multiple positions within the visual perception range (such as within the solid truncated pyramid formed by the near clipping surface 521 and the far clipping surface 522).
[0235] Optionally, if the visual perception range is determined from a third-person perspective, and the virtual camera is located outside the target virtual character, it can observe not only multiple virtual elements but also the target virtual character moving within the virtual environment. In this case, the visual perception range can be determined by the area captured by the virtual camera (where the virtual camera moves with the target virtual character, and the visual perception range is still related to the target virtual character); or, the visual perception range can be determined by the frustum whose position and orientation of the target virtual character are determined (see [reference]). Figure 5 As shown, taking the observation perspective of the target virtual character as the same as that of the virtual camera 510 as an example, the process of determining the visual perception range is not limited here.
[0236] In an optional embodiment, the element name of the virtual element is used as the first scene hot word; or, the action name of the interactive action corresponding to the virtual element is used as the first scene hot word.
[0237] Indicatively, each virtual element has a corresponding element name. The element name is used to represent the virtual element itself. There can be one element name or multiple element names, meaning that multiple element names can represent or indicate virtual elements. For example, the element name of virtual element 521 is "truck" or "car"; the element name of virtual element 522 is "grass" or "clump of grass," etc.
[0238] Optionally, after determining the virtual elements within the object's perception range, the element names of the virtual elements are summarized, and multiple element names are used as the first scene hot words corresponding to the first virtual scene, thus obtaining multiple first scene hot words.
[0239] Indicatively, during the game, players can command the main virtual character or non-player characters to interact with virtual elements, that is, there are interactive actions to realize the interaction process, and the names of the interactive actions are used as the first scene hot words.
[0240] For example, virtual element 521 is named a truck or car. Players can control the main virtual character or command non-player characters to interact with virtual element 521, such as getting on the truck, getting off the truck, hitting the truck, climbing on the truck, etc. This is the action name of the interaction. Or, virtual element 522 is named grass or bushes. Players can control the main virtual character or command non-player characters to interact with virtual element 522, such as hiding in the bushes, trimming the bushes, trampling the bushes, etc. This is the action name of the interaction.
[0241] Optionally, the names of actions that may interact with virtual elements are summarized, and at least one action name corresponding to each of the multiple virtual elements is used as the first scene hot words corresponding to the first virtual scene, thus obtaining multiple first scene hot words.
[0242] Step 430: Obtain the first scene type corresponding to the first virtual scene; obtain the set of hot words corresponding to the first scene type.
[0243] For illustrative purposes, the first scene type is the scene type corresponding to the first virtual scene. The scene type is used to represent the scene state of the virtual scene. For example, scene types include various types that describe the scene state, such as office type, battle type, kitchen type, bedroom type, classroom type, shop type, factory type, etc.
[0244] Indicatively, multiple scene types each correspond to a set of hot words. The set of hot words represents the set of scene-related words collected based on the scene type, and also represents the set of scene-related words collected based on at least one virtual scene under that scene type.
[0245] Optionally, after determining the first scene type corresponding to the first virtual scene, a set of hot words corresponding to the first scene type is obtained. The set of hot words is a collection of scene-related words collected based on the first scene type. For example, if the first scene type of the first virtual scene is combat, then the set of hot words corresponding to combat (set 2) is used as the set of hot words corresponding to the first scene type.
[0246] Step 431: Obtain the object perception range of the target virtual character in the first virtual scene.
[0247] Schematic illustration: The target virtual character includes at least one of a controlling virtual character and a non-player character. When obtaining the object perception range based on the target virtual character, if the target virtual character is a controlling virtual character, then the object perception range of the controlling virtual character in the first virtual scene is obtained; if the target virtual character is a non-player character, then the object perception range of the non-player character in the first virtual scene is obtained; or, if the target virtual character is both a controlling virtual character and a non-player character, then both the object perception range of the controlling virtual character in the first virtual scene and the object perception range of the non-player character in the first virtual scene are obtained.
[0248] Optionally, the object perception range includes at least one of the target virtual character's visual perception range, auditory perception range, olfactory perception range, and the range perceived by the target virtual character's perception skills or sensory tools.
[0249] Step 432: Based on the environmental perception information within the object perception range, at least two words are obtained from the hot word set as multiple first scene hot words.
[0250] Optionally, after determining the object perception range, the virtual elements within the object perception range are identified, and words that have element association relationships with the virtual elements are selected from the hot word set as multiple first scene hot words.
[0251] In illustrative terms, element relationships are used to represent the relationships that embody the states of virtual elements. For example: words containing the names of virtual elements are obtained from the hot word set as the first scene hot words; or, words used to trigger virtual elements are obtained from the hot word set as the first scene hot words; or, interactive words used to interact with virtual elements are obtained from the hot word set as the first scene hot words, etc.
[0252] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.
[0253] In summary, using the first virtual scene in which the target virtual character is located as a condition for analyzing natural language commands can provide more targeted textual information for the analysis process through the hot words of the first scene, reducing the ambiguity caused by poor quality of natural language commands. This not only facilitates a more accurate understanding of natural language commands but also makes the command analysis results more closely match the first virtual scene, improving the accuracy of the command analysis results. It greatly avoids ineffective or erroneous commands to non-player characters, providing players with a more efficient and comfortable operating environment and improving human-computer interaction efficiency. Using a small number of first scene hot words can also reduce the amount of data processing during command analysis, improving command analysis efficiency.
[0254] This application describes a method for obtaining hot words in a first scene. This method can obtain more hot words in a first scene by using the first scene type corresponding to the first virtual scene, or by using the object perception range of the target virtual character to obtain more precise hot words in a first scene. Alternatively, it can combine the first scene type and the object perception range to refine the hot word acquisition range while acquiring as many hot words as possible. This ensures that the hot words in a first scene do not deviate from the first virtual scene and can better represent the location of the target virtual character, improving the accuracy of hot word acquisition. This allows the hot words to better constrain the command analysis process, enhancing the correlation between the command analysis results and the first virtual scene, and ultimately improving the accuracy of the command analysis results.
[0255] In an optional embodiment, the virtual environment includes multiple virtual scenes. A first virtual scene can be determined based on the position of the target virtual character within the virtual environment. Then, multiple hot words for the first virtual scene are determined, so that natural language commands are converted into command analysis results under the constraints of these hot words. (Illustrative example, such as...) Figure 6 As shown above, Figure 3 The illustrated embodiment can also be implemented as follows: steps 610 to 640.
[0256] Step 610: Obtain natural language commands in speech form.
[0257] Natural language commands are used to instruct non-player characters.
[0258] As an illustration, the terminal acquires natural language commands in speech form based on the deployed microphone array, and the terminal performs instruction analysis based on the natural language commands itself; alternatively, the terminal sends the natural language commands to the server, which then performs instruction analysis based on the natural language commands.
[0259] In some embodiments, during the game running on the terminal, the presence of natural language commands in the form of voice for directing non-player characters in the physical environment is collected in real time; or, during the game running on the terminal, the presence of natural language commands in the form of voice for directing non-player characters in the physical environment is collected only after the player triggers the audio collection function.
[0260] In some embodiments, when there are multiple non-player characters, natural language commands can be either instructions for controlling multiple non-player characters simultaneously or instructions for directing some of the non-player characters. The command situation of natural language commands is not limited here.
[0261] Optionally, if the natural language command is an instruction to direct some non-player characters, the non-player character to be directed can be selected from multiple non-player characters based on subsequent analysis of the natural language command. For example, if the natural language command is in the form of voice, "Number 1, go and check the room upstairs," then even if the non-player characters include Number 1 and Number 2, this natural language command is only used to direct Number 1 to perform activities, which will not be elaborated here.
[0262] Step 620: Based on the position of the target virtual character in the virtual environment, determine the first virtual scene in which the target virtual character is located from multiple virtual scenes.
[0263] In some embodiments, the first virtual scene is a virtual scene among multiple virtual scenes in a virtual environment, and the virtual environment is a three-dimensional simulation environment in which a main virtual character and non-player characters participate in a virtual game.
[0264] Optionally, multiple virtual characters participating in the virtual game are in a virtual environment. The target virtual character includes at least one of a non-player character and a main virtual character. The virtual environment includes multiple virtual scenes, which together constitute the virtual environment. For example, the virtual environment includes virtual kitchen 1, virtual kitchen 2, virtual office, and virtual battle scene, etc.
[0265] Optionally, the target virtual character can be determined based on the player's choice; or, the target virtual character can be determined based on the default configuration, etc.
[0266] For illustration purposes, if the master virtual character is the target virtual character, then the first virtual scene is determined from multiple virtual scenes corresponding to the virtual environment based on the position of the master virtual character in the virtual environment.
[0267] For illustrative purposes, if the target virtual character is a non-player character, then the first virtual scene is determined from multiple virtual scenes corresponding to the virtual environment based on the non-player character's position in the virtual environment.
[0268] To illustrate, if both the controlling virtual character and the non-player character are considered as target virtual characters, then based on the controlling virtual character's first position in the virtual environment and the non-player character's second position in the virtual environment, at least one virtual scene is determined as the first virtual scene from among multiple virtual scenes corresponding to the virtual environment. For example, if the first position is in virtual scene 1 and the second position is in virtual scene 2, then virtual scene 1 and virtual scene 2 can both be considered as the first virtual scene; or, if the first position is in virtual scene 1 and the second position is also in virtual scene 1, then virtual scene 1 can be considered as the first virtual scene, and so on.
[0269] In other words, the target virtual character is in the first virtual scene.
[0270] In some embodiments, the position coordinates of the target virtual character in the virtual environment are obtained; the scene coordinate intervals corresponding to multiple virtual scenes are obtained respectively; and the first virtual scene in which the target virtual character is located is determined from multiple virtual scenes based on the relative position relationship between the position coordinates and the multiple scene coordinate intervals.
[0271] The scene coordinate range is used to characterize the location range of the virtual scene in the virtual environment.
[0272] In illustrative terms, the position coordinates of the target virtual character represent positional information as coordinate data. Optionally, the position coordinates of the target virtual character can be determined based on the world coordinate system corresponding to the virtual environment.
[0273] Indicatively, multiple virtual scenes within a virtual environment each correspond to a scene coordinate interval. The scene coordinate interval defines the three-dimensional spatial range of the virtual scene from three dimensions: horizontal, vertical, and depth. The scene coordinate intervals corresponding to different virtual scenes are independent and do not overlap. Optionally, the scene coordinate intervals corresponding to multiple virtual scenes are determined based on the world coordinate system of the virtual environment.
[0274] To illustrate, after obtaining the position coordinates, the position coordinates are matched with multiple scene coordinate intervals, and the virtual scene corresponding to the scene coordinate interval including the position coordinates is taken as the first virtual scene where the target virtual character is located. For example: the scene coordinate interval including the position coordinates is called the first scene coordinate interval, the first virtual scene corresponds to the first scene coordinate interval, and the object's coordinates are located within the first scene coordinate interval.
[0275] Optionally, if the master virtual character is the target virtual character, then the first position coordinates of the master virtual character in the virtual environment are obtained; based on the relative position between the first position coordinates and multiple scene coordinate intervals, the first virtual scene where the target virtual character is located is determined from multiple virtual scenes.
[0276] Optionally, if the non-player character is the target virtual character, then the first position coordinates of the non-player character in the virtual environment are obtained; based on the relative position between the first position coordinates and multiple scene coordinate intervals, the first virtual scene in which the target virtual character is located is determined from multiple virtual scenes.
[0277] In an optional embodiment, a first virtual scene corresponding to the first virtual game is determined based on the first virtual game in which the target virtual character participates.
[0278] Each virtual match corresponds to a virtual scene, and these virtual matches provide a simulated battle environment for the target virtual character.
[0279] As an illustration, besides determining the first virtual scene based on the location of the target virtual character in the virtual environment, the virtual scene corresponding to the first virtual game in which the target virtual character participates can also be used as the first virtual scene. For example, each virtual game corresponds to one virtual scene. If the virtual game in which the target virtual character participates is the first virtual game, then the virtual scene corresponding to the first virtual game is used as the first virtual scene.
[0280] Optionally, virtual matches can include various forms such as game missions (e.g., mission 1 corresponds to virtual scene 1, mission 2 corresponds to virtual scene 2, etc.) and game levels (e.g., game level 1 corresponds to virtual scene 1, game level 2 corresponds to virtual scene 2, etc.), which are not limited here.
[0281] Step 630: Obtain multiple hot words for the first virtual scene corresponding to the first virtual scene where the target virtual character is located.
[0282] Among them, many of the hot words in the first scene are scene-related words in the first virtual scene.
[0283] Optionally, the first scene hot words include the scene state of the first virtual scene; or, the first scene hot words include the element names of virtual elements within the first virtual scene; or, optionally, the first scene hot words include interactive words related to interacting with virtual elements within the first virtual scene, etc.
[0284] In an optional embodiment, the object perception range corresponding to the target virtual character is determined from the first virtual scene based on the position and orientation of the target virtual character in the first virtual scene.
[0285] Indicatively, the position is usually represented in coordinate form. To determine the position of the target virtual character in the first virtual scene, we need to determine the coordinates of the target virtual character in the first virtual scene, such as by combining the X-axis, Y-axis, and Z-axis coordinates in three-dimensional space.
[0286] Indicative, the orientation is used to represent the gaze direction of the target virtual character in the first virtual scene. It can usually be represented by an orientation vector, indicating the direction the target virtual character is facing.
[0287] Optionally, the object perception range of the target virtual character can be determined from the first virtual scene by taking into account the position and orientation. The object perception range is the three-dimensional spatial range in which the target virtual character perceives the existence of other scene elements.
[0288] As an illustration, the object's perception range is also affected by the target virtual character's visual perception ability, auditory perception energy, olfactory perception ability, and the perception skills it has acquired or the perception skills corresponding to the sensory tools. The determination of the object's perception range will not be elaborated here.
[0289] Optionally, the object perception range may also be affected by environmental factors such as virtual obstacles, virtual buildings, and virtual terrain in the first virtual scene. For example, the aforementioned environmental factors may affect the reception of visual perception, auditory perception, and other perceptual abilities, thereby affecting the determination of the object perception range.
[0290] Optionally, based on the target virtual character's position coordinates and orientation vector, combined with its visual perception capabilities and other perceptual abilities, as well as environmental factors, the range that the target virtual character can perceive in a specific direction is calculated. Illustratively, starting from the target virtual character's position, a line of sight is projected along its orientation vector within a certain distance to determine areas without obstacles and define the boundaries of the object's perceptual range.
[0291] In some embodiments, if the virtual environment is dynamic, or the position and / or orientation of the target virtual character changes, the calculation of the object perception range may need to be updated in real time.
[0292] Indicatively, by employing collision detection technology, light detection technology, and other techniques in a virtual environment to simulate real-time changes in object perception, a more accurate object perception range can be determined in real time, thereby helping to simulate the behavior and interaction of the target virtual character in the virtual environment.
[0293] In an optional embodiment, multiple first scene hot words are obtained based on environmental perception information within the object perception range.
[0294] As an illustration, after determining the object perception range corresponding to the target virtual character, multiple first scene hot words that are related to the first virtual scene can be obtained based on the object perception range.
[0295] In an optional embodiment, multiple scene hot words are obtained, wherein the multiple scene hot words correspond to at least two virtual scenes.
[0296] This is illustrative of the process of pre-acquiring scene hot words corresponding to multiple virtual scenes in a virtual environment. Optionally, based on the virtual elements in each virtual scene, at least one scene hot word corresponding to each of the multiple virtual scenes is acquired and stored.
[0297] Optionally, the element names of virtual elements in the virtual scene can be obtained as scene hot words; or, the action names of interactive actions that interact with virtual elements in the virtual scene can be obtained as scene hot words; or, the state names used to describe the state of virtual elements in the virtual scene can be obtained as scene hot words, etc.
[0298] In an optional embodiment, based on the first virtual scene in which the target virtual character is located, multiple first scene hot words corresponding to the first virtual scene are obtained from multiple scene hot words.
[0299] In illustrative terms, multiple scene hot words correspond to scene identifiers. The scene identifiers are used to represent the virtual scene when collecting scene hot words. For example, if scene hot words 1 and 2 are obtained based on virtual scene A, scene hot word 1 is marked with scene identifier a corresponding to virtual scene A, and scene hot word 2 is also marked with scene identifier a; if scene hot word 3 is obtained based on virtual scene B, scene hot word 3 is marked with scene identifier b corresponding to virtual scene B, and so on.
[0300] In some embodiments, based on the first virtual scene in which the target virtual character is located, multiple candidate scene hot words with a first scene identifier are obtained from multiple scene hot words. The first scene identifier is the scene identifier corresponding to the first virtual scene.
[0301] As an illustration, after determining the first virtual scene in which the target virtual character is located, and based on the fact that multiple scene hot words are each labeled with a scene identifier, scene hot words labeled with the first scene identifier can be selected from multiple scene hot words as candidate scene hot words.
[0302] In some embodiments, based on the object perception range of the target virtual character in the first virtual scene, at least two candidate scene hot words are selected as the first scene hot words from a plurality of candidate scene hot words.
[0303] As an illustration, after obtaining multiple candidate scene hot words based on the first virtual scene, since the object perception range is a part of the range in the first virtual scene, the multiple candidate scene hot words can be further filtered based on the object perception range, and at least two candidate scene hot words related to the object perception range can be selected as the first scene hot words.
[0304] Optionally, virtual elements within the object's perception range are identified as perceived virtual elements, i.e., perceived virtual elements are virtual elements that the target virtual character can perceive in the first virtual scene; the element names of perceived virtual elements are selected from multiple candidate scene hot words as first scene hot words; or, the action names of interactive actions corresponding to perceived virtual elements are selected from multiple candidate scene hot words as first scene hot words, etc.
[0305] Step 640: Convert natural language commands into text-based command analysis results based on multiple hot words from the first scenario.
[0306] In an optional embodiment, without generating command analysis results, in response to the target virtual character moving from the first virtual scene to the second virtual scene, multiple second scene hot words corresponding to the second virtual scene are obtained.
[0307] In a schematic representation, the virtual environment includes multiple virtual scenes, with the second virtual scene being a different virtual scene from the first. The target virtual character, as a participant in the virtual game, may change location within the virtual environment, such as moving from location A to location B. If, during the process of receiving a natural language command but before obtaining the command analysis result, the target virtual character moves from the first virtual scene to the second virtual scene, then it is necessary to determine the scene hot words used to constrain the generation of the command analysis result based on the most recently moved second virtual scene.
[0308] For example, if the target virtual character moves from the first virtual scene to the second virtual scene during the process of receiving a natural language command but before generating the command analysis result, that is, the target virtual character moves from the first virtual scene to the second virtual scene, then in order to analyze the situation of the virtual scene after the move more flexibly, multiple second scene hot words corresponding to the second virtual scene are obtained.
[0309] Optionally, a scene hot word with a second scene identifier is selected from multiple scene hot words as the second scene hot word; or, a scene hot word with a second scene identifier and associated with the current object's perception range is selected from multiple scene hot words as the second scene hot word.
[0310] The second scene identifier is the scene identifier corresponding to the second virtual scene. The current object perception range is the object perception range determined from the second virtual scene. For example, based on the target virtual character's perception capabilities (visual perception, auditory perception, olfactory perception, etc.), the current object perception range is determined from the second virtual scene.
[0311] In an optional embodiment, natural language commands are converted into command analysis results in text form based on multiple second-scenario hot words.
[0312] To illustrate, after acquiring multiple hot words for the second scenario, the natural language command is converted into a text-based command analysis result using a pre-trained natural language analysis model, constrained by these hot words.
[0313] Optionally, natural language commands and multiple second-scene hot words are used as input to the model. The natural language analysis model extracts the instruction feature representation of the natural language command and the hot word feature representation of the multiple second-scene hot words. The instruction feature representation and the hot word feature representation are fused to obtain a fused feature representation. The command analysis results in text form are output based on the fused feature representation.
[0314] In an alternative embodiment, non-player characters are directed to act within the virtual environment based on command analysis results.
[0315] Indicatively, the command analysis results are text-based content. Compared to natural language commands in speech form, command analysis results are more intuitive and allow for more targeted control of non-player characters.
[0316] Optionally, the forms of directing non-player character activities based on command analysis results include at least one of the following: (1) directing non-player character to move in the virtual environment; (2) directing non-player character to search for specified items; (3) directing non-player character to cooperate with the main virtual character to defend against attacks from enemy virtual characters; (4) directing non-player character to attack enemy virtual characters; (5) directing non-player character to change props, accessories, clothing, etc. No restrictions are placed on the form of activity here.
[0317] For example, if the command analysis result obtained after analyzing the natural language command is "Attack me", it means that a non-player character is instructed to launch a virtual attack on a virtual character that can be attacked in the virtual environment.
[0318] In some embodiments, if the server obtains the command analysis results, it can instruct a non-player character to move in the virtual environment and generate activity animation data based on the command analysis results. Then, the activity animation data is sent to the terminal so that the terminal can render and display the activity animation of the non-player character based on natural language commands.
[0319] In some embodiments, if the server obtains the command analysis results, it can send the command analysis results to the terminal. The terminal can then command the non-player character to move around in the virtual environment based on the command analysis results and generate activity animation data. The activity animation data is then rendered and displayed on the terminal interface, that is, the activity animation of the non-player character moving around based on natural language commands is displayed.
[0320] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.
[0321] In summary, using the first virtual scene in which the target virtual character is located as a condition for analyzing natural language commands can provide more targeted textual information for the analysis process through the hot words of the first scene, reducing the ambiguity caused by poor quality of natural language commands. This not only facilitates a more accurate understanding of natural language commands but also makes the command analysis results more closely match the first virtual scene, improving the accuracy of the command analysis results. It greatly avoids ineffective or erroneous commands to non-player characters, providing players with a more efficient and comfortable operating environment and improving human-computer interaction efficiency. Using a small number of first scene hot words can also reduce the amount of data processing during command analysis, improving command analysis efficiency.
[0322] In this embodiment, a method is described to determine a first virtual scene by the position of a target virtual character, and then determine the object perception range and obtain the content of the first scene hot words based on the position and orientation of the target virtual character. If multiple virtual characters are active in a virtual environment including multiple virtual scenes, the first virtual scene can be specifically determined from the virtual environment based on the position of the target virtual character. In this way, the object perception range can be determined in the first virtual scene by combining the orientation of the target virtual character. Even without generating command analysis results, the second scene hot words can be obtained based on the process of the target virtual character moving to the second virtual scene to constrain the analysis process of the command analysis results. This not only improves the flexibility of obtaining scene hot words used to constrain the analysis process, but also helps to further improve the flexibility of obtaining command analysis results, making it easier to give more accurate and flexible commands to non-player characters through command analysis results in text form.
[0323] In an optional embodiment, a pre-trained natural language analysis model is used to analyze natural language commands to obtain command analysis results in text form. (Illustrative example, such as...) Figure 7 As shown above, Figure 3 Step 230 shown can also be implemented as steps 710 to 730.
[0324] Step 710: Obtain the pre-trained natural language analysis model.
[0325] In illustration, a natural language analysis model is a pre-trained model used to analyze natural language commands in speech form and output the analysis results in text form.
[0326] The pre-trained natural language analysis model includes an acoustic network, a language network, and a pre-set dictionary.
[0327] Alternatively, the natural language analysis model can also be referred to as a natural language analysis system. In this case, the acoustic network, language network, and pre-defined dictionary in the natural language analysis model can also be referred to as the acoustic model, language model, and pre-defined dictionary in the natural language analysis system. Here, there is no limitation on the name of the natural language analysis model and its related networks.
[0328] Intuitively, an acoustic network, also known as an acoustic model, processes input in the form of speech to convert sound signals into potential speech units, such as phonemes or segments. A phoneme is the smallest unit of speech, a phonological unit in language, and is associated with the acoustic features of articulation. A phonological unit is a specific sound unit in the actual speech signal, describing the sound articulation in the actual speech and its acoustic details (such as duration, frequency, and intensity). By analyzing natural language commands in speech form through acoustic networks, possible sequences of speech units can be identified based on the acoustic features of the natural language commands (such as spectrum, tone, and speech rate). Acoustic networks can learn the mapping from acoustic features to speech units using machine learning techniques, such as deep neural networks (DNNs) or recurrent neural networks (RNNs).
[0329] Illustratively, language networks, also known as language models, are used to improve probability scores for highly probable sequences of language units when understanding text input. Based on grammatical rules and statistical information, language networks predict the plausibility and fluency of a given sequence of words or speech units. Language networks include n-gram models, recurrent neural network language models (RNNLMs), and transformer models, used to model and generate text sequences.
[0330] In illustrative terms, a pre-defined dictionary contains words that may be encountered in speech recognition or text understanding, along with their corresponding pronunciations or linguistic features. It provides a way to map text words to their pronunciations or linguistic features; in speech recognition, a pre-defined dictionary helps reduce potential pronunciation ambiguity and improves the system's recognition accuracy.
[0331] The preset dictionary includes multiple hot words for the first scenario, hot words for other scenarios, and general vocabulary.
[0332] For illustrative purposes, the first scene hot words are the scene hot words corresponding to the first virtual scene, the other scene hot words are the scene hot words corresponding to other virtual scenes besides the first virtual scene, and the general vocabulary are words frequently used in daily life, such as "we," "I," "here," "there," "big," "small," etc. The preset dictionary is a collection of multiple words collected in advance, including the first scene hot words, other scene hot words, and commonly used general vocabulary.
[0333] Step 720: Obtain multiple speech units corresponding to natural language commands through an acoustic network.
[0334] Among them, the phonetic unit is the basic building block of word pronunciation.
[0335] In some embodiments, instruction feature representations corresponding to natural language commands are extracted.
[0336] In a schematic way, the feature extraction process is used to represent dividing a continuous sound signal—a natural language command—into multiple short time periods and calculating acoustic features such as spectrum, pitch, and speech rate within these time periods. Each time period corresponds to one acoustic feature. The instruction feature representation can represent a sequence of acoustic features obtained by arranging multiple acoustic features according to the time periods, or it can represent a feature representation after fusing multiple acoustic features. This is not limited here.
[0337] In some embodiments, instruction feature representations are analyzed using acoustic networks.
[0338] To illustrate, using a command feature representation to represent multiple acoustic features, the goal of an acoustic network is to map a continuous sequence of acoustic features to a sequence of speech units, such as phonemes or segments. It learns to predict possible sequences of speech units (phonemes or segments) given specific acoustic features.
[0339] Optionally, the acoustic network outputs one or more possible sequences of speech units, each sequence containing multiple speech units. Different sequences of speech units may contain different speech units, or they may simply have different orders. These multiple sequences of speech units represent the arrangement of the speech units that the language network considers most likely when understanding the input natural language command.
[0340] In some embodiments, a context analysis network is additionally introduced during the analysis of instruction feature representations via acoustic networks. The context analysis network is used to analyze the context and, by considering the temporal dependencies between multiple acoustic features, helps to obtain a more reliable sequence of speech units.
[0341] Step 730: Analyze the sequence matching relationship between multiple speech units and selected words in a preset dictionary through language network analysis to obtain command analysis results in text form.
[0342] As an illustration, the preset dictionary includes multiple words, each corresponding to at least one phonetic unit. After obtaining at least one phonetic unit sequence, the at least one phonetic unit sequence can be matched with the words in the preset dictionary to analyze the at least one word represented by the phonetic unit sequence.
[0343] Optionally, considering the previous process of obtaining multiple first-scene hot words based on the first virtual scene, and since the first-scene hot words are used to constrain the command analysis results, at least multiple first-scene hot words can be used as selected words to be applied in the preset dictionary matching process when matching words through speech unit sequences. Furthermore, considering that speech recognition usually involves some commonly used general vocabulary, general vocabulary can also be used as selected words for word matching. That is, at least one speech unit sequence is matched with the first-scene hot words and general vocabulary in the selected vocabulary to analyze at least one word represented by the speech unit sequence. At least one word is more likely to include first-scene hot words, avoiding interference from other scene hot words. In other words, the selected vocabulary includes at least multiple first-scene hot words and general vocabulary.
[0344] In an optional embodiment, the selected vocabulary also includes other scenario-specific hot words.
[0345] For illustration purposes, if the selected words are the first scene hot words and general words, then at least one speech unit sequence can be matched with the first scene hot words and general words; if the selected words are the first scene hot words, other scene hot words and general words, then at least one speech unit sequence can be matched with the first scene hot words, other scene hot words and general words.
[0346] The vocabulary in the preset dictionary includes analysis weights. The first analysis weight of multiple first-scene hot words is higher than the second analysis weight of other scene hot words. The analysis weight refers to the degree of attention a word receives when participating in sequence matching.
[0347] Indicatively, each word in the pre-defined dictionary corresponds to an analysis weight. The higher the analysis weight, the greater the attention a word receives when participating in sequence matching, meaning it is easier to be noticed; conversely, the lower the analysis weight, the less attention a word receives when participating in sequence matching, meaning it is harder to be noticed. By adjusting the analysis weights, the degree of attention given to different words during sequence matching can be adjusted, so that words with higher analysis weights are prioritized for matching with word sequences, and words with higher analysis weights are also more likely to be selected as words in the command analysis results.
[0348] Optionally, multiple hot words for the first scene are obtained based on the first virtual scene in which the target virtual character is located. The hot words for the first scene can better reflect the situation of the first virtual scene and have a greater probability of being associated with natural language commands. Therefore, the first analysis weight of the hot words for the first scene is set higher than the second analysis weight of other hot words for the scene. This allows the focus to be placed more on the hot words for the first scene when words participate in sequence matching, and thus the selection of the hot words for the first scene as words to form the command analysis results is also more likely.
[0349] Optionally, the general term corresponds to the third analysis weight, where the third analysis weight is less than the first analysis weight and greater than the second analysis weight; or, the third analysis weight is equal to the first analysis weight and greater than the second analysis weight; or, the third analysis weight is less than the first analysis weight and equal to the second analysis weight; or, the third analysis weight is less than the first analysis weight and less than the second analysis weight, etc. The value of the third analysis weight is not limited here.
[0350] The first analysis weight, the second analysis weight, and the third analysis weight are usually values between 0 and 1. As long as the first analysis weight is greater than the second analysis weight, the specific values of the analysis weights are not limited here.
[0351] In an optional embodiment, the matching relationship between multiple speech units and selected words in a preset dictionary is analyzed to obtain multiple candidate word sequences.
[0352] This illustrative process illustrates that analyzing matching relationships represents analyzing the matching between selected words and speech units. Each selected word corresponds to at least one speech unit. By analyzing multiple speech units and the selected words, the conversion probability of a speech unit transforming into the selected word is determined. The conversion probability represents the probability of a match between the selected word and the speech unit. The higher the conversion probability, the closer the speech unit is to the selected word, meaning the greater the likelihood that the speech unit might express the selected word, and thus the stronger the matching relationship between the selected word and the speech unit.
[0353] Optionally, multiple speech units form at least one speech unit sequence. Based on the matching analysis process between the speech units and selected words, at least one candidate word sequence corresponding to each of the at least one speech unit sequence is determined, and multiple candidate word sequences are obtained.
[0354] Schematic, based on the analysis of natural language commands using a sound network, we obtain speech unit sequence 1 and speech unit sequence 2. Speech unit sequence 1 consists of speech unit a—speech unit b—speech unit c, and speech unit sequence 2 consists of speech unit a—speech unit d—speech unit c. These can be considered as multiple speech units including speech unit a, speech unit b, speech unit c, and speech unit d. After matching and analyzing these multiple speech units with selected words in a preset dictionary, the candidate word sequences obtained based on speech unit sequence 1 include candidate word sequence 11 and candidate word sequence 12. Candidate word sequence 11 is "Help me attack this place," and candidate word sequence 12 is "Help me attack the enemy." The candidate word sequences obtained based on speech unit sequence 2 include candidate word sequence 21 and candidate word sequence 22. Candidate word sequence 21 is "Help me supply this place," and candidate word sequence 22 is "Help me attack the enemy." Thus, we obtain multiple candidate word sequences including candidate word sequence 11, candidate word sequence 12, candidate word sequence 21, and candidate word sequence 22.
[0355] Among them, the multiple candidate word sequences include at least one word from multiple selected words, and the candidate word sequences are word sequences obtained based on the change relationships between multiple speech units.
[0356] For illustrative purposes, since the candidate word sequence is a word sequence obtained based on selected words in a preset dictionary, the words in the candidate word sequence are the words in the selected words; the candidate word sequence can be a sequence obtained from a single word or a sequence obtained from a combination of multiple words, without limitation here.
[0357] Since the candidate word sequence is a sequence obtained by matching multiple speech units with selected words, the speech unit transformation of at least one word in the candidate word sequence matches the transformation between multiple speech units. That is, the candidate word sequence is a word combination sequence predicted based on the changes between multiple speech units.
[0358] In an optional embodiment, the sequence semantics of multiple candidate word sequences are analyzed through a language network, and at least one candidate word sequence is obtained from the multiple candidate word sequences as a command analysis result in text form.
[0359] Indicatively, sequence semantics is used to represent the semantic information expressed by a candidate word sequence. It can reflect not only the semantic changes of at least one word in the candidate word sequence, but also the sentence semantics of the entire candidate word sequence.
[0360] Indicatively, the language network evaluates multiple candidate word sequences separately, and then selects at least one candidate word sequence that best fits the current semantic situation as the result of the command analysis.
[0361] Optionally, the probability of multiple candidate word sequences is evaluated through a language network to obtain the predicted probabilities of each candidate word sequence; the candidate word sequence with the highest predicted probability is taken as the command analysis result.
[0362] Indicatively, the prediction probability is used to characterize the probability estimation result of the candidate word sequence being correct under the constraints of multiple speech units. For example, the prediction probability of candidate word sequence 11 is 0.32, the prediction probability of candidate word sequence 12 is 0.8, the prediction probability of candidate word sequence 22 is 0.69, and the prediction probability of candidate word sequence 23 is 0.45. If the candidate word sequence with the highest prediction probability is taken as the command analysis result, then the output of candidate word sequence 12—"Help me attack the enemy"—is the command analysis result.
[0363] In some embodiments, after analyzing the sequence semantics of multiple candidate word sequences through a language network, probability distributions corresponding to the multiple candidate word sequences are generated; the multiple probability distributions and multiple speech units are input into a decoding network, and based on the constraints of the multiple speech units, at least one candidate word sequence is output as the command analysis result.
[0364] Indicatively, the probability distribution is used to describe the naturalness between at least one word in the candidate word sequence; the decoding network (or decoder) is used to combine multiple candidate word sequences, the corresponding probability distributions, and multiple language units to select the best candidate word sequence as the command analysis result.
[0365] Optionally, the decoding network may use beam search or other optimization algorithms to search among multiple candidate word sequences. Considering the probability distribution of the language network output and the constraints of multiple speech units (phonemes or segment sequences) of the acoustic network output, beam search can find the most likely candidate word sequence as the command analysis result in the potential sequence space, thereby improving the accuracy of speech recognition.
[0366] It is worth noting that the decoding network can be configured outside of the sound network and the language network, or it can be configured inside the language network. The preset dictionary can be configured inside the language network or outside the language network as an independent neural network layer. There are no restrictions on the network structure of the natural language analysis model here.
[0367] In some embodiments, the language network includes language subnetworks corresponding to at least two virtual scenes.
[0368] In a schematic representation, multiple virtual scenarios each correspond to a language subnetwork, with different language subnetworks used for targeted analysis of their respective virtual scenarios. For example, virtual scenario A corresponds to language subnetwork 1, and virtual scenario B corresponds to language subnetwork 2. When analysis is needed based on virtual scenario A, language subnetwork 1 is used for language analysis; when analysis is needed based on virtual scenario B, language subnetwork 2 is used for language analysis, and so on.
[0369] Optionally, if the natural language analysis model is referred to as a natural language analysis system, a language model can be configured to correspond to at least two virtual scenarios respectively, with multiple language models corresponding one-to-one with multiple virtual scenarios. Here, there are no restrictions on the model structure or system architecture; the above content is merely an illustrative description of the analysis process.
[0370] In some embodiments, based on the first virtual scene in which the target virtual character is located, a first language subnetwork corresponding to the first virtual scene is determined from at least two language subnetworks.
[0371] Indicatively, the first virtual scene is a virtual scene among at least two virtual scenes. In addition to determining the hot words of the first scene based on the first virtual scene in which the target virtual character is located, the first language sub-network corresponding to the first virtual scene will also be determined from at least two language sub-networks based on the first virtual scene. The first language sub-network is a sub-network layer that analyzes the candidate word sequence based on the first virtual scene.
[0372] In some embodiments, the sequence semantics of multiple candidate word sequences are analyzed through a first language sub-network, and at least one candidate word sequence is obtained from the multiple candidate word sequences as a command analysis result in text form.
[0373] Optionally, different language subnetworks are language subnetworks pre-trained based on corresponding virtual scenes, wherein the first language subnetwork is a language subnetwork trained based on the first virtual scene.
[0374] To illustrate, taking the training of the first language subnetwork through the first virtual scene as an example, the first scene hot words corresponding to the first virtual scene are obtained, and the first scene hot words and general vocabulary are used as sample data to train the first language subnetwork; or, multiple scene description words used to describe the first virtual scene are obtained, and the scene description words are used as sample data to train the first language subnetwork. The scene description words include multiple words describing the virtual scene state and describing the virtual elements therein. Other language subnetworks are also trained using corresponding virtual scene-related vocabulary (scene hot words, general vocabulary, scene description words, etc.), which will not be elaborated here.
[0375] Optionally, after obtaining multiple candidate word sequences under the constraint of hot words in the first scene, semantic analysis is performed on the multiple candidate word sequences through the first language sub-network corresponding to the first virtual scene to obtain the prediction probabilities corresponding to the multiple candidate word sequences respectively, and at least one candidate word sequence with the highest prediction probability is taken as the command analysis result in text form.
[0376] It is worth noting that the above process of determining the corresponding first language sub-network based on the first virtual scene in which the target virtual character is located is only an illustrative example. Multiple virtual scenes can also be configured to correspond to a sound network (such as a sound network trained based on the corresponding virtual scene). The corresponding first sound sub-network can be selected for analysis based on the first virtual scene, or the corresponding first sound sub-network and first language sub-network can be selected for analysis based on the first virtual scene. No limitation is imposed here.
[0377] In summary, using the first virtual scene in which the target virtual character is located as a condition for analyzing natural language commands can provide more targeted textual information for the analysis process through the hot words of the first scene, reducing the ambiguity caused by poor quality of natural language commands. This not only facilitates a more accurate understanding of natural language commands but also makes the command analysis results more closely match the first virtual scene, improving the accuracy of the command analysis results. It greatly avoids ineffective or erroneous commands to non-player characters, providing players with a more efficient and comfortable operating environment and improving human-computer interaction efficiency. Using a small number of first scene hot words can also reduce the amount of data processing during command analysis, improving command analysis efficiency.
[0378] In this embodiment, a process is described whereby, in addition to determining hot words for a first virtual scene based on the location of the target virtual character, a corresponding first language sub-network can be determined for targeted analysis based on the first virtual scene. This involves training corresponding language sub-networks for multiple virtual scenes, enabling the language sub-networks to perform more purposeful analysis of candidate word sequences under the constraints of the virtual scenes, avoiding the problem of blindly obtaining command analysis results. The process of obtaining candidate word sequences through the constraints of hot words for the first scene improves the accuracy of the candidate word sequences. Furthermore, combining this with the first language sub-network to further constrain the analysis process of the candidate word sequences helps to further improve the correlation between command analysis results and the first virtual scene, enhancing the accuracy of natural language command analysis. This facilitates more precise command control of non-player characters through more accurate command analysis results.
[0379] In an optional embodiment, the above-described virtual scene-based speech conversion method can be applied to a game scenario, where players issue natural language commands in voice form during gameplay, allowing the terminal to direct non-player characters based on these natural language commands. Therefore, the above-described virtual scene-based speech conversion method can also be called "an adaptive voice interaction method in a game scenario." (Illustrative example, such as...) Figure 8 The diagram shows a technical flowchart of a speech conversion method based on a virtual scene, which includes steps 810 to 880.
[0380] Step 810: The player enters the game level.
[0381] In a game, a level is a specific section or stage in a video game where players must complete specific objectives or tasks to advance to the next level. Levels are typically designed to progressively increase in difficulty to challenge players' skills and strategies. Each level may contain different environments, enemies, puzzles, and rewards, aiming to provide a diverse gaming experience.
[0382] Optionally, players can control a main virtual character to participate in virtual matches during gameplay. Players can also command other non-player characters to cooperate with the main virtual character in virtual matches. Non-player characters are virtual characters other than the main virtual character, such as virtual minions and virtual heroes deployed by default in the virtual match. For example, players can manually control the main virtual character to engage in virtual battles, and can also use voice commands to direct non-player characters to engage in virtual battles.
[0383] Taking the aforementioned target virtual character as the main control virtual character as an example, when determining the object perception range used to obtain the first scene hot words, the visual perception range (or object field of view) is used as the object perception range. Based on the movement of the main control virtual character in the virtual environment, the field of view of the main control virtual character can be updated in real time.
[0384] Step 820: Prepare a list of hot words based on virtual elements in the virtual scene of the game level.
[0385] In a typical game, multiple levels are provided, each corresponding to a virtual environment. Each virtual environment includes multiple virtual scenes. A game scene refers to the three-dimensional spatial area where the target virtual character resides within a specific level. Virtual scenes are not merely visual representations; they also include virtual elements, sound effects, and other details, all contributing to the player's gaming experience. A well-designed virtual scene can significantly enhance a game's immersion and enjoyment. Multiple virtual scenes may correspond to different scene types, which are not specified here.
[0386] Optionally, each scene type has its own unique virtual elements and atmosphere, which can enhance the game's storytelling and interactivity. For example, in a forest scene, players may need to avoid wild animals or find hidden paths; while in a futuristic city scene, players may need to use high-tech equipment to complete tasks.
[0387] Optionally, each game level may contain specific virtual elements, such as keys to unlock specific doors or treasure chests—doors leading to new areas or levels that require specific keys or completion of specific tasks. Virtual items obtained in specific game levels may have unique attack effects or abilities, such as providing additional defense or special resistances to help players deal with specific hostile virtual characters. Game levels contain unique virtual items, consumables, collectibles, gems / coins; additionally, game levels include puzzle items, specific vehicles, specific abilities (magic, skills, etc.); and items that interact with the environment, such as virtual explosives, virtual obstacles, and virtual buildings. Among these, the virtual elements unique to each specific level typically have unique functions or effects that can influence the player's gaming experience.
[0388] In some embodiments, a list of hot words corresponding to each of the multiple virtual scenes is prepared based on the virtual elements present in each of the multiple virtual scenes; or, a list of hot words corresponding to the multiple virtual scenes is prepared, which includes scene hot words corresponding to each of the multiple virtual scenes.
[0389] Optionally, to obtain a richer collection of scene hot words, this can be achieved by designing a hot word management system. This system preprocesses all game assets and can calculate the position and direction of target virtual characters during real-time player interaction. Based on the asset preprocessing, a hot word library (or hot word list) is generated, containing scene hot words for various virtual scenarios.
[0390] In some embodiments, virtual elements existing within the object perception range are determined from a first virtual scene where the target virtual character is located, based on the object perception range. That is, virtual elements within the object perception range are determined.
[0391] As an illustration, when the virtual scene changes, the hot word management system can automatically adjust the required scene hot words. Based on the current scene, it first selects candidate scene hot words from the predefined hot word library, and then performs a secondary selection based on the player's field of vision (object perception range).
[0392] Optionally, taking the field of view as the object's perceived range as an example, when determining the field of view of the target virtual character, the field of view is determined by the camera's frustum. The frustum is a spatial region defined by the camera's position and field of view (FOV). To improve efficiency, spatial segmentation techniques, such as quadtrees, octrees, or grids, can be used to organize virtual elements in the virtual scene. These data structures help quickly determine which virtual elements are within the field of view. After determining the spatial range, frustum culling is performed, calculating which virtual elements are within the frustum to determine which virtual elements need to be rendered.
[0393] To illustrate, six planes (top, bottom, left, right, front, and back) are extracted from the camera's view frustum to create the view frustum plane. For each virtual element, its bounding box is tested against the view frustum plane. If the bounding box is completely outside the view frustum, the virtual element's element name, action name, etc., do not need to be loaded into the first scene hotwords.
[0394] As an illustration, based on the results of view frustum culling, it is determined which virtual element names, action names, etc., need to be loaded into the first scene hot words that need to be used, and which virtual element hot words can be unloaded from memory. Optionally, thresholds for hot word loading and unloading can also be set to avoid frequent hot word loading and unloading operations.
[0395] Hierarchical Frustum Culling: Combining spatial segmentation techniques, it first unloads the element names and action names of virtual elements in large regions, and then unloads the element names and action names of virtual elements in small regions.
[0396] Among them, scenario-based hot words are words frequently used in specific application scenarios. By setting scenario-based hot words, the system can more accurately identify these words, thereby improving the overall recognition accuracy. These hot words either only appear in specific game levels of a certain game, or are words frequently used by players in specific game levels. These words are difficult to translate from speech to text based on pronunciation alone, or their usage frequency is very high, thus affecting the user experience.
[0397] Optionally, an initial list of trending terms can be compiled. This list of trending terms may include the following:
[0398] (1) Scene-specific item terms: such as “Area A”, “virtual stable”, “virtual crash site”, virtual “oil tanker”, etc.;
[0399] (2) Scene-specific character terms: such as “virtual hero 1”, “virtual soldier 2”, “virtual monster”, etc.;
[0400] (3) Common commands: such as “attack”, “defend”, “jump”, “run”, etc.
[0401] (4) Specific task terms: such as "open the door", "unlock the treasure chest", "dialogue", etc.;
[0402] (5) Environmental interaction words: such as “view map”, “use props”, “save game”, etc.
[0403] Step 830: Match the hot word list.
[0404] In a schematic way, after determining the virtual elements within the object's perception range, a list of hot words is matched to determine the scene hot words corresponding to the virtual elements as the first scene hot words.
[0405] Optionally, taking the example of multiple virtual scenes each corresponding to a hot word list, a first hot word list corresponding to the first virtual scene is determined based on the first virtual scene, and multiple first scene hot words are determined from the first hot word list based on virtual elements within the object's perception range. For example, the element name describing the virtual element can be used as the first scene hot word, or the action name of the interactive action that interacts with the virtual element can be used as the first scene hot word, etc.
[0406] Optionally, taking the example of multiple virtual scenes corresponding to a hot word list, multiple first scene hot words are determined from the hot word list based on virtual elements within the object's perception range. For example, the element name describing the virtual element can be used as the first scene hot word, or the action name of the interactive action that interacts with the virtual element can be used as the first scene hot word, etc.
[0407] Step 840, ASR system hot word weighting.
[0408] In illustrative terms, ASR (Automatic Speech Recognition) systems are commonly used in speech recognition technology, aiming to convert human speech instructions into text instructions. It involves processing, analyzing, and understanding sound signals, enabling computers to "understand" and perform corresponding operations.
[0409] The ASR system is the pre-trained natural language analysis model mentioned above. The ASR system includes the following key components.
[0410] (1) Acoustic model for speech recognition (acoustic network in the above natural language analysis model)
[0411] Acoustic models are used to map feature vectors of audio signals to basic speech units (such as phonemes). They are typically trained using machine learning techniques (such as deep neural networks) that use large amounts of labeled speech data to learn the relationship between speech signals and phonemes.
[0412] (2) Speech recognition language model (language network in the above natural language analysis model)
[0413] Language models are used to predict the probability of word sequences, helping the system select the most likely one from multiple possible word sequences. Typically, end-to-end networks already contain some language model information, and statistical methods (n-gram models) are used for decoding constraints to improve accuracy.
[0414] (3) Hot words for speech recognition (hot words for the first scene in the preset dictionary of the above natural language analysis model, or selected words including general vocabulary)
[0415] Hot words refer to words or phrases that require special attention and priority identification in specific virtual scenarios. For example, in a game scenario, "virtual stable" or "virtual crash site" are hot words. Hot words usually require higher recognition accuracy, so specialized models or algorithms may be used to process them. These hot words are often related to the virtual scenes in the game, and their corresponding words cannot be written based solely on pronunciation.
[0416] (4) Speech recognition decoder (the decoding network mentioned above)
[0417] The decoder is one of the core components of a speech recognition system. It combines the outputs of the acoustic model and the language model, using search algorithms (such as the Viterbi algorithm) to find the most likely sequence of words. The decoder's task is to convert the sequence of feature vectors into the final text output.
[0418] In an optional embodiment, a Weighted Finite-State Transducer (WFST) is used to handle the relationship between states and transitions. In speech recognition, states can represent different meanings based on different inputs, and transitions represent the changes between states. For example, when processing multiple speech units, each speech unit can be considered a state, and the transition between speech units (such as the transition from speech unit 1 to speech unit 2) is a state transition. Similarly, in the analysis of candidate word sequences, each word in the candidate word sequence can be considered a state, and the transition between words (such as the transition from word 1 to word 2) is a state transition, and so on.
[0419] To illustrate, WFST can assign a weight to each transition. The weight typically represents the cost or probability of the transition, such as the weight for transitioning from speech unit 1 to speech unit 2; or the weight for transitioning from word 1 to word 2, etc.
[0420] In some embodiments, WFST is used to represent various models and transformation relationships, such as acoustic models, language models, and pre-defined dictionaries. In speech recognition systems, WFST is used to combine different models (such as acoustic models, language models, and pre-defined dictionaries) into a unified decoding network, which can efficiently search for the most probable word sequences.
[0421] Schematic representation: the acoustic model represents the probability of phonemes or other basic speech units as a WFST; the language model represents the probability of candidate word sequences as a WFST; and the pre-defined dictionary represents the mapping between words and their corresponding phoneme sequences (speech unit sequences) as a WFST. The relationship between the acoustic model, language model, and WFST is as follows: in a speech recognition system, the acoustic model, language model, and dictionary are typically represented as independent WFSTs, and then combined into a comprehensive WFST through specific combination and optimization algorithms. This comprehensive WFST is used for decoding and recognition.
[0422] Optionally, the acoustic model (H) maps multiple acoustic features corresponding to natural language commands in speech form to phonemes or other basic speech units, resulting in a sequence of speech units. This mapping can be represented as a WFST, called H. Context-dependent models (C) are typically used within the acoustic model to obtain a sequence of speech units that is more closely related to the context.
[0423] The predefined dictionary (L) maps words to sequences of phonetic units, resulting in multiple candidate word sequences. This mapping can also be represented as a WFST, called L.
[0424] The language model (G) represents the probability of candidate word sequences as WFST, which is called G.
[0425] By combining H, C, L, and G, the corresponding HCLG for the speech recognition system is obtained, which generates a comprehensive WFST, called HCLG. This comprehensive WFST represents the complete mapping from feature vectors to word sequences. During decoding, the ASR system uses this comprehensive WFST, HCLG, to search for the most likely candidate word sequences as the command analysis result.
[0426] In a schematic manner, natural language commands in speech form are converted into multiple acoustic features. The acoustic model (H) and context-dependent model (C) in HCLG are used to score the multiple acoustic features and generate phoneme probabilities. Then, the phoneme probabilities are mapped to word sequences using a pre-defined dictionary (L) in HCLG. Finally, the command analysis results are obtained through language model (G), such as finding the most likely word sequence as the command analysis result through a decoder using a search algorithm (such as the Viterbi algorithm).
[0427] In an optional embodiment, it is necessary to prepare voice and text data related to the game levels. This data will be used to train the acoustic and language models. The voice data includes recorded and collected audio data of voice forms that players may use in the game; the text data includes text previously entered by the player and game level-related text such as level names and action commands associative input from the large language model.
[0428] Optionally, features are extracted from the speech-form natural language commands using Mel-Frequency Cepstral Coefficients (MFCCs) and fbank filter banks to train the acoustic model. After training the acoustic model using the extracted features, a Connectionist Temporal Classification (CTC) level model probability is obtained. The language model (LM) is trained using the data described above. Commonly used tools include the SRI Language Modeling Toolkit and the Ken Language Model Toolkit (KenLM). A pronunciation dictionary is constructed, containing all possible words and their corresponding phoneme sequences. The pronunciation dictionary format is one word per line and its phoneme representation. Then, graph construction tools such as arpa2fst, fstcompose, and fstreduced are used to construct the speech recognition decoding word graph, respectively.
[0429] Step 850: Perform input correction based on the large language model.
[0430] In one alternative embodiment, a large language model is used to generate sample data for training the ASR system, thereby training the ASR system with sample data before applying the ASR system; or, the ASR system is periodically trained with sample data during the application of the ASR system; or, the ASR system is trained in real time with sample data during the application of the ASR system, etc.
[0431] In illustration, a Large Language Model (LLM) refers to a highly complex and high-performance Natural Language Processing (NLP) model trained on massive amounts of text data. LLMs utilize deep learning techniques, particularly transformer architectures, to understand and generate natural language commands. LLMs excel in various NLP tasks, such as text generation, translation, question answering, and summarization. Large-scale data training: LLMs are typically trained using vast amounts of text data, potentially including billions or even trillions of words. This large dataset allows the model to learn richer linguistic features and patterns. Complex model architecture: These models often feature complex architectures, including multi-layer transformers and attention mechanisms, to capture subtle nuances in language. High performance: Due to the use of large amounts of data and complex model architectures, LLMs typically possess high understanding and generation capabilities, performing well across various language tasks.
[0432] Large language models employ prompts, which are the initial input text provided to the model when performing natural language processing tasks. The design and selection of prompts significantly impact the quality and accuracy of the model's output. Carefully designed prompts can guide the model to generate text that better reflects expectations. In other words, prompts provide contextual information, helping the model understand the user's intent and generate relevant text. For example, given the prompt "Write an article about climate change," the model will generate an article related to climate change. Prompts can also control the output style and format; they can contain specific instructions or formatting requirements to help control the model's output style and format. For instance, the prompt "Explain quantum mechanics in concise language" will guide the model to generate a concise and easy-to-understand explanation.
[0433] Furthermore, prompts can include previous dialogue content to help the model understand the context of the current conversation, thus generating more coherent responses. Effective prompt design requires consideration of several aspects: clear instructions; prompts should contain clear instructions, telling the model what task needs to be completed. For example, the prompt "Please summarize the main points of the following article" is clearer than the prompt "summarize". Prompts can also provide sufficient contextual information to help the model understand the task. For example, in dialogue systems, prompts can include previous dialogue content. Additionally, if there are specific formatting requirements for the output, these can be explicitly stated in the prompt. For example, the prompt "Please list the answers to the following questions in list format".
[0434] Based on player commands, dialogue, descriptions, etc. in the game. Provide context: Give the model enough contextual information so that it can generate reasonable voice input. Use examples: Guide the model to generate similar output by providing some examples. Provide explicit format: Ensure that the prompts are clearly stated.
[0435] In some embodiments, the process of generating sample data based on a large language model is described as follows.
[0436] To illustrate, suppose there is an adventure game where players can command non-player characters via natural language commands in the form of voice. Therefore, it is desirable for a large language model to generate some possible sample data to train the natural language analysis model in a more targeted manner.
[0437] Optionally, a pre-defined command template can be used to generate text corpora. For example, suppose the template is a string containing placeholders for inserting specific player command parameters. Placeholders can be represented by {}. For example, "{player_name}go to {direction}{location}{action}" means that the player's name (such as a non-player character's name) goes to a certain direction and location to perform a certain action. Suppose there are the following player command parameters: player_name = "teammate number one goes"; direction = "ahead"; location = "in the forest"; action = "throw a grenade"; concatenating the commands yields template = "teammate number one {player_name} is in {location}{action}", which means: teammate number one goes to the forest ahead and throws a virtual grenade.
[0438] Alternatively, if the player is in an adventure game, they can command a virtual character using natural language commands. Example commands include: 1. Move north; 2. Attack the enemy ahead; 3. Use a healing potion; 4. Open your inventory; 5. View the map. These commands allow the player to control the virtual character.
[0439] Based on this example instruction, sample data of some possible speech forms can be generated based on a large language model as data for training a natural language analysis model.
[0440] As an illustration, more contextual information can be provided to make the generated voice input more specific. For example, the prompts provided for the large language model are: "You are a player in an adventure game, your character is a warrior, and you are currently in the forest. You can command your character using voice input commands. Here are some example commands: 1. Walk north; 2. Attack the enemy in front; 3. Use a healing potion; 4. Open your inventory; 5. View the map; Current game state: Character: Warrior; Location: Forest; Quest: Find hidden treasure; Now, please generate some possible player voice inputs."
[0441] As an example, to make the generated voice input more diverse, more examples and different scenarios can be added to the prompts. For example, the prompts provided for the large language model could be: "You are a player in an adventure game, your character is a warrior, currently in a forest. You can use voice input commands to control your character; here are some example commands: 1. Walk north; 2. Attack the enemy ahead; 3. Use a healing potion; 4. Open your inventory; 5. View the map; 6. Pick up the gems on the ground; 7. Talk to the villagers; 8. Enter the cave; Current game state: warrior; Location: forest; Mission: Find hidden treasure; Now, please generate some possible player voice inputs," etc.
[0442] Optionally, based on the above prompts, the large language model may generate the following text: 1. Head east; 2. Use fireball to attack the enemy; 3. Check the quest log; 4. Drink a magic potion; 5. Search the nearby area; 6. "Trade with the merchant"; 7. "Enter the castle"; 8. "Equip a new sword", etc.
[0443] Through the above process, a large amount of text simulating player speech in a scene can be generated by the built-in human speaking habits of the large language model, and the text can be used as sample data; and / or, the text can be converted into speech to obtain speech input as sample data, thereby using richer sample data to train the natural language analysis model.
[0444] In an optional embodiment, text generated by a large language model is used as sample data, and the language model is updated using the sample data before being applied to the ASR system.
[0445] Step 860: Activate the language model for the current game level.
[0446] In an optional embodiment, the language model (G) in HCLG includes multiple language models, each corresponding to a multiple virtual scene. Based on the first virtual scene in which the target virtual character is located, the first language model corresponding to the first virtual scene is determined from the multiple language models (i.e., the first language sub-network corresponding to the first virtual scene is determined from the language network as described above).
[0447] Optionally, multiple language models are obtained by training on corresponding virtual scenes. During the process of training the language model using the text generated by the large language model as sample data, the text corresponding to each of the multiple virtual scenes is obtained, and the language model corresponding to the virtual scene is trained through the text.
[0448] Step 870: Select the language model during ASR system decoding.
[0449] To illustrate, when the ASR system performs decoding, it selects the first language model corresponding to the first virtual scene so that the content of the sound model input can be analyzed through the first language model.
[0450] Step 880, speech recognition and decoding.
[0451] Indicatively, multiple speech units corresponding to natural language commands are obtained through the acoustic model in the ASR system. Speech units are the basic building blocks of word pronunciation. The sequence matching relationship between multiple speech units and selected words in the preset dictionary is analyzed through the language model in the ASR system. The command analysis result in text form is the decoding result. The selected words include at least several hot words from the first scenario and general vocabulary.
[0452] In one optional embodiment, when the player-controlled target virtual character moves between different buildings, the spatial query logic provides complete location information to help the player quickly locate the enemy virtual character. For example, if the target virtual character is outside a motel and hears virtual gunfire coming from the second floor of the motel, the system will prompt "Virtual gunfire is coming from the second floor of the motel." When the player moves within the same building: When the player moves within the same building, the spatial query logic simplifies the feedback information, providing only information for sub-areas. For example, if the target virtual character is on the first floor of the motel and the enemy virtual character is on the second floor, the system will prompt "The enemy is on the second floor." When the player is in the same room: When the target virtual character and the enemy virtual character are in the same room, the spatial query logic provides more precise directional information to help the player react quickly. For example, if both the player and the enemy are on the second floor of the motel, the system will prompt "The enemy is to the right front." This hierarchical feedback helps players understand the enemy virtual character's location more intuitively, providing the most relevant location information based on the player's location.
[0453] In an optional embodiment, in game speech recognition, loading different language models and scene hot words according to different virtual scenes can significantly improve the accuracy and response speed of speech recognition. The following are the steps and methods to achieve this goal.
[0454] First, it is necessary to be able to identify the virtual scene in the current game. The loading of the language model decoding word map will be achieved in the following ways: (1) Game state detection: by detecting the game's state variables (such as player position, task progress, current activity, etc.), the first virtual scene in which the target virtual character is located is determined; (2) Event triggering: the scene is switched according to a specific event (such as entering a certain area or starting a certain task), thereby determining the first virtual scene in which the target virtual character is located; (3) Player input: the player can actively inform the first virtual scene mentioned above through voice or other input methods.
[0455] This involves training or selecting different language models for different virtual scenarios. For example, the language model for a combat scenario can focus more on vocabulary and phrases related to virtual combat, while the language model for an exploration scenario can focus more on vocabulary related to navigation and interaction.
[0456] This involves generating a set of scene hot words based on the virtual elements (such as virtual items and virtual characters) within the player's field of vision. These scene hot words are the most frequently triggered or most important words in the current player's scenario. For example, in a combat scenario, hot words might include "enemy," "attack," and "retreat," while in a trading scenario, hot words might include "buy," "sell," and "item."
[0457] As an example, based on the identified virtual scene, the corresponding language model and scene hot words are dynamically loaded.
[0458] Specifically, when a scene change is detected, the system switches to the language model corresponding to the virtual scene, which can be achieved by calling the speech recognition engine's API. In addition, when a scene change is detected, the system updates the list of currently used hot words so that the speech recognition engine can recognize these words more accurately.
[0459] In an optional embodiment, the decoding strategy of the speech recognition system can be dynamically adjusted based on the virtual scene in which the target virtual character is located, thereby improving the accuracy of speech recognition and the user experience. The following are some methods and steps.
[0460] As an illustration, the game status, player position, and current task are used to identify and classify the player's current scene.
[0461] Optionally, different decoding strategies can be defined according to different virtual scenarios. These strategies include: (1) vocabulary adjustment: adjust the vocabulary of speech recognition according to the virtual scenario, that is, the selected words in the preset dictionary mentioned above; (2) language model adjustment: use different language models to process different virtual scenarios; (3) confidence threshold adjustment: adjust the confidence threshold of speech recognition according to the virtual scenario, such as manually increasing the confidence threshold, and only when the predicted probability is higher than the confidence threshold will the candidate word sequence corresponding to the maximum preset probability be used as the instruction output result.
[0462] Among these features, by dynamically adjusting the decoder during game operation, the decoding strategy of the speech recognition device can be dynamically adjusted according to the virtual scene in which the target virtual character is located, thereby improving the flexibility of decoding.
[0463] In an optional embodiment, after the ASR system goes online, it can be continuously optimized and adjusted according to the actual usage of players to make it more adaptable to the needs of different virtual scenarios. Optimizing the language model includes the following operations: (1) Hot word adjustment: Based on player feedback and usage data, collect the text recognized by the ASR system and automatically rewrite it using the correction engine. The hot word list is dynamically adjusted through the rewritten player text to ensure that it covers the most commonly used and most important words so that the hot words of the first scenario can be obtained more accurately in the future; (2) Model update: Retrain the language model and / or the voice model and adjust the parameters of the model; (3) Testing and deployment: The updated model undergoes full comparative testing to ensure its performance and stability. Once the test is passed, the trained model can be deployed to the game; (4) Continuous iteration: Speech recognition is a continuous improvement process. Feedback is collected and analyzed regularly, and then the model is updated to adapt to changes in player behavior and the game. Tests are conducted in the actual game environment, data and feedback are collected, and the speech recognition system is continuously iterated and improved to improve the accuracy of command analysis results.
[0464] In an optional embodiment, the method of adjusting the language model of speech recognition according to the virtual scene helps to better understand and parse natural language commands in speech form. An appropriate language model is selected from a predefined language model library. When the scene changes, the language model for analysis can be automatically adjusted, and this process can be implemented through a preset language model management system. In addition, the system can also establish an effective feedback mechanism. Players can provide feedback and correct recognition errors or add new scene hot words through the cloud large language model. These feedbacks will update and improve the natural language analysis model online, and timely adjust and update the hot words and language models. In addition, to ensure the reliability and stability of the system, existing large language models can also be used for offline testing and evaluation. Combining system automatic testing and player feedback data in different scenarios, the hot words and language models are further optimized according to the test feedback data to meet the requirements of different scenarios. For example: when a player enters the interior of a city room, "turn on the light", "turn off the light", "raise the temperature", etc. may be the scene hot words of this virtual indoor scene, etc.
[0465] It should be noted that the above is only a schematic example, and the embodiments of this application are not limited thereto.
[0466] In an optional embodiment, taking the virtual minions including the main control virtual character and non-player virtual characters in the virtual environment as an example, the process of obtaining the command analysis result in text form for the natural language command in speech form is described as follows.
[0467] Schematically, as [[ID=IO]] Figure 9 shown, the virtual environment includes the main control virtual character 910 manually controlled by the player, the virtual minion 911 (referred to as No. 1) controlled by voice, and the virtual minion 912 (referred to as No. 2); if the main control virtual character 910 is used as the target virtual character.
[0468] After receiving the natural language command, it is determined that the first virtual scene where the main control virtual character 910 is located is the "virtual outdoor scene", and multiple first scene hot words corresponding to the first virtual scene are obtained, including "truck", "virtual grassland", "virtual house", "enter", "open the truck", "defend", "desert", "attack", "oak tree", "stable", etc.
[0469] Afterward, based on the multiple first scene hot words, the natural language command is converted into a command analysis result in text form, and the command analysis result is obtained as "No. 1, you go to the front Defense for a while", and this command analysis result can specifically command the virtual minion 911 to move forward and take a defensive posture to prevent the enemy virtual character from attacking the main control virtual character first; without the constraint of the first scene hot words, it is easy to recognize the natural language command as "No. 1, you go to the front SpeakingThe incorrect recognition result will instruct Player 1 to shout or make other vocalizations, thus affecting the player's virtual combat plan.
[0470] Alternatively, based on multiple hot words from the first scenario, the natural language command is converted into text form through command analysis, resulting in the command analysis result "Number 2, you go...". desert The command "Scout it out" can be analyzed to specifically direct virtual soldier 912 to move to the desert location to scout for enemy virtual characters or virtual treasure chests, etc. Without the constraint of the first scene's hot words, the natural language command could easily be identified as "Number 2, you go..." mountain range The "Investigate" command caused the virtual soldier 912 to move to the wrong location, resulting in low human-computer interaction efficiency.
[0471] Alternatively, based on multiple hot words from the first scenario, the natural language command is converted into text form through command analysis, resulting in the command analysis result "Number 1, Number 2, Give Me...". attack The command analysis results can specifically direct virtual soldiers 911 and 912 to launch virtual attacks on nearby attackable virtual characters; without the constraint of the first scene hot words, the natural language command is easily identified as "Number 1, Number 2, give me..." supply This causes virtual minions 911 and 912 to make erroneous behaviors that do not meet the player's expectations, failing to protect the main virtual character and making it easy for the enemy virtual character to win the virtual game.
[0472] Alternatively, based on multiple hot words from the first scene, the natural language command is converted into text form. The command analysis result is "Number 1, go find the nearby oak tree." This command analysis result can specifically instruct the virtual soldier 911 to search for oak trees in the nearby area so as to complete game tasks or find oak trees to avoid attacks. Without the constraint of the first scene hot words, the natural language command is easily identified as "Number 1, go find the nearby..." Project Proposal This causes the virtual soldier 911 to search for virtual elements in the vicinity that do not meet the player's expectations, thus failing to satisfy the player's virtual combat needs.
[0473] Alternatively, based on multiple hot words from the first scenario, the natural language command is converted into text form, and the command analysis result is obtained as "Number 1, Number 2, go over there". stable "Okay," the command analysis results can specifically instruct virtual soldiers 911 and 912 to search for nearby stables and move to their locations. Without the constraint of the first scene's hot words, the natural language command would easily be recognized as "Number 1, Number 2, go over there." Immediately"Okay," causing virtual minions 911 and 912 to mistakenly believe that the main virtual character wants to complete the virtual game on its own, thus failing to provide good game assistance to the main virtual character.
[0474] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.
[0475] Indicative, such as Figure 10 The process of command analysis based on the game terminal interface is explained as shown. The target virtual character is the main control virtual character 1010, and the first virtual scene in which it is located is a virtual battle scene, which also includes a non-player character 1020 (soldier A). The natural language command received is in the form of voice "xiaobingA geiwogongji". Based on the first virtual scene corresponding to multiple first scene hot words, the natural language command is analyzed by the constraints of the first scene hot words. The command analysis result is "soldier A, attack me", that is, "gongji" is determined to be "attack" rather than "supply", avoiding the analysis error of natural language command and obtaining a more accurate command analysis result, so as to better command the non-player character.
[0476] In summary, using the first virtual scene in which the target virtual character is located as a condition for analyzing natural language commands can provide more targeted textual information for the analysis process through the hot words of the first scene, reducing the ambiguity caused by poor quality of natural language commands. This not only facilitates a more accurate understanding of natural language commands but also makes the command analysis results more closely match the first virtual scene, improving the accuracy of the command analysis results. It greatly avoids ineffective or erroneous commands to non-player characters, providing players with a more efficient and comfortable operating environment and improving human-computer interaction efficiency. Using a small number of first scene hot words can also reduce the amount of data processing during command analysis, improving command analysis efficiency.
[0477] In this embodiment, the player's voice can be perceived more accurately within the game scenario, resulting in higher accuracy in recognizing player intent. This allows for more convenient interaction and operation within the game. Through voice commands, players can achieve faster and more natural operation, enhancing game smoothness and enjoyment. It also makes virtual characters and NPCs in the game more intelligent and realistic. Players can gain a more authentic gaming experience and enhance interactivity by conversing with in-game characters. Furthermore, players can complete various operations more quickly, such as switching virtual items and releasing skills. Compared to traditional button operations, faster response and execution, and more accurate perception of user intent improve game operation efficiency, giving game developers a competitive advantage. Simultaneously, this novel interaction process effectively differentiates itself from other games, helping to attract more players. In summary, the above technical solution can improve the game experience, enhance game interactivity, increase game operation efficiency, expand the game's audience, and enhance game competitiveness.
[0478] This application provides an exemplary embodiment of a method for controlling a virtual character. The method includes:
[0479] • Display at least one of the main virtual character and NPC located in the virtual environment;
[0480] For example, the main virtual character is a virtual character directly controlled by the user in the virtual environment. The virtual environment also contains one or more non-player characters (NPCs); NPCs and the main virtual character belong to the same virtual faction, and the NPCs are teammates of the main virtual character, following the main virtual character in activities within the virtual environment. Alternatively, NPCs can also be followers, pets, or other characters controlled by the main virtual character. Furthermore, NPCs can also be neutral characters, only cooperating with the main virtual character under specific conditions.
[0481] • Obtain natural language commands;
[0482] For example, natural language commands include natural semantics for directing NPCs; directing NPCs in a virtual scene through natural language commands can be done by direct input from the user or extracted from the user's voice input. This application does not limit the method of obtaining natural language commands.
[0483] Natural language commands include at least one of behavioral intent and scene entity; the behavioral intent is used to indicate the type of virtual activity performed by the NPC, such as indicating what type of virtual activity to perform in the virtual environment; the scene entity includes a first scene object, such as indicating which scene object in the virtual environment the virtual activity is performed on. The behavioral intent and scene entity included in the natural language command are part of the words in the natural language command.
[0484] In one example, natural language commands control the NPC from two aspects: what type of virtual activity to perform and what scene objects in the virtual environment to perform the virtual activity; it is a complex instruction for controlling the NPC.
[0485] • Respond to natural language commands and control NPCs to perform virtual activities;
[0486] For example, controlling an NPC to perform virtual activities is determined based on the environmental perception information of the controlling virtual character and / or the NPC. In some examples, the controlling virtual character and / or the NPC acquires environmental perception information within the perception range of visual, auditory, and movement trajectory perception methods. For example, virtual activities closely resemble the current perception of the virtual environment by the controlling virtual character and / or the NPC; the NPC is controlled with reference to the current perception of the virtual environment by the virtual character and / or the NPC. For example, the NPC possesses autonomous behavioral capabilities; the user merely directs the NPC, issuing commands, which the NPC understands and executes based on its own autonomous behavioral capabilities.
[0487] In summary, the method provided in this embodiment controls the complex instructions of the NPC from two aspects: what type of virtual activity to perform and what scene objects in the virtual environment to perform the virtual activity. It determines the virtual activity to be performed by referring to the environmental perception information of the main virtual character and / or the NPC. By referring to the environmental perception information of the main virtual character and / or the NPC, it ensures that the NPC and the main virtual character cooperate in performing virtual activities.
[0488] Figure 11 A flowchart illustrating a virtual character control method provided in an exemplary embodiment of this application is shown. The method includes:
[0489] Step 1101: Obtain the spatial dataset;
[0490] The spatial dataset of the virtual environment includes visual labels for scene objects. The visual labels are used to describe the visual features of scene objects in at least one dimension, such as material, transparency, color, shape, etc.
[0491] Step 1102: Display at least one of the main virtual character and NPC located in the virtual environment;
[0492] For example, the master virtual character is a virtual character that the user directly controls in the virtual environment. There are also one or more NPCs in the virtual environment; the NPCs and the master virtual character belong to the same virtual faction and are teammates of the master virtual character.
[0493] Step 1103: Obtain natural language commands in speech form;
[0494] Natural language commands include natural semantics for directing NPCs, and command information includes behavioral intent and scene entities; behavioral intent indicates the type of virtual activity performed by the NPC, and scene entities indicate the target entity targeted by the virtual activity;
[0495] Step 1104: Convert the natural language command in speech form into a natural language command in text form;
[0496] For example, automatic speech recognition (ASR) is performed on natural language commands in speech form to determine the natural language commands in text form. Speech-to-text processing typically involves calling components such as acoustic models and language models to recognize the pronunciation, vocabulary, and grammatical structure of natural language commands in speech form and convert them into natural language commands in text form.
[0497] Step 1105: Perform intent recognition on the natural language command to obtain the first category label;
[0498] For example, the first intent indication corresponding to the first category label has a command intent for a non-player character in one intent dimension; such as at least one of the following: indicating which non-player character among many non-player characters is the non-player character being commanded by the command text, indicating whether the virtual activity performed by the non-player character is related to launching a virtual attack, or indicating what kind of virtual activity the non-player character is performing.
[0499] Step 1106: Perform entity recognition on the natural language command to obtain the target entity;
[0500] The target entity is determined from the virtual environment by combining environmental perception information, which includes information perceived from the virtual environment by at least one of the main virtual character and non-player characters. The target entity may be an entity currently seen by the player, or it may be an entity that the player has seen in the past.
[0501] Step 1107: In response to a natural language command, control the NPC to perform a virtual activity according to the first intent corresponding to the first category label, or control the NPC to perform a virtual activity associated with the target entity, or control the NPC to perform a virtual activity associated with the target entity according to the first intent corresponding to the first category label.
[0502] For example, NPCs are controlled to perform virtual activities according to the instructions of the first category label and / or the target entity.
[0503] Step 1108: Broadcast the NPC's feedback information;
[0504] For example, a game application can generate and broadcast corresponding feedback information in real time based on the environmental awareness information of a non-player character. For instance, when the environmental awareness information of a non-player character triggers a broadcast condition, the game application can generate corresponding feedback information in real time based on the environmental awareness information that triggered the broadcast condition; or, when either a non-player character or the main virtual character triggers a broadcast condition, the game application can generate corresponding feedback information in real time based on the environmental awareness information of the non-player character.
[0505] The preprocessing stage in step 1101 can be implemented as follows:
[0506] Sub-step 1: Obtain the attribute text of scene objects in the virtual scene;
[0507] For example, attribute text is used to describe the inherent attributes of scene objects in a virtual scene. On the one hand, attribute text describes scene objects in a textual modality, providing semantic information for predicting the visual labels of scene objects while describing them. On the other hand, attribute text is a description of scene objects in a virtual scene, and given the large number of scene items and the existence of reusable object models, it can more accurately describe the inherent attributes of scene objects in the virtual scene.
[0508] In one alternative implementation, the attribute text includes at least one of the scene object's name in the virtual scene and its size in the virtual scene.
[0509] Sub-step 2: Obtain the appearance images of scene objects in the virtual scene;
[0510] For example, an appearance image is used to describe the style of scene objects; the appearance image carries the appearance style of scene objects, such as color, texture, and shape, as well as the relative positional relationships between the various sub-parts, in the image modality, and can comprehensively describe scene objects from the image modality (or visual modality). In an optional implementation, the appearance image of the scene object includes images obtained by observing the scene object from at least two perspectives.
[0511] Sub-step 3: Call the multimodal model to predict the attribute text and appearance image of scene objects to obtain the visual labels of scene objects;
[0512] For example, the multimodal model has the ability to perform model prediction on text information and image information of different modalities. In this embodiment, the input parameters of the multimodal model are the attribute text and appearance image of the scene object; the multimodal model predicts the visual labels of the scene object from two modalities, text modality and image modality, and the visual labels are used to describe the visual features of the scene object in at least one dimension.
[0513] Optionally, the multimodal model includes a visual question answering model; the question statement carries attribute text of scene objects; the question statement is used to guide the visual question answering model to convert the appearance image into visual labels of scene objects; on the other hand, while guiding the conversion of visual labels of scene objects, the question statement provides supplementary information about scene objects in text form.
[0514] Obtain the desired information about objects in the scene;
[0515] For example, the expectation information is used to indicate the descriptive dimensions of the scene object expected in the visual label, and / or the expected format of the visual label; in one example, the expectation information is used to indicate that the descriptive dimensions of the scene object expected in the visual label include, but are not limited to, at least one of the following: type description, material, transparency, color, surface features, and shape of the scene object. In another example, the expectation information is used to indicate that the expected format of the visual label is at least one of the following: Comma Separated Values (CSV), JavaScript Object Notation (JSON), and eXtensible Markup Language (XML).
[0516] Construct a query statement for the appearance image based on the expected information and attribute text;
[0517] For example, the first part of the question statement is supplementary information about the scene objects, carrying the attribute text of the scene objects; the second part of the question statement is the answer guidance statement for the visual question answering model, carrying the expected information.
[0518] Optionally, the spatial dataset includes matching labels for scene objects in the virtual scene; accordingly, the visual labels of the scene objects are rewritten in a pseudo-conversational style to obtain matching labels that conform to the spoken expression of natural language.
[0519] For example, as described above, visual tags are used to describe the visual features of scene objects in at least one dimension, and the appearance image of the scene object presents rich visual features. However, in spoken natural language, the description of scene objects cannot cover the visual features of scene objects in all dimensions. The purpose of performing pseudo-colloquial rewriting on the visual tags of scene objects is to obtain matching tags that are closer to spoken expressions; understandably, performing pseudo-colloquial rewriting can involve deleting parts of the visual tags, or changing the visual tags to tags with the same semantics but different textual expressions.
[0520] In one alternative implementation, pseudo-conversational rewriting is performed by invoking a natural language model. Visual labels of scene objects are input into the natural language model, which predicts matching labels that conform to spoken language expressions. This pseudo-conversational rewriting of visual labels is based on invoking the natural language model. For example, the natural language model carries prior knowledge of spoken language expressions. One example is a Large Language Model (LLM).
[0521] Optionally, the spatial dataset may also include spatial information of scene objects, and correspondingly, it may also include:
[0522] Obtain the spatial position of scene objects in the virtual scene, and use the spatial position of scene objects as auxiliary information to determine the visual labels of scene objects.
[0523] For example, the spatial position of a scene object in a virtual scene is used to indicate the deployment of the scene object in the virtual scene, and the spatial information is used to indicate the position, size, and other information of the scene object in the virtual scene after it is deployed.
[0524] Furthermore, obtain at least one of the following: the coordinate position, orientation information, bounding box information, and cover point information of the scene object in the virtual scene;
[0525] Among them, Location indicates the position of scene objects in the virtual scene, such as the coordinates of the center point or preset point of the scene object in the virtual scene; Rotation indicates the direction that scene objects face in the virtual scene, such as the direction that the front of the scene object faces in the virtual scene; Bounding Box indicates the size of scene objects in the virtual scene; Cover points indicate the recommended position points of the virtual character when the virtual character approaches the scene object, so that the scene object can cover the virtual character.
[0526] The intent recognition stage in step 1105 can be implemented as follows:
[0527] Sub-step 4: Call at least one hierarchical prediction network to identify the execution intent of the natural language command and obtain the first classification label of the natural language command intent to instruct the NPC;
[0528] For example, a hierarchical prediction network has the ability to predict the classification label corresponding to a command text. For example, a hierarchical prediction network includes at least two sub-networks, with an upper and lower network cascaded. The lower network further performs classification label prediction based on the prediction results output by the upper network. The at least two sub-networks are constructed based on a tree structure of multiple classification labels. The hierarchical prediction network constructs at least two sub-networks with corresponding hierarchical structures for the tree structure of classification labels, breaking down the prediction task of numerous classification labels into prediction sub-tasks performed by at least two sub-networks, thereby reducing the classification prediction complexity of each sub-network.
[0529] In one optional implementation, each of the at least one hierarchical prediction networks is used to predict the command intent of the command text for a non-player character in one intent dimension; the intent dimension includes at least one of the following: subject dimension, semantic dimension, and behavior dimension. For example, in the at least one hierarchical prediction network, the i-th layer subnetwork in each hierarchical prediction network is used to predict the first-level behavior label of the command text in one intent dimension, and the (i+1)-th layer subnetwork in the hierarchical prediction network is used to predict the second-level behavior label corresponding to the first-level behavior label of the command text, where i is a positive integer; the high-level behavior label (such as the first-level behavior label) predicted by the high-level subnetwork (such as the i-th layer subnetwork) in the hierarchical prediction network includes multiple sub-labels, or subordinate lower-level labels; the corresponding low-level subnetwork (such as the (i+1)-th layer subnetwork) needs to be called to further perform classification prediction (such as predicting the second-level behavior label). This breaks down the prediction task of various classification labels into prediction sub-tasks performed by at least two layers of subnetworks, thereby reducing the complexity of classification prediction for each layer of subnetwork.
[0530] Furthermore, the intent dimension includes a subject dimension, and the hierarchical prediction network includes a hierarchical subject prediction network with the ability to predict the subject type in a natural language command; the subject type is used to indicate the identity of the NPC commanded by the natural language command; for example, the subject prediction network is used to predict which of a number of non-player characters the non-player character commanded by the command text is, and the number of non-player characters commanded by the command text can be one or more.
[0531] And / or, the intent dimension includes the semantic dimension, and the hierarchical prediction network includes a hierarchical semantic prediction network with the ability to predict the semantic type in natural language commands; the semantic type is used to indicate the control method of the NPC initiating a virtual attack; for example, the subject prediction network is used to predict whether the virtual activities performed by a non-player character under the command text are related to initiating a virtual attack.
[0532] And / or, the intent dimension includes the behavior dimension, and the hierarchical prediction network includes a hierarchical behavior prediction network with the ability to predict the type of behavioral intent in natural language commands; the behavioral intent type is used to indicate how an NPC should behave in performing virtual activities. For example, the behavior prediction network is used to predict what virtual activities a non-player character should perform.
[0533] Accordingly, in step 1107, in response to a natural language command, the NPC is controlled to perform virtual activities according to the first intent corresponding to the first category label;
[0534] For example, the first intent indication corresponding to the first category label has a command intent for a non-player character in one intent dimension; such as at least one of the following: indicating which non-player character among many non-player characters is being commanded by the command text, indicating whether the virtual activity performed by the non-player character is related to launching a virtual attack, and indicating what kind of virtual activity the non-player character is performing. Following the indication of the first category label, the NPC is controlled to perform virtual activities.
[0535] The target entity recognition process in step 1106 can be implemented as follows:
[0536] Sub-step 5: Based on the target entity information indicated by the natural language command and the environmental perception information, query the target entity from the entity set in the virtual environment to obtain the target entity;
[0537] The target entity is determined from the virtual environment by combining environmental perception information, which includes information perceived by the main virtual character and / or non-player characters from the virtual environment.
[0538] Because the target entity information in natural language commands is often vague—for example, a natural language command might be "move to the red truck," where the target entity information is "red truck"—but there may be many red trucks in a virtual environment, the target entity cannot be accurately determined from the virtual environment based solely on the target entity information in the natural language command.
[0539] Therefore, this application provides a method for determining a target entity by combining target entity information and environmental perception information. In the previous example, the "red truck" expressed by the player in the natural language command should be a red truck that the player can see. Therefore, by combining the field of vision of the main virtual character, the red truck located within the field of vision of the main virtual character can be selected from multiple red trucks in the virtual environment. This red truck is the target entity indicated by the player in the natural language command.
[0540] Since natural language commands are issued by players based on their perception of the virtual environment, in order to accurately identify the target entity indicated in the natural language command, the game application combines the environmental perception information at the time the player issues the natural language command to identify the target entity in the natural language command. Based on the player's perception of the virtual environment, the entity in the virtual environment that the player can perceive and that is closest to the target entity information is deduced; this is the target entity.
[0541] For example, the virtual environment includes multiple candidate entities that match natural language commands, and the target entity is an entity selected from the multiple candidate entities based on the environmental awareness information of the controlling virtual character or non-player character. For instance, a game application first selects multiple candidate entities from the entity set that match the target entity information in the natural language command, and then selects the entity that the controlling virtual character or non-player character can perceive as the target entity based on the environmental awareness information.
[0542] For example, the environmental perception information of the main virtual character may include: the images or entities that the main virtual character can see when observing the virtual environment, the environmental sound effects that the main virtual character can hear, the direction of the source of the environmental sound effects, the type and volume of the environmental sound effects, the perception information obtained by the main virtual character through the use of perception skills (e.g., visual information, auditory information, sound wave information, light reflection information, etc.), and the perception information obtained by the main virtual character through the use of sensory tools (e.g., sensor signals, positioning signals of positioning tools, etc.).
[0543] The environmental perception information of non-player characters may include: the images or entities that non-player characters can see when observing the virtual environment, the environmental sound effects that non-player characters can hear, the direction of the source of the environmental sound effects, the type and volume of the environmental sound effects, the perception information obtained by non-player characters through the use of perception skills (e.g., visual information, auditory information, sound wave information, light reflection information, etc.), and the perception information obtained by non-player characters through the use of sensory tools (e.g., sensor signals, positioning signals of positioning tools, etc.).
[0544] It should be noted that environmental awareness information can include information perceived in real-time by the controlling virtual character and non-player characters upon receiving a natural language command, as well as information perceived historically by the controlling virtual character and non-player characters before receiving the natural language command. That is, the target entity may be an entity currently seen by the player, or it may be an entity the player has seen historically. Therefore, the game application needs to combine the real-time and historical environmental awareness information of the controlling virtual character and / or non-player characters to identify the target entity.
[0545] For example, when the natural language command is "Let's go back to the hotel we just passed", the game application needs to query the hotels that the main virtual character has passed by based on the historical environment awareness information of the main virtual character.
[0546] In one optional embodiment, the client retrieves the target entity from the entity set of the virtual environment based on the target entity information indicated by the natural language command and the environment-aware information. In another optional embodiment, the client reports a natural language command to the server, and the server retrieves the target entity from the entity set of the virtual environment based on the target entity information indicated by the natural language command and the environment-aware information.
[0547] Optionally, the game application infers the target entity based on target entity information, environmental awareness information, and pre-processed virtual environment data. The pre-processed virtual environment data includes a set of entities within the virtual environment.
[0548] The entity set includes entity information for each entity in the virtual environment. Entity information includes at least one of the following: name, type, location, features, text label, embedding vector of the text label, image, and embedding vector of the image.
[0549] The text label can be at least one word obtained by segmenting the features (e.g., descriptive text of an entity's appearance), and the embedding model is called to obtain the embedding vector corresponding to each text label.
[0550] The image can include images obtained by viewing a 3D model of an entity from at least one direction; for example, the image can include three views of the entity. The image embedding vector is the embedding vector of the image feature labels. The image feature labels are obtained by calling a multimodal model to perform feature recognition based on the entity's image and text labels.
[0551] For example, the entity set includes a first type of entity and a second type of entity. The first type of entity includes at least one of regions and buildings, and the second type of entity includes physical objects. The text labels of the first type of entities are obtained manually. For example, the text label for a specific room in a building is: a certain region, a certain building, a certain floor, and a room. The text labels of the second type of entities are automatically generated by calling a large language model.
[0552] The method for generating text labels for the second type of entity may include: acquiring at least one image of the entity; inputting the at least one image into a large language model to obtain descriptive text for the entity; segmenting the descriptive text to obtain at least one text label for the entity, each text label including a word obtained after segmentation; and performing vectorization on the text labels to obtain the embedding vectors corresponding to the text labels.
[0553] Embedded vectors are high-dimensional vector data of text tags. By pre-generating the embedded vectors of text tags, it is not necessary to repeatedly generate vectors for the text tags of each entity when matching target entities from the entity set based on target entity information, which can improve the matching efficiency between target entity information and entity information.
[0554] It's important to note that this method doesn't directly use the descriptive text as text labels. Instead, it uses the word segmentation results of the descriptive text as the text labels. This is because descriptive texts typically vary in length, and for longer descriptive texts, their embedding vectors perform poorly when comparing similarity in actual searches. For example, when searching for "a rusty blue truck" using "truck," "a rusty blue truck" is likely to rank after "a car" in the recall results. However, for the search term "truck," regardless of how complex its additional description is, "truck" should be prioritized in the search. Therefore, when performing similarity searches, the comparison should be based on individual feature words, not the descriptive text itself.
[0555] Furthermore, in order to further extract the image features of the entity, the method provided in this application embodiment can also further extract the hidden visual features in the image of the entity based on the text label and image of the entity. Taking the first entity as an example, the method includes: obtaining at least one view image of the first entity; obtaining at least one text label of the first entity; calling a multimodal model to extract the image features of the first entity based on at least one view image and at least one text label of the first entity, and obtaining the image embedding vector of the first entity.
[0556] Since text labels are either manually annotated or generated by a large language model based on human language characteristics, and text labels extracted based on human language habits may overlook some features of the entity. For example, when manually labeling an oil drum, the labeler might only include more noticeable text labels like "metal" or "rusty," while ignoring detailed features such as "rusty," "blue," "yellow," or "paint chipping in the lower right corner." Therefore, the method provided in this application also uses a multimodal model, based on the entity's text labels and entity image, to extract more image features; and performs image similarity matching with the target entity based on these image features, thereby improving the accuracy of target entity recognition.
[0557] The training method for the multimodal model can be as follows: Input sample entity images and sample labels into a pre-trained multimodal model. Based on the predicted labels output by the pre-trained multimodal model and the loss of the sample labels, fine-tune the pre-trained multimodal model to obtain a multimodal model that can output image feature labels based on the input text labels and images. The image feature labels include more detailed descriptive text extracted from the image. For example, the text description "a metal oil drum" can only be associated with the two text labels "metal" and "oil drum". However, the Clip (multimodal) model can also identify hidden visual information in the image of the oil drum, such as image feature labels "rusty", "blue", and "yellow". These image features are features that do not appear in the text labels. Therefore, combining the Clip model for image feature search can further improve the accuracy of entity search.
[0558] In one alternative embodiment, the game application may query the target entity using the following method.
[0559] (1) Parse natural language commands to obtain target entity information; the target entity information includes at least one of the following: entity type, entity name, entity location, entity features.
[0560] Optionally, the game application calls a large language model to parse the natural language commands and obtain the target entity information. It then calls an embedding model to vectorize the target entity information, obtaining the target embedding vector of the target entity information.
[0561] For example, if the natural language command is "Come to the blue truck in front", the target entity information that the large language model can parse from it includes: the location information "in front" and the entity name "blue truck".
[0562] For example, named entity recognition technology from the field of natural language processing can also be used to extract target entity information from natural language commands. Named entity recognition technology is used to identify entities with specific meanings in text, such as names of people, place names, locative words, adjectives, etc.
[0563] For example, a natural language command like "Go to the motel and find a cardboard box behind the red sofa on the first floor" can be identified using named entity recognition technology as follows: architectural noun "motel first floor"; object nouns "sofa" and "cardboard box"; locative word "behind"; and adjective "red". After logical construction, a hierarchical scene query can be formed with search type, content, locative words, and constraint information. The search type narrows the data retrieval scope based on similarity matching, the search content is a specific entity description, the adjective is incorporated into the description text for the query, and the locative word forms the constraint on the search scope. The floor constraint is determined by the player's location; in indoor scenes, the vertical range needs to be narrowed to avoid finding entities that are invisible across floors.
[0564] (2) Calculate the similarity between the target entity information and the information of each entity in the entity set.
[0565] For example, the entity set includes entity information for at least one entity, and the entity information for each entity may include at least one of the text embedding vector of a text label and the embedding vector of an image feature. Then, the target entity information can be compared with the text embedding vector and the image embedding vector respectively to calculate similarity, and the entity with higher similarity is identified as the target entity.
[0566] The methods for similarity matching with text embedding vectors and similarity matching with image embedding vectors are given below.
[0567] 1) Similarity includes the text similarity between the target entity information and the text embedding vector.
[0568] Taking the first entity in the entity set as an example, the game application (client or server) segments the target entity information to obtain at least one target entity label; converts the at least one target entity label into at least one target embedding vector; obtains the entity information of the first entity, which includes a text embedding vector, which is an embedding vector obtained based on the text label of the first entity; calculates the text parent similarity between at least one target embedding vector and the text embedding vector respectively, and obtains at least one text parent similarity corresponding to at least one target embedding vector; and determines the sum of at least one text parent similarity as the text similarity between the target entity information and the entity information of the first entity.
[0569] For example, if there are two target entity labels, target entity label 1 and target entity label 2, and the entity set also contains two entities, entity 1 and entity 2, with entity 1 corresponding to text embedding vector 1 and entity 2 corresponding to text embedding vector 2, then we calculate the similarity 1 between target entity label 1 and text embedding vector 1, and the similarity 2 between target entity label 2 and text embedding vector 1. The sum of similarity 1 and similarity 2 is determined as the text similarity between the target entity information and entity 1. Next, we calculate the similarity 3 between target entity label 1 and text embedding vector 2, and the similarity 4 between target entity label 2 and text embedding vector 2. The sum of similarity 3 and similarity 4 is determined as the text similarity between the target entity information and entity 2.
[0570] For example, the entity information of the first entity includes at least one text embedding vector; at least one target embedding vector includes a first target embedding vector; then the game application calculates the text sub-similarity between the first target embedding vector and at least one text embedding vector respectively to obtain at least one text sub-similarity; and determines the highest value among the at least one text sub-similarity as the text parent similarity corresponding to the first target embedding vector.
[0571] For example, the target entity label has one element: target entity label 1. The entity set also contains one entity: entity 1. Entity 1 corresponds to text embedding vector 1 and text embedding vector 3. Then, the similarity between target entity label 1 and text embedding vector 1 is calculated as 1, and the similarity between target entity label 1 and text embedding vector 3 is calculated as 5. The larger of the similarity values 1 and 5 is determined as the similarity between the target entity information (target entity label 1) and entity 1.
[0572] For example, given the input target entity information "blue car", the target entity information is first segmented into the target entity labels "blue" and "car", representing that the target entity has two features. Then, the similarity score of each feature is calculated with each entity in the entity set, and the maximum similarity score of each feature is taken. Finally, the similarity scores corresponding to all query features are summed to represent the text similarity between the target entity information and the query entity.
[0573] For example, the target entity information contains two target entity labels, "blue" and "vehicle". These two target entity labels are converted into two target embedding vectors. Then, the entity information in the entity set is obtained. For example, the entity set includes entity 1 and entity 2. The text labels of entity 1 include "metal", "old", "vehicle", "truck", "blue", and "damaged". The text labels of entity 2 include "metal", "scratches", "vehicle", "damaged", "armored truck", and "black".
[0574] Calculate the text similarity between the target entity information and entity 1: First, calculate the similarity between the target entity label "blue" and each text label of entity 1, and take the maximum value. For example, the similarity between the target entity label "blue" and the text label "blue" of entity 1 is 1 (1 is the maximum value). Then, calculate the similarity between the target entity label "car" and each text label of entity 1, and take the maximum value. For example, the similarity between the target entity label "car" and the text label "car" of entity 1 is 1. Then, sum the two similarities between the target entity labels "car" and "blue" to get the final similarity between the target entity information and entity 1, which is 2.
[0575] Similarly, to calculate the text similarity between the target entity information and entity 2: first, calculate the similarity between the target entity label "blue" and each text label of entity 2, and take the maximum value. For example, the similarity between the target entity label "blue" and the text label "black" of entity 1 is 0.91. Then, calculate the similarity between the target entity label "car" and each text label of entity 1, and take the maximum value. For example, the similarity between the target entity label "car" and the text label "car" of entity 2 is 1 (1 is the maximum value). Then, sum the two similarities between the target entity labels "car" and "blue" to get the final similarity between the target entity information and entity 2, which is 1.91.
[0576] It can be seen that the similarity 2 between the target entity information and entity 1 is higher than the similarity 1.91 between the target entity information and entity 2.
[0577] 2) Similarity includes the image similarity between the target entity information and the image embedding vector.
[0578] Taking the first entity in the entity set as an example, the game application (client or server) segments the target entity information to obtain at least one target entity label; converts the at least one target entity label into at least one target embedding vector; obtains the entity information of the first entity, which includes an image embedding vector, which is an embedding vector extracted from the image of the first entity; calculates the image parent similarity between at least one target embedding vector and the text embedding vector respectively, and obtains at least one image parent similarity corresponding to at least one target embedding vector; and determines the sum of at least one image parent similarity as the image similarity between the target entity information and the entity information of the first entity.
[0579] For example, if there are two target entity labels, target entity label 1 and target entity label 2, and the entity set also contains two entities, entity 1 and entity 2, with entity 1 corresponding to image embedding vector 1 and entity 2 corresponding to image embedding vector 2, then we calculate the similarity 1 between target entity label 1 and image embedding vector 1, and the similarity 2 between target entity label 2 and image embedding vector 1. The sum of similarity 1 and similarity 2 is determined as the image similarity between the target entity information and entity 1. Next, we calculate the similarity 3 between target entity label 1 and image embedding vector 2, and the similarity 4 between target entity label 2 and image embedding vector 2. The sum of similarity 3 and similarity 4 is determined as the image similarity between the target entity information and entity 2.
[0580] For example, the entity information of the first entity includes at least one image embedding vector; at least one target embedding vector includes a first target embedding vector; then the game application calculates the image sub-similarity between the first target embedding vector and at least one image embedding vector respectively to obtain at least one image sub-similarity; and determines the highest value among the at least one image sub-similarity as the image parent similarity corresponding to the first target embedding vector.
[0581] For example, the target entity label has one element: target entity label 1. The entity set also contains one entity: entity 1. Entity 1 corresponds to image embedding vector 1 and image embedding vector 3. Then, the similarity 1 between target entity label 1 and image embedding vector 1 is calculated, and the similarity 5 between target entity label 1 and image embedding vector 3 is calculated. The larger of similarity 1 and similarity 5 is determined as the similarity between the target entity information (target entity label 1) and entity 1.
[0582] In an optional embodiment, the calculation of image similarity can be performed using the target image embedding vector, that is, the target image embedding vector is used instead of the target embedding vector in the above method. The target image embedding vector can be obtained by: segmenting the target entity information into words to obtain at least one target entity label; inputting the at least one target entity label into a multimodal model to obtain the target image embedding vector.
[0583] 3) Similarity includes text similarity and image similarity.
[0584] When a game application has both text similarity and image similarity with a target entity, it determines the average of the text similarity and image similarity as the similarity between the first entity and the target entity.
[0585] Alternatively, if the first entity and the target entity have text similarity and image similarity, the larger value between the text similarity and the image similarity is determined as the similarity between the first entity and the target entity.
[0586] Alternatively, if the first entity and the target entity have text similarity and image similarity, the sum of the text similarity and image similarity is determined as the similarity between the first entity and the target entity.
[0587] (3) Determine the target entity from the entity set based on similarity and environmental perception information.
[0588] In one optional embodiment, the game application first performs a similarity search based on the target entity information to obtain at least one candidate entity with high similarity, and sorts them according to similarity; then, based on the location, orientation, and other information of the main virtual character and / or non-player characters, it performs perceptual filtering and sorting to finally filter out the target entity.
[0589] For example, the game application determines the target perception range based on target entity information and environmental perception information; and selects the target entity from the entities perceived within the target perception range based on similarity.
[0590] Alternatively, the game application identifies the x most similar entities from the entity set based on similarity, forming a candidate entity list; it then determines the target perception range based on the target entity information and environmental perception information, and identifies the entities within the target perception range from the x entities as the target entities. x is a positive integer.
[0591] For example, a game application might identify the entity with the highest similarity within its perception range as the target entity.
[0592] It should be noted that when multiple entities with equal similarity exist within the target's perception range, they can be sorted according to their distance from the controlling virtual character. For example, if there are at least two entities with the highest similarity within the perception range, the entity with the highest similarity and closest distance to the controlling virtual character within the perception range will be identified as the target entity.
[0593] In one alternative embodiment, the game application identifies the entity with the highest similarity within its perception range and exceeding a threshold as the target entity. Exemplarily, the number of perception ranges is at least one. For example, the perception range includes at least one of the following: the field of vision of the main virtual character, the auditory range of the main virtual character, the perception range of the main virtual character for virtual items, the field of vision of a non-player character, the auditory range of a non-player character, and the perception range of a non-player character for virtual items.
[0594] For example, the game application can sequentially traverse at least one of the aforementioned perception ranges, or the game application can sequentially traverse at least one perception range determined according to the natural language command. For example, if the natural language command is "Can you see the truck ahead? Move to the truck," then the game application can first traverse the field of vision of the main virtual character, and then traverse the field of vision of the non-player characters, matching the entity with the highest similarity that is above a threshold.
[0595] For example, when at least two perception ranges exist, the traversal order of the at least two perception ranges can be preset, or the traversal order of the at least two perception ranges can be determined according to the intent indicated in the natural language command. For example, when the intent is to perform a search for an item, the visual range is traversed first. When the intent is to find a combat scene, the auditory range is traversed first.
[0596] If a natural language command contains multi-level nested instructions, the first target entity will be searched first, and then the search for the next target entity will continue based on the position of the previous target entity. If the natural language command includes a second target entity determined based on the position of the first target entity, the first target entity is retrieved from the entity set of the virtual environment based on the first target entity information indicated by the natural language command and the environment-aware information; then, the second target entity is retrieved from the entity set of the virtual environment based on the second target entity information, the first target entity, and the environment-aware information.
[0597] Responding to natural language commands, control NPCs to perform virtual activities associated with the target entity.
[0598] For example, a large language model is invoked to identify the behavioral intent of natural language commands, and control instructions or sequences of control instructions are generated based on the behavioral intent. Based on the control instructions or sequences of control instructions, a non-player character is controlled to complete activities based on the target entity.
[0599] The broadcast of NPC feedback information in step 1108 can be implemented as follows:
[0600] Sub-step 6: Broadcast feedback information from non-player characters. The feedback text corresponding to the feedback information is non-fixed text generated based on the environmental perception information of non-player characters.
[0601] For example, a game application can generate and broadcast corresponding feedback information in real time based on the environmental awareness information of a non-player character. For instance, when the environmental awareness information of a non-player character triggers a broadcast condition, the game application can generate corresponding feedback information in real time based on the environmental awareness information that triggered the broadcast condition; or, when either a non-player character or the main virtual character triggers a broadcast condition, the game application can generate corresponding feedback information in real time based on the environmental awareness information of the non-player character.
[0602] For example, the feedback text is obtained by inferring from the static entity data of the 3D virtual environment and the dynamic environmental perception information of the non-player character using a large language model. The large language model will generate different feedback texts based on the different perception situations of the non-player character.
[0603] For example, feedback text can also be derived from static entity data of the 3D virtual environment and dynamic environmental perception information of the non-player character using a large language model, and can be tailored to the personality traits of the non-player character. For instance, when the non-player character is a robust, middle-aged man, the feedback text can be written in a bold and unrestrained tone; when the non-player character is a reporter, the feedback text can be written in a news reporting tone.
[0604] It should be noted that due to the randomness of the results generated by the large language model, in the same scene, when the environmental perception information of non-player characters is the same, the generated feedback text may also be different.
[0605] The environmental perception information can include information perceived in real time by non-player characters, or information perceived in the past by non-player characters.
[0606] For example, the feedback information is inferred by the game application based on the environmental perception information of non-player characters and static entity data of the 3D virtual environment. The feedback information can be in the form of voice or text. The static entity data includes data on relatively unchanging entities in the 3D virtual environment, such as model information of various buildings, terrains, and vehicles in the 3D virtual environment.
[0607] For example, feedback information may include at least one of the following: immediate feedback, execution feedback, and dynamic feedback. Immediate feedback and execution feedback are both feedback information generated in response to natural language commands, while dynamic feedback is feedback information spontaneously generated by a non-player character.
[0608] 1. Real-time feedback is feedback information generated instantly upon receiving a natural language command. Real-time feedback can include response broadcasts and real-time broadcasts. Response broadcasts are used to respond to queries in natural language commands; real-time broadcasts are used to provide feedback on the reception status of natural language commands.
[0609] For example, a client can instruct a non-player character to move within a 3D virtual environment based on the intent of a natural language command. Upon receiving a natural language command, the client can generate immediate feedback indicating that the command has been received, or it can provide an immediate response to the command.
[0610] In one optional embodiment, the feedback information includes a response broadcast. The response broadcast includes a reply to the natural language command posed by the player. For example, if the natural language command asks for the location of a target entity, the corresponding response broadcast should include the query result for the location of the target entity.
[0611] When the terminal receives a natural language command whose behavioral intent is an inquiry, it broadcasts the response from the non-player character. The response includes the response content generated based on the non-player character's environmental perception information of the 3D virtual environment in response to the natural language command.
[0612] Natural language commands are used to request information from non-player characters. These commands can be called query natural language commands. A query natural language command is a natural language command intended to ask a question. A query natural language command contains at least one question.
[0613] Natural language commands are instructions conveyed using human natural language. The terminal receives natural language commands from the player and directs non-player characters according to the intent behind those commands. For example, the terminal receives the player's voice audio, converts it into text, and obtains the natural language command. Alternatively, the terminal receives the text of a natural language command input by the player.
[0614] Players can issue natural language commands using spoken or written language. The game application will then use a large language model to understand these commands, extract the intent expressed within them, and generate control instructions based on that intent to direct the actions of non-player characters.
[0615] For example, a natural language command could be "Come here." The game application can then parse the intention of the natural language command as "Control the non-player character to move to the location of the main virtual character." The game application can then obtain the location of the main virtual character, generate a navigation route for the non-player character to move to the location of the main virtual character, and control the non-player character to move to the location of the main virtual character according to the navigation route.
[0616] For example, a natural language command could be "go pick up the treasure chest". The game application can then parse the command to determine that its intent is "to control a non-player character to move to the location of the treasure chest and pick it up". The game application can then obtain the location of the treasure chest, control the non-player character to move to the location of the treasure chest, and perform the operation of picking up the treasure chest after arriving at the location.
[0617] For example, when a natural language command is used to inquire about the intelligence of a target entity, a first response is broadcast based on the non-player character's perception of the target entity; the first response includes the non-player character's perception of the intelligence of the target entity.
[0618] Alternatively, when the natural language command query includes a question about the location of the target entity, a second response is broadcast based on the non-player character's perception of the target entity. The second response includes a description of the target entity's location.
[0619] For example, a large language model is invoked to parse the received natural language command, obtaining its intent. When the intent is a query, the command is determined to be a query command. When the command is a query, the large language model can infer the response text based on the intent, static entity data, and environmental awareness. In other words, the large language model parses the received natural language command, outputting the intent as a query and the response text. Following the query intent processing logic, the game application invokes a text-to-speech service to convert the response text into an audio message, which is then sent to the client for playback.
[0620] For example, when the large language model recognizes that the intent of the natural language command is a question, since there may be hundreds or thousands of questions from users, it is impossible to store the response to each question locally on the client. Therefore, the game application will call the text-to-speech service to generate a response broadcast in real time based on the response text returned by the large language model, so that the game application can generate the corresponding response broadcast to reply to the player's question in an instant.
[0621] For example, a game application (client or server) calls a large language model to parse the query natural language command, and infers the response text based on the static environmental data and environmental perception information of the three-dimensional virtual environment; based on the response text, it generates human voice audio and obtains the response broadcast.
[0622] For example, the client reports the received query natural language command to the server. The server calls a large language model to parse the query natural language command, and infers the response text based on the static environmental data and environmental perception information of the 3D virtual environment. Based on the response text, it generates human voice audio and obtains the response broadcast. The server sends the response broadcast to the client. The client then broadcasts the response.
[0623] 2. Execution feedback is the feedback generated when a task is executed according to the intent indicated by a natural language command after receiving the command. Execution feedback can also be called execution broadcast.
[0624] For example, when natural language commands are used to instruct a non-player character to perform a target task, the client can generate execution feedback during the execution of the target task, which indicates the execution status of the target task; or, the client can generate execution feedback after the non-player character completes the target task, which indicates the execution result of the target task.
[0625] In one optional embodiment, the feedback information includes feedback broadcasts (also referred to as execution feedback). Feedback broadcasts include broadcasts of the task execution status and results generated in response to the player's natural language command for the task. For example, if the task's natural language command includes moving to a target entity, the corresponding feedback broadcast might include: moving towards the target entity, or having reached the target entity.
[0626] When the terminal receives a natural language command whose behavioral intent is to perform a task, it broadcasts a feedback broadcast. The feedback broadcast includes the execution status after executing the natural language command based on the non-player character's environmental perception information of the three-dimensional virtual environment.
[0627] Natural language commands are used to instruct non-player characters to perform tasks. These commands can be called task natural language commands. Task natural language commands include a description of the task, which may include at least one of the following: task name, the action required of the non-player character, the method used to perform the task, the task objective, and the location of the task objective.
[0628] For example, if the intent of a task's natural language command includes finding a target entity, a first feedback broadcast is given, which includes the search results for the target entity from a non-player character.
[0629] Alternatively, if the intent of the task's natural language command includes the use of virtual props, a second feedback broadcast is given to indicate that the virtual props are ready.
[0630] Alternatively, if the intent of the task's natural language command includes the use of virtual props, a third feedback broadcast is given to indicate the result of using the virtual props.
[0631] Alternatively, when the intent of a task's natural language command includes controlling the movement of a non-player character, a fourth feedback broadcast is given, which indicates the result of the movement.
[0632] Alternatively, if the intent of the task's natural language command includes performing an interaction with the target entity, a fifth feedback broadcast is given, which indicates at least one of the non-player character's search results for the target entity and the interaction results between the non-player character and the target entity.
[0633] Alternatively, if the intent of the task's natural language command includes item interaction, a sixth feedback broadcast is given, which is used to instruct a non-player character to perform the item interaction result.
[0634] For example, a large language model is invoked to parse the received natural language command to obtain its intent. When the intent is a task, the command can be identified as a task-oriented natural language command. When the command is indeed a task-oriented command, the large language model can infer the task instruction sequence based on the intent, static entity data, and environmental awareness information. This sequence is used to control a non-player character to execute the task. Upon receiving the task instruction sequence, the game application controls the non-player character's activities to execute the task according to the instructions within the sequence. During this process, the application can obtain and broadcast feedback on the execution status based on the feedback logic corresponding to different task instructions.
[0635] For example, since the commands that non-player characters can execute in a 3D virtual environment are limited and traversable, the client can store the feedback broadcasts corresponding to each character's commands locally. When a non-player character executes a corresponding command, the client can read the local feedback broadcasts and broadcast them via voice.
[0636] Alternatively, when there is variable content in the feedback broadcast corresponding to a certain task instruction, such as the target entity in the feedback broadcast being variable, or the location information in the feedback broadcast being variable, the client can also generate broadcast voice with variable content based on the target entity or location information returned by the large language model, and concatenate it with the broadcast voice with immutable content stored locally to obtain the final feedback broadcast.
[0637] For example, when the task's natural language command is "Find nearby treasure chests," parsing this command using a large language model reveals that the intent is a task, the task objective is to find a target entity, and the target entity is a treasure chest. Based on static entity data and environmental awareness information, inference can be made that a treasure chest exists to the left front of the main virtual character. If the target entity can be found, the feedback broadcast corresponding to this task objective is "Find a at x," where x is the location of the target entity and a is the name of the target entity. The game application can then use the location information "to the left front" from the inference result to generate location speech using online text-to-speech technology; and use the name of the target entity, "treasure chest," to generate name speech using online text-to-speech technology. Finally, the location speech and name speech are concatenated into the feedback broadcast to obtain the final feedback broadcast.
[0638] For example, the game application (client or server) calls a large language model to parse the natural language commands for the task, and infers the task execution instructions based on the static environmental data and environmental perception information of the three-dimensional virtual environment; controls the non-player character to perform the task according to the task execution instructions; generates feedback text according to the execution status of the non-player character; and generates human voice audio based on the feedback text to obtain feedback broadcast.
[0639] For example, the client reports the task's natural language command to the server; the server parses the task's natural language command, and based on the static environmental data and environmental perception information of the 3D virtual environment, infers the task execution instructions; controls the non-player character to execute the task according to the task execution instructions; generates feedback text based on the execution status of the non-player character's task; generates human voice audio based on the feedback text, and obtains feedback broadcast; the server sends the feedback broadcast to the client; the client receives and broadcasts the feedback broadcast.
[0640] 3. Dynamic feedback is feedback information that is spontaneously generated based on the environmental perception information of non-player characters. Dynamic feedback can also be called dynamic broadcast.
[0641] For example, without receiving natural language commands, the client can spontaneously generate dynamic feedback based on its perception of the 3D virtual environment. This dynamic feedback is used to indicate abnormal situations discovered by non-player characters within the 3D virtual environment. For example, abnormal situations may include: detecting hostile virtual characters, detecting changes in the status of friendly virtual characters, detecting dangerous situations, detecting signs of combat or looting, etc.
[0642] In one alternative embodiment, the feedback information includes dynamic announcements (also referred to as "dynamic feedback"). Dynamic announcements include announcements of abnormal situations perceived by non-player characters. For example, when a non-player character perceives an attack, an announcement is made that an attack is in progress; when a non-player character perceives a dangerous situation ahead, an announcement is made that there is danger ahead.
[0643] When the non-player character's environmental perception information of the 3D virtual environment meets the dynamic broadcasting conditions, the terminal broadcasts dynamic announcements.
[0644] Dynamic broadcasts are announcements generated by game applications based on environmental awareness information from non-player characters, identifying abnormal situations that require reporting. Dynamic broadcasts include alerts about the abnormal situation. For example, abnormal situations may include at least one of the following: detection of an enemy virtual character, detection of a change in the status of a friendly virtual character, or detection of new traces.
[0645] For example, when a non-player character senses a hostile virtual character, a first dynamic announcement is broadcast; the first dynamic announcement includes the location of the hostile virtual character sensed by the non-player character.
[0646] Alternatively, a second dynamic announcement may be broadcast when a non-player character perceives a dangerous situation; the second dynamic announcement is used to alert the player to the dangerous situation.
[0647] Alternatively, when a non-player character perceives a change in the status of a friendly virtual character, a third dynamic announcement can be broadcast; the third dynamic announcement is used to notify the friendly virtual character of a change in status.
[0648] Alternatively, if a non-player character discovers new traces in the 3D virtual environment, a fourth dynamic announcement will be broadcast; the fourth dynamic announcement is used to notify of the new traces. These new traces can be combat traces and / or looting traces.
[0649] For example, since the number of dynamic announcements that non-player characters can trigger in a 3D virtual environment is finite and traversable, the client can store dynamic announcement voices locally. When a corresponding dynamic announcement voice is triggered, the client can read the voice from the local storage and play it. Similarly, some dynamic announcement voices may contain variable content. The client can generate variable content in real time using text-to-speech technology and concatenate it with the dynamic announcement template to obtain the final dynamic announcement.
[0650] In one alternative embodiment, when generating feedback information, it may be necessary to query the location of the target entity in the 3D virtual environment, or to determine the target entity indicated in the natural language command from the 3D virtual environment. In this case, an entity query method is required to query the target entity.
[0651] Before executing a target entity query, static entity data for a 3D virtual environment needs to be pre-constructed. This static entity data construction serves both the reasoning of the large language model and the real-time feedback text generation on the game application side after the reasoning results are returned.
[0652] In one alternative embodiment, entities in the 3D virtual environment are traversed, and entity information for each entity is exported from the game engine to obtain static entity data.
[0653] For example, for in-game scenes, an editor tool was developed to perform full-scene StaticMesh traversal, special Actor traversal (such as interactive doors, which are not within the scope of StaticMesh), and vegetation traversal, exporting the position, orientation, and bounding box size of entities as static entity data for subsequent use by the large language model for labeling and inference.
[0654] StaticMesh is a static geometry resource type in Unreal Engine 4 (UE4) game engine, used to represent immutable 3D models such as buildings and props. Actor is a basic object class in Unreal Engine 4 (UE4) game engine, representing an entity in the game world, such as a character, object, or light source. An Actor can contain multiple components to achieve different functionalities.
[0655] For example, static entity data can also include spatial data, which is manually labeled. Spatial data is more complex than static entities because it contains multiple floors and nested layers within floors (e.g., a motel area contains a second floor, and the second floor contains room 201). For this type of data, manual labeling is used. Add a Volume to the scene in the editor, divide the required area of the entire image, and label it accordingly. A Volume is a special Actor class in the Unreal Engine 4 (UE4) game engine, representing a 3D region with a specific function, such as a trigger area or an audio area.
[0656] Once the static entity data is obtained, entities or spaces can be queried based on the static entity data during game operation.
[0657] For example, when a player issues a natural language command, the command, along with pre-built static entity data and real-time captured runtime data (including environmental awareness information from non-player characters and / or the main virtual character), is used to make a reasoning request to the large language model. Furthermore, when an NPC needs to provide dynamic feedback based on runtime data, entity queries are also performed to generate feedback text based on the current state.
[0658] When a natural language command contains a target entity, the game application queries the target entity based on the current position and orientation of the main virtual character, the position and orientation of teammates or enemies, and selects the target entity that matches the target entity information from at least one candidate entity located within the perception range of the main virtual character and / or non-player character, based on the target entity information described in the natural language command and the environmental perception information of the main virtual character and / or non-player character.
[0659] Optionally, natural language commands are used to instruct non-player characters to perform activities related to a target entity. For example, the target entity could be the non-player character's destination, an object the non-player character needs to observe, a target the non-player character needs to attack, or an object the non-player character needs to interact with.
[0660] A target entity is a virtual object existing in a 3D virtual environment. A target entity is an entity described in a natural language command. An entity can refer to a fixed virtual object in the 3D virtual environment; for example, a target entity can be a virtual building, virtual terrain, virtual vehicle, virtual plant, virtual prop, virtual item, etc. An entity can also refer to a virtual object in the 3D virtual environment that can interact with a virtual character; for example, a target entity can be an interaction point (e.g., a door, window, cabinet, cellar, etc.), virtual prop, virtual character, virtual light source, etc.
[0661] For example, a natural language command could be "Please open the door for me," in which case the target entity could be "door," and the non-player character would need to perform the door-related activity "open the door." Alternatively, a natural language command could be "Is the kitchen safe?" in which case the target entity could be "kitchen," and the non-player character would need to perform the kitchen-related activity "check if there are any enemy virtual characters or other dangerous situations in the kitchen."
[0662] Optionally, the natural language command may include target entity information. Target entity information can be descriptive text about the target entity. For example, target entity information may include at least one of the following: the target entity's name, type, location, and characteristics. The game application can use the target entity information in the natural language to locate the target entity from among many entities in the 3D virtual environment.
[0663] The game application parses natural language commands, extracts target entity information, determines the target entity from the entity set in the three-dimensional virtual environment based on the target entity information, and then controls the non-player character to perform activities related to the target entity according to the intention of the natural language commands.
[0664] The target entity is determined from the three-dimensional virtual environment by combining environmental perception information, which includes information perceived from the three-dimensional virtual environment by at least one of the main virtual character and non-player characters.
[0665] Because the target entity information in natural language commands is often vague—for example, a natural language command might be "move to the red truck," where the target entity information is "red truck"—but there may be many red trucks in a 3D virtual environment, the target entity cannot be accurately determined from the 3D virtual environment based solely on the target entity information in the natural language command.
[0666] Therefore, this application provides a method for determining a target entity by combining target entity information and environmental perception information. In the previous example, the "red truck" expressed by the player in the natural language command should be a red truck that the player can see. Therefore, by combining the field of vision of the main virtual character, the red truck located within the field of vision of the main virtual character can be selected from multiple red trucks in the three-dimensional virtual environment. This red truck is the target entity indicated by the player in the natural language command.
[0667] Since natural language commands are issued by players based on their perception of the 3D virtual environment, in order to accurately identify the target entity indicated in the natural language command, the game application combines the environmental perception information at the time the player issues the natural language command to identify the target entity in the natural language command. Based on the player's perception of the 3D virtual environment, the entity that the player can perceive in the 3D virtual environment and that is closest to the target entity information is deduced; this is the target entity.
[0668] For example, the 3D virtual environment includes multiple candidate entities that match natural language commands. The target entity is selected from these candidate entities based on environmental perception information from the controlling virtual character or a non-player character. For instance, a game application first selects multiple candidate entities from the entity set that match the target entity information in the natural language command, and then selects the entity that the controlling virtual character or a non-player character can perceive as the target entity based on the environmental perception information.
[0669] For example, the perception range of the main virtual character is determined based on its location and orientation. This perception range might be a fan-shaped area of a certain field of view in front of the main virtual character; entities outside this fan-shaped area are filtered out. Then, the target entities are sorted according to the similarity between their information and that of the candidate entities, and entities with higher similarity are selected as the target entities. When two entities within the perception range have the highest and equal similarity scores—for example, if entities 2 and 3 both have a similarity score of 2—then entity 2, which is closer to the main virtual character, is selected as the target entity.
[0670] When feedback information needs to be generated based on the location description of the target entity, the game application can perform a spatial query to generate a location description based on the location relationship between the main virtual character and the target entity.
[0671] Spatial query is primarily used when you need to obtain the location of the main virtual character, non-player characters, enemy virtual characters, gunshots, etc. For example, it can be used to broadcast announcements containing spatial information such as "Enemy spotted on the second floor of a motel".
[0672] The spatial query logic defines the concepts of parent and child regions. A parent region refers to a large area in a 3D virtual environment encompassing multiple buildings or isolated spaces, such as a motel with front and back yards and multiple rooms and a basement. Child regions refer to independent spaces within the parent region; these can be enclosed spaces or spaces with special functions, such as motel guest rooms or kitchens.
[0673] To make the location descriptions returned by spatial queries more similar to human expressions, when the main virtual character and the queried location are in different parent regions, the query result returns complete text information (e.g., if the main virtual character is outside a motel and the queried location is on the second floor of the motel, the result returns "The enemy is on the second floor of the motel"). When the main virtual character and the queried location are in the same parent region but different child regions, the query result returns only the information of the child region (e.g., if the main virtual character is on the first floor of the motel and the queried location is on the second floor, the result returns "The enemy is on the second floor"). When the main virtual character and the queried location are in the same parent region and the same child region, the query result only returns the position relative to the player's perspective (e.g., if both the main virtual character and the queried location are on the second floor of the motel, the result returns "The enemy is to the right front").
[0674] For example, when the controlling virtual character and the target entity are located in different parent spaces, the location description includes the parent space and child space of the target entity. When the controlling virtual character and the target entity are located in different child spaces of the same parent space, the location description includes the child space of the target entity. When the controlling virtual character and the target entity are located in the same child space of the same parent space, the location description includes the relative orientation of the target entity and the controlling virtual character; wherein, a parent space includes at least one child space.
[0675] Regarding the broadcast of NPC feedback information in step 1108, the feedback information is implemented as spatial human voice audio processed with spatial sound effects. The spatial sound effects processing can be implemented as follows:
[0676] Sub-step 7: Obtain the human voice audio corresponding to the feedback text;
[0677] In this embodiment of the application, the process of generating the feedback text is as described in sub-step 6 above, and will not be repeated here.
[0678] In some embodiments, the process of obtaining the corresponding human voice audio based on the feedback text can be implemented as follows: obtaining the feedback text and character feature information corresponding to the non-player character; generating human voice audio that matches the character feature information based on the feedback text.
[0679] Schematic illustration: Character feature information is used to describe the attributes of a non-player character; that is, character feature information describes the current state and characteristics of the non-player character. In some embodiments, the attributes of a non-player character include at least one of basic attributes and contextual performance attributes. Basic attributes are attributes pre-configured for the non-player character; that is, attributes that do not change due to the 3D virtual environment or the current context, such as the non-player character's age, gender, or faction. Contextual performance attributes are attributes that are determined in real-time within the context of the 3D virtual environment; that is, attributes associated with the current context, such as the non-player character's emotional information and behavioral information in the current context.
[0680] Optionally, the character characteristic information includes at least one of the following: basic character information, character emotional information, and character behavioral information. Basic character information indicates the non-player character's fundamental attributes, such as age, gender, and faction. Character emotional information indicates the non-player character's emotional state in the dialogue context corresponding to the feedback text, representing their situational performance attributes, such as excitement, anger, or sadness. Character behavioral information indicates the actions performed by the non-player character in the dialogue context, representing their situational performance attributes, such as running, attacking, or healing from injury.
[0681] In some embodiments, a pre-trained speech generation model is used to generate human voice audio. This speech generation model generates speech that matches the character feature information of a non-player character. Illustratively, feedback text and character feature information are input into the pre-trained speech generation model to obtain human voice audio, which is then used as the human voice audio.
[0682] Optionally, the above speech generation model can be implemented using neural network models such as convolutional neural networks, feedforward neural networks, residual networks, Transformers, and multimodal large language models (MLLM), without specific limitations.
[0683] In some embodiments, the speech generation model described above includes a text encoder, a character encoder, and a decoder. The text encoder encodes the input feedback text; that is, it inputs the feedback text into the text encoder to obtain a text-encoded representation. The character encoder encodes the input character feature information; that is, it inputs the character feature information into the character encoder to obtain a character-encoded representation.
[0684] To illustrate, after obtaining the text encoding representation and the role encoding representation, the text encoding representation and the role encoding representation are fused to obtain the encoding representation for input to the decoder. That is, the text encoding representation and the role encoding representation are fused to obtain the joint encoding representation, which is then input into the decoder to generate human voice audio.
[0685] In some embodiments, the character encoder includes at least one of a first sub-encoder, a second sub-encoder, and a third sub-encoder. The first sub-encoder is used for feature encoding based on the basic attributes of a non-player character; that is, when the character feature information includes basic character information, the basic character information is input into the first sub-encoder to obtain a first character encoding representation. The second sub-encoder is used for feature encoding based on the emotional state of a non-player character; that is, when the character feature information includes emotional character information, the emotional character information is input into the second sub-encoder to obtain a second character encoding representation. The third sub-encoder is used for feature encoding based on the behavioral state of a non-player character; that is, when the character feature information includes behavioral character information, the behavioral character information is input into the third sub-encoder to obtain a third character encoding representation.
[0686] Sub-step 8: Based on the relative orientation of the non-player character and the main virtual character in the three-dimensional virtual environment, obtain the sound effect parameters corresponding to the non-player character;
[0687] The illustrative relative positions of the non-player character and the main virtual character in the three-dimensional virtual environment are used to indicate the relative positions between the first position corresponding to the non-player character and the second position corresponding to the main virtual character in the three-dimensional virtual environment.
[0688] Optionally, the aforementioned relative orientation relationship includes at least one of positional distance, positional direction, and spatial enclosure between the non-player character and the main virtual character. The positional distance is used to indicate the distance between the non-player character and the main virtual character. The positional direction is used to indicate the angular relationship between the first position of the non-player character and the orientation direction of the main virtual character. The spatial enclosure is used to indicate the surrounding situation of the spatial elements forming the virtual space in the three-dimensional virtual environment around the main virtual character, and the propagation effect of space on audio when the non-player character emits audio in the virtual space.
[0689] In some embodiments, determining the relative positional relationship between the non-player character and the main virtual character can be achieved by: obtaining first position information of the non-player character in the three-dimensional virtual environment, obtaining second position information of the main virtual character in the three-dimensional virtual environment, and determining the relative positional relationship between the non-player character and the main virtual character based on the first position information and the second position information.
[0690] In some embodiments, the first position information of the non-player character and the second position information of the main virtual character are position information determined based on the same preset coordinate system. Optionally, the preset coordinate system can be implemented as the world coordinate system corresponding to the three-dimensional virtual environment, or the preset coordinate system can be implemented as a coordinate system established with the main virtual character as the origin.
[0691] In some embodiments, when the sound effect parameters corresponding to the non-player character are pre-generated and stored in a database, the sound effect parameters corresponding to the non-player character are retrieved from the database. In other embodiments, the sound effect parameters corresponding to the non-player character can also be generated in real time, that is, based on relative positional relationships, the sound effect parameters corresponding to the non-player character are generated.
[0692] Optionally, the sound effect parameters for non-player characters may include at least one of the following types:
[0693] The first type is distance type.
[0694] Schematic, distance-type audio parameters are used to indicate adjustments to the sound effects of human voice audio based on the distance between the non-player character and the controlling virtual character. In some embodiments, distance-type audio parameters can adjust the audio volume of the human voice audio to simulate the effect of distance on sound volume; and / or, distance-type audio parameters can adjust the audio delay of the human voice audio to simulate the effect of distance on sound propagation time.
[0695] The second type is the direction type.
[0696] Schematic, directional audio parameters are used to indicate adjustments to the sound effects of human voice audio based on the direction of the non-player character relative to the main virtual character. In some embodiments, directional audio parameters can adjust the audio volume of the human voice audio to simulate whether the main virtual character is facing scene audio elements; and / or, directional audio parameters can adjust the volume parameters of the human voice audio in different channels to simulate the direction in which scene sound source elements emit audio.
[0697] The third type is spatial special effects.
[0698] Indicatively, spatial effect type audio parameters are used to indicate adjustments to sound effects within a virtual space formed by spatial elements. Specifically, these parameters indicate the impact of the virtual space on the sound effects emitted by non-player characters. In some embodiments, spatial effect type audio parameters can add echo effects and reverberation effects to simulate the effects of sound within a space.
[0699] Optionally, the sound effect parameters for spatial effects include at least one of the following: reverberation time parameter, pre-delay parameter, wet / dry mix parameter, spatial width parameter, and distance effect parameter. Specifically, the reverberation time parameter indicates the rate attenuation of audio in the virtual space; the pre-delay parameter indicates the time difference between the direct arrival of audio on the main virtual character and the arrival of audio on the first reflection; the wet / dry mix parameter indicates the ratio between the directly propagated audio and the audio propagated through reflection; the spatial width parameter indicates the degree of audio diffusion on the horizontal plane in the virtual space; and the distance effect parameter indicates the attenuation of audio as it travels through propagation distance.
[0700] The fourth type is custom types.
[0701] The illustrative example shows custom-type sound effect parameters, which are user-defined parameters used for sound effect adjustments. Optionally, users can customize the overall volume of the scene audio, whether to add background music, and the volume of different types of audio, etc.
[0702] In some embodiments, the client provides the user with a custom interface for sound effect parameters, through which the user configures custom sound effect parameters.
[0703] In some embodiments, a parameter prediction model is configured in the user's custom sound effect parameter server to perform personalized learning of the custom sound effect parameters. Illustratively, custom sound effect parameters of multiple candidate accounts are obtained, and these multiple custom sound effect parameters are input into the parameter prediction model to be trained. The parameter prediction model is iteratively trained to obtain the parameter prediction model, which is then used to optimize the system-generated sound effect parameters (e.g., the distance-type, direction-type, and spatial effect-type sound effect parameters mentioned above).
[0704] Sub-step 9: Adjust the human voice audio based on the sound effect parameters, generate and broadcast the spatial human voice audio corresponding to the non-player character;
[0705] Among them, spatial human voice audio is used to represent the perceived audio effect of the main virtual character on the non-player character under relative orientation.
[0706] In this embodiment, the vocal audio is adjusted according to the sound effect parameters corresponding to the non-player character to generate spatial vocal audio corresponding to the non-player character. That is, the spatial vocal audio is the audio data corresponding to the non-player character obtained by adjusting the vocal audio with sound effect parameters. In some embodiments, the spatial vocal audio includes the audio adjusted with sound effect parameters, and at least one of the following data: audio duration information, audio playback start timestamp, audio playback end timestamp, audio playback condition information, and audio corresponding element identification information.
[0707] Optionally, when the sound effect parameters indicate adjustment of the volume of the non-player character's voice audio, the volume of the non-player character's voice audio on at least one channel is adjusted based on the sound effect parameters. Optionally, when the sound effect parameters indicate adjustment of the playback delay of the non-player character's voice audio, the playback start time corresponding to the non-player character's voice audio is adjusted based on the sound effect parameters. Optionally, when the sound effect parameters indicate adjustment of the tempo of the non-player character's voice audio, the playback speed of the non-player character's voice audio is adjusted based on the sound effect parameters. Optionally, when the sound effect parameters indicate adjustment of the pitch of the non-player character's voice audio, the frequency in the spectrum of the non-player character's voice audio is adjusted based on the sound effect parameters. When the sound effect parameters indicate adjustment of the timbre of the non-player character's voice audio, a filter corresponding to the non-player character is determined based on the sound effect parameters, and the timbre corresponding to the voice audio is adjusted through the filter.
[0708] As an example, the client plays spatial human voice audio to broadcast feedback information from non-player characters.
[0709] For the speech-to-text stage in step 1104, it can be implemented as follows:
[0710] Sub-step 10: Obtain multiple hot words for the first virtual scene corresponding to the first virtual scene where the target virtual character is located;
[0711] The target virtual character includes at least one of an NPC and a controlling virtual character.
[0712] In illustrative terms, the main virtual character is the virtual character commanded by the player, while an NPC is a virtual character that assists the main virtual character in virtual matches. Typically, a virtual match includes one main virtual character and at least one NPC. There are certain differences between the main virtual character and the NPC.
[0713] Optionally, the main virtual character is the virtual character primarily controlled by the player, while the NPC is a virtual character that the player chooses to command. For example, the main virtual character is a virtual character that the player manually controls during the game, and the player's manual operations on the terminal interface are used to control the main virtual character; the NPC is a virtual character that the player commands via voice, and the player occasionally issues natural language commands in the form of voice to command the NPC.
[0714] In some embodiments, a first virtual scene is determined based on a target virtual character, wherein the target virtual character is at least one of an NPC and a controlling virtual character, that is, the process of determining the first virtual scene is implemented in at least one of the following ways.
[0715] (1) If the main virtual character is selected as the target virtual character, then the virtual scene where the main virtual character is located is taken as the first virtual scene;
[0716] (2) If an NPC is chosen as the target virtual character, the virtual scene where the NPC is located is taken as the first virtual scene; if there are multiple NPCs, the virtual scene containing the most NPCs can be taken as the first virtual scene, or the NPC closest to the main virtual character can be taken as the target virtual character and the virtual scene where it is located can be taken as the first virtual scene, or the NPC with the largest character attribute value (such as at least one of virtual health, virtual mana, virtual defense, etc.) can be taken as the target virtual character and the virtual scene where it is located can be taken as the first virtual scene;
[0717] (3) If you choose to use both the main virtual character and the NPC as the target virtual character, you can use the virtual scene where the main virtual character and the NPC are located as the first virtual scene, etc.
[0718] In some embodiments, the target virtual character is determined...
Claims
1. A voice conversion method based on a virtual scene, characterized by, The method comprises: acquiring a natural language command in a voice form, the natural language command being used to instruct a non-player character, a target virtual character being in a first virtual scene, the target virtual character comprising at least one of the non-player character and a host virtual character; acquiring an object perception range of the target virtual character in the first virtual scene, the object perception range being a three-dimensional spatial range in which the target virtual character has a perception of other scene elements; the object perception range comprising at least one of a visual perception range of the target virtual character, an auditory perception range of the target virtual character, an olfactory perception range of the target virtual character, a perception skill or a perception device possessed by the target virtual character, and a range perceived by the perception device; acquiring a plurality of first scene hotwords based on environmental perception information in the object perception range, the plurality of first scene hotwords being scene-related words of the first virtual scene; converting the natural language command into a command analysis result in a text form based on the plurality of first scene hotwords.
2. The method of claim 1, wherein, The acquiring of the plurality of first scene hotwords based on the environmental perception information in the object perception range comprises: acquiring a first scene type corresponding to the first virtual scene, acquiring a hotword set corresponding to the first scene type, acquiring the object perception range of the target virtual character in the first virtual scene, and acquiring at least two words from the hotword set as the plurality of first scene hotwords based on the environmental perception information in the object perception range.
3. The method of claim 1, wherein, The acquiring of the object perception range of the target virtual character in the first virtual scene comprises: in a case where the object perception range comprises a visual perception range, generating a preset visual cone as the visual perception range, the preset visual cone being a three-dimensional region for measuring the visual perception range, with a position of the target virtual character as a starting point and an orientation of the target virtual character as a cone center line.
4. The method of claim 1, wherein, The environmental perception information comprises virtual elements. The acquiring of the plurality of first scene hotwords based on the environmental perception information in the object perception range comprises: determining the virtual elements in the first virtual scene that are in the object perception range; taking an element name of the virtual element as the first scene hotword, or taking an action name of an interactive action corresponding to the virtual element as the first scene hotword.
5. The method according to any one of claims 1 to 4, characterized in that, The first virtual scene is one of a plurality of virtual scenes in a virtual environment. After the acquiring of the natural language command in the voice form, the method further comprises: based on a position of the target virtual character in the virtual environment, determining the first virtual scene in which the target virtual character is located from the plurality of virtual scenes.
6. The method of claim 5, wherein, The method further comprises: in a case where the command analysis result is not generated, in response to the target virtual character moving from the first virtual scene to a second virtual scene, acquiring a plurality of second scene hotwords corresponding to the second virtual scene; converting the natural language command into the command analysis result in the text form based on the plurality of second scene hotwords.
7. The method according to any one of claims 1 to 4, characterized in that, The method further comprises: determine the first virtual scene corresponding to the first virtual game session in which the target virtual character participates; a plurality of virtual game sessions correspond to a virtual scene respectively, and the plurality of virtual game sessions are simulation battle environments provided for the target virtual character.
8. The method of claim 1, wherein, The method further comprises: obtaining a plurality of scene hotwords, the plurality of scene hotwords corresponding to at least two virtual scenes, the plurality of scene hotwords corresponding to scene identifiers, and the scene identifiers being used to represent virtual scenes when the scene hotwords are collected; obtaining a plurality of candidate scene hotwords with first scene identifiers from the plurality of scene hotwords based on the first virtual scene in which the target virtual character is located, the first scene identifiers being scene identifiers corresponding to the first virtual scene; selecting at least two candidate scene hotwords from the plurality of candidate scene hotwords as the first scene hotwords based on the environmental perception information within the object perception range of the target virtual character in the first virtual scene.
9. The method according to any one of claims 1 to 4, characterized in that, The conversion of the natural language command into the command analysis result in text form based on the plurality of first scene hotwords comprises: obtaining a pre-trained natural language analysis model, the pre-trained natural language analysis model comprising an acoustic network, a language network, and a preset dictionary, the preset dictionary comprising the plurality of first scene hotwords, other scene hotwords, and general vocabulary; obtaining a plurality of phonetic units corresponding to the natural language command through the acoustic network, the phonetic unit being a basic unit of vocabulary pronunciation; analyzing a sequence matching relationship between the plurality of phonetic units and selected vocabulary in the preset dictionary through the language network to obtain the command analysis result in text form, the selected vocabulary comprising at least the plurality of first scene hotwords and the general vocabulary.
10. The method of claim 9, wherein, The selected vocabulary further comprises the other scene hotwords. The vocabulary in the preset dictionary comprises an analysis weight, a first analysis weight of the plurality of first scene hotwords being higher than a second analysis weight of the other scene hotwords, and the analysis weight being a degree of attention of vocabulary participating in sequence matching.
11. The method of claim 9, wherein, The analysis of the sequence matching relationship between the plurality of phonetic units and the selected vocabulary in the preset dictionary through the language network to obtain the command analysis result in text form comprises: analyzing a matching relationship between the plurality of phonetic units and the selected vocabulary in the preset dictionary to obtain a plurality of candidate vocabulary sequences, the plurality of candidate vocabulary sequences comprising at least one vocabulary in the plurality of selected vocabulary, and the candidate vocabulary sequence being a vocabulary sequence obtained based on a change relationship between the plurality of phonetic units; analyzing sequence semantics of the plurality of candidate vocabulary sequences through the language network to obtain at least one candidate vocabulary sequence from the plurality of candidate vocabulary sequences as the command analysis result in text form.
12. The method of claim 11, wherein, The language network comprises language sub-networks corresponding to at least two virtual scenes respectively. The analyzing the sequence semantics of the plurality of candidate word sequences through the language network comprises: determining, based on the first virtual scene in which the target virtual character is located, a first language sub-network corresponding to the first virtual scene from at least two language sub-networks; analyzing the sequence semantics of the plurality of candidate word sequences through the first language sub-network, and obtaining at least one candidate word sequence from the plurality of candidate word sequences as the command analysis result in the form of text.
13. A voice conversion apparatus based on a virtual scene, characterized by, The apparatus comprises: an obtaining module configured to obtain a natural language command in a voice form, the natural language command being used to instruct a non-player character, a target virtual character being located in a first virtual scene, the target virtual character comprising at least one of the non-player character and a master virtual character; the obtaining module is further configured to obtain an object perception range of the target virtual character in the first virtual scene, the object perception range being a three-dimensional spatial range in which the target virtual character has perception of other scene elements; the object perception range comprising at least one of a visual perception range of the target virtual character, an auditory perception range of the target virtual character, an olfactory perception range of the target virtual character, a range perceived by a perception skill or a perception device possessed by the target virtual character, and a plurality of first scene hot words obtained based on environmental perception information in the object perception range, the plurality of first scene hot words being scene-related words of the first virtual scene; an analyzing module configured to convert the natural language command into a command analysis result in a text form based on the plurality of first scene hot words.
14. A computer device, comprising: The computer device comprises a processor and a memory, the memory storing at least one program, the at least one program being loaded and executed by the processor to implement the virtual scene-based voice conversion method according to any one of claims 1 to 12.
15. A computer readable storage medium characterized by: The storage medium stores at least one program, the at least one program being loaded and executed by the processor to implement the virtual scene-based voice conversion method according to any one of claims 1 to 12.
16. A computer program product, characterised in that, The computer program product comprises computer instructions, the computer instructions being executed by the processor to implement the virtual scene-based voice conversion method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Method and system for driving virtual character to act by real-time voice
CN111939558A
Speech recognition method and device, electronic equipment and storage medium
CN112562659A