Scene object processing method, device, equipment and storage medium
By combining multimodal models and natural language models, visual labels of virtual scene objects are generated and optimized, which solves the problem of inefficient object labeling in virtual scenes and achieves efficient and accurate visual label generation.
Patent Information
- Application Number
- CN202411097944.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-09
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-08-09
AI Technical Summary
In virtual scenes, the visual labeling of scene objects is inefficient, making it difficult to efficiently batch process a large number of scene objects.
By obtaining the attribute text and appearance images of scene objects, a multimodal model is used for prediction to generate visual labels for scene objects, and the labels are optimized through a natural language model to improve labeling efficiency and accuracy.
It achieves efficient batch annotation of visual labels for scene objects, ensures the accuracy and comprehensiveness of the labels, and improves the efficiency of object feature description in virtual scenes.
Smart Images

Figure CN119091340B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and storage medium for processing scene objects. Background Art
[0002] A large number of scene objects are deployed in an application program including a virtual scene to provide a virtual space for the virtual activities of virtual characters.
[0003] In related technologies, it is usually necessary to record the visual tags of scene objects when drawing or writing scene object material files; or after completing the deployment of the virtual scene, mark the visual tags of scene objects one by one, and then search for scene objects in the virtual scene based on the visual tags.
[0004] However, there are a large number of scene objects in the virtual scene, and the above marking method is inefficient. Summary of the Invention
[0005] This application provides a method, apparatus, device, and storage medium for processing scene objects. The technical solutions are as follows:
[0006] According to one aspect of the present application, a method for processing scene objects is provided, the method comprising:
[0007] Obtaining a first modality image, a second modality image, and a classification label of the second modality image;
[0008] Inputting the first modality image and the second modality image into the image translation model to obtain a first translated image, where the first translated image is an image generated when the image translation model predicts the first modality image based on the second modality image;
[0009] Inputting the first translated image into a classification network to obtain a first predicted probability, and inputting the second modality image into the classification network to obtain a second predicted probability, wherein the first predicted probability is a probability that the image content of the first translated image belongs to the target tissue, and the second predicted probability is a probability that the image content of the second modality image belongs to the target tissue;
[0010] Calculating a structural consistency error based on the first predicted probability, the second predicted probability, and the classification label, where the structural consistency error is used to represent the difference between the first predicted probability and the second predicted probability and the classification label respectively;
[0011] Based on the structural consistency error, backward error propagation training is performed on the image translation model to obtain a trained image translation model.
[0012] According to another aspect of the present application, a device for processing scene objects is provided, the device comprising:
[0013] An acquisition module, configured to acquire attribute text of a scene object in a virtual scene, wherein the attribute text is used to introduce inherent attributes of the scene object in the virtual scene;
[0014] The acquisition module is further configured to acquire an appearance image of the scene object in the virtual scene, wherein the appearance image is used to describe the style of the scene object;
[0015] A processing module is used to call a multimodal model to perform prediction on the attribute text and the appearance image of the scene object to obtain a visual label of the scene object, where the visual label is used to describe the visual features of the scene object in at least one dimension.
[0016] In an optional design of the present application, the multimodal model includes a visual question answering model, and the processing module is further configured to:
[0017] Constructing a question sentence for the appearance image, wherein the question sentence carries the attribute text of the scene object;
[0018] The appearance image of the scene object and the question statement are input into the visual question answering model to obtain an answer statement, and the answer statement is used as the visual label of the scene object.
[0019] In an optional design of the present application, the question statement includes at least two sub-statements; and the processing module is further configured to:
[0020] Inputting the appearance image of the scene object and the first sentence of the at least two sub-sentences into the visual question answering model to obtain a first answer sub-sentence;
[0021] Repeat the above steps until at least two answer sub-sentences corresponding one-to-one to the at least two sub-sentences are obtained, wherein the at least two sub-sentences are used to inquire about the visual features of the appearance image from multiple dimensions;
[0022] Sentence aggregation is performed on the at least two answer sub-sentences to extract the visual label of the scene object.
[0023] In an optional design of the present application, the processing module is further configured to:
[0024] Acquiring desired information of the scene object, where the desired information is used to indicate desired description dimensions of the scene object in the visual tag and / or desired format of the visual tag;
[0025] constructing the question sentence of the appearance image according to the expected information and the attribute text;
[0026] Among them, the first sub-part of the question statement is the supplementary introduction information of the scene object, carrying the attribute text of the scene object; the second sub-part of the question statement is the answer guidance statement for the visual question answering model, carrying the expected information.
[0027] In an optional design of the present application, the multimodal model includes an image description model, and the processing module is further configured to:
[0028] Inputting the appearance image of the scene object into the image description model to predict a description text of the scene object;
[0029] Sentence aggregation is performed on the description text and the attribute text to extract the visual label of the scene object.
[0030] In an optional design of the present application, the attribute text of the scene object includes at least one of the name of the scene object in the virtual scene and the size of the scene object in the virtual scene;
[0031] And / or, the appearance image of the scene object includes images obtained by observing the scene object from at least two viewing angles.
[0032] In an optional design of the present application, the processing module is further configured to:
[0033] Performing pseudo-colloquial rewriting on the visual label of the scene object to obtain a matching label that conforms to the spoken expression of natural language.
[0034] In an optional design of the present application, the processing module is further configured to:
[0035] The visual label of the scene object is input into a natural language model to predict the matching label that conforms to the spoken expression of the natural language, and the natural language model carries prior knowledge of the spoken expression of the natural language.
[0036] In an optional design of the present application, the processing module is further configured to:
[0037] Obtaining a first sample label pair, the first sample label pair including a first label before being rewritten into a pseudo-colloquial language and a second label obtained after being rewritten into a pseudo-colloquial language;
[0038] constructing a rewriting guidance sentence based on the first sample label pair and the visual label, wherein the rewriting guidance sentence has natural semantics of rewriting the visual label with reference to the first sample label pair;
[0039] The rewriting guide sentence is input into a natural language model to predict the matching label that conforms to the spoken expression of natural language.
[0040] In an optional design of the present application, the acquisition module is further configured to:
[0041] The spatial position of the scene object in the virtual scene is acquired, and the spatial position of the scene object is determined as auxiliary information of the visual label of the scene object.
[0042] In an optional design of the present application, the processing module is further configured to:
[0043] Obtaining at least one of the coordinate position, orientation information, bounding box information, and cover point information of the scene object in the virtual scene;
[0044] Among them, the coordinate position is used to indicate the position of the scene object in the virtual scene, the orientation information is used to indicate the direction facing the scene object in the virtual scene, the bounding box information is used to indicate the size of the scene object in the virtual scene, and the cover point information indicates the recommended virtual character standing position when the virtual character approaches the scene object.
[0045] In an optional design of the present application, the processing module is further configured to construct a spatial information library of the virtual scene based on visual labels of a plurality of scene objects in the virtual scene;
[0046] The acquisition module is further used to acquire natural language commands;
[0047] The processing module is further configured to filter out a first scene object corresponding to the natural language command from a plurality of scene objects in the virtual scene based on a similarity between the natural language command and the visual tags in the spatial information library;
[0048] The processing module is further configured to execute a virtual activity for the first scene object based on the natural language command.
[0049] In an optional design of the present application, the natural language command includes multiple input words; the processing module is further configured to:
[0050] determining a similarity score between each word in the plurality of input words and the visual label of the first scene object;
[0051] determining a sum of a plurality of similarity scores as a similarity between the first scene object and the natural language command, the plurality of similarity scores corresponding one-to-one to the plurality of input words;
[0052] When the similarity between the first scene object and the natural language command exceeds the similarity between other scene objects in the spatial information library and the natural language command, the first scene object is determined as the object corresponding to the natural language command.
[0053] In an optional design of the present application, the processing module is further configured to construct a spatial information library of the virtual scene based on appearance images of a plurality of scene objects in the virtual scene;
[0054] The acquisition module is further used to acquire natural language commands;
[0055] The processing module is further configured to call a similarity matching network to predict a degree of similarity between the natural language command and the appearance images of the plurality of scene objects in the spatial information library, and obtain a second scene object among the plurality of scene objects having the highest similarity to the natural language command;
[0056] The processing module is further configured to execute a virtual activity for the second scene object based on the natural language command.
[0057] According to another aspect of the present application, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method for processing scene objects as described above.
[0058] According to another aspect of the present application, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the method for processing scene objects as described above.
[0059] According to another aspect of the present application, a computer program product is provided, which includes computer instructions stored in a computer-readable storage medium. A processor reads and executes the computer instructions from the computer-readable storage medium to implement the above-mentioned method for processing scene objects as described above.
[0060] The beneficial effects of the technical solution provided by this application include at least:
[0061] By calling the multimodal model to perform predictions on attribute text and appearance images, the visual labels of scene objects are predicted based on the descriptive information of scene objects in the text dimension and the image dimension. In the process of predicting the visual labels of scene objects, the descriptive information of scene objects in multiple dimensions is used to ensure that the characteristics of scene objects are fully described from the natural language semantics in the text dimension and the visual information in the image dimension, thereby ensuring the accuracy of the visual labels. Calling the multimodal model to perform predictions on attribute text and appearance images can realize batch annotation of visual labels of scene objects, thereby improving the efficiency of obtaining visual labels of scene objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0063] Figure 1 is a schematic diagram of controlling a non-player character in a virtual environment provided by an exemplary embodiment of the present application;
[0064] Figure 2 is a structural block diagram of a computer system provided by an exemplary embodiment of the present application;
[0065] Figure 3 is a schematic diagram of a method for processing scene objects provided by an exemplary embodiment of the present application;
[0066] Figure 4 is a flow chart of a method for processing scene objects provided by an exemplary embodiment of the present application;
[0067] Figure 5 is a flow chart of a method for processing scene objects provided by an exemplary embodiment of the present application;
[0068] Figure 6 is a schematic diagram of a visual label of a scene object provided by an exemplary embodiment of the present application;
[0069] Figure 7 is a flow chart of a method for processing scene objects provided by an exemplary embodiment of the present application;
[0070] Figure 8 is a flow chart of a method for processing scene objects provided by an exemplary embodiment of the present application;
[0071] Figure 9 This is a schematic diagram of information about scene objects provided by an exemplary embodiment of the present application;
[0072] Figure 10 is a flow chart of a method for processing scene objects provided by an exemplary embodiment of the present application;
[0073] Figure 11 is a schematic diagram of a natural language command provided by an exemplary embodiment of the present application;
[0074] Figure 12 is a schematic diagram of an appearance image of a scene object provided by an exemplary embodiment of the present application;
[0075] Figure 13 is a schematic diagram of a spatial information library of a virtual scene provided by an exemplary embodiment of the present application;
[0076] Figure 14 is a flow chart of a method for controlling a virtual character provided by an exemplary embodiment of the present application;
[0077] Figure 15 is a structural block diagram of a scene object processing device provided by an exemplary embodiment of the present application;
[0078] Figure 16 It is a structural block diagram of a terminal provided by an exemplary embodiment of the present application.
[0079] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application. DETAILED DESCRIPTION
[0080] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0081] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0082] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0083] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the attribute text, appearance images, and other information involved in this application are all obtained with full authorization.
[0084] It should be understood that although the terms first, second, etc. may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, a first parameter may also be referred to as a second parameter, and similarly, a second parameter may also be referred to as a first parameter without departing from the scope of this disclosure. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0085] The embodiments of the present application provide a solution for users to use natural language commands in the form of voice to command non-player characters (NPCs) in a virtual environment. In this solution, in addition to the main virtual character controlled by the user, the user usually has one or more NPCs as teammates. The user can use a relatively casual rather than mechanical human conversation method to command the NPC to make the desired feedback to the scene objects in the virtual environment, thereby commanding the NPC to complete the task in collaboration with the main virtual character. For example, the NPC has the ability to act autonomously, and the user only commands the NPC. The user gives commands, and the NPC understands and executes the commands based on its own autonomous behavior capabilities. Exemplary reference Figure 1 , the user controls the main virtual character 10 to play the game, and the main virtual character 10 has an NPC teammate 20. There is a car 30 within the field of view of the main virtual character 10. The user speaks a natural voice command in the form of speech: "No. 2, you go to ambush behind car 30", and then the NPC teammate 20 will automatically move to the back of car 30 to ambush. It should be noted that Figure 1 There may be multiple cars in the scene shown, and the NPC teammate 20 will accurately understand that the car the user is talking about is the car 30 located within the user's field of view. In other words, the NPC teammate 20 has a relatively intelligent natural language understanding ability.
[0086] On the one hand, compared with the traditional technology of using mechanical fixed instructions to command NPC, the embodiment of the present application uses relatively random natural language commands to command NPC, which can provide users with more natural, more complex and more flexible language command capabilities.
[0087] On the other hand, the intelligence of the natural language understanding capability provided by the embodiment of the present application is reflected in the following: when understanding natural language commands, NPC teammates not only consider the literal information of the natural language commands. NPC teammates will also combine the main virtual character and / or the NPC's own environmental perception information in the virtual environment to assist in understanding the semantics in the natural language commands. That is, the NPC not only considers the information of the natural language command modality, but also considers the perception information of other modalities such as vision, hearing, radar, etc. to assist in the understanding and execution of natural language commands. Since natural language commands in the form of voice are spoken commands rather than written commands, there may be multiple candidate understanding methods or unclear places when considering the literal information of natural language commands alone. NPC teammates, combined with the main virtual character and / or the NPC's own environmental perception information in the virtual environment, can determine a reasonable understanding method from multiple candidate understanding methods, or eliminate doubts about unclear places, thereby achieving a more intelligent natural language understanding capability. The above-mentioned environmental perception information includes but is not limited to at least one of the following:
[0088] Visual information perceived within the field of view of the controlling virtual character;
[0089] Auditory information perceived within the hearing range of the controlling avatar;
[0090] The information perceived by the control avatar's sensory skills or sensory gadgets;
[0091] Visual information perceived within the NPC's field of view;
[0092] Auditory information perceived within the NPC's hearing range;
[0093] Information perceived by the NPC's perception skills or perception tools.
[0094] Combined with reference Figure 1 The program includes at least one of the following five stages:
[0095] Phase 1: Preprocessing of spatial data;
[0096] The server will pre-process the scene objects in the virtual scene and construct the spatial data of the virtual scene.
[0097] The spatial data of the virtual scene includes the visual labels of the scene objects and the spatial positions of the scene objects. In some examples, the spatial data also includes the appearance images of the scene objects. Scene objects are any objects that appear in the virtual scene, such as Figure 1 Cars, walls, boxes, etc.
[0098] First, we introduce the process of pre-acquiring visual labels for scene objects. This involves acquiring attribute information for the virtual scene objects, such as at least one of their names and dimensions. Furthermore, we acquire appearance images of the scene objects. These images include images of the scene objects observed from at least two perspectives. Observing the scene objects from multiple perspectives ensures that the appearance images contain comprehensive information about the scene objects.
[0099] The server calls a multimodal model to predict the attribute information and appearance images of scene objects. Based on the multimodal model's ability to process text-dimensional and image-dimensional information, it extracts hidden features from the text-dimensional attribute information and the image-dimensional appearance images, and decodes the hidden features to predict the visual labels of the scene objects. The visual labels are used to describe the scene objects in at least one dimension in a textual manner; for example, describing the scene objects in dimensions such as material, transparency, color, and shape. The natural language model is called to perform label optimization on the visual labels to obtain the visual labels of the scene objects. The natural language model has the ability to generate text and rewrites the visual labels input into the natural language model in a manner consistent with the spoken expression of natural language to achieve label optimization of the visual labels and obtain optimized visual labels for the scene objects. The optimized visual labels conform to the spoken expression of natural language and describe the scene objects in at least one dimension in a manner close to spoken expression.
[0100] For example, when describing a virtual bed in a virtual scene, people often focus on the color, material, and placement of the virtual bed, but often neglect to pay attention to the craftsmanship used on the headboard and side panels, such as carving or embossing. The goal of performing pseudo-colloquial rewriting on the visual labels of scene objects is to obtain matching labels that are closer to the spoken description. For example, by enriching the labels of dimensions such as color, material, and placement that are of interest in spoken description with multiple words with the same semantics, and removing the labels of the surface features of the headboard and side panels that are less concerned in spoken description.
[0101] Next, we'll discuss the spatial position of scene objects. This includes at least one of the following: location, such as the coordinates of the scene object's center or a preset point in the virtual scene; rotation, such as the direction the front of the scene object faces in the virtual scene; a bounding box, which indicates the size of the scene object in the virtual scene; and cover points, which indicate recommended locations for the avatar when it approaches the scene object, enabling the scene object to mask the avatar.
[0102] Next, we will introduce the process of pre-acquiring the appearance image of the scene object. The appearance image is acquired during the process of predicting the visual label of the scene object. The appearance image includes images obtained by observing the scene object from at least two perspectives.
[0103] Optionally, the spatial data of each scene object in the entire virtual environment may be pre-acquired and organized into a data set for use in a subsequent query process.
[0104] Phase 2: Speech recognition and intent recognition;
[0105] When the user controls the main virtual character, the terminal receives the natural language command input by the user in the form of voice; the server performs voice recognition and intent recognition on the natural language command to analyze the user's input instruction;
[0106] Speech recognition is to perform text conversion on the natural language instructions input by the user through voice, obtain the instruction text corresponding to the natural language instructions, and convert the voice information into text information. Call the encoding network to perform feature encoding on the instruction text to obtain the feature representation of the natural language instructions in the latent space; call the text segmentation network to perform segmentation on the feature representation to obtain multiple clauses in the instruction text, and then perform intent recognition on each of the multiple clauses separately. In the case that the natural language instructions input by the user through voice are long sentences, the instruction text corresponding to the natural language instructions can be segmented to meet the needs of analyzing the user's input instructions in the long sentence scenario.
[0107] The following introduces the intent recognition of each clause; the user's input instruction is obtained through intent recognition, and the user's input instruction includes an entity, semantic type, subject type and intent type in at least one clause.
[0108] On the one hand, a Conditional Random Field (CRF) approach is used to identify entities in clauses, such as identifying entities in clauses as scene objects such as buildings, virtual items, and virtual vegetation in virtual scenes. Entities in clauses are used to indicate the scene objects targeted by virtual activities. For example, if the entity in a clause is a virtual building, and the clause carries the semantics of instructing an NPC to launch a virtual attack, then the virtual attack is carried out against the virtual building, that is, a virtual attack is launched against the virtual building. On the other hand, a prediction network is called to predict the semantic type, subject type, and intent type in the clause. The semantic type includes whether to instruct an NPC to launch a virtual attack; the subject type is used to indicate whether the subject of the clause is one or more NPCs, or a master virtual character controlled by the user; and the intent type is the desired virtual activity indicated by the clause, such as moving, launching a virtual attack, etc.
[0109] Phase 3: Query of spatial data;
[0110] For the entity in the input command, it is necessary to determine the spatial location of the entity in the virtual environment. If there are multiple candidate entities in the virtual environment that match the entity in the input command, the target entity that matches the input command can be uniquely determined from the multiple candidate entities by combining the environmental perception information of the controlling virtual character and / or NPC.
[0111] For example, assuming that the environmental perception information is the field of view of the controlling avatar, the spatial data in the virtual environment is used as the query scope, and the real-time position and orientation of the controlling avatar is used as the query reference conditions to find the spatial position of the target entity in the virtual environment corresponding to the input command. This allows the NPC to be subsequently controlled to move to the vicinity of the target entity, or to perform the action indicated by the input command on the target entity.
[0112] However, since the main virtual character and / or NPC may have multiple environmental perception information, multiple environmental perception information can be integrated to assist in the determination process of the above-mentioned target entity, or multiple environmental perception information can be used according to priority to assist in the determination process of the above-mentioned target entity, and there is no limitation on this.
[0113] Stage 4: NPC voice feedback;
[0114] The NPC will provide voice feedback to the user. The types of voice feedback include: immediate feedback, command execution feedback, and dynamic feedback. Among them, immediate feedback means that after the user issues a natural language command to the NPC, the NPC immediately provides voice feedback on the natural language command, such as "received", "start execution", "OK", etc.; execution command feedback means that after the NPC executes the control command indicated by the intention of the natural language command, the NPC provides voice feedback on the execution result of the control command, such as "successfully executed", "failed to execute successfully", etc.; dynamic feedback means that when the user does not issue a natural language command to the NPC, the NPC provides voice feedback based on real-time environmental perception.
[0115] The voice feedback is generated by converting text content into speech. This text content is inferred using a large language model combined with the aforementioned spatial data and dynamic environmental perception information. Using a large language model to dynamically generate text content for voice feedback can overcome the monotony, mechanical, and repetitive nature of fixed, templated voice feedback. It also eliminates the need to pre-store numerous pre-created template voices, thus reducing the overall client data size.
[0116] Stage 5: NPC behavior control;
[0117] Taking the user's input as input, and the spatial data of the virtual scene and the runtime information corresponding to the master virtual character as reference, the AI model is used to determine the target entity in the virtual space and the control instructions for the NPC;
[0118] Control instructions are instructions that can be executed by an NPC's behavior tree. The subsequent actions a control instruction instructs the NPC to perform can be a single action or a sequence of actions. In the case of a multi-action sequence instruction, the multi-action sequence instruction is used to instruct the NPC to execute a sequence of actions. The data structure storing multi-action sequence instructions is in list form, caching the pending instructions in the sequence instruction. When the NPC's behavior tree completes the current instruction, the cache is called back and the pending instructions in the cache are sent for execution.
[0119] Hot word update function:
[0120] For phase 2, this solution also introduces a hot word system to improve the accuracy of speech recognition.
[0121] The hot word system is used to provide scene hot words related to the scene involved in the natural language instruction during the speech recognition process, so that when a natural language instruction input by voice is received, the hot word library can be used as the instruction analysis result based on the currently running virtual scene, so that the matching degree between the instruction analysis result and the virtual scene is higher.
[0122] The virtual scene includes at least one of a plurality of virtual scenes such as a virtual battle scene, a virtual transaction scene, a virtual office scene, and a virtual kitchen scene. Frequently used words in a plurality of virtual scenes, specific virtual elements existing in the virtual scene, etc. (such as virtual buildings and virtual characters that only exist in a certain type of scene) are pre-analyzed as scene hot words, and a scene hot word library consisting of a plurality of scene hot words is obtained, and the plurality of scene hot words correspond to virtual scenes respectively. In addition, different scene language models are pre-trained for different virtual scenes. For example, the scene language model in the battle scene focuses more on words related to battle, and the scene language model in the transaction scene focuses more on words related to transaction; the scene language model can perform more targeted analysis of natural language commands received in the virtual scene.
[0123] For example: when a user controls a non-player character, a natural language command is received. Based on the game state, object position, and current task completion status, the first virtual scene in which the target virtual character (at least one of the main virtual character and the non-player character) is currently located is determined to be a virtual battle scene, and a battle scene language model corresponding to the virtual battle scene is obtained. In addition, based on the field of view of the target virtual character, the first scene hot words in the virtual battle scene are obtained, including: "truck", "virtual grass", "virtual house", "enter", "open truck", "defense", "desert", "attack", "oak tree", "stable", "gun", etc. Under the constraints of the first scene hot words, the natural language command is decoded and processed by a pre-trained natural language analysis model (including acoustic network, language network, preset dictionary, etc.), and the instruction analysis result is output to accurately control the non-player character. The instruction analysis result is illustrated as follows.
[0124] (1) The command analysis result is "No. 1, you go to the front Defense The command analysis result can be used to specifically control No. 1 to move forward and take a defensive posture to prevent the enemy virtual character from attacking the main virtual character first; if there is no constraint of the hot words in the first scene, it is easy to recognize the natural language command as "No. 1, go to the front and make a sound", thereby affecting the player's virtual combat plan.
[0125] (2) The result of the command analysis is "No. 2, you go desert The command analysis result can be used to specifically control No. 2 to move to the desert to explore whether there are virtual elements such as enemy virtual characters or virtual treasure chests; if there is no constraint of the hot words in the first scene, it is easy to recognize the natural language command as "No. 2, you go Mountains "Explore it", which causes No. 2 to move to the wrong position, and the human-computer interaction efficiency is low.
[0126] (3) The command analysis result is "No. 1, No. 2, give me attack ", the command analysis result can be used to specifically control No. 1 and No. 2 to attack nearby attackable virtual characters; if there is no constraint of the hot words in the first scene, it is easy to recognize the natural language command as "No. 1 No. 2, give me supply ”, causing No. 1 and No. 2 to make incorrect behaviors that do not meet the players’ expectations, which can neither protect the main virtual character nor easily allow the enemy virtual character to win the virtual game.
[0127] (4) The command analysis result is “No. 1, go find a nearby oak tree”. This command analysis result can be used to specifically control No. 1 to find an oak tree in the nearby area so that he can complete the game task or find an oak tree that can avoid attacks. If there is no constraint of the hot words in the first scene, it is easy to recognize the natural language command as “No. 1, go find a nearby oak tree”. Project Book ", causing No. 1 to search nearby for virtual elements that do not meet the player's expectations and fail to meet the player's virtual combat needs.
[0128] (5) The command analysis result is "No. 1, No. 2, go over there stable OK", the command analysis result can be used to control No. 1 and No. 2 to search for the nearby stables and move to the location of the stables in a targeted manner; if there is no constraint of the hot words in the first scene, it is easy to recognize the natural language command as "No. 1 No. 2, go over there , right away OK", thus making No. 1 and No. 2 mistakenly believe that the main virtual character currently wants to complete the virtual game on its own, and therefore unable to provide better game assistance to the main virtual character.
[0129] (6) The command analysis result is "No. 2, pick up the one in front firearms ", the command analysis result can be used to specifically control No. 2 to search for firearms from the front and pick up the firearms; if there is no constraint of the hot words in the first scene, it is easy to recognize the natural language command as "No. 2, pick up the firearms in front Pickled Crab ", which makes it impossible for No. 2 to accurately find the virtual element that the main virtual character in front wants to pick up. It may prompt the player that "the crab cannot be found", or it may cause the player to pick up the wrong virtual element "crab", which will interfere with the player's virtual combat process.
[0130] Ambient sound effect function:
[0131] In addition, the solution also provides a spatial audio enhancement solution for the ambient sound effects of the entire virtual scene.
[0132] The audio played by the terminal includes ambient audio and NPC audio. Ambient audio is generated based on the characteristics of the virtual scene in which the main virtual object is currently located, giving the user an immersive experience. NPC audio is generated based on the characteristics of the NPC, allowing the user to intuitively experience the character's emotions and physical state through the sounds they hear.
[0133] Ambient Audio: This system identifies the elements of the virtual scene in which the master virtual object resides, generates appropriate elemental sound effects in real time based on these elements, or selects appropriate elemental sound effects from a sound effects library. These elemental sound effects are then synthesized to create the ambient audio. For example, if the virtual scene is a forest at night, the elements might include trees, owls, and insects, and the ambient audio might include the rustling of leaves, the calls of owls, and the chirping of insects.
[0134] NPC Audio: Identify the NPCs included in the virtual scene and generate corresponding character voices based on their character type, current emotion, and current behavior. For example, if the NPC is a middle-aged man running, the generated character voice will have a deep male voice tone and the sound of running breaths.
[0135] When playing ambient audio and / or NPC audio to the user, the terminal performs audio enhancement processing on the ambient audio and / or NPC audio to improve the realism of the audio.
[0136] Figure 2 The computer system 100 includes a first terminal 110 , a server 120 , and a second terminal 130 .
[0137] The first terminal 110 has a client 111 installed and running that supports a virtual environment. This client 111 can be a multiplayer online battle program. When the first terminal runs client 111, the user interface of client 111 is displayed on the screen of the first terminal 110. Client 111 can be any of a battle royale shooting game, a virtual reality (VR) application, an augmented reality (AR) program, a three-dimensional map program, a virtual reality game, an augmented reality game, a casual game, a party game, a first-person shooter (FPS), a third-person shooter (TPS), a multiplayer online tactical competitive game (MOBA), or a simulation game (SLG). In this embodiment, an FPS game is used as an example for explanation. First terminal 110 is used by first user 112. First user 112 uses first terminal 110 to control a first virtual object in a virtual environment to perform activities. The first virtual object may be referred to as the virtual object of first user 112. Activities of the first virtual object include, but are not limited to, at least one of: moving, jumping, teleporting, performing skills, using props, adjusting body posture, crawling, walking, running, riding, flying, jumping, driving, picking up, shooting, attacking, and throwing. Illustratively, the first virtual object is, for example, a simulated human character or an animated character.
[0138] The second terminal 130 has a client 131 installed and running that supports a virtual environment. The client 131 can be a multiplayer online battle program. When the second terminal 130 runs the client 131, the user interface of the client 131 is displayed on the screen of the second terminal 130. The client can be any of a battle royale shooting game, a VR application, an AR program, a three-dimensional map program, a virtual reality game, an augmented reality game, an FPS, a TPS, a MOBA, and a SLG. In this embodiment, the client is a MOBA game as an example. The second terminal 130 is a terminal used by a second user 132. The second user 132 uses the second terminal 130 to control a second virtual object in the virtual environment to perform activities. The second virtual object can be referred to as the virtual object of the second user 132. Schematically, the second virtual object is a second virtual object, such as a simulated human character or an anime character.
[0139] Optionally, the first virtual object and the second virtual object are in the same virtual environment. Optionally, the first virtual object and the second virtual object may belong to the same camp, the same team, the same organization, have a friendship relationship, or have temporary communication permissions. Optionally, the first virtual object and the second virtual object may belong to different camps, different teams, different organizations, or have a hostile relationship.
[0140] Optionally, the client installed on the first terminal 110 and the second terminal 130 is the same, or the client installed on the two terminals is the same type of client on different operating system platforms (Android or iOS). The first terminal 110 can generally refer to one of multiple terminals, and the second terminal 130 can generally refer to another of the multiple terminals. This embodiment only uses the first terminal 110 and the second terminal 130 as an example. The first terminal 110 and the second terminal 130 can be of the same or different device types, including at least one of a smartphone, a tablet computer, an e-book reader, an MP3 player, an MP4 player, a laptop computer, and a desktop computer.
[0141] Figure 2 Only two terminals are shown, but in different embodiments, multiple other terminals 140 can access the server 120. Optionally, one or more terminals 140 are terminals corresponding to developers. A development and editing platform for a client that supports a virtual environment is installed on the terminal 140. The developer can edit and update the client on the terminal 140 and transmit the updated client installation package to the server 120 via a wired or wireless network. The first terminal 110 and the second terminal 130 can download the client installation package from the server 120 to update the client.
[0142] The first terminal 110 , the second terminal 130 , and the other terminals 140 are connected to the server 120 via a wireless network or a wired network.
[0143] Server 120 includes at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 120 provides backend services for clients supporting a 3D virtual environment. Optionally, server 120 performs primary computing tasks, while the client performs secondary computing tasks; alternatively, server 120 performs secondary computing tasks, while the client performs primary computing tasks; alternatively, server 120 and the client utilize a distributed computing architecture for collaborative computing.
[0144] In an illustrative example, the server 120 includes a processor 122, a user account database 123, a battle service module 124, and a user-facing input / output interface (I / O interface) 125. The processor 122 is configured to load instructions stored in the server 120 and process data in the user account database 123 and the battle service module 124. The user account database 123 is configured to store user account data used by the first terminal 110, the second terminal 130, and other terminals 140, such as user account avatars, user account nicknames, user account combat power indexes, and the service areas where the user accounts are located. The battle service module 124 is configured to provide multiple battle rooms for users to engage in battles, such as 1v1 battles, 3v3 battles, 5v5 battles, etc. The user-facing I / O interface 125 is configured to establish communication and exchange data with the first terminal 110 and / or the second terminal 130 via a wireless network or a wired network.
[0145] The method provided in this application can be applied to, but is not limited to, at least one of the following scenarios: virtual reality applications, three-dimensional map programs, casual games, party games, first-person shooter games (FPS), third-person shooter games (TPS), multiplayer online battle arena games (MOBA), multiplayer gunfight survival games, etc. The following embodiments are illustrated by their application in games. The method provided in this application can be executed entirely by the client, such as in a stand-alone game; it can also be executed entirely by the server, such as in a cloud game, where the server performs all the calculations and the client is only responsible for displaying and collecting the user's operation behavior; the client and server can also collaborate, with the client responsible for part and the server responsible for part.
[0146] Figure 3A schematic diagram of a method for processing scene objects provided by an exemplary embodiment of the present application is shown.
[0147] Question 302 reads: "Below is a rendering of a 3D model named A, with dimensions B. Please use the model name, image, and dimensions to understand the 3D model, summarize its overall characteristics, and output the visual label in JSON format." The first half of question 302, separated by the semicolon ";," is the first subsection, which contains supplementary information about the scene object. This supplementary information includes the acquired attribute text of the scene object, such as the name of the scene object in the virtual scene being A and the dimensions of the scene object in the virtual scene being B. The second half, the second subsection, is the answer guidance statement for the visual question answering model.
[0148] The appearance image 304 of the scene object includes images of the scene object from nine perspectives. The question sentence 302 and the appearance image 304 of the scene object are input into the multimodal model 310 to obtain a visual label 315 for the scene object. For example, the visual label 315 for the scene object includes: "Type description: wooden double bed; material: [wood]; transparency: opaque; color: [brown]; surface feature: [carving]; shape: [rectangular]."
[0149] The visual labels 315 of the scene objects are input into a natural language model 320, which predicts matching labels 325 that match the spoken natural language expression. Natural language model 320 is implemented as a large language model (LLM). Natural language model 320 carries prior knowledge of spoken natural language expressions and, based on this prior knowledge, performs a pseudo-colloquial rewrite of the visual labels 315 of the scene objects to obtain rewritten labels (also called matching labels 325) that are more similar to the spoken natural language expression. The rewritten labels include: bed, brown, wood, rectangle.
[0150] The spatial position 327 of the scene object in the virtual scene is obtained. The spatial position 327 includes the coordinate position, orientation information, bounding box information, and cover point information of the scene object in the virtual scene; it is used to indicate the position, size, and other information of the scene object in the virtual scene. The matching label 325 and the spatial position 327 are used together as the label of the scene object in the virtual scene to facilitate searching for the scene object in the virtual scene. For example, Figure 3 The embodiment shown is Figure 1 Phase 1 pointed out in: A way to implement the preprocessing of spatial data.
[0151] Figure 4A flowchart of a method for processing scene objects provided by an exemplary embodiment of the present application is shown. The method can be executed by a computer device. The method includes:
[0152] Step 510: Obtain attribute text of scene objects in the virtual scene;
[0153] Exemplarily, attribute text is used to introduce the inherent attributes of scene objects in a virtual scene; on the one hand, attribute text realizes the description of scene objects in a textual mode, and while describing the scene objects, it provides semantic information for the prediction of the visual labels of the scene objects. On the other hand, attribute text is a description of scene objects in a virtual scene. In the face of the large number of scene objects in the virtual scene and the existence of reused object models, it can more accurately describe the inherent attributes of scene objects in the virtual scene. For example, in a virtual environment, in order to avoid redesigning the object model of the scene object, the object model of the virtual table is subjected to at least one of the transformation operations such as scaling, stretching, and rotation, and the transformed object model is deployed in the virtual environment as a virtual bench; in the face of the above-mentioned model reuse, attribute text can describe the inherent characteristics of scene objects in the virtual scene.
[0154] In one optional implementation, the attribute text includes at least one of the name and size of the scene object in the virtual scene. Exemplarily, the size of the scene object in the virtual scene indicates the virtual space occupied by the scene object in the virtual scene, thereby preventing the impact on the description of the scene object caused by changing the size of the object model when the object model is reused. The name in the virtual scene can reflect the purpose of the scene object in the virtual scene, thereby preventing the impact on the description of the scene object when deploying two different scene objects based on the same object model when the object model is reused.
[0155] Step 520: Acquire appearance images of scene objects in the virtual scene;
[0156] For example, appearance images are used to describe the style of scene objects. Based on the richness of natural language, players can describe scene objects in virtual scenes from multiple semantic perspectives. When constructing visual labels for scene objects, it is necessary to obtain comprehensive description information of the scene objects. Appearance images carry the color, texture, shape and other appearance styles of scene objects in the picture modality, as well as the relative positional relationships between the various sub-parts, and can comprehensively describe scene objects from the picture modality (or visual modality).
[0157] In one optional implementation, the appearance image of a scene object includes images obtained by observing the scene object from at least two perspectives. Observing the scene object from at least two perspectives in the appearance image can avoid occlusion caused by the scene object's spatial structure, which can prevent the object's entire appearance from being rendered in a single perspective.
[0158] It should be noted that the appearance image can be obtained by the scene object in the virtual scene, such as capturing the appearance image of the scene object in the scene object; the appearance image can also be obtained outside the virtual scene, for example, during the development and design process of the scene object, capturing the appearance image of the scene object in the development interface.
[0159] Step 530: Calling the multimodal model to perform prediction on the attribute text and appearance image of the scene object to obtain the visual label of the scene object;
[0160] Exemplarily, the multimodal model has the ability to perform model predictions on text and image information in different modalities. In this embodiment, the input parameters of the multimodal model are the attribute text and appearance image of the scene object; the multimodal model predicts the visual label of the scene object from both the text modality and the image modality. The visual label is used to describe the visual characteristics of the scene object in at least one dimension.
[0161] Exemplarily, the multimodal model includes an artificial neural network (ANN), and the network of the artificial neural network includes but is not limited to at least one of the following: Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Temporal Convolutional Networks (TCN), Long Short-Term Memory (LSTM), Multilayer Perceptron (MLP), Support Vector Machine (SVM). This application does not limit the network structure of the multimodal network. Exemplarily, the multimodal model has the ability to extract visual features of scene objects from input information of text modality and image modality to obtain visual labels of scene objects. Exemplarily, the multimodal model can be a classification model that assigns scene objects to preset type labels; it can also be a prediction model that predicts labels for scene objects that can describe the visual characteristics of scene objects.
[0162] To sum up, the method provided in this embodiment predicts the visual labels of scene objects based on the description information of scene objects in the text dimension and the image dimension by calling a multimodal model to perform predictions on attribute text and appearance images; in the process of predicting the visual labels of scene objects, the description information of scene objects in multiple dimensions is used to ensure that the characteristics of scene objects are fully described from the natural language semantics in the text dimension and the visual information in the image dimension, thereby ensuring the accuracy of the visual labels; calling a multimodal model to perform predictions on attribute text and appearance images can realize batch annotation of visual labels of scene objects, thereby improving the efficiency of obtaining visual labels of scene objects.
[0163] Next, the multimodal model is further introduced;
[0164] Figure 5 The flowchart of a method for processing scene objects provided by an exemplary embodiment of the present application is shown. The method can be executed by a computer device. Figure 4 In the illustrated embodiment, step 530 may be implemented as steps 532 and 534:
[0165] Step 532: Constructing a question statement for the appearance image;
[0166] Exemplarily, the question statement carries attribute text of the scene objects; on the one hand, the question statement is used to guide the visual question answering model to convert the appearance image into the visual label of the scene object; on the other hand, while guiding the conversion of the visual label of the scene object, the question statement provides supplementary information of the scene object in text form.
[0167] In an optional implementation, step 532 in this embodiment may be implemented as the following sub-steps:
[0168] Obtain the desired information about scene objects;
[0169] Exemplarily, the desired information is used to indicate desired description dimensions of scene objects in the visual tag and / or the desired format of the visual tag. In one example, the desired information is used to indicate that the desired description dimensions of scene objects in the visual tag include, but are not limited to, at least one of the following: type description, material, transparency, color, surface features, and shape of the scene object. In another example, the desired information is used to indicate that the desired format of the visual tag is at least one of the following: Comma Separated Values (CSV), JavaScript Object Notation (JSON), Extensible Markup Language (XML), and Excel.
[0170] Constructing a question statement about the appearance image based on the expected information and attribute text;
[0171] Exemplarily, the first subpart of the question statement is supplementary introduction information of the scene object, carrying the attribute text of the scene object; the second subpart of the question statement is an answer guidance statement for the visual question answering model, carrying expected information.
[0172] Figure 6 A schematic diagram of a visual label of a scene object provided by an exemplary embodiment of the present application is shown. In one example, the question statement 402 is: "Below is a rendering of a three-dimensional model named A, with a size of B; please understand what the three-dimensional model is through the model name, image, and size information, and summarize the overall features, and finally output the visual label in Json format.". The first half is the first sub-part, which is the supplementary introduction information of the scene object, divided by the semicolon ";" in the question statement 402. The supplementary introduction information of the scene object includes the attribute text of the scene object obtained from the data set 400, such as the name of the scene object in the virtual scene is A, and the size of the scene object in the virtual scene is B.
[0173] The second half is the second sub-section, which is the answer guidance statement for the visual question answering model. Furthermore, the question statement includes a preset statement template, and the attribute text of the scene object is filled into the preset statement template to obtain the question statement. The attribute text of the scene object is filled in the position where the word "A" and word "B" in the question statement are located; the name of the scene object in the attribute text in the virtual scene is marked as the asset name (asset name), and the size of the scene object in the attribute text in the virtual scene is recorded as the asset size (asset size); filling it into the preset statement template to obtain the question statement.
[0174] Step 534: Input the appearance image of the scene object and the question sentence into the visual question answering model to obtain the answer sentence, and use the answer sentence as the visual label of the scene object;
[0175] For example, the appearance image of the scene object and the question sentence are input into the visual question answering model, and the question sentence prompts the guidance information of the visual label generated by the visual question answering model in text form, and at the same time provides supplementary information of the scene object to further describe the scene object. Figure 6The appearance image 404 of the scene object includes images of the scene object from nine viewing angles. The question statement 402 and the appearance image 404 of the scene object are input into the visual question answering model 410 to obtain a visual label 415 for the scene object. For example, the visual label 415 of the scene object includes: "Type description: wooden double bed; material: [wood]; transparency: opaque; color: [brown]; surface feature: [carving]; shape: [rectangular]."
[0176] Furthermore, visual labels 415 for scene objects are used to supplement dataset 400. After supplementation with visual labels 415, dataset 400 includes appearance images 404 of scene objects, visual labels 415, and spatial locations 420 of scene objects in the virtual scene. Spatial locations 420 include coordinate positions, orientation information, bounding box information, and cover point information, which will be described in a separate embodiment below. In various embodiments of the present application, dataset 400 is also referred to as a spatial information library for the virtual scene.
[0177] In this embodiment, the visual question answering model (VQA) can be implemented as at least one of the following: Large Language Model with Vision Ability (LLAVA), Mini Generative Pre-trained Transformer 4 (MiniGPT4), Cognitive Visual Language Model (CogVLM), Generative Pre-trained Transformer 4 with Vision (GPT4V), and Gemini model.
[0178] In an optional implementation, step 534 in this embodiment may be implemented as the following sub-steps:
[0179] Inputting an appearance image of a scene object and the first of at least two sub-sentences into a visual question answering model to obtain a first answer sub-sentence;
[0180] Exemplarily, the visual question answering model takes an appearance image and a first statement as input parameters, and the resulting first answer sub-statement is an answer statement guided by the first statement. Exemplarily, at least two sub-statements are used to inquire about the visual features of the appearance image from multiple dimensions; the first answer sub-statement is used to answer the visual features of the appearance image inquired about in one dimension corresponding to the first statement.
[0181] Repeat the above steps until at least two answer sub-sentences corresponding to at least two sub-sentences are obtained;
[0182] Each sub-sentence of the at least two sub-sentences and the appearance image are input into the visual question answering model to obtain a corresponding answer sub-sentence, that is, at least two answer sub-sentences corresponding to the at least two sub-sentences are obtained.
[0183] Perform sentence aggregation on at least two answer sub-sentences to extract visual labels of scene objects;
[0184] Exemplarily, sentence aggregation is performed on the at least two answer sub-sentences, extracting part or all of the content from the at least two answer sub-sentences to obtain visual labels for scene objects. Exemplarily, sentence aggregation can be performed on the input answer sub-sentences based on an artificial neural network, or can be performed by extracting words at predetermined positions in the answer sub-sentences to achieve sentence aggregation.
[0185] In summary, the method provided in this embodiment predicts the attribute text and appearance image by calling the visual question answering model to perform predictions on the attribute text and appearance image, thereby realizing the prediction of the visual label of the scene object based on the descriptive information of the scene object in the text dimension and the image dimension; the question statement for the appearance image includes the attribute text, which guides the visual question answering model to generate the answer statement while providing supplementary information for the scene object; in the process of predicting the visual label of the scene object, the descriptive information of the scene object in multiple dimensions is used, which can ensure that the characteristics of the scene object are fully described from the natural language semantics in the text dimension and the visual information in the image dimension, thereby ensuring the accuracy of the visual label and improving the efficiency of obtaining the visual label of the scene object.
[0186] Figure 7 The flowchart of a method for processing scene objects provided by an exemplary embodiment of the present application is shown. The method can be executed by a computer device. Figure 4 In the illustrated embodiment, step 530 may be implemented as steps 536 and 538:
[0187] Step 536: Input the appearance image of the scene object into the image description model to predict the description text of the scene object;
[0188] Exemplarily, the image captioning model has the ability to describe the appearance image of the input model in text form. By calling the image captioning model, the appearance image of the graphical modality is converted into a description text in the text modality. Exemplarily, the image captioning model includes but is not limited to at least one of the following: the image captioning model (Show and Tell), the display attention and narration model (Show, Attend and Tell, SAAT), the bootstrap latent integer picture model (BLIP), and the transformer-based image captioning model (TICM).
[0189] Step 538: performing sentence aggregation on the description text and attribute text to extract visual labels of scene objects;
[0190] Exemplarily, part or all of the content in the description text and attribute text is determined as the visual label of the scene object obtained in the image area; exemplary, sentence aggregation is performed on the description text and attribute text, which can be based on an artificial neural network to perform sentence aggregation on the input description text and attribute text; or words can be intercepted at preset positions in the description text and attribute text to achieve sentence aggregation.
[0191] To sum up, the method provided in this embodiment, by calling the image description model to perform prediction on the appearance image, realizes the description text of the scene objects in the text dimension and the image dimension, and converts it into the description text in the text dimension; further aggregates the description text and attribute text predicted by the image description model, and obtains the description information from multiple dimensions of the description text from the image dimension and the attribute text from the text dimension; it can ensure that the characteristics of the scene objects are fully described from the natural language semantics in the text dimension and the visual information in the image dimension, thereby ensuring the accuracy of the visual labels and improving the efficiency of obtaining the visual labels of the scene objects.
[0192] In an optional implementation, a pseudo-colloquial rewriting is performed on the visual label. The pseudo-colloquial rewriting is described as follows:
[0193] Figure 8 The flowchart of a method for processing scene objects provided by an exemplary embodiment of the present application is shown. The method can be executed by a computer device. Figure 4 Based on the embodiment shown, the process further includes steps 540 and 545:
[0194] Step 540: Performing pseudo-colloquial rewriting on the visual labels of the scene objects to obtain matching labels that conform to the spoken expressions of natural language;
[0195] Exemplarily, as introduced above, visual labels are used to describe the visual features of scene objects in at least one dimension, and the appearance images of scene objects present rich visual features of scene objects. However, in the spoken expression of natural language, the description of scene objects cannot cover the visual features of scene objects in all dimensions. For example, when describing a virtual bed in a virtual scene in spoken expression, people usually pay attention to the color, material, placement, etc. of the virtual bed, but usually neglect to pay attention to the craftsmanship used for the headboard and side panels of the bed, such as carving, embossing, etc. The purpose of performing pseudo-spoken rewriting on the visual labels of scene objects is to obtain matching labels that are closer to spoken expressions; it can be understood that performing pseudo-spoken rewriting can be to delete part of the content in the visual label, or to change the visual label to a label with the same semantics but different textual expression.
[0196] In an optional implementation, the pseudo-colloquial rewriting is performed by calling a natural language model; accordingly, step 540 in this embodiment can be implemented as follows:
[0197] Input the visual labels of scene objects into the natural language model to predict matching labels that match the spoken expressions in natural language;
[0198] In this implementation, the pseudo-colloquial rewriting of visual labels is implemented by invoking a natural language model. For example, the natural language model carries prior knowledge of spoken natural language expressions. For example, the natural language model is implemented as a large language model (LLM).
[0199] Furthermore, executing the call to the natural language model to perform the pseudo-spoken rewriting includes:
[0200] Obtain a first sample label pair;
[0201] Constructing a rewritten guide sentence based on the first sample label pair and the visual label;
[0202] Input the rewritten guide sentence into the natural language model to predict matching labels that conform to the spoken expression of natural language;
[0203] In this example, the first sample label pair includes a first label before pseudo-colloquial rewriting and a second label obtained after pseudo-colloquial rewriting. Taking a virtual bed in a virtual scene as an example, the first label in the first sample label pair is: "Type description: wooden double bed; material: [wood]; transparency: opaque; color: [brown]; surface feature: [carved]; shape: [rectangular]." The second label is: "bed, brown, wooden, rectangular."
[0204] The rewriting guide sentence has the natural semantics of rewriting the visual label with reference to the first sample label pair. In one example, the visual label that needs to be rewritten in a pseudo-colloquial manner is: "Type description: green plant pot; material: [plant]; transparency: opaque; color: [brown, green]; surface feature: [smooth]; shape: [irregular shape]". Accordingly, based on the first sample label pair and the visual label, the resulting guide sentence is: Example: Rewrite "Type description: wooden double bed; material: [wood]; transparency: opaque; color: [brown]; surface feature: [carved]; shape: [rectangular]" to "bed, brown, wooden, rectangular". Based on the above example, "Type description: green plant pot; material: [plant]; transparency: opaque; color: [brown, green]; surface feature: [smooth]; shape: [irregular shape]" is rewritten.
[0205] Step 545: obtaining the spatial position of the scene object in the virtual scene, and determining the spatial position of the scene object as auxiliary information of the visual label of the scene object;
[0206] Exemplarily, the spatial position of the scene object in the virtual scene is used to indicate the deployment of the scene object in the virtual scene, and the spatial information is used to indicate the position, size, and other information of the scene object in the virtual scene after the scene object is deployed in the virtual scene.
[0207] In an optional implementation, step 545 in this embodiment may be implemented as follows:
[0208] Obtaining at least one of the coordinate position, orientation information, bounding box information, and cover point information of a scene object in the virtual scene;
[0209] Among them, the coordinate position (Location) is used to indicate the position of the scene object in the virtual scene, such as the coordinate information of the center point or preset point of the scene object in the virtual scene; the orientation (Rotation) is used to indicate the direction facing the scene object in the virtual scene, such as the direction facing the front of the scene object in the virtual scene, and the bounding box (BoundingBox) is used to indicate the size of the scene object in the virtual scene; the cover points (Cover points) are used to indicate the recommended virtual character position points when the virtual character approaches the scene object, so that the scene object can mask the virtual character.
[0210] Figure 9 The figure shows a schematic diagram of scene object information provided by an exemplary embodiment of the present application. Various types of scene object information are stored in a table, such as Figure 9As shown, the first column of the table records the asset name of the scene object, indicating the name of the scene object's texture resource and skeleton model resource. The second column records the entity name of the scene object, which is used to indicate the name of the scene object in the virtual scene. The third column is the image storage path of the scene object's appearance image. The fourth column is the spatial location of the appearance image, such as coordinates and orientation, also known as tactical information. The fifth column is the asset scale ratio of the scene object, which is used to indicate the size of the asset. The sixth column is the image name of the scene object's appearance image. The seventh column is the appearance image of the scene object.
[0211] In summary, the method provided in this embodiment predicts attribute text and appearance images by calling a multimodal model, thereby realizing the prediction of visual labels of scene objects based on the description information of scene objects in the text dimension and the image dimension; in the process of predicting the visual labels of scene objects, the description information of scene objects in multiple dimensions is used, which can ensure that the characteristics of scene objects are fully described from the natural language semantics in the text dimension and the visual information in the image dimension, thereby ensuring the accuracy of the visual labels. Further, a pseudo-colloquial rewriting is performed on the visual labels to obtain matching labels that are closer to spoken expressions, thereby improving the accuracy of matching labels to describe scene objects in spoken expression scenes. Further, the spatial position of the scene objects in the virtual scene is obtained, and the position, size, etc. of the scene in the virtual scene are described, providing more basis for screening scene objects in the virtual scene.
[0212] Next, determining the first scene object in the virtual scene is introduced. Figure 10 A flowchart of a method for processing scene objects provided by an exemplary embodiment of the present application is shown. The method can be executed by a computer device. Figure 4 Based on the embodiment shown, steps 552 to 558 and steps 562 to 568 are also included:
[0213] Step 552: constructing a spatial information library of the virtual scene based on the visual labels of the multiple scene objects in the virtual scene;
[0214] For example, please refer to the above embodiments for the introduction of visual tags, which will not be repeated here. The spatial information library includes visual tags of multiple scene objects in the virtual scene. Figure 8 In the paper, a method of rewriting visual labels into a pseudo-colloquial language is introduced. Figure 8 When combined with the present embodiment to form a new embodiment, the spatial information database is constructed based on the rewritten labels.
[0215] Step 554: Obtain natural language commands;
[0216] For example, a natural language command is a command used to direct a non-player character in a virtual scene. A user may input a natural language command in text or voice. The natural language command typically includes multiple words to indicate a virtual action that the non-player character needs to be directed to perform in the virtual scene.
[0217] Step 556: Based on the similarity between the natural language command and the visual tags in the spatial information library, a first scene object corresponding to the natural language command is selected from a plurality of scene objects in the virtual scene;
[0218] For example, based on the similarity between some or all words in the natural language command and the visual tags in the spatial information library, a first scene object corresponding to the natural language command is selected from multiple scene objects in the virtual scene. As described above, natural language commands are used to direct non-player characters in the virtual scene; natural language commands not only include language describing the virtual scene object to which the virtual activity is directed (for example, if a natural language command instructs to move to a virtual building, the virtual activity of moving is performed on the virtual building), but also include words indicating the identity of the non-player character performing the virtual activity.
[0219] In an optional implementation, the natural language command includes multiple input words; accordingly, step 556 can be implemented as follows:
[0220] Determining a similarity score between each of the plurality of input words and the visual label of the first scene object;
[0221] refer to Figure 11 , the multiple input words included in the natural language command 600 include "blue" and "car"; the virtual scene includes two scene objects; it can be understood that in different embodiments, the virtual scene may include more scene objects. Figure 11 In the example, scene object A 601 has visual tags including metal, old, car, truck, blue, and damaged; scene object B 602 has visual tags including metal, scratched, car, damaged, armored truck, and black. Similarity scores are calculated for each of the input words included in natural language command 600 and the visual tags of the scene objects.
[0222] Determining the sum of the multiple similarity scores as the similarity between the first scene object and the natural language command;
[0223] Exemplarily, multiple similarity scores correspond to multiple input words one by one; each input word has a similarity score with the visual label of the first scene object, and the similarity score is the maximum value of the cosine similarity between the input word and each visual label of the first scene object. Exemplarily, the cosine similarity can be implemented as the cosine similarity between the input word and the visual label in the natural semantic latent space. Figure 11 , multiple similarity scores correspond one-to-one to multiple input words. For the input word "blue" in natural language command 600, the similarity score with scene object A 601 is 1.0. For the input word "car" in natural language command 600, the similarity score with scene object A 601 is also 1.0. The similarity between scene object A 601 and natural language command 600 is 1.0 + 1.0 = 2.0. Similarly, the similarity between scene object B 602 and natural language command 600 is calculated to be 1.0 + 0.91 = 1.91.
[0224] If the similarity between the first scene object and the natural language command exceeds the similarity between other scene objects in the spatial information library and the natural language command, determining the first scene object as the object corresponding to the natural language command;
[0225] Exemplarily, the first scene object is the scene object in the virtual scene that has the greatest similarity to the natural language command. Figure 11 In the example, the similarity between scene object A 601 and the natural language command 600 is greater than the similarity between scene object B 602 and the natural language command 600, and scene object A 601 is determined as the first scene object.
[0226] Step 558: Executing a virtual activity on the first scene object based on the natural language command;
[0227] Exemplarily, the virtual activity indicated by the natural language command is performed on the first scene object, such as launching a virtual attack on the first scene object, moving to the location of the first scene object, turning on or off the first scene object, picking up the first scene object, etc.
[0228] Step 562: constructing a spatial information library of the virtual scene based on the appearance images of multiple scene objects in the virtual scene;
[0229] For example, for the introduction of the appearance image, please refer to the various embodiments above, which will not be repeated here. The spatial information library includes appearance images of multiple scene objects in the virtual scene.
[0230] Figure 12A schematic diagram illustrating the appearance of a scene object provided by an exemplary embodiment of the present application is shown. First, object traversal 452 is performed to obtain each scene object (Actor) 454 in the virtual scene. Assets referenced by each scene object are identified and extracted 456. These assets include, but are not limited to, 3D static meshes and texture resources. For each identified scene object, multiple perspectives 458 are traversed around the identified asset; for example, nine perspectives are switched to observe the scene object from each of these perspectives.
[0231] For example, by using a multi-view scene capture tool, scene objects can be observed from nine perspectives, resulting in a nine-square grid of appearance images of the scene objects, converting the 3D model structure of the scene objects into 2D image information. For example, by traversing the assets referenced by each scene object in the virtual scene, the 3D model structure of each asset is captured from multiple angles and ultimately stored as 2D image information. This allows for convenient acquisition of appearance images of scene objects, thereby constructing a game asset dataset.
[0232] The segmentation channels 460 are traversed, and information is captured in different channels 462. The segmentation channels are first traversed to capture information in each segmentation channel of the same viewpoint, and then the viewpoints are traversed to obtain information from different viewpoints.
[0233] The multiple segmentation channels are described as follows: First, a virtual camera model is used to capture the asset. The viewing angle is selected based on a preset set of angles to ensure comprehensive coverage of the asset from all directions. Second, the asset's positional transformation information within the virtual scene is recorded, including its position and rotation, to ensure accurate restoration of its position and orientation within the scene. Third, the asset's scale information is recorded for analysis and processing to determine the actual size and proportion of scene objects. Third, to describe the spatial occupancy of scene objects within the virtual scene, the bounding box size of the scene objects is calculated and recorded.
[0234] Step 564: Obtain natural language commands;
[0235] For example, a natural language command is a command used to direct a non-player character in a virtual scene. A user may input a natural language command in text or voice. The natural language command typically includes multiple words to indicate a virtual action that the non-player character needs to be directed to perform in the virtual scene.
[0236] Step 566: Calling a similarity matching network to predict the similarity between the natural language command and the appearance images of the multiple scene objects in the spatial information library, and obtaining a second scene object with the highest similarity to the natural language command among the multiple scene objects;
[0237] For example, the similarity matching network has the ability to predict the degree of similarity between image information and text information. As described above, natural language commands can be input in text or voice. When natural language commands are input in voice, automatic speech recognition (ASR) is used to convert the voice information into text format.
[0238] Exemplarily, the similarity matching network can be implemented as a contrastive language-image pretraining model (CLIP). The network parameters of CLIP carry rich visual information. Based on the appearance images of multiple scene objects in the spatial information library and the natural language information in the natural language commands, the similarity relationship between image information and text information is predicted.
[0239] Step 568: Executing a virtual activity for the second scene object based on the natural language command;
[0240] Exemplarily, the virtual activity indicated by the natural language command is performed on the second scene object, such as launching a virtual attack on the second scene object, moving to the location of the second scene object, opening or closing the second scene object, picking up the second scene object, etc.
[0241] It should be noted that the second scene object in this step and the first scene object in step 558 may be the same or different, and there is no limitation on this.
[0242] It should be noted that steps 552 to 558 in this embodiment can be Figure 4 The steps in the embodiment are combined into a new embodiment and implemented separately; Steps 562 to 568 in this embodiment can be combined with Figure 4 The various steps in the embodiment are combined into new embodiments and implemented separately; this application does not limit this.
[0243] In an optional example, the spatial information library of the virtual scene is based on appearance images and visual labels of multiple scene objects in the virtual scene; and also includes identity identifiers of the scene objects to index the scene objects. Figure 13A schematic diagram of a spatial information library for a virtual scene provided by an exemplary embodiment of the present application is shown. For example, a virtual scene includes five objects. The first object 621 is identified by SM_xxx01. The spatial information library includes appearance images of the first object 621 observed from nine perspectives. The visual labels for the first object 621 are: Type Description: Small Building with a Loft; Material: Concrete, Wood, Glass; Transparency: Opaque; Color: Brown, Gray; Surface Features: Old, Damaged; Shape: Rectangle, Triangle. Figure 13 Also shown are the identification marks, appearance images, and visual labels of the second object 622 to the fifth object 625 .
[0244] First, execute the similarity based on the natural language command and the visual tags in the spatial information library shown in step 556, and screen out the first scene object corresponding to the natural language command from the multiple scene objects in the virtual scene; if there are multiple scene objects with similar similarities in step 556, and they are the multiple scene objects with the highest similarity in the virtual scene (there are multiple similar scene objects that cannot be distinguished); or if the similarity of the first scene object is lower than the preset confidence threshold (the confidence of the scene object determined according to the similarity of the visual tags is insufficient), execute the calling of the similarity matching network prediction shown in step 566.
[0245] In summary, the method provided in this embodiment predicts the visual labels of scene objects based on the descriptive information of scene objects in the text and image dimensions by invoking a multimodal model to perform predictions on attribute text and appearance images. This ensures the accuracy of the visual labels and improves the efficiency of obtaining visual labels for scene objects. Two similarity evaluation methods are provided: the similarity between natural language commands and visual labels, and the similarity between natural language commands and appearance images. By identifying a first scene object in a virtual scene, it is possible to determine whether the natural language command is intended to instruct the execution of a virtual activity directed at the first scene object.
[0246] An exemplary embodiment of the present application provides a method for controlling a virtual character. The method includes:
[0247] Displaying at least one of a control avatar and an NPC in a virtual environment;
[0248] For example, the controlling avatar is the avatar that the user directly controls in the virtual environment. The virtual environment also contains one or more non-player characters (NPCs). NPCs and the controlling avatar belong to the same virtual faction and are teammates of the controlling avatar, following the controlling avatar in the virtual environment. NPCs can also be followers or pets controlled by the controlling avatar. Furthermore, NPCs can be neutral characters, acting in concert with the controlling avatar only under certain conditions.
[0249] Obtain natural language commands;
[0250] Exemplarily, natural language commands include natural semantics for commanding NPCs; commanding NPCs in virtual scenes through natural language commands, natural language commands can be directly input by users or extracted from voice information input by users. This application does not limit the method of obtaining natural language commands.
[0251] A natural language command includes at least one of a behavioral intent and a scene entity. The behavioral intent is used to indicate the type of virtual activity to be performed by the NPC, such as indicating which type of virtual activity to be performed in the virtual environment. The scene entity includes a first scene object, such as indicating which scene object in the virtual environment the virtual activity is performed on. The behavioral intent and scene entity included in the natural language command are partial words in the natural language command.
[0252] In one example, the natural language command controls the NPC in terms of what type of virtual activity to perform and for which scene objects in the virtual environment the virtual activity is performed; it is a complex instruction for controlling the NPC.
[0253] Controlling NPCs to perform virtual activities in response to natural language commands;
[0254] Exemplarily, controlling an NPC to perform a virtual activity is determined based on environmental perception information from the controlling virtual character and / or NPC. In some examples, the controlling virtual character and / or NPC obtains environmental perception information within the perceptual range of visual, auditory, and motion trajectory perception. Exemplarily, the virtual activity closely matches the controlling virtual character and / or NPC's current perception of the virtual environment, and the NPC is controlled based on the virtual character's and / or NPC's current perception of the virtual environment. Exemplarily, the NPC has autonomous behavioral capabilities, and the user merely directs the NPC. The user gives commands, and the NPC understands and executes the commands based on its own autonomous behavioral capabilities.
[0255] To sum up, the method provided in this embodiment determines the virtual activities that need to be performed by referring to the environmental perception information of the master virtual character and / or NPC in response to complex instructions for controlling the NPC in terms of what type of virtual activities to perform and for what scene objects in the virtual environment to perform the virtual activities; and uses the environmental perception information of the master virtual character and / or NPC as a reference to ensure that the NPC and the master virtual character perform virtual activities in collaboration.
[0256] Figure 14 A flowchart of a method for controlling a virtual character provided by an exemplary embodiment of the present application is shown. The method includes:
[0257] Step 201: Acquire a spatial dataset;
[0258] The spatial dataset of the virtual environment includes visual labels of scene objects, which are used to describe the visual characteristics of scene objects in at least one dimension, such as describing scene objects in terms of material, transparency, color, shape, etc.
[0259] Step 202: Display at least one of the main virtual character and the NPC in the virtual environment;
[0260] For example, the main virtual character is a virtual character that the user directly controls in the virtual environment. The virtual environment also contains one or more NPCs; the NPCs and the main virtual character belong to the same virtual camp and are teammates of the main virtual character.
[0261] Step 203: Acquire a natural language command in voice form;
[0262] Natural language commands include natural semantics for directing NPCs, and command information includes behavioral intentions and scene entities. Behavioral intentions are used to indicate the type of virtual activity to be performed by the NPC, and scene entities are used to indicate the target entity of the virtual activity.
[0263] Step 204: converting the natural language command in voice form into a natural language command in text form;
[0264] For example, Automatic Speech Recognition (ASR) processing is performed on the natural language command in the form of speech to determine the natural language command in the form of text. Speech-to-text processing generally involves invoking components such as acoustic models and language models to recognize the pronunciation, vocabulary, and grammatical structure of the natural language command in the form of speech and convert it into the natural language command in the form of text.
[0265] Step 205: Perform intent recognition on the natural language command to obtain a first classification label;
[0266] Exemplarily, the first intention indication corresponding to the first classification label includes a command intention directed at a non-player character on an intention dimension; for example, at least one of the following: indicating which of the multiple non-player characters the non-player character is commanded by the command text, indicating whether the virtual activity performed by the non-player character is related to launching a virtual attack, and indicating what kind of virtual activity the non-player character performs.
[0267] Step 206: Perform entity recognition on the natural language command to obtain a target entity;
[0268] The target entity is determined from the virtual environment using environmental perception information, which includes information perceived by at least one of the controlling avatar and non-player characters. The target entity can be the entity currently visible to the player or an entity the player has seen previously.
[0269] Step 207: In response to the natural language command, controlling the NPC to perform a virtual activity according to the first intention corresponding to the first classification label, or controlling the NPC to perform a virtual activity associated with the target entity, or controlling the NPC to perform a virtual activity associated with the target entity according to the first intention corresponding to the first classification label;
[0270] Exemplarily, the NPC is controlled to perform a virtual activity according to the first classification tag and / or the instruction of the target entity.
[0271] Step 208: reporting the NPC's feedback information;
[0272] For example, the game application can generate corresponding feedback information in real time based on the environmental perception information of the non-player character and broadcast it. For example, when the environmental perception information of the non-player character triggers a broadcast condition, the game application can generate corresponding feedback information in real time based on the environmental perception information that triggered the broadcast condition; or when the non-player character or the master virtual character triggers a broadcast condition, the game application can generate corresponding feedback information in real time based on the environmental perception information of the non-player character.
[0273] The preprocessing stage of step 201 can be implemented as follows:
[0274] Sub-step 1: Obtain attribute text of scene objects in the virtual scene;
[0275] For example, attribute text is used to describe the inherent properties of scene objects in a virtual scene. On the one hand, attribute text describes scene objects in a textual manner, providing semantic information for predicting their visual labels. On the other hand, attribute text describes scene objects in a virtual scene. Given the large number of objects in a virtual scene and the presence of reused object models, it can more accurately describe the inherent properties of scene objects in the virtual scene.
[0276] In an optional implementation, the attribute text includes at least one of a name of the scene object in the virtual scene and a size of the scene object in the virtual scene.
[0277] Sub-step 2: obtaining an appearance image of a scene object in the virtual scene;
[0278] Exemplarily, an appearance image is used to describe the style of a scene object. An appearance image, in a picture modality, carries the scene object's appearance style, such as color, texture, and shape, as well as the relative positional relationships between its sub-components, enabling a comprehensive description of the scene object from a picture modality (or visual modality). In one optional implementation, the appearance image of a scene object includes images obtained by observing the scene object from at least two perspectives.
[0279] Sub-step 3: Call the multimodal model to perform prediction on the attribute text and appearance image of the scene object to obtain the visual label of the scene object;
[0280] Exemplarily, the multimodal model has the ability to perform model predictions on text and image information in different modalities. In this embodiment, the input parameters of the multimodal model are the attribute text and appearance image of the scene object; the multimodal model predicts the visual label of the scene object from both the text modality and the image modality. The visual label is used to describe the visual characteristics of the scene object in at least one dimension.
[0281] Optionally, the multimodal model includes a visual question answering model; the question statement carries attribute text of the scene object; on the one hand, the question statement is used to guide the visual question answering model to convert the appearance image into the visual label of the scene object; on the other hand, while guiding the conversion of the visual label of the scene object, the question statement provides supplementary information of the scene object in text form.
[0282] Get the expected information of scene objects;
[0283] Exemplarily, the desired information is used to indicate desired description dimensions of scene objects in the visual tag and / or the desired format of the visual tag. In one example, the desired information is used to indicate that the desired description dimensions of scene objects in the visual tag include, but are not limited to, at least one of the following: type description, material, transparency, color, surface features, and shape of the scene object. In another example, the desired information is used to indicate that the desired format of the visual tag is at least one of the following: Comma Separated Values (CSV), JavaScript Object Notation (JSON), or Extensible Markup Language (XML).
[0284] Construct a question statement of the appearance image based on the expected information and attribute text;
[0285] Exemplarily, the first subpart of the question statement is supplementary introduction information of the scene object, carrying the attribute text of the scene object; the second subpart of the question statement is an answer guidance statement for the visual question answering model, carrying expected information.
[0286] Optionally, the spatial dataset includes matching labels of scene objects in the virtual scene; accordingly: performing pseudo-colloquial rewriting on the visual labels of the scene objects to obtain matching labels that conform to spoken expressions of natural language.
[0287] For example, as described above, visual labels are used to describe the visual features of scene objects in at least one dimension, and the appearance image of the scene object presents rich visual features of the scene object. However, in the spoken expression of natural language, the description of the scene object cannot cover the visual features of the scene object in all dimensions. The purpose of performing a pseudo-colloquial rewriting on the visual label of the scene object is to obtain a matching label that is closer to the spoken expression; it can be understood that performing a pseudo-colloquial rewriting can be to delete part of the content in the visual label, or to change the visual label to a label with the same semantics but different textual expression.
[0288] In one optional implementation, the pseudo-colloquial rewriting is performed by invoking a natural language model. The visual labels of scene objects are input into the natural language model, and matching labels that match the spoken expressions of natural language are predicted. The pseudo-colloquial rewriting of the visual labels is performed based on invoking the natural language model. Exemplarily, the natural language model carries prior knowledge of spoken natural language expressions. For example, the natural language model is implemented as a Large Language Model (LLM).
[0289] Optionally, the spatial dataset also includes spatial information of scene objects, and accordingly, also includes:
[0290] The spatial position of the scene object in the virtual scene is obtained, and the spatial position of the scene object is determined as auxiliary information of the visual label of the scene object.
[0291] Exemplarily, the spatial position of the scene object in the virtual scene is used to indicate the deployment of the scene object in the virtual scene, and the spatial information is used to indicate the position, size, and other information of the scene object in the virtual scene after the scene object is deployed in the virtual scene.
[0292] Further, obtaining at least one of the coordinate position, orientation information, bounding box information, and cover point information of the scene object in the virtual scene;
[0293] Among them, the coordinate position (Location) is used to indicate the position of the scene object in the virtual scene, such as the coordinate information of the center point or preset point of the scene object in the virtual scene; the orientation (Rotation) is used to indicate the direction facing the scene object in the virtual scene, such as the direction facing the front of the scene object in the virtual scene; the bounding box (BoundingBox) is used to indicate the size of the scene object in the virtual scene; the cover points (Cover points) are used to indicate the recommended virtual character position points when the virtual character approaches the scene object, so that the scene object can mask the virtual character.
[0294] The intention recognition phase of step 205 can be implemented as follows:
[0295] Sub-step 4: Call at least one hierarchical prediction network to perform intent recognition on the natural language command to obtain a first classification label of the natural language command intent to command the NPC;
[0296] Exemplarily, the hierarchical prediction network has the ability to predict the classification label corresponding to the command text. Exemplarily, the hierarchical prediction network includes at least two layers of sub-networks, with the upper and lower layers being cascaded, and the lower layer further performing classification label prediction based on the prediction results output by the upper layer. At least two layers of sub-networks are constructed based on a tree structure of multiple classification labels. The hierarchical prediction network constructs at least two layers of sub-networks corresponding to the tree structure of the classification labels, splitting the prediction task of a wide variety of classification labels into prediction sub-tasks performed by at least two layers of sub-networks, thereby reducing the complexity of classification prediction in each layer of sub-network.
[0297] In an optional implementation, each hierarchical prediction network in at least one hierarchical prediction network is used to predict the command intention of the command text for a non-player character on an intention dimension; the intention dimension includes at least one of the subject dimension, the semantic dimension, and the behavior dimension. Exemplarily, in at least one hierarchical prediction network, the i-th layer sub-network in each hierarchical prediction network is used to predict the first-level behavior label of the command text on an intention dimension, and the i+1-th layer sub-network in the hierarchical prediction network is used to predict the second-level behavior label corresponding to the command text in the first-level behavior label, where i is a positive integer; the high-level behavior label (such as the first-level behavior label) predicted by the high-level sub-network (such as the i-th layer sub-network) in the hierarchical prediction network includes multiple sub-labels, or subordinate lower-level labels; the corresponding low-level sub-network (such as the i+1-th layer sub-network) needs to be called to further perform classification prediction (such as predicting the second-level behavior label). The prediction task of a wide variety of classification labels is split into prediction sub-tasks performed by at least two layers of sub-networks to reduce the complexity of classification prediction of each layer of sub-network.
[0298] Furthermore, the intent dimension includes a subject dimension, and the hierarchical prediction network includes a hierarchical subject prediction network, which has the ability to predict the subject type in a natural language command; the subject type is used to indicate the identity of the NPC commanded by the natural language command; exemplarily, the subject prediction network is used to predict which of the many non-player characters the non-player character commanded by the command text is, and the number of non-player characters commanded by the command text can be one or more.
[0299] And / or, the intent dimension includes a semantic dimension, the hierarchical prediction network includes a hierarchical semantic prediction network, and the semantic prediction network has the ability to predict the semantic type in natural language commands; the semantic type is used to indicate the NPC's control method for initiating a virtual attack; exemplarily, the subject prediction network is used to predict whether the virtual activity performed by the non-player character commanded by the command text is related to initiating a virtual attack.
[0300] Alternatively, the intent dimension includes a behavior dimension, the hierarchical prediction network includes a hierarchical behavior prediction network, and the behavior prediction network is capable of predicting the behavior intent type in a natural language command; the behavior intent type is used to indicate how the NPC will behave in performing a virtual activity. Exemplarily, the behavior prediction network is used to predict what virtual activity the non-player character will perform.
[0301] Accordingly, in step 207 , in response to the natural language command, the NPC is controlled to perform a virtual activity according to the first intention corresponding to the first classification label;
[0302] For example, the first intent corresponding to the first category tag indicates a command intent directed at a non-player character on an intent dimension, such as at least one of the following: indicating which non-player character is being commanded by the command text among a plurality of non-player characters, indicating whether the virtual action performed by the non-player character is related to launching a virtual attack, and indicating what virtual action the non-player character is to perform. According to the instruction of the first category tag, the NPC is controlled to perform the virtual action.
[0303] The target entity identification process of step 206 can be implemented as follows:
[0304] Sub-step 5: querying and obtaining the target entity from the entity set of the virtual environment according to the target entity information of the target entity indicated by the natural language command and the environment perception information;
[0305] The target entity is determined from the virtual environment in combination with environmental perception information, and the environmental perception information includes information perceived by the controlling virtual character and / or the non-player character from the virtual environment.
[0306] Since the target entity information in natural language commands is usually vague, for example, a natural language command can be "move to the red truck", where the target entity information is "red truck", and there may be many red trucks in the virtual environment. It is impossible to accurately determine the target entity in the virtual environment based solely on the target entity information in the natural language command.
[0307] Therefore, embodiments of the present application provide a method for determining a target entity by combining target entity information and environmental perception information. For the previous example, the "red truck" expressed by the player in the natural language command should be a red truck that the player can see. Therefore, based on the field of view of the controlling virtual character, the red truck within the field of view of the controlling virtual character can be selected from multiple red trucks in the virtual environment. This red truck is the target entity indicated by the player in the natural language command.
[0308] Because natural language commands are issued based on the player's perception of the virtual environment, the game application uses the player's environmental perception information at the time of the command to accurately identify the target entity. Based on the player's perception of the virtual environment, the game application infers the entity in the virtual environment that is closest to the target entity and is considered the target entity.
[0309] Exemplarily, the virtual environment includes multiple candidate entities that match the natural language command, and the target entity is an entity selected from the multiple candidate entities based on the environmental perception information of the controlling virtual character or the non-player character. For example, the game application first selects multiple candidate entities from the entity set that match the target entity information in the natural language command, and then selects an entity that can be perceived by the controlling virtual character or the non-player character from the multiple candidate entities as the target entity based on the environmental perception information.
[0310] For example, the environmental perception information of the main virtual character may include: the pictures or entities that the main virtual character can see when observing the virtual environment, the environmental sound effects that the main virtual character can hear, the source direction of the environmental sound effects, the type and volume of the environmental sound effects, the perception information obtained by the main virtual character by using perception skills (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and the perception information obtained by the main virtual character by using sensory props (for example, sensor signals of sensors, positioning signals of positioning props, etc.).
[0311] The environmental perception information of the non-player character may include: the images or entities that the non-player character can see when observing the virtual environment, the environmental sounds that the non-player character can hear, the source direction of the environmental sounds, the type and volume of the environmental sounds, the perception information obtained by the non-player character through the use of perception skills (for example, visual information, auditory information, sound wave information, light reflection information, etc.), and the perception information obtained by the non-player character through the use of sensory props (for example, sensor signals from sensors, positioning signals from positioning props, etc.).
[0312] It should be noted that environmental perception information can include both the real-time perception of the controlling avatar and non-player characters upon receiving a natural language command, and can also include historical perception information of the controlling avatar and non-player characters prior to receiving the natural language command. In other words, the target entity may be an entity currently visible to the player, or an entity the player has previously seen. Therefore, the game application needs to combine the real-time and historical environmental perception information of the controlling avatar and / or non-player characters to identify the target entity.
[0313] For example, when the natural language command is "Let's go back to the hotel we just passed by", the game application needs to query the hotels that the main virtual character has passed by based on the historical environmental perception information of the main virtual character.
[0314] In one optional embodiment, the client queries and obtains the target entity from a set of entities in the virtual environment based on the target entity information indicated by the natural language command and the environment perception information. In another optional embodiment, the client reports the natural language command to the server, and the server queries and obtains the target entity from a set of entities in the virtual environment based on the target entity information indicated by the natural language command and the environment perception information.
[0315] Optionally, the game application infers the target entity based on the target entity information, the environment perception information, and the pre-processed virtual environment data, wherein the pre-processed virtual environment data includes a set of entities in the virtual environment.
[0316] The entity set includes entity information of each entity in the virtual environment. The entity information includes at least one of the following: name, type, location, features, text label, text label embedding vector, image, and image embedding vector.
[0317] The text label may be at least one word obtained after segmenting a feature (for example, a descriptive text of an entity appearance), and the embedding model is called to obtain an embedding vector corresponding to each text label.
[0318] The image may include an image obtained by observing a three-dimensional model of the entity from at least one direction. For example, the image may include three views of the entity. The image embedding vector is the embedding vector of the image feature label. The image feature label is obtained by using a multimodal model to perform feature recognition based on the image and text label of the entity.
[0319] For example, the entity set includes first-category entities and second-category entities. The first-category entities include at least one of regions and buildings, and the second-category entities include physical objects. The text labels for the first-category entities are manually labeled. For example, the text labels for a room in a building might be: region, building, floor, room. The text labels for the second-category entities are automatically generated by invoking a large language model.
[0320] A method for generating text labels for the second type of entity may include: obtaining at least one image of the entity, inputting the at least one image into a large language model to obtain a description text of the entity; segmenting the description text to obtain at least one text label for the entity, where each text label includes a word obtained after the segmentation process; and performing a vectorization operation on the text label to obtain an embedding vector corresponding to the text label.
[0321] Embedding vectors are high-dimensional vector data of text labels. Pre-generating embedding vectors for text labels eliminates the need to repeatedly generate vectors for each entity's text label when matching target entities from an entity set based on target entity information, thereby improving the matching efficiency between target entity information and entity information.
[0322] It's important to note that this method doesn't use the description text directly as the text label. Instead, it uses the word segmentation results of the description text as the text label. This is because description texts often vary in length, and the embedding vectors of longer description texts are less effective when performing actual searches and comparing similarities. For example, when searching for "a small blue truck with rusty body" using the word "truck," "a small blue truck with rusty body" will likely be ranked after "a car" in the recall results. However, in reality, the search for "truck" should prioritize searching for "truck" regardless of the complexity of its additional description. Therefore, when performing similarity searches, the comparison should be based on individual feature words, not the description text.
[0323] In addition, in order to further extract the image features of the entity, the method provided in the embodiment of the present application can also further extract hidden visual features in the image of the entity based on the text label and image of the entity. Taking the first entity as an example, the method includes: obtaining at least one perspective image of the first entity; obtaining at least one text label of the first entity; calling the multimodal model to extract the image features of the first entity based on at least one perspective image and at least one text label of the first entity, and obtain the image embedding vector of the first entity.
[0324] Since text labels are manually annotated or generated by a large language model based on human language characteristics, and text labels extracted based on human language habits may ignore some features of the entity. For example, when manually annotating a text label for an oil drum, only more eye-catching text labels such as "metal" and "rusty" may be annotated, while ignoring detailed features such as "rusty", "blue", "yellow", and "paint peeling in the lower right corner". Therefore, the method provided in the embodiment of the present application also uses a multimodal model based on the entity's text label and entity image, so that the multimodal model extracts more image features from it; based on the image features, image similarity matching with the target entity is performed to improve the recognition accuracy of the target entity.
[0325] The multimodal model training method can include: inputting sample entity images and sample labels into a pre-trained multimodal model, and fine-tuning the pre-trained multimodal model based on the loss of the predicted labels and sample labels output by the pre-trained multimodal model. This results in a multimodal model that can output image feature labels based on the input text labels and image. The image feature labels include more detailed descriptive text extracted from the image. For example, the text description of "a metal oil drum" can only be associated with the two text labels "metal" and "oil drum." However, the Clip (multimodal) model can also identify hidden visual information in the image of the oil drum, such as the image feature labels "rusty," "blue," and "yellow." These image features are not present in the text labels. Therefore, combining the Clip model with image feature search can further improve the accuracy of entity search.
[0326] In an optional embodiment, the game application may query the target entity using the following method.
[0327] (1) Parsing the natural language command to obtain target entity information; the target entity information includes at least one of the following: entity type, entity name, entity location, and entity features.
[0328] Optionally, the game application calls a large language model to parse the natural language command to obtain target entity information of the target entity, and calls an embedding model to vectorize the target entity information to obtain a target embedding vector of the target entity information.
[0329] For example, if the natural language command is "Come to the blue truck in front", the target entity information that can be parsed by the large language model includes: the direction information "front" and the entity name "blue truck".
[0330] For example, named entity recognition technology in the field of natural language processing can also be used to extract target entity information from natural language commands. Named entity recognition technology is used to identify entities with specific meanings in text, such as names of people, places, locative words, adjectives, etc.
[0331] For example, if the natural language command is "Go to the motel on the first floor and find a cardboard box behind the red sofa," named entity recognition technology can identify the following: the architectural noun "motel first floor"; the object nouns "sofa" and "cardboard box"; the locative word "behind"; and the adjective "red." After logical construction, a hierarchical scenario query call is formed, complete with search type, content, locative words, and constraint information. The search type narrows the data retrieval scope based on similarity matching, the search content is a specific entity description, and adjectives are incorporated into the description text for querying. The locative words constrain the search scope. The floor constraint is determined by the player's location. In indoor scenes, the range must be narrowed to avoid searching for entities that are not visible across floors.
[0332] (2) Calculate the similarity between the target entity information and each entity information in the entity set.
[0333] For example, the entity set includes entity information for at least one entity, and the entity information for each entity may include at least one of a text embedding vector for a text label and an embedding vector for an image feature. Similarity can be calculated between the target entity information and the text embedding vector and the image embedding vector, respectively, and the entity with the highest similarity is determined as the target entity.
[0334] The following are methods for similarity matching with text embedding vectors and similarity matching with image embedding vectors.
[0335] 1) Similarity includes the text similarity between the target entity information and the text embedding vector.
[0336] Taking the first entity in the entity set as an example, the game application (client or server) segments the target entity information to obtain at least one target entity tag; converts the at least one target entity tag into at least one target embedding vector; obtains the entity information of the first entity, the entity information includes a text embedding vector, and the text embedding vector is an embedding vector obtained based on the text tag conversion of the first entity; calculates the text parent similarity between the at least one target embedding vector and the text embedding vector respectively, and obtains at least one text parent similarity corresponding to the at least one target embedding vector; and determines the sum of the at least one text parent similarity as the text similarity between the target entity information and the entity information of the first entity.
[0337] For example, there are two target entity labels, target entity label 1 and target entity label 2, and there are also two entities in the entity set, entity 1 and entity 2. Entity 1 corresponds to text embedding vector 1, and entity 2 corresponds to text embedding vector 2. Similarity 1 is calculated between target entity label 1 and text embedding vector 1, and similarity 2 is calculated between target entity label 2 and text embedding vector 1. The sum of similarity 1 and similarity 2 is determined as the text similarity between the target entity information and entity 1. Similarity 3 is calculated between target entity label 1 and text embedding vector 2, and similarity 4 is calculated between target entity label 2 and text embedding vector 2. The sum of similarity 3 and similarity 4 is determined as the text similarity between the target entity information and entity 2.
[0338] Exemplarily, the entity information of the first entity includes at least one text embedding vector; at least one target embedding vector includes a first target embedding vector; the game application calculates the text sub-similarity between the first target embedding vector and the at least one text embedding vector respectively to obtain at least one text sub-similarity; and the highest value of the at least one text sub-similarity is determined as the text parent similarity corresponding to the first target embedding vector.
[0339] For example, the number of target entity tags is 1: target entity tag 1. The number of entities in the entity set is also 1: entity 1. Entity 1 corresponds to text embedding vector 1 and text embedding vector 3. Similarity 1 is calculated between target entity tag 1 and text embedding vector 1, and similarity 5 is calculated between target entity tag 1 and text embedding vector 3. The larger value of similarity 1 or similarity 5 is determined as the similarity between the target entity information (target entity tag 1) and entity 1.
[0340] For example, for the input target entity information "blue car," we first perform word segmentation on the target entity information, breaking it down into the target entity labels "blue" and "car," indicating that the query target entity has two features. We then calculate similarity between each feature and each entity in the entity set. The maximum similarity score for each feature is taken, and finally, the sum of the similarity scores corresponding to all query features is taken, representing the textual similarity between the target entity information and the query entity.
[0341] For example, the target entity information contains two target entity labels, "blue" and "car." The two target entity labels are converted into two target embedding vectors. Then, the entity information in the entity set is obtained. For example, the entity set includes entity 1 and entity 2. The text labels of entity 1 include "metal," "old," "car," "truck," "blue," and "damaged." The text labels of entity 2 include "metal," "scratches," "car," "damaged," "armored truck," and "black."
[0342] Calculate the text similarity between the target entity information and entity 1: first calculate the similarity between the target entity label "blue" and each text label of entity 1, and take the maximum value. For example, the similarity between the target entity label "blue" and the text label "blue" of entity 1 is 1 (1 is the maximum value); then calculate the similarity between the target entity label "car" and each text label of entity 1, and take the maximum value. For example, the similarity between the target entity label "car" and the text label "car" of entity 1 is 1. Then sum the two similarities of the target entity labels "car" and "blue" to obtain the final similarity between the target entity information and entity 1, which is 2.
[0343] Similarly, calculate the text similarity between the target entity information and entity 2: first calculate the similarity between the target entity label "blue" and each text label of entity 2, and take the maximum value. For example, the similarity between the target entity label "blue" and the text label "black" of entity 1 is 0.91; then calculate the similarity between the target entity label "car" and each text label of entity 1, and take the maximum value. For example, the similarity between the target entity label "car" and the text label "car" of entity 2 is 1 (1 is the maximum value). Then sum the two similarities of the target entity labels "car" and "blue" to obtain the final similarity between the target entity information and entity 2, which is 1.91.
[0344] It can be seen that the similarity between the target entity information and entity 1 is 2, which is higher than the similarity between the target entity information and entity 2, which is 1.91.
[0345] 2) Similarity includes the image similarity between the target entity information and the image embedding vector.
[0346] Taking the first entity in the entity set as an example, the game application (client or server) segments the target entity information to obtain at least one target entity tag; converts the at least one target entity tag into at least one target embedding vector; obtains the entity information of the first entity, the entity information includes an image embedding vector, and the image embedding vector is an embedding vector obtained based on the image extraction of the first entity; calculates the image parent similarity of the at least one target embedding vector and the text embedding vector respectively, and obtains at least one image parent similarity corresponding to the at least one target embedding vector; and determines the sum of the at least one image parent similarity as the image similarity between the target entity information and the entity information of the first entity.
[0347] For example, there are two target entity labels, target entity label 1 and target entity label 2, and there are also two entities in the entity set, entity 1 and entity 2. Entity 1 corresponds to image embedding vector 1, and entity 2 corresponds to image embedding vector 2. Similarity 1 is calculated between target entity label 1 and image embedding vector 1, and similarity 2 is calculated between target entity label 2 and image embedding vector 1. The sum of similarity 1 and similarity 2 is determined as the image similarity between the target entity information and entity 1. Similarity 3 is calculated between target entity label 1 and image embedding vector 2, and similarity 4 is calculated between target entity label 2 and image embedding vector 2. The sum of similarity 3 and similarity 4 is determined as the image similarity between the target entity information and entity 2.
[0348] Exemplarily, the entity information of the first entity includes at least one image embedding vector; the at least one target embedding vector includes a first target embedding vector; the game application calculates the image sub-similarity between the first target embedding vector and the at least one image embedding vector respectively to obtain at least one image sub-similarity; and the highest value of the at least one image sub-similarity is determined as the image parent similarity corresponding to the first target embedding vector.
[0349] For example, the number of target entity tags is 1: target entity tag 1. The number of entities in the entity set is also 1: entity 1. Entity 1 corresponds to image embedding vector 1 and image embedding vector 3. Similarity 1 is calculated between target entity tag 1 and image embedding vector 1, and similarity 5 is calculated between target entity tag 1 and image embedding vector 3. The larger value of similarity 1 or similarity 5 is determined as the similarity between the target entity information (target entity tag 1) and entity 1.
[0350] In an optional embodiment, the target image embedding vector can be used to calculate image similarity, that is, the target image embedding vector is used instead of the target embedding vector in the above method. The target image embedding vector can be obtained by: segmenting the target entity information to obtain at least one target entity label; and inputting the at least one target entity label into a multimodal model to obtain the target image embedding vector.
[0351] 3) Similarity includes text similarity and image similarity.
[0352] When the first entity and the target entity have corresponding text similarity and image similarity, the game application determines an average value of the text similarity and the image similarity as the similarity between the first entity and the target entity.
[0353] Alternatively, when the first entity and the target entity have corresponding text similarity and image similarity, the larger value of the text similarity and the image similarity is determined as the similarity between the first entity and the target entity.
[0354] Alternatively, when the first entity and the target entity have corresponding text similarity and image similarity, the sum of the text similarity and the image similarity is determined as the similarity between the first entity and the target entity.
[0355] (3) Determine the target entity from the entity set based on similarity and environmental perception information.
[0356] In an optional embodiment, the game application first performs a similarity search based on the target entity information, obtains at least one candidate entity with a high similarity, and sorts them according to the similarity; then, based on the position, position, orientation and other information of the main virtual character and / or the non-player character, the game application performs perceptual filtering and sorting, and finally filters out the target entity.
[0357] Exemplarily, the game application determines a target perception range based on target entity information and environmental perception information; and obtains the target entity by filtering entities perceived within the target perception range based on similarity.
[0358] Alternatively, the game application determines x entities with the highest similarity from the entity set based on similarity to form a candidate entity list; determines a target perception range based on the target entity information and the environmental perception information, and determines the entity within the target perception range among the x entities as the target entity, where x is a positive integer.
[0359] Exemplarily, the game application determines the entity with the highest similarity in the perception range as the target entity.
[0360] It should be noted that when multiple entities with equal similarity exist within the target perception range, they can be sorted by their distance from the controlling virtual character. For example, if there are at least two entities with the highest similarity within the perception range, the entity with the highest similarity within the perception range and the closest distance to the controlling virtual character will be determined as the target entity.
[0361] In an optional embodiment, the game application determines the entity with the highest similarity within the perception range and above a threshold as the target entity. Exemplarily, the number of perception ranges is at least one. For example, the perception range includes at least one of the following: the visual range of the controlling virtual character, the auditory range of the controlling virtual character, the perception range of the controlling virtual character's virtual props, the visual range of non-player characters, the auditory range of non-player characters, and the perception range of non-player characters' virtual props.
[0362] For example, the game application may sequentially traverse at least one of the aforementioned perception ranges, or the game application may sequentially traverse at least one perception range determined based on a natural language command. For example, if the natural language command is "Can you see the truck in front? Move to the truck?", the game application may prioritize traversing the control virtual character's field of view, followed by the field of view of non-player characters, and match entities with the highest similarity that exceeds a threshold.
[0363] For example, when there are at least two perception ranges, the order in which the at least two perception ranges are traversed can be predetermined, or the order in which the at least two perception ranges are traversed can be determined based on the intent indicated in the natural language command. For example, when the intent is to search for an object, the visual range is prioritized. When the intent is to search for combat scenes, the auditory range is prioritized.
[0364] If there are multiple levels of nested instructions in the natural language command, the first target entity will be searched first, and then the next target entity will be searched based on the position of the previous target entity. In the case where the natural language command includes a second target entity determined based on the position of the first target entity, the first target entity is queried from the entity set of the virtual environment based on the first target entity information and environmental perception information of the first target entity indicated by the natural language command; then, the second target entity is queried from the entity set of the virtual environment based on the second target entity information, the first target entity, and the environmental perception information.
[0365] In response to natural language commands, the NPC is controlled to perform a virtual activity associated with a target entity.
[0366] Exemplarily, a large language model is called to identify the behavioral intention of natural language commands, and a control instruction or a control instruction sequence is generated according to the behavioral intention. According to the control instruction or the control instruction sequence, a non-player character is controlled to complete activities based on the target entity.
[0367] The broadcasting of the NPC feedback information in step 208 can be implemented as follows:
[0368] Sub-step 6: reporting feedback information of the non-player character, wherein the feedback text corresponding to the feedback information is a non-fixed text generated based on the environmental perception information of the non-player character;
[0369] For example, the game application can generate corresponding feedback information in real time based on the environmental perception information of the non-player character and broadcast it. For example, when the environmental perception information of the non-player character triggers a broadcast condition, the game application can generate corresponding feedback information in real time based on the environmental perception information that triggered the broadcast condition; or when the non-player character or the master virtual character triggers a broadcast condition, the game application can generate corresponding feedback information in real time based on the environmental perception information of the non-player character.
[0370] For example, the feedback text is derived by inferring the large language model based on the static entity data of the 3D virtual environment and the dynamic environmental perception information of the non-player characters. The large language model generates different feedback texts based on the different perception conditions of the non-player characters.
[0371] For example, the feedback text can be derived by a large language model based on static entity data of the 3D virtual environment and the dynamic environmental perception information of the non-player character, and can be tailored to the personality traits of the non-player character. For example, if the non-player character is a strong and handsome man, the feedback text can adopt a heroic tone; if the non-player character is a reporter, the feedback text can adopt a news report tone.
[0372] It should be noted that due to the randomness of the results generated by the large language model, in the same scene, when the environmental perception information of the non-player character is the same, the generated feedback text may also be different.
[0373] The environmental perception information may include information perceived by the non-player character in real time, or may include information perceived by the non-player character in history.
[0374] For example, the feedback information is derived by the game application based on the non-player character's environmental perception information and the static entity data of the 3D virtual environment. The feedback information can be broadcast in the form of voice or text. The static entity data includes data on relatively unchanging entities in the 3D virtual environment, such as model information of various buildings, terrain, and vehicles in the 3D virtual environment.
[0375] For example, the feedback information may include at least one of the following: immediate feedback, execution feedback, and dynamic feedback. Instant feedback and execution feedback are both feedback information generated in response to natural language commands, while dynamic feedback is feedback information spontaneously generated by non-player characters.
[0376] 1. Instant feedback is the feedback information generated immediately upon receiving a natural language command. Instant feedback can include reply announcements and instant announcements. Reply announcements are used to reply to the inquiry in the natural language command; instant announcements are used to provide feedback on the reception of the natural language command.
[0377] For example, a client can direct a non-player character to move in a 3D virtual environment based on the intent of a natural language command. Upon receiving a natural language command, the client can generate immediate feedback corresponding to the natural language command. The immediate feedback indicates receipt of the natural language command, or provides an immediate response to the natural language command.
[0378] In an optional embodiment, the feedback information includes a reply broadcast. The reply broadcast includes a response to a natural language query command issued by the player. For example, if the natural language query command includes a query for the location of a target entity, the corresponding reply broadcast should include the query result for the location of the target entity.
[0379] When the behavioral intention of the received natural language command is an inquiry, the terminal broadcasts a reply message of the non-player character; the reply message includes a reply content for the natural language command generated based on the environmental perception information of the non-player character of the three-dimensional virtual environment.
[0380] The natural language command is used to request information from a non-player character. This natural language command can be referred to as a query natural language command. A query natural language command is a natural language command intended to be a query. The query natural language command contains at least one question.
[0381] Natural language commands are instructions conveyed using natural human language. The terminal receives natural language commands from players and directs non-player characters based on the intent of the natural language commands. For example, the terminal receives the player's voice audio and converts it into text to generate the natural language commands. Alternatively, the terminal receives the text of natural language commands input by the player.
[0382] Players can issue natural language commands using spoken or written language. The game application uses a large language model to understand the natural language commands, extract the intent expressed in the natural language commands, and generate control instructions based on the intent to direct the activities of non-player characters.
[0383] For example, the natural language command may be "Come to me", then the game application can parse out that the intention of the natural language command is "control the non-player character to move to the location of the controlling virtual character", then the game application obtains the location of the controlling virtual character, generates a navigation route for the non-player character to move to the location of the controlling virtual character, and controls the non-player character to move to the location of the controlling virtual character according to the navigation route.
[0384] For another example, the natural language command may be "go pick up the treasure chest", then the game application can parse out that the intention of the natural language command is "control the non-player character to move to the location of the treasure chest and pick up the treasure chest", then the game application can obtain the location of the treasure chest, control the non-player character to move to the location of the treasure chest, and perform the operation of picking up the treasure chest after arriving at the location of the treasure chest.
[0385] For example, when the natural language command includes an intelligence inquiry about the target entity, a first reply report is broadcast based on the non-player character's perception of the target entity; the first reply report includes the non-player character's intelligence perception result about the target entity.
[0386] Alternatively, in a case where the natural language command includes an inquiry about the location of the target entity, a second reply broadcast is broadcast based on the non-player character's perception of the target entity, and the second reply broadcast includes a description of the location of the target entity.
[0387] Exemplarily, a large language model is called to parse the received natural language command to obtain the intent of the natural language command. When the intent is of the inquiry type, the natural language command can be determined to be an inquiry natural language command. When the natural language command is an inquiry natural language command, the large language model can infer the intent of the natural language command based on the static entity data and environmental perception information to obtain the reply text. That is, the large language model is called to parse the received natural language command, and the intent of the natural language command is output as an inquiry, as well as the reply text of the natural language command. The game application calls the text-to-speech service according to the processing logic of the inquiry intent, converts the reply text into a reply broadcast, and the reply broadcast is an audio broadcast, and the reply audio is sent to the client for broadcast.
[0388] For example, when the large language model recognizes that the intention of the natural language command is to ask a question, since the user may ask hundreds of questions and it is impossible to store the answer to each question locally on the client, the game application will call the text-to-speech service to generate a reply broadcast in real time based on the reply text returned by the large language model, so that the game application can instantly generate a corresponding reply broadcast to respond to the player's question.
[0389] For example, a game application (client or server) calls a large language model to parse natural language commands, and infers reply text based on the static environmental data and environmental perception information of the three-dimensional virtual environment; generates human voice audio based on the reply text to obtain a reply broadcast.
[0390] For example, the client reports the received natural language query command to the server, the server calls the large language model to parse the natural language query command, and infers the reply text based on the static environmental data and environmental perception information of the three-dimensional virtual environment; generates human voice audio based on the reply text to obtain a reply broadcast; the server sends the reply broadcast to the client; and the client broadcasts the reply broadcast.
[0391] 2. Execution feedback is the feedback generated when a task is executed according to the intent indicated by a natural language command after receiving it. Execution feedback can also be called execution broadcast.
[0392] For example, when a natural language command is used to instruct a non-player character to perform a target task, the client can generate execution feedback during the non-player character's performance of the target task, and the execution feedback is used to indicate the execution status of the target task; or, the client can generate execution feedback after the non-player character completes the target task, and the execution feedback is used to indicate the execution result of the target task.
[0393] In an optional embodiment, the feedback information includes a feedback report (also referred to as execution feedback). The feedback report includes a report on the task execution status and results generated in response to the task natural language command issued by the player. For example, if the task natural language command includes moving to a target entity, the corresponding feedback report may include: moving to the target entity, or having reached the target entity.
[0394] When the behavioral intention of the received natural language command is to execute a task, the terminal broadcasts a feedback broadcast; the feedback broadcast includes the execution status after the natural language command is executed according to the environmental perception information of the non-player character of the three-dimensional virtual environment.
[0395] The natural language command is used to instruct the non-player character to perform a task. This natural language command may be referred to as a task natural language command. The task natural language command includes a description of the task. For example, the description may include at least one of the following: the task name, the action to be performed by the non-player character, the method to be used to perform the task, the task objective, and the location of the task objective.
[0396] For example, when the intention of the task natural language command includes searching for a target entity, a first feedback announcement is broadcast, and the first feedback announcement includes the search result of the non-player character for the target entity.
[0397] Alternatively, when the intention of the task natural language command includes using a virtual prop, a second feedback announcement is broadcast, and the second feedback announcement is used to indicate that the virtual prop is ready.
[0398] Alternatively, when the intention of the task natural language command includes the use of virtual props, a third feedback announcement is broadcast, and the third feedback announcement is used to indicate the result of the use of the virtual props.
[0399] Alternatively, when the intention of the task natural language command includes controlling the movement of the non-player character, a fourth feedback announcement is announced, and the fourth feedback announcement is used to indicate the movement result of the movement.
[0400] Alternatively, when the intention of the task natural language command includes performing an interaction with a target entity, a fifth feedback announcement is broadcast, and the fifth feedback announcement is used to indicate at least one of the search result of the non-player character for the target entity and the interaction result between the non-player character and the target entity.
[0401] Alternatively, when the intention of the task natural language command includes item interaction, a sixth feedback announcement is broadcast, and the sixth feedback announcement is used to indicate the item interaction result of the non-player character performing the item interaction.
[0402] Exemplarily, a large language model is called to parse the received natural language command to obtain the intent of the natural language command. When the intent is a task, the natural language command can be determined to be a task natural language command. When the natural language command is a task natural language command, the large language model can reason according to the intent of the natural language command, based on static entity data and environmental perception information, to obtain a task instruction sequence. The task instruction sequence is used to control the non-player character to perform the task. When the game application receives the task instruction sequence, it controls the non-player character's activities to perform the task according to the task instructions in the task instruction sequence. While controlling the non-player character according to the task instructions, the corresponding execution status can be obtained and feedback broadcasted according to the feedback broadcast logic corresponding to different task instructions.
[0403] For example, since the commands that can be executed by non-player characters in a three-dimensional virtual environment are finite and traversable, the client can locally store the feedback broadcast corresponding to each character's command. When the non-player character executes the corresponding command, the client can read the local feedback broadcast and perform voice broadcast.
[0404] Alternatively, when there is variable content in the feedback broadcast corresponding to a certain task instruction, for example, the target entity in the feedback broadcast is variable, or the location information in the feedback broadcast is variable; the client can also generate a broadcast voice with variable content based on the target entity or location information returned by the large language model, and splice it with the broadcast voice with immutable content stored locally to obtain the final feedback broadcast.
[0405] For example, if the task natural language command is "Find a nearby treasure chest," parsing it with the large language model reveals the task intent, the target entity, and the treasure chest. Reasoning based on static entity data and environmental perception information reveals the presence of a treasure chest to the left and front of the main control virtual character. If the target entity can be found, the corresponding feedback announcement for this task objective is "Found aaa at xxx," where xxx is the target entity's location and aaa is its name. The game application can then use online text-to-speech technology to generate a location audio based on the inference result, "front left." Based on the target entity's name, "treasure chest," the application can also use online text-to-speech technology to generate a name audio. The location audio and name audio are then concatenated into the feedback announcement for the final feedback announcement.
[0406] Exemplarily, a game application (client or server) calls a large language model to parse natural language commands for a task, and infers task execution instructions based on the static environmental data and environmental perception information of the three-dimensional virtual environment; controls a non-player character to perform the task according to the task execution instructions; generates feedback text based on the execution status of the non-player character performing the task; and generates human voice audio based on the feedback text to obtain a feedback broadcast.
[0407] For example, the client reports a task natural language command to the server; the server parses the task natural language command and, based on the static environmental data and environmental perception information of the three-dimensional virtual environment, infers the task execution instructions; the non-player character is controlled to perform the task according to the task execution instructions; feedback text is generated based on the execution status of the non-player character performing the task; human voice audio is generated based on the feedback text to obtain a feedback broadcast; the server sends the feedback broadcast to the client; the client receives and broadcasts the feedback broadcast.
[0408] 3. Dynamic feedback is spontaneously generated based on the non-player character's environmental perception information. Dynamic feedback can also be called dynamic broadcast.
[0409] For example, the client can spontaneously generate dynamic feedback based on its perception of the 3D virtual environment without receiving natural language commands. This dynamic feedback is used to indicate abnormal conditions detected by the non-player character in the 3D virtual environment. For example, abnormal conditions may include: detecting a hostile virtual character, detecting a change in the status of a friendly virtual character, detecting a dangerous situation, or detecting traces of combat or looting.
[0410] In an optional embodiment, the feedback information includes dynamic notifications (also referred to as "dynamic feedback"). Dynamic notifications include notifications of abnormal situations perceived by non-player characters. For example, when a non-player character perceives being attacked, a notification of "being attacked" is given; when a non-player character perceives a dangerous situation ahead, a notification of "danger ahead" is given.
[0411] The terminal broadcasts a dynamic broadcast when the environmental perception information of the non-player character on the three-dimensional virtual environment meets the dynamic broadcast conditions.
[0412] Dynamic notifications are generated by the game application based on the environmental perception information of non-player characters and when it determines that an abnormal situation requires notification. Dynamic notifications include alerts about abnormal situations. For example, abnormal situations may include at least one of the following: the discovery of a hostile virtual character, the discovery of a change in the status of a friendly virtual character, or the discovery of new traces.
[0413] For example, when a non-player character perceives a hostile virtual character, a first dynamic broadcast is broadcast; the first dynamic broadcast includes the position of the hostile virtual character perceived by the non-player character.
[0414] Alternatively, when the non-player character senses a dangerous situation, a second dynamic broadcast is broadcast; the second dynamic broadcast is used to prompt the dangerous situation.
[0415] Alternatively, when the non-player character senses a change in the state of the friendly virtual character, a third dynamic broadcast is broadcast; the third dynamic broadcast is used to prompt that the state of the friendly virtual character has changed.
[0416] Alternatively, when the non-player character discovers new traces in the three-dimensional virtual environment, a fourth dynamic broadcast is broadcasted; the fourth dynamic broadcast is used to indicate the new traces, wherein the new traces may be combat traces and / or looting traces.
[0417] For example, since the number of dynamic announcements that non-player characters can trigger in a 3D virtual environment is limited and traversable, the client can locally store dynamic announcement voices. When a corresponding dynamic announcement voice is triggered, the client can read the dynamic announcement voice from the local storage and announce it. Similarly, some dynamic announcement voices can contain variable content. The client can generate variable content in real time using text-to-speech technology and splice it with the dynamic announcement template to produce the final dynamic announcement.
[0418] In an optional embodiment, when generating feedback information, it may be necessary to query the location of a target entity in the 3D virtual environment, or to determine the target entity indicated in a natural language command from the 3D virtual environment. In this case, an entity query method is required to query the target entity.
[0419] Before executing a target entity query, static entity data for the 3D virtual environment must be pre-built. This static entity data serves both the inference of the large language model and the real-time feedback text generation on the game application side after the inference results are returned.
[0420] In an optional embodiment, entities in the three-dimensional virtual environment are traversed, and entity information of each entity is exported from a game engine to obtain static entity data.
[0421] For example, for in-game scenes, editor tools have been developed to perform StaticMesh traversal of the entire scene, special Actor traversal (such as interactive doors, which do not fall into the StaticMesh category), vegetation traversal, and export the entity's position, orientation, and bounding box size as static entity data for subsequent labeling and inference by the large language model.
[0422] StaticMesh is a static geometry resource type in the Unreal Engine 4 (UE4) game engine, used to represent immutable 3D models such as buildings and props. Actor is a basic object class in the Unreal Engine 4 (UE4) game engine, representing an entity in the game world, such as a character, object, or light source. Actors can contain multiple components to implement different functions.
[0423] For example, static entity data can also include spatial data, which is obtained by manual labeling. Spatial data is more complex than static entities because there are multiple floors in the space and multiple layers of nesting within the floors (such as the motel area includes the second floor, and the second floor area includes room 201). For this type of data, manual calibration is used to construct it. Add Volume to the scene in the editor, divide the area required for the entire map and label it accordingly. Volume is a special Actor class in the UnrealEngine4 (UE4) game engine, which represents a three-dimensional area with specific functions, such as a trigger area, audio area, etc.
[0424] After obtaining the static entity data, you can query entities or spaces based on the static entity data during game operation.
[0425] For example, when a player issues a natural language command, it is accompanied by pre-built static entity data and real-time runtime data (including environmental perception information of non-player characters and / or the controlling avatar) to make an inference request to the large language model. Furthermore, when an NPC needs to provide dynamic feedback based on runtime data, an entity query is also performed to generate feedback text based on the current state.
[0426] When a natural language command contains a target entity, the game application will query the target entity based on the position and orientation of the current controlling virtual character, the position and orientation of teammates or enemies, and based on the target entity information of the target entity described in the natural language command and the environmental perception information of the controlling virtual character and / or non-player character, filter out the target entity that matches the target entity information from at least one candidate entity within the perception range of the controlling virtual character and / or non-player character.
[0427] Optionally, natural language commands are used to instruct the non-player character to perform actions related to a target entity. For example, the target entity can be a destination for the non-player character to move to, an object that the non-player character needs to observe, a target that the non-player character needs to attack, or an object that the non-player character needs to interact with.
[0428] A target entity is a virtual object that exists in a 3D virtual environment. A target entity is an entity described in a natural language command. An entity can refer to a virtual object that is stationary in a 3D virtual environment. For example, a target entity can be a virtual building, virtual terrain, virtual vehicle, virtual plant, virtual prop, virtual item, etc. An entity can also refer to a virtual object in a 3D virtual environment that can interact with a virtual character. For example, a target entity can be an interaction point in a 3D virtual environment (e.g., a door, window, cabinet, cellar, etc.), a virtual prop, a virtual character, a virtual light source, etc.
[0429] For example, a natural language command could be "Please open the door for me," then the target entity could be "door," and the non-player character would need to perform a door-related action, "Open the door." Alternatively, a natural language command could be "Is the kitchen safe?" then the target entity could be "kitchen," and the non-player character would need to perform a kitchen-related action, "Check the kitchen for enemy avatars or other dangerous situations."
[0430] Optionally, the natural language command includes target entity information of the target entity. The target entity information may be text describing the target entity. For example, the target entity information may include at least one of the following: the name, type, location, and characteristics of the target entity. Based on the target entity information in the natural language, the game application can locate the target entity from among multiple entities in the three-dimensional virtual environment.
[0431] The game application parses the natural language commands, extracts the target entity information therein, determines the target entity from the entity set of the three-dimensional virtual environment based on the target entity information, and then controls the non-player character to perform activities related to the target entity according to the intention of the natural language commands.
[0432] The target entity is determined from the three-dimensional virtual environment in combination with environmental perception information, and the environmental perception information includes information perceived by at least one of the main virtual character and the non-player character from the three-dimensional virtual environment.
[0433] Since the target entity information in natural language commands is usually vague, for example, a natural language command can be "move to the red truck", where the target entity information is "red truck", and there may be many red trucks in the three-dimensional virtual environment. It is impossible to accurately determine the target entity in the three-dimensional virtual environment based solely on the target entity information in the natural language command.
[0434] Therefore, embodiments of the present application provide a method for determining a target entity by combining target entity information and environmental perception information. For the previous example, the "red truck" expressed by the player in the natural language command should be a red truck that the player can see. Therefore, based on the field of view of the controlling virtual character, the red truck within the field of view of the controlling virtual character can be selected from multiple red trucks in the three-dimensional virtual environment. This red truck is the target entity indicated by the player in the natural language command.
[0435] Because natural language commands are issued based on the player's perception of the 3D virtual environment, the game application uses the player's environmental perception information at the time the command is issued to accurately identify the target entity. Based on the player's perception of the 3D virtual environment, the game application infers the entity in the 3D virtual environment that is closest to the target entity's information. This entity is then identified as the target entity.
[0436] Exemplarily, a three-dimensional virtual environment includes multiple candidate entities that match a natural language command, and a target entity is an entity selected from the multiple candidate entities based on environmental perception information of a controlling virtual character or a non-player character. For example, a game application first selects multiple candidate entities from a set of entities that match the target entity information in the natural language command, and then selects an entity that can be perceived by the controlling virtual character or the non-player character from the multiple candidate entities as the target entity based on the environmental perception information.
[0437] For example, the control virtual character's perception range is determined based on its location and orientation. For example, the perception range is a sector-shaped area of a certain visual field in front of the control virtual character. Entities outside this sector are filtered out. The target entity information and candidate entity information are then sorted based on their similarity, with the entity with the highest similarity selected as the target entity. If two entities within the perception range have the highest and equal similarity scores—for example, if entity 2 and entity 3 both have a similarity score of 2—entity 2, which is closer to the control virtual character, is selected as the target entity.
[0438] When feedback information needs to be generated based on the position description of the target entity, the game application can perform a spatial query to generate the position description based on the positional relationship between the master virtual character and the target entity.
[0439] Spatial queries are primarily used to obtain the locations of the main virtual character, non-player characters, hostile virtual characters, gunshots, and more. For example, a report containing spatial information can be broadcast, such as "Enemy spotted on the second floor of a motel."
[0440] Spatial query logic defines the concepts of parent regions and child regions. A parent region refers to a large area within a 3D virtual environment that encompasses multiple buildings or isolated spaces, such as a motel with front and back yards, multiple rooms, and a basement. A child region refers to an independent space within a parent region. These can be enclosed spaces or specialized spaces, such as a motel guest room or kitchen.
[0441] In order to make the location description returned by the spatial query closer to human expression, when the controlling virtual character and the queried location are in different parent areas, the query result returns the complete text information (for example, if the controlling virtual character is outside the motel and the queried location is on the second floor of the motel, then the query result returns "The enemy is on the second floor of the motel"). When the controlling virtual character and the queried location are in the same parent area but different child areas, the query result only returns the child area information (for example, if the controlling virtual character is on the first floor of the motel and the queried location is on the second floor, then the query result returns "The enemy is on the second floor"). When the controlling virtual character and the queried location are in the same parent area and the same child area, the query result only returns the direction relative to the main perspective player (for example, if the controlling virtual character and the queried location are both on the second floor of the motel, then the query result returns "The enemy is in the right front position").
[0442] Exemplarily, when the controlling virtual character and the target entity are located in different parent spaces, the position description includes the parent space and child space where the target entity is located. When the controlling virtual character and the target entity are located in different child spaces of the same parent space, the position description includes the child space where the target entity is located. When the controlling virtual character and the target entity are located in the same child space of the same parent space, the position description includes the relative position of the target entity and the controlling virtual character; wherein, a parent space includes at least one child space.
[0443] Regarding the broadcasting of the NPC feedback information in step 208, the feedback information is implemented as spatial human voice audio processed with spatial sound effects. The spatial sound effect processing process can be implemented as follows:
[0444] Sub-step 7: Obtain the human voice audio corresponding to the feedback text;
[0445] In the embodiment of the present application, the process of generating the feedback text is as shown in the above-mentioned sub-step 6, which will not be repeated here.
[0446] In some embodiments, the process of obtaining corresponding human voice audio based on feedback text can be implemented as follows: obtaining feedback text and character feature information corresponding to a non-player character; and generating human voice audio matching the character feature information based on the feedback text.
[0447] Illustratively, character feature information is used to describe the attributes of a non-player character. Specifically, character feature information is information used to describe the current state and characteristics of a non-player character. In some embodiments, the attributes of a non-player character include at least one of basic attributes and situational attributes. Basic attributes are pre-configured attributes for a non-player character, i.e., attributes that do not change due to the 3D virtual environment or the current situation. Examples include the non-player character's age, gender, or martial arts school. Scenario-specific attributes are attributes determined in real time within the context of the 3D virtual environment. These attributes are associated with the current situation, such as the non-player character's emotions and behavior in the current situation.
[0448] Optionally, the character characteristic information includes at least one of basic character information, character emotion information, and character behavior information. The basic character information is used to indicate the basic characteristics of the non-player character, and is information about the basic attributes of the non-player character, such as the non-player character's age, gender, or martial arts background. The character emotion information is used to indicate the emotional state of the non-player character in the dialogue context corresponding to the feedback text, and is a situational attribute of the non-player character, such as indicating whether the non-player character is excited, angry, or sad. The character behavior information is used to indicate the actions performed by the non-player character in the dialogue context, and is a situational attribute of the non-player character, such as whether the non-player character is running, attacking, or receiving injury treatment.
[0449] In some embodiments, human voice audio is generated using a pre-trained speech generation model. The speech generation model is used to generate speech that matches the character characteristics of the non-player character. Illustratively, feedback text and character characteristics are input into the pre-trained speech generation model to generate human voice audio, which is then used as the human voice audio.
[0450] Optionally, the above-mentioned speech generation model can be implemented by a neural network model such as a convolutional neural network, a feedforward neural network, a residual network, a Transformer, a multimodal large language model (MLLM), etc., which is not specifically limited here.
[0451] In some embodiments, the speech generation model includes a text encoder, a role encoder, and a decoder. The text encoder is configured to perform text encoding on the input feedback text, i.e., the feedback text is input into the text encoder to obtain a text encoding representation. The role encoder is configured to perform feature encoding on the input role feature information, i.e., the role feature information is input into the role encoder to obtain a role encoding representation.
[0452] In an illustrative manner, after obtaining the text coding representation and the role coding representation, the text coding representation and the role coding representation are fused to obtain the coding representation input to the decoder, that is, the text coding representation and the role coding representation are fused to obtain a joint coding representation, and the joint coding representation is input into the decoder to generate human voice audio.
[0453] In some embodiments, the character encoder includes at least one of a first sub-encoder, a second sub-encoder, and a third sub-encoder. The first sub-encoder is configured to perform feature encoding based on the basic attributes of the non-player character. Specifically, if the character feature information includes basic character information, the basic character information is input into the first sub-encoder to obtain a first character encoding representation. The second sub-encoder is configured to perform feature encoding based on the emotional state of the non-player character. Specifically, if the character feature information includes character emotional information, the character emotional information is input into the second sub-encoder to obtain a second character encoding representation. The third sub-encoder is configured to perform feature encoding based on the behavioral state of the non-player character. Specifically, if the character feature information includes character behavioral information, the character behavioral information is input into the third sub-encoder to obtain a third character encoding representation.
[0454] Sub-step 8: obtaining sound effect parameters corresponding to the non-player character based on the relative position relationship between the non-player character and the controlling virtual character in the three-dimensional virtual environment;
[0455] Illustratively, the relative position relationship between the non-player character and the controlling virtual character in the 3D virtual environment is used to indicate the relative relationship between the first position corresponding to the non-player character and the second position corresponding to the controlling virtual character in the 3D virtual environment.
[0456] Optionally, the above-mentioned relative orientation relationship includes at least one of position distance, position direction, and spatial encirclement between the non-player character and the main virtual character, wherein the position distance is used to indicate the distance relationship between the non-player character and the main virtual character, the position direction is used to indicate the angular relationship between the first position of the non-player character and the orientation direction of the main virtual character, and the spatial encirclement is used to indicate the surrounding situation of the main virtual character by the spatial elements forming the virtual space in the three-dimensional virtual environment, and the propagation effect of the space on the audio when the non-player character emits audio in the virtual space.
[0457] In some embodiments, the determination of the relative orientation relationship between the non-player character and the controlling virtual character can be implemented as follows: obtaining first position information of the non-player character in the three-dimensional virtual environment, obtaining second position information of the controlling virtual character in the three-dimensional virtual environment, and determining the relative orientation relationship between the non-player character and the controlling virtual character based on the first position information and the second position information.
[0458] In some embodiments, the first position information of the non-player character and the second position information of the controlling virtual character are position information determined based on the same preset coordinate system. Optionally, the preset coordinate system can be implemented as a world coordinate system corresponding to the three-dimensional virtual environment, or as a coordinate system established with the controlling virtual character as the origin.
[0459] In some embodiments, when the sound effect parameters corresponding to the non-player characters are pre-generated and stored in a database, the sound effect parameters corresponding to the non-player characters are retrieved from the database. In other embodiments, the sound effect parameters corresponding to the non-player characters may also be generated in real time, i.e., based on relative position relationships.
[0460] Optionally, the sound effect parameters corresponding to the non-player character may include at least one of the following types:
[0461] The first type is distance type
[0462] Illustratively, a distance-based sound effect parameter is used to indicate adjustments to the sound effects of the human voice audio based on the distance between the non-player character and the controlling virtual character. In some embodiments, the distance-based sound effect parameter can adjust the volume of the human voice audio to simulate the effect of distance on the volume of the sound; and / or the distance-based sound effect parameter can adjust the audio delay of the human voice audio to simulate the effect of distance on the sound propagation time.
[0463] Second, direction type
[0464] Illustratively, directional sound effect parameters are used to indicate adjustments to the sound effects of the human voice audio based on the direction of the non-player character relative to the controlling avatar. In some embodiments, directional sound effect parameters can adjust the volume of the human voice audio to simulate whether the controlling avatar is facing a scene audio element; and / or directional sound effect parameters can adjust the volume parameters of the human voice audio in different channels to simulate the direction of the audio emitted by a scene sound source element.
[0465] The third type is spatial special effects
[0466] Illustratively, spatial effect-type sound effect parameters are used to indicate adjustments to the sound effects within the virtual space based on the virtual space formed by spatial elements. Specifically, spatial effect-type sound effect parameters indicate the impact of the virtual space on the audio performance of audio emitted by non-player characters. In some embodiments, spatial effect-type sound effect parameters can include echo and reverberation effects to simulate the effect of sound within a space.
[0467] Optionally, the spatial effect type sound effect parameters include at least one of a reverberation time parameter, a pre-delay parameter, a wet / dry mix parameter, a space width parameter, and a distance effect parameter. The reverberation time parameter indicates the rate at which audio decays in the virtual space; the pre-delay parameter indicates the time difference between the audio directly reaching the controlling virtual character and the audio reaching the controlling virtual character after the first reflection; the wet / dry mix parameter indicates the ratio of directly transmitted audio to reflected audio; the space width parameter indicates the degree of audio diffusion on the horizontal plane in the virtual space; and the distance effect parameter indicates the attenuation of audio as it propagates over distance.
[0468] Fourth, custom type
[0469] Schematically, the custom type of sound effect parameters are user-defined parameters for adjusting sound effects. Optionally, the user can customize the overall volume of the scene audio, whether to add background music, customize the volume of different types of audio, etc.
[0470] In some embodiments, the client provides the user with a customization interface for sound effect parameters, and the user configures the customized sound effect parameters through the customization interface.
[0471] In some embodiments, a parameter prediction model for personalized learning of user-defined sound effect parameters is configured in the server. Illustratively, the customized sound effect parameters of multiple candidate accounts are obtained, the multiple customized sound effect parameters are input into the parameter prediction model to be trained, the parameter prediction model to be trained is iteratively trained to obtain a parameter prediction model, and the parameter prediction model is used to optimize the sound effect parameters generated by the system (e.g., the sound effect parameters of the aforementioned distance-type, direction-type, and spatial-type).
[0472] Sub-step 9: Adjusting the vocal audio based on the sound effect parameters, generating and broadcasting spatial vocal audio corresponding to the non-player character;
[0473] Among them, spatial human voice audio is used to express the perceptual effect of the audio produced by the main virtual character to the non-player character under the relative orientation relationship.
[0474] In an embodiment of the present application, human voice audio is adjusted based on the sound effect parameters corresponding to the non-player character, thereby generating spatial human voice audio corresponding to the non-player character. Specifically, spatial human voice audio is audio data corresponding to the non-player character obtained by adjusting the human voice audio using the sound effect parameters. In some embodiments, the spatial human voice audio includes the audio after the sound effect parameters have been adjusted, as well as at least one of the following data: audio duration information, audio playback start timestamp, audio playback end timestamp, audio playback condition information, and element identification information corresponding to the audio.
[0475] Optionally, when the sound effect parameters indicate an adjustment to the volume of the non-player character's vocal audio, the volume of the non-player character's vocal audio on at least one channel is adjusted based on the sound effect parameters. Optionally, when the sound effect parameters indicate an adjustment to the playback delay of the non-player character's vocal audio, the start time of the audio playback corresponding to the non-player character's vocal audio is adjusted based on the sound effect parameters. Optionally, when the sound effect parameters indicate an adjustment to the speed of the non-player character's vocal audio, the playback speed of the non-player character's vocal audio is adjusted based on the sound effect parameters. Optionally, when the sound effect parameters indicate a pitch of the non-player character's vocal audio, the frequencies in the spectrum of the non-player character's vocal audio are adjusted based on the sound effect parameters. When the sound effect parameters indicate a timbre of the non-player character's vocal audio, a filter corresponding to the non-player character is determined based on the sound effect parameters, and the timbre of the vocal audio is adjusted using the filter.
[0476] Schematically, the client plays the spatial human voice audio to realize the broadcast of non-player character feedback information.
[0477] The speech-to-text conversion phase of step 204 can be implemented as follows:
[0478] Sub-step 10: obtaining a plurality of first scene hot words corresponding to the first virtual scene where the target virtual character is located;
[0479] The target virtual character includes at least one of an NPC and a master virtual character.
[0480] For illustration, the controlling avatar is the avatar commanded by the player, while the NPC is the avatar that assists the controlling avatar in virtual games. Typically, a virtual game involves one controlling avatar and at least one NPC. There are certain differences between the controlling avatar and the NPC.
[0481] Optionally, the main virtual character is the virtual character mainly controlled by the player, and the NPC is the virtual character chosen by the player to command. For example, the main virtual character is the virtual character manually controlled by the player during the game, and the player's manual operations on the terminal interface are used to control the main virtual character; the NPC is the virtual character commanded by the player in voice form, and the player occasionally issues natural language commands in voice form to command the NPC.
[0482] In some embodiments, the first virtual scene is determined based on the target virtual character, and the target virtual character is at least one of an NPC and a master virtual character, which means that the process of determining the first virtual scene is implemented as at least one of the following.
[0483] (1) If the master virtual character is selected as the target virtual character, the virtual scene where the master virtual character is located is used as the first virtual scene;
[0484] (2) If an NPC is selected as the target virtual character, the virtual scene where the NPC is located is the first virtual scene; if there are multiple NPCs, the virtual scene including the most NPCs can be used as the first virtual scene. Alternatively, the NPC closest to the main virtual character can be selected as the target virtual character, and the virtual scene where the NPC is located can be used as the first virtual scene. Alternatively, the NPC with the largest character attribute value (such as at least one of the attribute values such as virtual health, virtual mana, and virtual defense) can be selected as the target virtual character, and the virtual scene where the NPC is located can be used as the first virtual scene.
[0485] (3) If the main virtual character and the NPC are selected as the target virtual characters, the virtual scene in which the main virtual character and the NPC are located can be used as the first virtual scene, etc.
[0486] In some embodiments, the target virtual character is determined based on the player's selection; alternatively, the target virtual character is a system default setting.
[0487] In some embodiments, a plurality of first scene hot words are determined based on the first virtual scene.
[0488] The plurality of first scene hot words are scene-related words of the first virtual scene, that is, there is an association relationship between the first scene hot words and the first virtual scene, and the first scene hot words are words that describe the first virtual scene.
[0489] Optionally, the first scene hot words include the scene status of the first virtual scene. For example, if the first virtual scene is a virtual restaurant, the first scene hot words include business status, rest status, closed status, etc.
[0490] Optionally, the first scene hot words include element names of virtual elements in the first virtual scene, such as: the first virtual scene is a virtual restaurant, and the first scene hot words include virtual tables, virtual chairs, virtual kitchen utensils, virtual plates, etc.
[0491] Optionally, the first scene hot words include interactive vocabulary that interacts with virtual elements in the first virtual scene. For example, if the first virtual scene is a virtual battlefield, the first scene hot words include attack, attack virtual character, defense, building barriers, etc.
[0492] In some embodiments, after receiving a natural language command, multiple first scene hot words are collected based on the first virtual scene; or, after receiving a natural language command, multiple first scene hot words corresponding to the first virtual scene are filtered from multiple scene hot words acquired in advance based on the first virtual scene, etc.
[0493] In an optional embodiment, the object perception range of the target virtual character in the first virtual scene is obtained.
[0494] Among them, the object perception range is the three-dimensional spatial range in which the target virtual character perceives the existence of other scene elements.
[0495] Optionally, when the target virtual character is the master virtual character, the object perception range of the master virtual character in the first virtual scene is obtained; when the target virtual character is an NPC, the object perception range of the NPC in the first virtual scene is obtained; or, when the target virtual character is the master virtual character and the NPC, the object perception range of the master virtual character in the first virtual scene and the object perception range of the NPC in the first virtual scene are obtained.
[0496] In some embodiments, the object perception range includes at least one of the following.
[0497] (1) The visual perception range of the target virtual character.
[0498] For illustration, visual perception range refers to the spatial range that an observer can visually perceive or understand. It is often used to describe the area that an organism or technological device can visually cover or perceive. The visual perception range of a target avatar is the three-dimensional spatial range perceived by the target avatar as the observer.
[0499] Optionally, based on the position and visual observation ability of the target virtual character in the first virtual scene, the visual perception range of the target virtual character is acquired.
[0500] (2) The auditory perception range of the target virtual character.
[0501] For illustration, auditory perception range refers to the spatial range that a listener can perceive or understand. It is typically used to describe the range of sounds that an organism or technological device can hear or perceive. The auditory perception range of a target avatar is the three-dimensional spatial range perceived by the target avatar through its auditory perception capabilities.
[0502] Optionally, based on the position and auditory perception ability of the target virtual character in the first virtual scene, the auditory perception range of the target virtual character is acquired.
[0503] (3) The olfactory perception range of the target virtual character.
[0504] Schematically, the olfactory perception range refers to the spatial range of the odor perceived or recognized by the sense of smell. The olfactory perception range of the target virtual character is the three-dimensional spatial range that the target virtual character can perceive through the sense of smell.
[0505] Optionally, the olfactory perception range of the target virtual character is acquired based on the position of the target virtual character in the first virtual scene and the olfactory perception ability.
[0506] (4) The target virtual character’s perceptual skills or the range of perception of the perceptual tool.
[0507] Illustratively, the perception skill is the ability of the target virtual character to perceive the first virtual scene, and the perception skill includes at least one of a perception acquisition form and a perception enhancement form.
[0508] Optionally, the "perception acquisition" mode is used to represent the process of acquiring a previously unpossessed perceptual skill. Once a target virtual character acquires a perceptual skill through the "perception acquisition" mode, they can use the perceptual skill to conduct targeted scene perception of the first virtual scene. Alternatively, the "perception enhancement" mode is used to represent the process of enhancing a certain perceptual ability through the acquisition of a perceptual skill.
[0509] Illustratively, the perception prop is a virtual prop used to perceive the first virtual scene, and the perception prop includes at least one of a perception acquisition form and a perception enhancement form.
[0510] Optionally, different perception skills or perception tools each correspond to a preset effective range, which represents a range that can be perceived when the perception skill or perception tool is applied, and the preset effective range represents the object perception range.
[0511] In some embodiments, the object perception range is a spherical range, an annular range, an irregular three-dimensional space range, etc. The shape of the object perception range is not limited here.
[0512] In an optional embodiment, a plurality of first scene hot words are acquired based on environmental perception information within the object perception range.
[0513] Illustratively, environmental perception information represents the environmental information perceived and acquired by the target virtual character within the object's perception range. Optionally, environmental perception information includes virtual elements identified by raycasting, and may also include regional structural information determined by analyzing the geometric structure and layout information of the virtual environment.
[0514] Schematically, after determining the object perception range corresponding to the target virtual character, based on the range of the object perception range in the first virtual scene, and the environmental perception information determined based on perception of the target virtual character within the object perception range, multiple first scene hot words representing the state of the target virtual character are obtained.
[0515] In an optional embodiment, the environmental perception information includes virtual elements; and the virtual elements in the first virtual scene that are within the object perception range are determined.
[0516] Schematically, the object perception range is a part of the three-dimensional space range in the first virtual scene; the first virtual scene includes a large number of virtual elements, and the virtual elements are elements that constitute the virtual scene, such as virtual elements including virtual ground, virtual buildings, virtual trees, virtual characters, virtual stones, etc.
[0517] In some embodiments, when the object perception range is determined, virtual elements within the object perception range are determined from a plurality of virtual elements corresponding to the first virtual scene.
[0518] In some embodiments, the element name of the virtual element is used as the first scene hot word; or, the action name of the interactive action corresponding to the virtual element is used as the first scene hot word.
[0519] Sub-step 11: converting the natural language command in voice form into text form based on the multiple first scenario hot words;
[0520] Indicatively, the hot words based on the first scene are words that are highly associated with the first virtual scene. Therefore, when hoping to command the NPC to better cooperate with the activities of the main virtual character through natural language commands, the hot words of the first scene can be used as constraints on the basis of the first virtual scene. In this way, in the process of converting the natural language commands into text form, the text content of the command analysis results is more consistent with the first virtual scene, avoiding obtaining command analysis results that are incompatible with the first virtual scene.
[0521] In an optional embodiment, when a natural language command is analyzed by a natural language analysis model, multiple speech units corresponding to the natural language command are acquired through an acoustic network.
[0522] Among them, phonetic units are the basic building blocks of vocabulary pronunciation.
[0523] In some embodiments, a command feature representation corresponding to a natural language command is extracted; and the command feature representation is analyzed through an acoustic network.
[0524] For example, using a command feature representation representing multiple acoustic features, the acoustic network aims to map a continuous sequence of acoustic features to a sequence of speech units, such as phonemes or segments. It learns to predict the likely sequence of speech units (phonemes or segments) given the acoustic features.
[0525] Optionally, the acoustic network outputs one or more possible speech unit sequences, each of which includes multiple speech units. Different speech unit sequences may contain different speech units, or may simply differ in the order of the speech units. The multiple speech unit sequences represent the most likely arrangements of speech units considered by the language network when understanding the input natural language command.
[0526] The natural language analysis model includes an acoustic network, a preset dictionary, and language sub-networks corresponding to at least two virtual scenes.
[0527] Schematically, the virtual environment includes at least two virtual scenes. For example, the virtual environment is a large virtual world, which includes a virtual kitchen 1, a virtual kitchen 2, a virtual office, etc., and each of the virtual kitchen 1, the virtual kitchen 2, the virtual office, etc. can be regarded as a virtual scene.
[0528] In an optional embodiment, based on a first virtual scene in which the target virtual character is located, a first language sub-network corresponding to the first virtual scene is determined from at least two language sub-networks.
[0529] Schematically, multiple virtual scenes each correspond to a language subnetwork, and different language subnetworks are used to analyze the corresponding virtual scenes in a targeted manner. For example, virtual scene A corresponds to language subnetwork 1, and virtual scene B corresponds to language subnetwork 2. When analysis is required based on virtual scene A, language subnetwork 1 is used for language analysis; when analysis is required based on virtual scene B, language subnetwork 2 is used for language analysis, and so on.
[0530] Schematically, the first virtual scene is a virtual scene among at least two virtual scenes. In addition to determining the first scene hot words based on the first virtual scene in which the target virtual character is located, the first language subnetwork corresponding to the first virtual scene is also determined from at least two language subnetworks based on the first virtual scene. The first language subnetwork is a subnetwork layer that performs semantic analysis on natural language commands based on the first virtual scene.
[0531] In an optional embodiment, the first language sub-network analyzes the sequence matching relationship between multiple speech units and selected words in a preset dictionary to convert the natural language command in speech form into text form.
[0532] Optionally, different language sub-networks are language sub-networks pre-trained based on corresponding virtual scenes, wherein the first language sub-network is a language sub-network pre-trained based on the first virtual scene.
[0533] In some embodiments, matching relationships between multiple speech units and selected words in a preset dictionary are analyzed to obtain multiple candidate word sequences.
[0534] The selected words include at least a plurality of first-scene hot words and common words.
[0535] Illustratively, the process of analyzing matching relationships represents the process of analyzing the matching between selected vocabulary and speech units. Illustratively, each selected vocabulary corresponds to at least one speech unit. By analyzing multiple speech units and the selected vocabulary, the conversion probability of converting a speech unit to the selected vocabulary is determined. The conversion probability represents the probability of matching between the selected vocabulary and the speech unit. A greater conversion probability indicates that the speech unit is closer to the selected vocabulary, i.e., a greater likelihood that the speech unit can express the selected vocabulary, and thus a stronger matching relationship between the selected vocabulary and the speech unit.
[0536] Optionally, multiple speech units form at least one speech unit sequence, and based on a matching analysis process between the speech units and the selected vocabulary, at least one candidate vocabulary sequence corresponding to the at least one speech unit sequence is determined, and multiple candidate vocabulary sequences are obtained.
[0537] In some embodiments, the sequence semantics of multiple candidate vocabulary sequences are analyzed by the first language sub-network, and at least one candidate vocabulary sequence is obtained from the multiple candidate vocabulary sequences as a text form for natural language command conversion.
[0538] Schematically, sequence semantics is used to represent the semantic information expressed by the candidate vocabulary sequence, which can not only reflect the semantic changes of at least one word in the candidate vocabulary sequence, but also reflect the sentence semantics of the entire candidate vocabulary sequence.
[0539] Schematically, the natural language command in text form is called the command analysis result; multiple candidate vocabulary sequences are evaluated separately by the first language sub-network, so that at least one candidate vocabulary sequence that is more in line with the current semantic situation is taken as the command analysis result, that is, the natural language command in speech form is converted into text form.
[0540] It should be noted that the above are merely illustrative examples and are not limited to the embodiments of the present application.
[0541] Those skilled in the art will appreciate that the above embodiments may be implemented independently, or the above embodiments may be freely combined to form new embodiments to implement the scene object processing method of the present application.
[0542] Figure 15 The following is a block diagram of a scene object processing device provided by an exemplary embodiment of the present application. The device includes:
[0543] An acquisition module 810 is configured to acquire attribute text of a scene object in a virtual scene, wherein the attribute text is used to describe inherent attributes of the scene object in the virtual scene;
[0544] The acquisition module 810 is further configured to acquire an appearance image of the scene object in the virtual scene, wherein the appearance image is used to describe the style of the scene object;
[0545] The processing module 820 is used to call a multimodal model to perform prediction on the attribute text and the appearance image of the scene object to obtain a visual label of the scene object, where the visual label is used to describe the visual features of the scene object in at least one dimension.
[0546] In an optional implementation of this embodiment, the multimodal model includes a visual question answering model, and the processing module 820 is further configured to:
[0547] Constructing a question sentence for the appearance image, wherein the question sentence carries the attribute text of the scene object;
[0548] The appearance image of the scene object and the question statement are input into the visual question answering model to obtain an answer statement, and the answer statement is used as the visual label of the scene object.
[0549] In an optional implementation of this embodiment, the question statement includes at least two sub-statements; the processing module 820 is further configured to:
[0550] Inputting the appearance image of the scene object and the first sentence of the at least two sub-sentences into the visual question answering model to obtain a first answer sub-sentence;
[0551] Repeat the above steps until at least two answer sub-sentences corresponding one-to-one to the at least two sub-sentences are obtained, wherein the at least two sub-sentences are used to inquire about the visual features of the appearance image from multiple dimensions;
[0552] Sentence aggregation is performed on the at least two answer sub-sentences to extract the visual label of the scene object.
[0553] In an optional implementation of this embodiment, the processing module 820 is further configured to:
[0554] Acquiring desired information of the scene object, where the desired information is used to indicate desired description dimensions of the scene object in the visual tag and / or desired format of the visual tag;
[0555] constructing the question sentence of the appearance image according to the expected information and the attribute text;
[0556] Among them, the first sub-part of the question statement is the supplementary introduction information of the scene object, carrying the attribute text of the scene object; the second sub-part of the question statement is the answer guidance statement for the visual question answering model, carrying the expected information.
[0557] In an optional implementation of this embodiment, the multimodal model includes an image description model, and the processing module 820 is further configured to:
[0558] Inputting the appearance image of the scene object into the image description model to predict a description text of the scene object;
[0559] Sentence aggregation is performed on the description text and the attribute text to extract the visual label of the scene object.
[0560] In an optional implementation of this embodiment, the attribute text of the scene object includes at least one of a name of the scene object in the virtual scene and a size of the scene object in the virtual scene;
[0561] And / or, the appearance image of the scene object includes images obtained by observing the scene object from at least two viewing angles.
[0562] In an optional implementation of this embodiment, the processing module 820 is further configured to:
[0563] Performing pseudo-colloquial rewriting on the visual label of the scene object to obtain a matching label that conforms to the spoken expression of natural language.
[0564] In an optional implementation of this embodiment, the processing module 820 is further configured to:
[0565] The visual label of the scene object is input into a natural language model to predict the matching label that conforms to the spoken expression of the natural language, and the natural language model carries prior knowledge of the spoken expression of the natural language.
[0566] In an optional implementation of this embodiment, the processing module 820 is further configured to:
[0567] Obtaining a first sample label pair, the first sample label pair including a first label before being rewritten into a pseudo-colloquial language and a second label obtained after being rewritten into a pseudo-colloquial language;
[0568] constructing a rewriting guidance sentence based on the first sample label pair and the visual label, wherein the rewriting guidance sentence has natural semantics of rewriting the visual label with reference to the first sample label pair;
[0569] The rewriting guide sentence is input into a natural language model to predict the matching label that conforms to the spoken expression of natural language.
[0570] In an optional implementation of this embodiment, the acquisition module 810 is further configured to:
[0571] The spatial position of the scene object in the virtual scene is acquired, and the spatial position of the scene object is determined as auxiliary information of the visual label of the scene object.
[0572] In an optional implementation of this embodiment, the processing module 820 is further configured to:
[0573] Obtaining at least one of the coordinate position, orientation information, bounding box information, and cover point information of the scene object in the virtual scene;
[0574] Among them, the coordinate position is used to indicate the position of the scene object in the virtual scene, the orientation information is used to indicate the direction facing the scene object in the virtual scene, the bounding box information is used to indicate the size of the scene object in the virtual scene, and the cover point information indicates the recommended virtual character standing position when the virtual character approaches the scene object.
[0575] In an optional implementation of this embodiment, the processing module 820 is further configured to construct a spatial information library of the virtual scene based on the visual labels of the plurality of scene objects in the virtual scene;
[0576] The acquisition module 810 is further configured to acquire a natural language command;
[0577] The processing module 820 is further configured to select a first scene object corresponding to the natural language command from a plurality of scene objects in the virtual scene based on a similarity between the natural language command and the visual tags in the spatial information library;
[0578] The processing module 820 is further configured to execute a virtual activity on the first scene object based on the natural language command.
[0579] In an optional implementation of this embodiment, the natural language command includes multiple input words; the processing module 820 is further configured to:
[0580] determining a similarity score between each word in the plurality of input words and the visual label of the first scene object;
[0581] Determining a sum of a plurality of similarity scores as a similarity between the first scene object and the natural language command, the plurality of similarity scores corresponding one-to-one to the plurality of input words;
[0582] When the similarity between the first scene object and the natural language command exceeds the similarity between other scene objects in the spatial information library and the natural language command, the first scene object is determined as the object corresponding to the natural language command.
[0583] In an optional implementation of this embodiment, the processing module 820 is further configured to construct a spatial information library of the virtual scene based on appearance images of multiple scene objects in the virtual scene;
[0584] The acquisition module 810 is further configured to acquire a natural language command;
[0585] The processing module 820 is further configured to invoke a similarity matching network to predict a degree of similarity between the natural language command and the appearance images of the plurality of scene objects in the spatial information library, and obtain a second scene object among the plurality of scene objects having the highest similarity to the natural language command;
[0586] The processing module 820 is further configured to execute a virtual activity on the second scene object based on the natural language command.
[0587] It should be noted that the device provided in the above embodiment only uses the division of the above-mentioned functional modules as an example to implement its functions. In actual applications, the above-mentioned functions can be assigned to different functional modules according to actual needs, that is, the content structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0588] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method; the technical effects achieved by each module performing operations are the same as the technical effects in the embodiment of the method, and will not be elaborated here.
[0589] An embodiment of the present application also provides a computer device, which includes: a processor and a memory, wherein a computer program is stored in the memory; the processor is used to execute the computer program in the memory to implement the scene object processing method provided by the above-mentioned method embodiments.
[0590] Figure 16The following is a block diagram of a terminal structure provided by an exemplary embodiment of the present application. Terminal 1900 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Terminal 1900 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other similar names.
[0591] Typically, terminal 1900 includes: a processor 1901 and a memory 1902. Processor 1901 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 1901 may be implemented in at least one of the following hardware forms: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). Processor 1901 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, processor 1901 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the display screen. In some embodiments, processor 1901 may also include an AI (Artificial Intelligence) processor, which is responsible for processing computing operations related to machine learning.
[0592] Memory 1902 may include one or more computer-readable storage media, which may be non-transitory. Memory 1902 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in memory 1902 is used to store at least one instruction, which is executed by processor 1901 to implement the scene object processing method provided in the method embodiment of the present application.
[0593] In some embodiments, terminal 1900 may optionally include a peripheral device interface 1903 and at least one peripheral device. The processor 1901, memory 1902, and peripheral device interface 1903 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 1903 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 1904, a touchscreen display 1905, a camera assembly 1906, an audio circuit 1907, and a power supply 1908.
[0594] The peripheral device interface 1903 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 1901 and the memory 1902. In some embodiments, the processor 1901, the memory 1902, and the peripheral device interface 1903 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1901, the memory 1902, and the peripheral device interface 1903 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0595] RF circuit 1904 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. RF circuit 1904 communicates with communication networks and other communication devices via electromagnetic signals. RF circuit 1904 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. RF circuit 1904 may optionally include an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. RF circuit 1904 may communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, RF circuit 1904 may also include circuitry related to Near Field Communication (NFC), although this application does not limit this.
[0596] The touch screen display 1905 is used to display a UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. The touch screen display 1905 also has the ability to collect touch signals on or above the surface of the touch screen display 1905. The touch signal can be input as a control signal to the processor 1901 for processing. At this time, the touch screen display 1905 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there can be one touch screen display 1905, set on the front panel of the terminal 1900; in other embodiments, there can be at least two touch screen displays 1905, respectively set on different surfaces of the terminal 1900 or in a folding design; in still other embodiments, the touch screen display 1905 can be a flexible display, set on a curved surface or a folding surface of the terminal 1900. Even more, the touch screen display 1905 can be set into a non-rectangular irregular shape, that is, a special-shaped screen. The touch screen display 1905 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0597] The camera assembly 1906 is used to capture images or videos. Optionally, the camera assembly 1906 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting with VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 1906 may also include a flash. The flash can be a monochrome temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.
[0598] The audio circuit 1907 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 1901 for processing, or input into the radio frequency circuit 1904 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 1900. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 1901 or the radio frequency circuit 1904 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as distance measurement. In some embodiments, the audio circuit 1907 may also include a headphone jack.
[0599] Power supply 1908 is used to power various components in terminal 1900. Power supply 1908 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 1908 includes a rechargeable battery, the rechargeable battery can be wired or wirelessly rechargeable. A wired rechargeable battery is charged via a wired line, while a wireless rechargeable battery is charged via a wireless coil. The rechargeable battery can also support fast charging technology.
[0600] In some embodiments, the terminal 1900 further includes one or more sensors 1909 . The one or more sensors 1909 include, but are not limited to, an acceleration sensor 1910 , a gyroscope sensor 1911 , a pressure sensor 1912 , an optical sensor 1913 , and a proximity sensor 1914 .
[0601] The accelerometer 1910 can detect the magnitude of acceleration on the three coordinate axes of the coordinate system established by the terminal 1900. For example, the accelerometer 1910 can be used to detect the components of gravity acceleration on the three coordinate axes. The processor 1901 can control the touch screen display 1905 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer 1910. The accelerometer 1910 can also be used to collect game or user motion data. The gyroscope sensor 1911 can detect the body orientation and rotation angle of the terminal 1900. The gyroscope sensor 1911 can work with the accelerometer 1910 to collect the user's 3D movements on the terminal 1900. Based on the data collected by the gyroscope sensor 1911, the processor 1901 can implement the following functions: motion sensing (such as changing the UI based on the user's tilt operation), image stabilization during shooting, game control, and inertial navigation.
[0602] The pressure sensor 1912 can be provided on the side frame of the terminal 1900 and / or below the touch screen display 1905. When the pressure sensor 1912 is provided on the side frame of the terminal 1900, it can detect the user's gripping signal of the terminal 1900, and the processor 1901 can perform left and right hand recognition or shortcut operations based on the gripping signal collected by the pressure sensor 1912. When the pressure sensor 1912 is provided below the touch screen display 1905, the processor 1901 controls the operable controls on the UI interface based on the user's pressure operation on the touch screen display 1905. Operable controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0603] Optical sensor 1913 is used to detect ambient light intensity. In one embodiment, processor 1901 can control the display brightness of touchscreen display 1905 based on the ambient light intensity detected by optical sensor 1913. Specifically, when the ambient light intensity is high, the display brightness of touchscreen display 1905 is increased; when the ambient light intensity is low, the display brightness of touchscreen display 1905 is decreased. In another embodiment, processor 1901 can also dynamically adjust the shooting parameters of camera assembly 1906 based on the ambient light intensity detected by optical sensor 1913.
[0604] Proximity sensor 1914, also known as a distance sensor, is typically located on the front panel of terminal 1900. Proximity sensor 1914 is used to detect the distance between the user and the front of terminal 1900. In one embodiment, when proximity sensor 1914 detects that the distance between the user and the front of terminal 1900 is gradually decreasing, processor 1901 controls touchscreen display 1905 to switch from the screen-on state to the screen-off state. When proximity sensor 1914 detects that the distance between the user and the front of terminal 1900 is gradually increasing, processor 1901 controls touchscreen display 1905 to switch from the screen-off state to the screen-on state.
[0605] Those skilled in the art will appreciate that the above structure does not limit the terminal 1900 , and the terminal 1900 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0606] In an exemplary embodiment, a chip is further provided. The chip includes a programmable logic circuit and / or program instructions. When the chip runs on a computer device, it is used to implement the method for processing scene objects described in the above aspects.
[0607] In an exemplary embodiment, a computer program product is also provided. The computer program product includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions to implement the scene object processing methods provided in the above-described method embodiments.
[0608] In an exemplary embodiment, a computer-readable storage medium is further provided, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the scene object processing methods provided by the above-mentioned method embodiments.
[0609] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0610] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0611] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A method for processing scene objects, characterized in that: The method comprises: Acquire attribute text of a scene object in a virtual scene, wherein the attribute text is used to introduce inherent attributes of the scene object in the virtual scene, and the attribute text of the scene object includes at least one of a name of the scene object in the virtual scene and a size of the scene object in the virtual scene; Acquire an appearance image of the scene object in the virtual scene, where the appearance image is used to describe the style of the scene object; Acquiring desired information of the scene object, where the desired information is used to indicate a desired description dimension of the scene object in a visual tag and / or a desired format of the visual tag; Constructing a question sentence about the appearance image based on the desired information and the attribute text, wherein the first subpart of the question sentence is supplementary introduction information about the scene object and carries the attribute text of the scene object; and the second subpart of the question sentence is an answer guidance sentence for the visual question answering model and carries the desired information; Inputting the appearance image of the scene object and the question sentence into the visual question answering model to obtain an answer sentence, and using the answer sentence as the visual label of the scene object, wherein the visual label is used to describe the visual features of the scene object in at least one dimension; Acquiring a spatial position of the scene object in the virtual scene, and determining the spatial position of the scene object as auxiliary information of a visual label of the scene object, wherein the spatial position is used to indicate a deployment status of the scene object in the virtual scene; In which, the spatial position includes at least one of the coordinate position, orientation information, bounding box information, and cover point information of the scene object in the virtual scene, the coordinate position is used to indicate the position of the center point or preset point on the scene object in the virtual scene, the orientation information is used to indicate the direction facing the scene object in the virtual scene, the bounding box information is used to indicate the size of the scene object in the virtual scene, and the cover point information indicates the recommended virtual character standing position when the virtual character approaches the scene object.
2. The method according to claim 1, characterized in that The question statement includes at least two sub-statements; inputting the appearance image of the scene object and the question statement into the visual question answering model to obtain an answer statement, and using the answer statement as the visual label of the scene object includes: Inputting the appearance image of the scene object and the first sentence of the at least two sub-sentences into the visual question answering model to obtain a first answer sub-sentence; Repeat the above steps until at least two answer sub-sentences corresponding one-to-one to the at least two sub-sentences are obtained, wherein the at least two sub-sentences are used to inquire about the visual features of the appearance image from multiple dimensions; Sentence aggregation is performed on the at least two answer sub-sentences to extract the visual label of the scene object.
3. The method according to claim 1, characterized in that The method also includes a picture description model, and the method further includes: Inputting the appearance image of the scene object into the image description model to predict a description text of the scene object; Sentence aggregation is performed on the description text and the attribute text to extract the visual label of the scene object.
4. The method according to any one of claims 1 to 3, characterized in that: The appearance image of the scene object includes images obtained by observing the scene object from at least two viewing angles.
5. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Performing pseudo-colloquial rewriting on the visual label of the scene object to obtain a matching label that conforms to the spoken expression of natural language.
6. The method according to claim 5, characterized in that The performing pseudo-colloquial rewriting on the visual label of the scene object to obtain a matching label that conforms to the spoken expression of a natural language includes: The visual label of the scene object is input into a natural language model to predict the matching label that conforms to the spoken expression of the natural language, and the natural language model carries prior knowledge of the spoken expression of the natural language.
7. The method according to claim 6, characterized in that Inputting the visual label of the scene object into a natural language model to predict the matching label that conforms to the spoken expression of the natural language includes: Obtaining a first sample label pair, the first sample label pair including a first label before being rewritten into a pseudo-colloquial language and a second label obtained after being rewritten into a pseudo-colloquial language; constructing a rewriting guidance sentence based on the first sample label pair and the visual label, wherein the rewriting guidance sentence has natural semantics of rewriting the visual label with reference to the first sample label pair; The rewriting guide sentence is input into a natural language model to predict the matching label that conforms to the spoken expression of natural language.
8. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Constructing a spatial information library of the virtual scene according to visual labels of a plurality of scene objects in the virtual scene; Get natural language commands; Based on the similarity between the natural language command and the visual tags in the spatial information library, screening out a first scene object corresponding to the natural language command from a plurality of scene objects in the virtual scene; Based on the natural language command, a virtual activity is performed on the first scene object.
9. The method according to claim 8, characterized in that The natural language command includes a plurality of input words; and the selecting, based on the similarity between the natural language command and the visual tags in the spatial information library, a first scene object corresponding to the natural language command from a plurality of scene objects in the virtual scene, comprises: determining a similarity score between each word in the plurality of input words and the visual label of the first scene object; determining a sum of a plurality of similarity scores as a similarity between the first scene object and the natural language command, the plurality of similarity scores corresponding one-to-one to the plurality of input words; When the similarity between the first scene object and the natural language command exceeds the similarity between other scene objects in the spatial information library and the natural language command, the first scene object is determined as the object corresponding to the natural language command.
10. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Constructing a spatial information library of the virtual scene according to appearance images of a plurality of scene objects in the virtual scene; Get natural language commands; calling a similarity matching network to predict the similarity between the natural language command and the appearance images of the plurality of scene objects in the spatial information library, and obtaining a second scene object having the highest similarity to the natural language command among the plurality of scene objects; Based on the natural language command, a virtual activity is performed on the second scene object.
11. A device for processing scene objects, characterized in that: The device comprises: an acquisition module, configured to acquire attribute text of a scene object in a virtual scene, wherein the attribute text is used to introduce inherent attributes of the scene object in the virtual scene, and the attribute text of the scene object includes at least one of a name of the scene object in the virtual scene and a size of the scene object in the virtual scene; The acquisition module is further configured to acquire an appearance image of the scene object in the virtual scene, wherein the appearance image is used to describe the style of the scene object; a processing module, configured to obtain desired information of the scene object, wherein the desired information is used to indicate desired description dimensions of the scene object in the visual tag and / or desired format of the visual tag; Constructing a question sentence about the appearance image based on the desired information and the attribute text, wherein the first subpart of the question sentence is supplementary introduction information about the scene object and carries the attribute text of the scene object; and the second subpart of the question sentence is an answer guidance sentence for the visual question answering model and carries the desired information; Inputting the appearance image of the scene object and the question sentence into the visual question answering model to obtain an answer sentence, and using the answer sentence as the visual label of the scene object, wherein the visual label is used to describe the visual features of the scene object in at least one dimension; Acquiring a spatial position of the scene object in the virtual scene, and determining the spatial position of the scene object as auxiliary information of a visual label of the scene object, wherein the spatial position is used to indicate a deployment status of the scene object in the virtual scene; In which, the spatial position includes at least one of the coordinate position, orientation information, bounding box information, and cover point information of the scene object in the virtual scene, the coordinate position is used to indicate the position of the center point or preset point on the scene object in the virtual scene, the orientation information is used to indicate the direction facing the scene object in the virtual scene, the bounding box information is used to indicate the size of the scene object in the virtual scene, and the cover point information indicates the recommended virtual character standing position when the virtual character approaches the scene object.
12. A computer device, characterized in that: The computer device includes: a processor and a memory, wherein the memory stores at least one program; the processor is used to execute the at least one program in the memory to implement the scene object processing method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that The readable storage medium stores executable instructions, and the executable instructions are loaded and executed by the processor to implement the scene object processing method according to any one of claims 1 to 10.
14. A computer program product, characterized in that The computer program product includes computer instructions, which are stored in a computer-readable storage medium. The processor reads and executes the computer instructions from the computer-readable storage medium to implement the scene object processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Static label identification model training method, static label identification method and device
CN116051921A
Display device and visual question and answer method
CN116301337A