Virtual character interaction method and device
By identifying the intention of the user's question and obtaining multimodal data, and processing these data using a large language model, virtual characters can more accurately answer questions involving the user's environment or state, solving the problem of limited interaction experience in the prior art.
Patent Information
- Application Number
- CN202510174167.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-02-17
AI Technical Summary
In the prior art, the interaction between people and virtual characters is carried out in the form of voice Q&A, and questions that require visual observation environment cannot be answered accurately, which affects the user's interactive experience.
By identifying the user's intent to ask questions, determine whether it is an intention related to the user's environment or status, obtain corresponding multimodal data (such as images, sounds, actions, etc.), and use a large language model to process these data to generate more accurate responses.
It realizes that virtual characters answer questions involving user environment or status more realistically and accurately, improves the interactive experience, and avoids unnecessary waste of data collection and processing resources.
Smart Images

Figure CN120105009A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a virtual character interaction method and device. Background Art
[0002] Virtual characters are appearing in more and more fields, and the interaction between people and virtual characters in more and more scenarios has become a new mode of human-computer interaction. At present, the interaction between people and virtual characters is carried out in the form of voice questions and answers. In the process of interaction between users and virtual characters, the virtual characters cannot answer some questions of users or answer them inaccurately, which affects the user's interactive experience. Summary of the invention
[0003] In view of this, the present disclosure proposes a virtual character interaction method, device, electronic device, storage medium and computer program product.
[0004] According to one aspect of the present disclosure, a virtual character interaction method is provided, the method comprising:
[0005] In response to a question message sent by a user to a virtual character running in an application, determining an intention corresponding to the question message;
[0006] Determine whether the intention corresponding to the question information is a preset intention, and if the intention corresponding to the question information is the preset intention, trigger the application to obtain target information matching the preset intention; wherein the preset intention is related to the environment in which the user is located and / or the state of the user;
[0007] The target multimodal data is processed by the target model to generate response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data;
[0008] The response information is broadcasted to the user through the virtual character.
[0009] In a possible implementation, the target information includes one or more of image information of the user, image information of the user's environment, sound information of the user's environment, movement information of the user, tactile information of the user, and physiological information of the user.
[0010] In a possible implementation, the preset intention is related to the state of the user;
[0011] The triggering the application to obtain target information matching the preset intention includes:
[0012] Start the camera and control it to collect the user's image information;
[0013] and / or,
[0014] At least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device is obtained.
[0015] In a possible implementation, the preset intention is related to the environment in which the user is located;
[0016] The triggering the application to obtain target information matching the preset intention includes:
[0017] Start the camera and control the camera to collect image information of the user's environment;
[0018] and / or,
[0019] Start the microphone and control it to collect sound information of the user's environment.
[0020] In a possible implementation, the processing of the target multimodal data by using the target model also includes:
[0021] A preset model matching the preset intent is selected from a plurality of preset models as the target model; wherein different preset models are used to process information matching different intents.
[0022] In a possible implementation, the processing of the target multimodal data by the target model to generate response information includes:
[0023] The question information and / or the target information are processed by the target model to extract the attribute characteristics of the user; and the question information and the target information are processed to extract multimodal characteristics; wherein the attribute characteristics include: age and / or emotion;
[0024] The response information is generated based on the attribute feature and the multimodal feature.
[0025] In a possible implementation, the controlling the camera to collect image information of the user includes:
[0026] Control the camera to take at least one shot and perform target detection on the shot image until it is detected that the target body part involved in the question information is included in the currently shot image, wherein, when it is detected that the currently shot image does not include the target body part, adjust the shooting parameters of the camera, and / or prompt the user to adjust the posture of himself or the camera.
[0027] In a possible implementation manner, determining the intention corresponding to the question information includes:
[0028] Obtain the user's historical questions stored in the application;
[0029] In combination with the historical questions, the intention corresponding to the question information is determined.
[0030] In a possible implementation, the broadcasting of the response information to the user by the virtual character includes:
[0031] Determining a target state of the virtual character when making a broadcast according to the target information;
[0032] The virtual character is controlled to broadcast the response information to the user in the target state.
[0033] According to another aspect of the present disclosure, a virtual character interaction device is provided, the device comprising:
[0034] A questioning module, configured to determine the intention corresponding to the questioning information in response to the questioning information asked by the user to the virtual character running in the application program;
[0035] an acquisition module, used to determine whether the intention corresponding to the question information is a preset intention, and if the intention corresponding to the question information is the preset intention, trigger the application to acquire target information matching the preset intention; wherein the preset intention is related to the environment in which the user is located and / or the state of the user;
[0036] A response module, used to process the target multimodal data through a target model to generate response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data;
[0037] The reporting module is used to report the response information to the user through the virtual character.
[0038] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0039] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0040] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0041] Through various aspects of the present disclosure, in response to the question information of the user asking the virtual character running in the application, the intention corresponding to the question information is determined; it is judged whether the intention corresponding to the question information is a preset intention, and when the intention corresponding to the question information is the preset intention, the application is triggered to obtain the target information matching the preset intention; wherein the preset intention is related to the environment in which the user is located and / or the state of the user; the target multimodal data is processed by the target model to generate the response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data; the response information is broadcast to the user by the virtual character. In this way, based on the multimodal interaction technology and the intention recognition technology, the target information can be obtained when the intention corresponding to the user's question information is an intention related to the environment in which the user is located and / or the state of the user; and then the large language model is used to process the multimodal data, and the complementarity and correlation between different modal data are used to improve the virtual character's ability to understand the user's question information, so that the questions related to the environment in which the user is located or the state of the user can be answered more realistically and accurately, and more accurate, natural and reasonable response information can be generated, so as to realize the immersive interaction of hyper-realistic virtual characters. At the same time, unlike directly acquiring multimodal data, the target information is acquired only when the intention corresponding to the user's question information is the preset intention, thereby avoiding the waste of data collection and processing resources caused by directly acquiring the target information.
[0042] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0044] Figure 1 A schematic diagram showing a scenario of interaction between a user and a virtual character according to an embodiment of the present disclosure.
[0045] Figure 2 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown.
[0046] Figure 3 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown.
[0047] Figure 4 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown.
[0048] Figure 5 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown.
[0049] Figure 6 A flow chart of a method for virtual character interaction according to an embodiment of the present disclosure is shown.
[0050] Figure 7 A structural diagram of a virtual character interaction device according to an embodiment of the present disclosure is shown.
[0051] Figure 8 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0052] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.
[0053] References to "one embodiment" or "some embodiments" etc. described in this specification mean that one or more embodiments of the present disclosure include specific features, structures or characteristics described in conjunction with the embodiment. Thus, the phrases "exemplary", "in one embodiment", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0054] In the present disclosure, "at least one" means one or more, and "plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: including the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or plural.
[0055] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.
[0056] Figure 1 A schematic diagram showing a scene of interaction between a user and a virtual character according to an embodiment of the present disclosure is shown; Figure 1 As shown, the electronic device is configured with a display screen, and the user can interact with the virtual character on the display screen. During the interaction, the user asks questions to the virtual character, and the electronic device generates response information after understanding the question information, and feeds back the response information to the user in real time through the virtual character on the display screen.
[0057] In the related art, the interaction between users and virtual characters is carried out in the form of voice questions and answers; however, the virtual characters are unable to answer or answer inaccurately some of the users' questions. For example, when users ask questions such as "How do you think I'm wearing today?" or "Is the light in the room a little dim?" which require visual observation of the environment to answer accurately, the virtual characters are unable to give true and accurate answers, which affects the interactive experience between users and the virtual characters.
[0058] In order to solve the above technical problems, an embodiment of the present disclosure provides a virtual character interaction method (see below for detailed description). This virtual character interaction method is based on multimodal interaction technology and intention recognition technology, and can answer questions related to the user's environment or the user's status more realistically and accurately, thereby realizing immersive interaction with hyper-realistic virtual characters.
[0059] Exemplarily, the virtual character interaction method provided in the embodiment of the present disclosure can be configured in an electronic device in the form of software and / or hardware so that the electronic device can perform the virtual character interaction function. In the following embodiments, the execution subject is an electronic device as an example for explanation. Among them, the electronic device can be any device with computing power, for example, a personal computer (PC), a mobile terminal, a server, etc., and the mobile terminal can be a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, etc., with various operating systems, touch screens and / or display screens; the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and basic cloud computing services such as big data and artificial intelligence platforms.
[0060] It should be noted that the above-mentioned application scenarios described in the embodiments of the present disclosure are for the purpose of more clearly illustrating the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. A person of ordinary skill in the art may know that the technical solutions provided by the embodiments of the present disclosure are also applicable to similar technical problems in response to the emergence of other similar or new scenarios, such as live broadcast rooms, metaverses, etc.
[0061] Figure 2 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown. Figure 2 As shown, the method may include the following steps:
[0062] Step 201: In response to a question message sent by a user to a virtual character running in an application, determine an intention corresponding to the question message.
[0063] Among them, the application can be any computer program that can realize the operation of a virtual character and interact with a user; the application can be in the form of an application that can be independently installed in an electronic device and run, or it can be a small program that depends on other applications or operating environments, or it can be a plug-in. The specific form of the application is not limited in the embodiments of the present disclosure. For example, the application can be various types of applications such as video applications and e-commerce platforms, which can respond to the virtual character's awakening conditions to awaken the virtual character to run, such as the user triggering a control to awaken the virtual character, or automatically awakening the virtual character on certain pages.
[0064] Among them, virtual characters can be digital people or other non-human images, and there is no limitation on this.
[0065] It is understandable that one or more rounds of dialogues may be conducted during the interaction between a user and a virtual character running in an application, wherein during each round of dialogue, the electronic device may obtain the user's latest question information and determine the intention corresponding to the question information.
[0066] Exemplarily, the user's question information may be the user's voice, which may also be referred to as question voice. For example, the electronic device may obtain the user's voice through a configured microphone, so that the user can ask questions to the virtual character running in the application in the form of voice; or the user's question information may be text input by the user, which may also be referred to as question text. For example, the electronic device may obtain the text input by the user through a configured display screen, so that the user can ask questions to the virtual character running in the application in the form of text. The user's question information may also be provided in other forms such as pictures and documents, which are not limited in the present disclosure.
[0067] For example, the user's question information can be processed by a trained intent recognition model based on a machine learning algorithm to determine the intent corresponding to the user's question information. For example, the user's question information can be input into a pre-trained intent recognition model, and the intent recognition model determines the intent corresponding to the question information in a predefined intent category based on the extracted features. The intent can reflect the user's real intention and needs. In other words, the intent reflects the user's idea of achieving a certain purpose by asking questions. Among them, the intent recognition model and the predefined intent categories can be set according to needs, and there is no limitation on this. The intent recognition model can be trained through existing training methods. For example, the question information sample can be input into the intent recognition model to obtain the predicted intent category, and the loss between the predicted intent category and the intent category label corresponding to the question information sample is calculated. The parameters of the intent recognition model are adjusted based on the loss to obtain the trained intent recognition model; as an example, the predefined intent categories may include: chat intent, professional field intent, intent related to the user's environment, intent related to the user's status, etc.; chat intent means that the user wants to chat with the virtual character. For example, the user's question information is "Tell me a joke", and the corresponding intent is the chat intent; professional field intent means that the user wants to obtain specific fields (such as law) from the virtual character. , physics, biology, finance, chemistry, mathematics, literature, etc.), for example, the user's question information is "What is the expression of the law of universal gravitation", the corresponding intention is the physics field intention; the intention related to the user's environment means that the user wants to communicate with the virtual character regarding the user's current environment. For example, the user's question information is "Is it a little dim to read under the current light?", which involves the intensity of light in the user's current environment, and the corresponding intention is the intention related to the user's environment; the intention related to the user's status means that the user wants to communicate with the virtual character regarding the user's current status (such as the user's own clothing, makeup, height, weight, blood pressure, movements, etc.), for example, the user's question information is "How about my outfit?", which involves the user's own clothing, and the corresponding intention is the intention related to the user's status.
[0068] Exemplarily, the input data of the trained intent recognition model can be data in text format. When the user's question information is in the form of question voice, the electronic device can use automatic speech recognition technology (Automatic Speech Recognition, ASR) to convert the user's question information in the form of question voice into question text, and then input the question text into the trained intent recognition model to determine the intention corresponding to the question voice.
[0069] In a possible implementation, determining the intent corresponding to the question information includes: obtaining the user's historical questions stored in the application; and determining the intent corresponding to the question information in combination with the historical questions. Considering that different users have different language expression habits, such as using inverted sentences, having accents, swallowing words, using dialects or slang, and the user's historical questions can reflect the user's language expression habits, therefore, in combination with the user's historical questions, the intent corresponding to the user's question information can be more accurately determined.
[0070] For example, the language expression habits that users are used to using can be analyzed from the user's historical questions. For example, inverted sentences are frequently used in historical questions, and it can be known that users are used to using inverted sentences. When determining the intention corresponding to the user's question information, if it is determined that there is an inverted sentence, the inverted sentence can be adjusted to a normal sentence and then input into the intention recognition model to more accurately determine the intention corresponding to the question information. For another example, if swallowing words frequently occurs in historical questions, when determining the intention corresponding to the user's question information, if it is determined that there is a swallowing word phenomenon, the question information can be completed and input into the intention recognition model, so as to more accurately determine the intention corresponding to the question information based on comprehensive sentences.
[0071] Step 202: determine whether the intention corresponding to the question information is a preset intention, and if the intention corresponding to the question information is the preset intention, trigger the application to obtain target information matching the preset intention; wherein the preset intention is related to the user's environment and / or the user's status.
[0072] Exemplarily, the preset intention may be an intention related to the user's environment or an intention related to the user's status in the above-mentioned predefined intention categories; thereby, it can be determined whether the intention corresponding to the question information determined from the above-mentioned predefined intention categories is the preset intention.
[0073] Among them, when the preset intention is an intention related to the user's state, that is, when the preset intention is related to the user's state, if the intention corresponding to the user's question information is the preset intention, the application needs to combine the user's current state to accurately and comprehensively understand the user's question information; for example, the user's question information is "How is my outfit?", the intention corresponding to the question information is an intention related to the user's state, and it is necessary to obtain the user's actual clothing situation in order to more accurately answer the question information; for another example, the user's question information is "Is my blood pressure high?", the intention corresponding to the question information is an intention related to the user's state, and it is necessary to obtain the user's blood pressure data in order to more accurately answer the question information. Similarly, when the preset intention is an intention related to the user's environment, that is, when the preset intention is related to the user's environment, if the intention corresponding to the question information is the preset intention, the application needs to combine the user's current environment to accurately and comprehensively understand the user's question information; for example, the user's question information is "Is it dark to read in the current light?", the intention corresponding to the question information is an intention related to the user's environment, and it is necessary to obtain the light intensity in the user's environment in order to more accurately answer the question information.
[0074] In a possible implementation, the target information includes one or more of image information of the user, image information of the user's environment, sound information of the user's environment, movement information of the user, tactile information of the user, and physiological information of the user.
[0075] Among them, the image information of the user represents an image containing the user's body parts, such as the user's face image, hand image, eyebrow image, etc.; the image information of the user's environment represents an image containing the user's environment, such as the image of the user's room, etc.; the sound information of the user's environment represents the background sound of the user's environment, such as music played in the environment, wind, rain, traffic noise, etc.; the action information of the user represents the user's body action information; the physiological information of the user represents the information reflecting the user's vital signs, such as blood pressure, heart rate, breathing rate, sleep duration, etc.; the tactile information of the user represents the tactile feedback information of the user's palm, fingers and other body parts. Exemplarily, the image information of the user and the image information of the user's environment can be obtained by shooting with a camera, wherein the camera can be the camera of the electronic device running the application, or a camera connected to the electronic device; the sound information of the user's environment can be collected by a microphone, wherein the microphone can be the microphone of the electronic device running the application, or a microphone connected to the electronic device; the physiological information, action information or tactile information of the user can be obtained through a wearable device worn by the user (such as a smart watch, a fitness ring, a game glove, etc.).
[0076] Since the target information and the user's question information have different information types or information sources, the target information and the question information form multimodal data.
[0077] It can be understood that the user's image information, the user's motion information, the user's tactile information, and the user's physiological information can all reflect the user's state and match the intention related to the user's state. The image information of the user's environment and the sound information of the user's environment can reflect the user's current environment and match the intention related to the user's environment. As an example, the preset intention is related to the user's state; the triggering of the application to obtain the target information matching the preset intention includes: starting the camera and controlling the camera to collect the user's image information; and / or obtaining at least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device. Exemplarily, when the camera is started and the camera is controlled to collect the user's image information, the built-in camera of the electronic device or the external camera connected to the electronic device can be controlled to start and shoot once or continuously to collect the user's image information, so that the electronic device can "visually" perceive the user's current state. At the same time, when the intention corresponding to the user's question information is related to the user's state, the camera is started to shoot, thereby avoiding the waste of image acquisition and processing resources caused by directly collecting and processing image information. Exemplarily, when acquiring at least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device, information on the user's corresponding status collected by the wearable device in a previous period of time can be acquired; or, when it is determined that the preset intention is related to the user's status, the electronic device can trigger the user's wearable device to start up to collect information on the user's corresponding status.
[0078] Exemplarily, in a scenario where the preset intention is related to the user's state, when the application is triggered to obtain target information matching the preset intention, the electronic device may further analyze the question information to determine which aspect of the user's state the question information specifically relates to. For example, keywords may be preset for different aspects of the state, and based on the preset keywords contained in the question information, it may be determined that the question information relates to one or more aspects of the user's action, touch, vision, and physiology, thereby obtaining information on the corresponding aspects of the state; wherein, if the question information includes preset keywords for the visual aspect of the state, the user's image information is obtained; if the question information includes preset keywords for the action aspect of the state, the user's action information is obtained; if the question information includes preset keywords for the tactile aspect of the state, the user's tactile information is obtained; and if the question information includes preset keywords for the physiological aspect of the state, the user's physiological information is obtained.
[0079] As another example, the preset intention is related to the environment in which the user is located; the triggering of the application to obtain target information matching the preset intention includes: starting the camera and controlling the camera to collect image information of the environment in which the user is located, and / or starting the microphone and controlling the microphone to collect sound information of the environment in which the user is located. Exemplarily, when the camera is started and controlled to collect image information of the environment in which the user is located, the built-in camera of the electronic device or the external camera connected to the electronic device can be controlled to start and take one or continuous shots to collect image information of the environment in which the user is located, so that the electronic device can "visually" perceive the user's current environment. At the same time, the camera is started to shoot only when the intention corresponding to the user's question information is related to the environment in which the user is located, thereby avoiding the waste of data collection and processing resources caused by directly collecting and processing image information. Exemplarily, when the microphone is started and controlled to collect sound information of the user's environment, the built-in microphone of the electronic device or the external microphone connected to the electronic device can be controlled to start and collect sound, so that the electronic device can perceive the user's current environment from the "hearing" sense. At the same time, the microphone is started to collect sound only when the intention corresponding to the user's question information is related to the user's environment, thereby avoiding the waste of data collection and processing resources caused by directly collecting and processing sound information.
[0080] Exemplarily, in a scenario where the preset intention is related to the user's environment, when the application is triggered to obtain target information matching the preset intention, the electronic device may further analyze the question information to determine what type of data (such as visual data, auditory data, etc.) the question information specifically involves. For example, keywords may be preset for different types of data, and the type of data involved in the question information may be determined based on the preset keywords contained in the question information, so that the corresponding type of environmental information may be obtained. Among them, if the question information includes visual preset keywords, image information of the user's environment is obtained, and if the question information includes auditory preset keywords, sound information of the user's environment is obtained.
[0081] It should be noted that the user's image, sound, physiological and other data involved in this disclosure are obtained with the user's consent and authorization, and the collection, use and processing of relevant data comply with relevant laws, regulations and standards of relevant countries and regions.
[0082] In a possible implementation, when the electronic device controls the camera to collect the image information of the user and when the electronic device controls the camera to collect the image information of the environment where the user is located, the shooting parameters of the camera may be different; the electronic device may set the shooting parameters of the camera as needed, or select from a preset combination of multiple shooting parameters to meet the needs of collecting image information types. Among them, the shooting parameters of the camera may include: viewing angle, exposure, white balance, frame rate, aperture, etc. For example, when the camera collects the image information of the user, the relative position of the user and the camera is usually close, and the camera may use a narrower viewing angle to collect more facial or body details of the user; and when the camera collects the image information of the environment where the user is located, a wide-angle lens or a wider viewing angle may be used so that the viewing angle of the camera can cover as much environment as possible. For another example, when the camera collects the image information of the user, the white balance may be set according to the skin color of the user to ensure that the skin color of the user is natural and avoid color cast, so that when the user's question information involves the skin color of the user, the user's question can be understood more accurately; and when the camera collects the image information of the environment where the user is located, the white balance may be set according to the color of the ambient light source to ensure that the color of each object in the collected image is accurate, so that the objects in the environment can be distinguished more accurately. For another example, when the camera collects image information of the user, the exposure can be adjusted according to the lighting conditions of the user's face or skin to avoid overexposure or underexposure, so that the user's face or body parts can be clearly visible; when the camera collects image information of the user's environment, the exposure can use HDR mode, so that high-contrast objects in the environment can be clearly displayed. For another example, when the camera collects image information of the user, a relatively large aperture and a high frame rate can be used, so that the user's expression or action details can be captured more clearly; when the camera collects image information of the user's environment, a relatively small aperture and a low frame rate can be used, so that more objects in the environment can be captured.
[0083] Step 203: Process the target multimodal data through the target model to generate response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data.
[0084] Among them, a large language model (LLMs) is a deep learning model trained with a large amount of text data, the purpose of which is to generate text, understand natural language or complete related language processing tasks. Large language models usually have a large number of parameters and can capture complex patterns and dependencies of language; illustratively, the target model can be a large language model for processing multimodal data built on the basis of Tongyi Qianwen, pre-trained generative transformer (GPT) model, Meta AI large language model (LLaMa), etc., or a large language model with multi-model data processing capabilities; the training process of the target model can be based on the existing technology and will not be repeated here. Since the user's question information and target information are data of different modes, the target model processes the user's question information and target information matching the preset intention and other data of different modes, so that the multimodal data can be comprehensively analyzed, and the complementarity and correlation between different modal data can be used to improve the understanding of the user's question information, thereby generating more accurate, natural and reasonable response information.
[0085] In a possible implementation, the processing of the target multimodal data by the target model also includes: selecting a preset model that matches the preset intent from multiple preset models as the target model; wherein different preset models are used to process information that matches different intents.
[0086] Exemplarily, based on a large language model trained on a huge corpus and capable of understanding general text, it is possible to further fine-tune information matching different intentions, thereby obtaining multiple preset models for processing information matching different intentions. For example, on the basis of a large language model for processing multimodal data, the question information and one or more of the user's image information, the image information of the environment, the action information, the tactile information, and the physiological information can be used as multimodal training data to further train and fine-tune the large language model, thereby obtaining a preset model for processing information matching different intentions. In this way, each preset model not only has the general feature capability of learning a language, but can also better process information matching the corresponding intention, thereby improving the ability to understand the question information of the corresponding intention, so as to generate the best response information.
[0087] Furthermore, considering that users of different ages, or the same user in different emotions, may have different acceptable response information for the same question information; for example, for users of different ages, adults usually expect to obtain more comprehensive response information, while minors are more likely to accept response information that is easy to understand; for another example, when users are happy, they are more tolerant of the content of response information, while when users are angry or sad, they are often more sensitive to the content of response information. Therefore, the target model can comprehensively consider attributes such as the user's age and / or emotion, so as to generate response information that is easier for users to accept.
[0088] In a possible implementation, the processing of target multimodal data through a target model to generate response information includes: processing the question information and / or the target information through the target model to extract attribute characteristics of the user; and processing the question information and the target information to extract multimodal characteristics; wherein the attribute characteristics include: age and / or emotion; and generating the response information based on the attribute characteristics and the multimodal characteristics.
[0089] The target model has functions such as age recognition, emotion recognition, and multimodal recognition. Exemplarily, the target model may be configured with multiple attribute recognition sub-models such as age recognition and emotion recognition that are pre-trained, and a pre-trained multimodal recognition sub-model that is pre-trained.
[0090] Exemplarily, age space and emotion space can be pre-constructed. For example, age space can include minors, adults, the elderly, etc., and emotion space can include: anger, sadness, happiness, tension, confidence, pain, loss, contempt, gratitude, excitement, etc. Then, in the constructed age space and emotion space, the corresponding attribute recognition sub-model is trained to obtain a trained age recognition sub-model and emotion recognition sub-model, and the trained age recognition sub-model and / or emotion recognition sub-model are configured in the target model. In this way, the question information and / or target information are processed by each attribute sub-model of the target model to extract the corresponding attribute features; at the same time, the question information can be processed by the trained multimodal recognition sub-model to extract multimodal features; then based on the attribute features and multimodal features, response information that conforms to the user attributes is generated. The attribute sub-model and the multimodal recognition sub-model can be implemented based on relevant technologies, and the present disclosure does not limit their specific implementation methods.
[0091] In some scenarios, the target information matching the preset intent can be processed by the target model to extract the user's attribute features. As an example, the user's facial image can be processed by the target model to extract the user's age features, and then based on the age features, combined with the multimodal features extracted from the user's question information and target information, response information that meets the age can be generated. For example, when an adult is talking to a virtual character, the adult's facial image can be captured by the camera, and then the current user can be predicted to be an adult through the facial image, so that response information that meets the adult's needs can be generated, such as response information with more comprehensive and detailed content; for another example, when a child is talking to a virtual character, the electronic device captures the child's facial image through the camera, and then the current user can be predicted to be a minor through the facial image, so that response information that meets the needs of minors can be generated, such as response information with easy-to-understand content. As another example, the user's image can be processed through the target model to extract the user's emotional features, and then based on the emotional features, combined with the multimodal features extracted from the user's question information and the target information, response information that matches the emotion can be generated. For example, when the user's mouth corners are upturned, or the user is dancing, it means that the user is happy, and response information that matches happiness can be generated, such as response information containing humorous words; for another example, when the user's expression is frowning, the corners of the mouth are drooping, or the head and shoulders are drooping, it means that the user is disappointed, and response information that matches disappointment can be generated, such as response information containing encouraging words; for another example, when the user covers his face with his hands and his body trembles slightly, it means that the user is sad, and response information that matches sadness can be generated, such as response information containing comforting words; for another example, when the user's body is straight and he holds his chest and head high, it means that the user is confident, and response information that matches confidence can be generated, such as response information containing praise words.
[0092] In some scenarios, the user's question information can be processed through the target model to extract the user's attribute characteristics. For example, the user's emotional characteristics can be extracted through the user's question voice, and then the target model can be used to generate response information that conforms to the emotion based on the user's emotional characteristics and combined with the multimodal features extracted from the question voice and target information. Considering the different emotions of the speaker, the acoustic characteristics of the spoken voice will be different. For example, the speaker's tone will be different with different emotions. The tone may be higher when happy and lower when sad. For example, the speaking speed may be faster when excited, and the speaking speed may be slower when frustrated; for example, when nervous, there may be more pauses when speaking, so the user's current emotional characteristics can be extracted by analyzing the tone, speed, volume, pauses, etc. of the question voice.
[0093] Exemplarily, the target model can also be configured in a server in the cloud, and the electronic device can also send the multimodal data to the cloud, so that based on the computing resources in the cloud, the target model is called to process the multimodal data to generate response information. For example, the electronic device can convert the multimodal data into data in the format of the Object Storage Service (OSS), and then send it to the cloud server, so that the server uses the target model to process the target multimodal data and generate response information.
[0094] Step 204: announcing the response information to the user via the virtual character on the display screen.
[0095] A virtual character refers to a digital or virtualized character or individual that exists in a virtual environment, and may be an animated character, a virtual assistant, or a character generated by artificial intelligence (AI). As an example, a virtual character may be displayed in a holographic warehouse on a display screen. For example, multiple virtual characters may be pre-designed, and users may select a virtual character of a corresponding image according to their preferences.
[0096] In one possible implementation, the response information includes a response text; and broadcasting the response information to the user through the virtual character on the display screen includes: generating a response voice and a response action based on the response text; driving the physical movement of the virtual character based on the response action, and synchronously playing the response voice.
[0097] For example, the response text can be converted into the response voice by using TTS (Text To Speech) technology. Considering that the response information generated by the target model is usually information in text form, that is, the response text, the text-to-speech technology can be used to convert the response text into the response voice, so that the response voice can be broadcasted to the user by driving the virtual character, thereby realizing the voice dialogue between the user and the virtual character and improving the user's immersive interactive experience.
[0098] In some scenarios, the response text can be converted into a response voice of the target timbre. The target timbre can be a timbre selected by the user, a pre-set timbre of a different virtual character image, or a timbre of the user's preference determined based on the user's historical interaction data. The corresponding TTS model can be pre-trained for different timbres. After determining the target timbre, the TTS model corresponding to the target timbre is selected, and the response text is input into the TTS model. The TTS model can generate a response voice of the target timbre. The model structure of the TTS model is not limited, and the training process of the TTS model can be carried out in an existing manner.
[0099] Exemplarily, the response text generated by the target model can be analyzed and processed based on a pre-trained action generation model to generate a corresponding response action. The specific structure and training process of the action generation model can be carried out in an existing manner. Furthermore, the existing animation generation technology can be used to drive the physical movement of the virtual character based on the response action, thereby generating the action animation of the virtual character, and the response voice and the action animation can be synchronized through the timestamp to achieve synchronous playback of the response voice. As an example, the virtual character can have a three-dimensional digital body, and while broadcasting the response information, the three-dimensional digital body of the virtual character can be driven to move synchronously based on the response action.
[0100] For example, by designing the facial expressions, skin texture, eye luster and other details of the virtual character, the human image in the real world can be realistically and meticulously restored, and the audience can visually get closer to the feeling of talking to a real person, thereby providing users with a face-to-face hyper-realistic immersive interactive experience. For example, the user's preferences can be determined through the user's historical interaction data, and the facial expressions, skin texture, corresponding character image, etc. of the virtual character can be determined based on the user's preferences.
[0101] In a possible implementation, the broadcasting of the response information to the user through the virtual character includes: determining the target state of the virtual character when broadcasting according to the target information; and controlling the virtual character to broadcast the response information to the user in the target state. In this way, the virtual character broadcasts the response information to the user in the target state determined by the target information; so that the broadcasting state of the virtual character remains integrated with the user's environment and / or matches the user's state, making the user feel more intimate and considerate, and enhancing the user's willingness to interact with the virtual character; at the same time, it can help the user better understand the meaning expressed by the response information. Exemplarily, the target state may include expression, sound, brightness, action, etc.
[0102] As an example, the target state of the virtual character when broadcasting can be determined based on the user's image information; illustratively, the user's emotions can be analyzed through the user's image information, and the facial expression of the virtual character when broadcasting can be determined based on the emotions; for example, when the user is happy, the virtual character's eyes can be controlled to be smiling, the virtual character's mouth can be controlled to be a happy state with upturned corners of the mouth, and so on; and when the user is sad, the virtual character's mouth can be controlled to be a sad state with downturned corners of the mouth, and so on; in this way, in response to questions asked by the user with different emotions, the virtual character can make corresponding facial expressions when answering, so that the emotion of the virtual character when broadcasting the response information matches the user's emotion.
[0103] As another example, the target state of the virtual character when making a report can be determined based on the image information of the user's environment; illustratively, the light intensity of the environment, the scene the user is in, etc. can be analyzed through the image information of the user's environment, so as to determine the brightness, sound, action, etc. of the virtual character when making a report; for example, when the intensity of the ambient light is strong, the brightness of the virtual character can be increased; for another example, when the user is identified in an outdoor scene such as a street, the voice of the virtual character can be increased, and when the user is identified in a bedroom, the voice of the virtual character can be lowered; for another example, when the user is identified in the snow, the virtual character can be controlled to "hug the whole body tightly and tremble" while reporting. In this way, in response to the user's questions in different environments, the brightness, sound, action, etc. of the virtual character are adjusted accordingly to blend in with the environment.
[0104] As another example, the target state of the virtual character's broadcast can be determined based on the sound information of the user's environment; for example, the noise level of the environment can be analyzed through the sound information of the user's environment, so as to determine the voice of the virtual character when broadcasting; for example, when the environment is noisy, the voice of the virtual character can be increased; for another example, when the background sound of the scene where the user is located is loud through sound information analysis, the voice of the virtual character can be increased. In this way, the voice of the virtual character is adjusted accordingly to the user's questions in different environments to blend in with the environment.
[0105] In the disclosed embodiment, in response to the question information of the user to the virtual character running in the application, the intention corresponding to the question information is determined; it is judged whether the intention corresponding to the question information is a preset intention, and when the intention corresponding to the question information is the preset intention, the application is triggered to obtain the target information matching the preset intention; wherein the preset intention is related to the environment in which the user is located and / or the state of the user; the target multimodal data is processed by the target model to generate the response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data; the response information is broadcast to the user by the virtual character. In this way, based on the multimodal interaction technology and the intention recognition technology, the target information can be obtained when the intention corresponding to the user's question information is an intention related to the environment in which the user is located and / or the state of the user; and then the large language model is used to process the multimodal data, and the complementarity and correlation between different modal data are used to improve the virtual character's ability to understand the user's question information, so that the questions related to the environment in which the user is located or the state of the user can be answered more realistically and accurately, and more accurate, natural and reasonable response information can be generated, so as to realize the immersive interaction of hyper-realistic virtual characters. At the same time, unlike directly acquiring multimodal data, the target information is acquired only when the intention corresponding to the user's question information is the preset intention, thereby avoiding the waste of data collection and processing resources caused by directly acquiring the target information.
[0106] The following takes the scenario of voice interaction between a user and a virtual character (i.e., the user's question information is the question voice, and the virtual character broadcasts the answer voice), the preset intention is related to the user's status, and the target message is the user's image information as an example to exemplify the above virtual character interaction method.
[0107] Figure 3 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown. Figure 3 As shown, the following steps are included:
[0108] Step 301: In response to a question voice from a user to a virtual character running in an application, determine an intention corresponding to the question voice.
[0109] This step 301 can be used as the above Figure 2 A possible implementation of step 201 in FIG.
[0110] As an example, the user's question voice can be received through a microphone configured in the electronic device, so that the user's question can be perceived from the "hearing" sense.
[0111] Exemplarily, the question voice can be converted into question text, and then the question text is input into a trained intent recognition model to determine the intent corresponding to the question voice.
[0112] Step 302: determine whether the intention corresponding to the question voice is an intention related to the user's status, and if the intention corresponding to the question voice is an intention related to the user's status, trigger the application to obtain user image information that matches the intention related to the user's status.
[0113] This step 302 can be used as the above Figure 2 A possible implementation of step 202 in FIG.
[0114] Among them, the question voice involves the user's visual state; for example, the user's question voice can be "How about my outfit?", "Are the dark circles under my eyes serious?" and other questions involving the user's visual state.
[0115] For example, when the intention corresponding to the question voice is an intention related to the user's state, the preset keywords involved in the question voice can be analyzed to determine that the question voice involves the user's visual state, so as to trigger the application to obtain the user's image information. For example, the preset keywords can be words such as "outfit", "appearance", "eyes", "hair" and other entities that require visual observation.
[0116] Exemplarily, the triggering application obtains the user's image information that matches the intention related to the user's state, including: starting the camera, and controlling the camera to collect the user's image information; thereby being able to perceive the user's current state from a "visual" perspective.
[0117] In a possible implementation, the controlling the camera to collect the image information of the user includes: controlling the camera to take at least one shot, and performing target detection on the shot image until it is detected that the target body part involved in the question information is included in the currently shot image, wherein when it is detected that the currently shot image does not include the target body part, adjusting the shooting parameters of the camera, and / or prompting the user to adjust the posture of himself or the camera. The number of target body parts can be one or more, so that it is ensured that the collected image information of the user contains the target body part involved in the user's question information, so that the user's question can be understood and answered more accurately.
[0118] Exemplarily, the electronic device can analyze the user's question information, for example, it can analyze the keywords of different body parts contained in the question information (such as face, hands, legs, arms, neck, mouth, eyebrows, hair, etc.) to determine the target body part involved in the user's question information, and then after the electronic device controls the camera to take each shot, it performs target detection on the image taken this time to determine whether the target body part is included in the image taken this time. If it is detected that the target body part is included in the image taken this time, the electronic device controls the camera to end image acquisition and uses the image taken this time as the image information of the user collected; if it is detected that the image taken this time does not contain the target body part, the electronic device adjusts the shooting parameters of the camera and controls the camera to continue the next shooting with the adjusted shooting parameters, or the electronic device can prompt the user to adjust his or her own posture or the posture of the camera, and after the user has adjusted his or her own posture or the posture of the camera, controls the camera to continue the next shooting.
[0119] As an example, when adjusting the shooting parameters of the camera, the camera's viewing angle, exposure, white balance, resolution, aperture, etc. can be adjusted. For example, when the image captured this time only contains part of the user's face and the user's hair, the camera's viewing angle can be lowered and the next shot can be taken to capture the user's entire face or more other body parts. For another example, if the user's skin is too dark or underexposed in the image captured this time, the exposure can be increased or the white balance can be adjusted, and the next shot can be taken to capture the user's eyebrows, hair and other body parts. For another example, if the body part is blurred and indistinguishable in the image captured this time, the aperture or resolution can be increased, and the next shot can be taken to capture a clearer body part. As another example, when prompting the user to adjust his or her posture, a prompt voice can be issued to the user, such as, "Please stand back a little, let me see your feet", "Please look at me", etc., to prompt the user to adjust his or her posture. As another example, when prompting the user to adjust the position of the camera, a prompt voice may be issued to the user, for example, "point the phone down a little", "turn the camera left", etc., to prompt the user to adjust the position of the camera.
[0120] Step 303: Process the target multimodal data through the target model to generate a response text; wherein the target multimodal data includes the question voice and the image information of the user.
[0121] This step 303 can be used as the above Figure 2 A possible implementation of step 203 in FIG.
[0122] Step 304: Generate a response voice and a response action according to the response text, and drive the physical movement of the virtual character based on the response action, and synchronously play the response voice.
[0123] This step 304 can be used as the above Figure 2 A possible implementation of step 204 in FIG.
[0124] In the disclosed embodiment, the user can initiate a voice question to the electronic device. Based on the multimodal interaction technology and the intention recognition technology, the electronic device obtains the user's image information when the intention corresponding to the user's voice question is an intention related to the user's state and the voice question involves the user's visual state, and then uses a large language model to process the user's image information and the data of different modes such as the voice question, and utilizes the complementarity and correlation between the data of different modes to improve the virtual character's understanding ability of the questions related to the user's visual state, and generate more accurate, natural and reasonable response information; at the same time, the virtual character's body movement can be driven during the interactive feedback of the voice question, and the response voice can be played synchronously, so as to realize the immersive interaction of the hyper-realistic virtual character and significantly improve the interactive experience between the user and the virtual character. In addition, unlike directly obtaining multimodal data, the user's image information is obtained only when it is determined that the intention corresponding to the user's voice question is an intention related to the user's state, thereby avoiding the waste of data collection and processing resources caused by directly obtaining the user's image information.
[0125] The following takes the scenario of voice interaction between a user and a virtual character (i.e., both the question information and the answer information are messages in voice form), the preset intention is related to the user's status, and the target message is one or more of the user's action information, the user's tactile information, and the user's physiological information as an example to exemplify the above-mentioned virtual character interaction method.
[0126] Figure 4 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown. Figure 4 As shown, the following steps are included:
[0127] Step 401: In response to a question voice from a user to a virtual character running in an application, determine an intention corresponding to the question voice.
[0128] This step 401 is similar to the above Figure 3 The step 301 is the same as that in the above step and will not be described again.
[0129] Step 402: determine whether the intention corresponding to the questioning voice is an intention related to the user's state, and if the intention corresponding to the questioning voice is an intention related to the user's state, trigger the application to obtain one or more of the user's action information, user's tactile information, and user's physiological information that match the intention related to the user's state.
[0130] This step 402 can be used as the above Figure 2 A possible implementation of step 202 in FIG.
[0131] The question voice involves at least one of the user's actions, touch, and physiology; exemplarily, when the intention corresponding to the question voice is an intention related to the user's state, the preset keywords involved in the question voice can be analyzed to determine the action, touch, physiology, and other aspects of the question voice involved in order to trigger the application to obtain the user's corresponding state information. For example, if the question voice involves the user's tactile state, the application is triggered to obtain the user's tactile information. The preset keywords corresponding to the states of different aspects can be set as needed. For example, the preset keywords corresponding to the action state can be words such as "posture", "action", etc. related to the user's actions, the preset keywords corresponding to the tactile state can be words such as "vibration", "tactile", etc. related to the user's tactile sense, and the preset keywords corresponding to the physiological state can be words such as "blood pressure", "sleep", "heart rate", etc. related to the user's physiological indicators.
[0132] Exemplarily, the application is triggered to obtain one or more of the user's motion information, the user's tactile information, and the user's physiological information that match the intention related to the user's state, including: obtaining at least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device, so as to perceive the user's current state from different aspects.
[0133] As an example, when a user is wearing a smart watch with a blood pressure sensor, the user can ask the electronic device a question, "Is my blood pressure high?", and the electronic device can obtain the user's blood pressure data collected by the smart watch. For example, if the smart watch is already connected to the electronic device, the electronic device can directly request the blood pressure data from the smart watch, or obtain the blood pressure data that the smart watch has uploaded to the electronic device. If the electronic device and the smart watch are not connected, the connection can be started (such as Bluetooth connection), or the user can be prompted to connect the electronic device and the smart watch.
[0134] As another example, when a user is wearing a fitness ring for fitness, the user can ask the electronic device a question, "Is my posture correct?", and the electronic device can obtain the user's motion information collected by the fitness ring. For example, if the fitness ring is already connected to the electronic device, the electronic device can directly request the user's motion information from the fitness ring, or obtain the user's motion information that the fitness ring has uploaded to the electronic device. If the electronic device and the fitness ring are not connected, the connection can be started or the user can be prompted to connect the electronic device and the fitness ring.
[0135] As another example, when a user is playing a game wearing a gaming glove with a tactile feedback function, he or she may ask the electronic device a question such as "Why is the vibration in my hand suddenly stronger?", and the electronic device may obtain the user's tactile information collected by the gaming glove. For example, if the gaming glove is already connected to the electronic device, the electronic device may directly request the user's tactile information from the gaming glove, or obtain the user's tactile information that the gaming glove has uploaded to the electronic device. If the electronic device and the gaming glove are not connected, the connection may be initiated or the user may be prompted to connect the electronic device and the gaming glove.
[0136] Step 403: Process the target multimodal data through the target model to generate a response text; wherein the target multimodal data includes the question voice and one or more of the user's motion information, the user's tactile information, and the user's physiological information.
[0137] This step 403 can be used as the above Figure 2 A possible implementation of step 203 in FIG.
[0138] Step 404: Generate a response voice and a response action according to the response text, and drive the physical movement of the virtual character based on the response action, and synchronously play the response voice.
[0139] This step 404 can be used as the above Figure 2 A possible implementation of step 204 in FIG.
[0140] In the disclosed embodiment, the user can initiate a voice question to the electronic device. The electronic device, based on the multimodal interaction technology and the intention recognition technology, obtains the user's action information, tactile information or physiological information when the intention corresponding to the user's voice question is an intention related to the user's state and the voice question involves the user's action, touch or physiological state, and then uses a large language model to process the question voice and the user's action information, tactile information or physiological information and other different modal data, and utilizes the complementarity and correlation between different modal data to improve the virtual character's understanding ability of the problem involving the user's action, touch or physiological state, and generate more accurate, natural and reasonable response information; at the same time, the virtual character's body movement can be driven during the interactive feedback of the question voice, and the response voice can be played synchronously, so as to realize the immersive interaction of the hyper-realistic virtual character and significantly improve the interactive experience between the user and the virtual character. In addition, unlike directly obtaining multimodal data, the user's action information, tactile information or physiological information is obtained only when it is determined that the intention corresponding to the user's voice question is an intention related to the user's state, thereby avoiding the waste of data collection and processing resources caused by directly obtaining the user's action information, tactile information or physiological information.
[0141] The following is an example of a scenario in which a user interacts with a virtual character through voice, the preset intention is related to the user's environment, and the target message is image information of the user's environment and / or sound information of the user's environment, to exemplify the above virtual character interaction method.
[0142] Figure 5 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown. Figure 5 As shown, the following steps are included:
[0143] Step 501: In response to a question voice from a user to a virtual character running in an application, determine an intention corresponding to the question voice.
[0144] This step 501 is similar to the above Figure 3 The step 301 is the same as that in the above step and will not be described again.
[0145] Step 502: determine whether the intention corresponding to the question voice is an intention related to the user's environment, and if the intention corresponding to the question voice is an intention related to the user's environment, trigger the application to obtain image information of the user's environment and / or sound information of the user's environment that matches the intention related to the user's environment.
[0146] This step 502 can be used as the above Figure 2 A possible implementation of step 202 in FIG.
[0147] For example, the user's voice question may be "Is it a little dim to read in the current light?", "What kind of flower is this next to me?", "What music is playing on TV?" and other questions related to the user's environment.
[0148] Exemplarily, triggering the application to obtain image information of the user's environment that matches the intention related to the user's environment includes: starting a camera, and controlling the camera to collect image information of the user's environment.
[0149] In a possible implementation, the controlling the camera to collect image information of the environment of the user includes: controlling the camera to shoot at least once, and performing target detection on the shot image until the target object involved in the question information is detected in the currently shot image, wherein, when it is detected that the currently shot image does not contain the target object, adjusting the shooting parameters of the camera, and / or prompting the user to adjust the posture of the camera or the camera. The number of target objects can be one or more, so that it is ensured that the image information of the user's environment collected contains the target object involved in the user's question information, so that the user's question can be understood and answered more accurately.
[0150] Exemplarily, the electronic device can analyze the user's question information, for example, it can analyze the keywords of different objects contained in the question information (such as TV, sofa, flower, floor, lamp, cup, etc.) to determine the target object involved in the user's question information, and then after the electronic device controls the camera to take each shot, it performs target detection on the image taken this time to determine whether the target object is included in the image taken this time. If it is detected that the target object is included in the image taken this time, the electronic device controls the camera to end image acquisition and uses the image taken this time as the image information of the user's environment; if it is detected that the image taken this time does not contain the target object, the electronic device adjusts the shooting parameters of the camera and controls the camera to continue the next shooting with the adjusted shooting parameters, or the electronic device can prompt the user to adjust his or her own posture or the posture of the camera, and after the user has adjusted his or her own posture or the posture of the camera, controls the camera to continue the next shooting.
[0151] As an example, when adjusting the shooting parameters of the camera, the camera's viewing angle, exposure, white balance, resolution, aperture, etc. can be adjusted. For example, if the image captured this time only contains the user and a few objects, the camera's viewing angle can be increased or switched to a wide-angle lens, and the next shot can be taken to capture more objects in the environment. For another example, if the object is blurred and indistinguishable in the image captured this time, the aperture or resolution can be increased, and the next shot can be taken to capture a clearer object. As another example, when prompting the user to adjust their posture, a prompt voice can be issued to the user, such as, "Please stand a little to the left, don't block the TV", etc., to prompt the user to adjust their posture. As another example, when prompting the user to adjust the posture of the camera, a prompt voice can be issued to the user, such as, "The phone is facing down a little", "Turn the camera to the left", etc., to prompt the user to adjust the posture of the camera.
[0152] Exemplarily, triggering the application to obtain the sound information of the user's environment that matches the intention related to the user's environment includes: starting a microphone and controlling the microphone to collect the sound information of the user's environment. For example, the microphone can be controlled to collect music played on a TV, wind, rain, traffic noise, etc.
[0153] Step 503: Process the target multimodal data through the target model to generate a response text; wherein the target multimodal data includes the question voice and image information of the user's environment and / or sound information of the user's environment.
[0154] This step 503 can be used as the above Figure 2 A possible implementation of step 203 in FIG.
[0155] Step 504: Generate a response voice and a response action according to the response text, and drive the physical movement of the virtual character based on the response action, and synchronously play the response voice.
[0156] This step 504 can be used as the above Figure 2 A possible implementation of step 204 in FIG.
[0157] In the disclosed embodiment, the user can initiate a voice question to the electronic device. The electronic device, based on the multimodal interaction technology and the intention recognition technology, obtains the image information of the user's environment and / or the sound information of the user's environment when the intention corresponding to the user's voice question is an intention related to the user's environment, and then uses a large language model to process the image information of the user's environment and / or the sound information of the user's environment, and the question voice and other different modal data, and uses the complementarity and correlation between different modal data to improve the virtual character's understanding ability of the problem related to the user's environment, and generate more accurate, natural and reasonable response information; at the same time, the virtual character's body movement can be driven during the interactive feedback of the question voice, and the response voice can be played synchronously, so as to realize the immersive interaction of the hyper-realistic virtual character and significantly improve the interactive experience between the user and the virtual character. In addition, unlike directly obtaining multimodal data, the image information of the user's environment and / or the sound information of the user's environment are obtained only when it is determined that the intention corresponding to the user's voice question is an intention related to the user's environment, thereby avoiding the waste of data collection and processing resources caused by directly obtaining the image information of the user's environment and / or the sound information of the user's environment.
[0158] Figure 6 A flowchart of a method for virtual character interaction according to an embodiment of the present disclosure is shown as follows: Figure 6 As shown in the figure, first, the user can trigger the interaction with the virtual character. In each round of dialogue, the electronic device obtains the user's question information and determines the intention corresponding to the question information (i.e., execution) from the chat intention, professional field intention, intention related to the user's environment, and intention related to the user's status through the intention recognition model. Figure 2 The preset intention is pre-set as an intention related to the user's environment and an intention related to the user's state. Then, when the intention corresponding to the question information is an intention related to the user's environment, the image information of the user's environment and / or the sound information of the user's environment are obtained accordingly. When the intention corresponding to the question information is an intention related to the user's state, the image information, action information, tactile information or physiological information of the user is obtained accordingly (i.e., the execution Figure 2 Step 202); then, the above-obtained information is processed by calling the corresponding large language model to generate response information (i.e., executing Figure 2 Step 203); Finally, the virtual character broadcasts the response information to the user (i.e., executes Figure 2 Step 204); thereby achieving ultra-realistic immersive interaction.
[0159] Based on the same inventive concept of the above method embodiment, an embodiment of the present disclosure further provides a virtual character interaction device, which can be used to execute the technical solution described in the above method embodiment.
[0160] Figure 7 A structural diagram of a virtual character interaction device according to an embodiment of the present disclosure is shown as follows: Figure 7 As shown, the device may include: a questioning module 701, which is used to determine the intention corresponding to the questioning information in response to the questioning information asked by the user to the virtual character running in the application; an acquisition module 702, which is used to determine whether the intention corresponding to the questioning information is a preset intention, and when the intention corresponding to the questioning information is the preset intention, trigger the application to obtain target information matching the preset intention; wherein the preset intention is related to the environment in which the user is located and / or the state of the user; a response module 703, which is used to process the target multimodal data through a target model to generate response information; wherein the target multimodal data includes the questioning information and the target information; the target model is a pre-built large language model for processing multimodal data; a broadcasting module 704, which is used to broadcast the response information to the user through the virtual character.
[0161] In the disclosed embodiment, in response to the question information of the user to the virtual character running in the application, the intention corresponding to the question information is determined; it is judged whether the intention corresponding to the question information is a preset intention, and when the intention corresponding to the question information is the preset intention, the application is triggered to obtain the target information matching the preset intention; wherein the preset intention is related to the environment in which the user is located and / or the state of the user; the target multimodal data is processed by the target model to generate the response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data; the response information is broadcast to the user by the virtual character. In this way, based on the multimodal interaction technology and the intention recognition technology, the target information can be obtained when the intention corresponding to the user's question information is an intention related to the environment in which the user is located and / or the state of the user; and then the large language model is used to process the multimodal data, and the complementarity and correlation between different modal data are used to improve the virtual character's ability to understand the user's question information, so that the questions related to the environment in which the user is located or the state of the user can be answered more realistically and accurately, and more accurate, natural and reasonable response information can be generated, so as to realize the immersive interaction of hyper-realistic virtual characters. At the same time, unlike directly acquiring multimodal data, the target information is acquired only when the intention corresponding to the user's question information is the preset intention, thereby avoiding the waste of data collection and processing resources caused by directly acquiring the target information.
[0162] In a possible implementation, the target information includes one or more of image information of the user, image information of the user's environment, sound information of the user's environment, movement information of the user, tactile information of the user, and physiological information of the user.
[0163] In one possible implementation, the preset intention is related to the user's state; the acquisition module 702 is also used to: start the camera and control the camera to collect image information of the user; and / or obtain at least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device.
[0164] In one possible implementation, the preset intention is related to the environment in which the user is located; the acquisition module 702 is also used to: start the camera and control the camera to collect image information of the user's environment; and / or start the microphone and control the microphone to collect sound information of the user's environment.
[0165] In a possible implementation, the processing of the target multimodal data by the target model also includes: selecting a preset model that matches the preset intent from multiple preset models as the target model; wherein different preset models are used to process information that matches different intents.
[0166] In a possible implementation, the response module 703 is further used to: process the question information and / or the target information through the target model to extract the user's attribute characteristics; and process the question information and the target information to extract multimodal characteristics; wherein the attribute characteristics include: age and / or emotions; and generate the response information based on the attribute characteristics and the multimodal characteristics.
[0167] In one possible implementation, the acquisition module 702 is also used to: control the camera to take at least one shot, and perform target detection on the shot image until it is detected that the target body part involved in the question information is included in the currently shot image, wherein when it is detected that the currently shot image does not contain the target body part, the shooting parameters of the camera are adjusted, and / or the user is prompted to adjust the posture of himself or the camera.
[0168] In a possible implementation, the questioning module 701 is further used to: obtain historical questions of the user stored in the application; and determine the intention corresponding to the question information in combination with the historical questions.
[0169] In a possible implementation, the reporting module 704 is further used to: determine the target state of the virtual character when reporting according to the target information; and control the virtual character to report the response information to the user in the target state.
[0170] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0171] The embodiment of the present disclosure also provides a computer-readable storage medium on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0172] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0173] The embodiments of the present disclosure also provide a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0174] Figure 8 1 is a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 may be provided as a server or a terminal device. Figure 8 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.
[0175] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™ or the like.
[0176] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.
[0177] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0178] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.
[0179] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0180] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0181] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0182] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0183] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0184] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.
[0185] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A virtual character interaction method, characterized in that: The method comprises: In response to a question message sent by a user to a virtual character running in an application, determining an intention corresponding to the question message; Determine whether the intention corresponding to the question information is a preset intention, and if the intention corresponding to the question information is the preset intention, trigger the application to obtain target information matching the preset intention; wherein the preset intention is related to the environment in which the user is located and / or the state of the user; The target multimodal data is processed by the target model to generate response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data; The response information is broadcasted to the user through the virtual character.
2. The method according to claim 1, characterized in that: The target information includes one or more of the following: image information of the user, image information of the user's environment, sound information of the user's environment, action information of the user, tactile information of the user, and physiological information of the user.
3. The method according to claim 1 or 2, characterized in that: The preset intention is related to the user's state; The triggering the application to obtain target information matching the preset intention includes: Start the camera and control it to collect the user's image information; and / or, At least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device is obtained.
4. The method according to claim 1 or 2, characterized in that: The preset intention is related to the environment in which the user is located; The triggering the application to obtain target information matching the preset intention includes: Start the camera and control the camera to collect image information of the user's environment; and / or, Start the microphone and control it to collect sound information of the user's environment.
5. The method according to claim 1, characterized in that The target multimodal data is processed by the target model, and the process also includes: A preset model matching the preset intent is selected from a plurality of preset models as the target model; wherein different preset models are used to process information matching different intents.
6. The method according to claim 1, characterized in that The target multimodal data is processed by the target model to generate response information, including: The question information and / or the target information are processed through the target model to extract the user's attribute characteristics; and the question information and the target information are processed to extract multimodal characteristics; wherein the attribute characteristics include: age and / or emotion; based on the attribute characteristics and the multimodal characteristics, the response information is generated.
7. The method according to claim 3, characterized in that The controlling camera to collect image information of the user includes: Control the camera to take at least one shot and perform target detection on the shot image until it is detected that the target body part involved in the question information is included in the currently shot image, wherein, when it is detected that the currently shot image does not include the target body part, adjust the shooting parameters of the camera, and / or prompt the user to adjust the posture of himself or the camera.
8. The method according to claim 1, characterized in that The determining the intention corresponding to the question information includes: Obtain the user's historical questions stored in the application; In combination with the historical questions, the intention corresponding to the question information is determined.
9. The method according to claim 1, characterized in that: The step of broadcasting the response information to the user through the virtual character includes: Determining a target state of the virtual character when making a broadcast according to the target information; The virtual character is controlled to broadcast the response information to the user in the target state.
10. A virtual character interaction device, characterized in that: The device comprises: A questioning module, used for asking question information to a virtual character running in an application program, and determining the intention corresponding to the question information; an acquisition module, used to determine whether the intention corresponding to the question information is a preset intention, and if the intention corresponding to the question information is the preset intention, trigger the application to acquire target information matching the preset intention; wherein the preset intention is related to the environment in which the user is located and / or the state of the user; A response module, used to process the target multimodal data through a target model to generate response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data; The reporting module is used to report the response information to the user through the virtual character.
11. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the method described in any one of claims 1 to 9 when executing the instructions stored in the memory.
12. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 9 is implemented.
13. A computer program product, comprising a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, wherein when the computer-readable code is executed in a processor of an electronic device, the processor in the electronic device executes the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Virtual object interaction method and device, electronic equipment and storage medium
CN117950492A
Context environment-based dialogue interaction processing method and device
CN119202332A
Question answering method and apparatus, device, and storage medium
WO2024188242A1
Question answering method and apparatus, and device and storage medium
WO2024227415A1
Cited By
AI agent digital human interaction system and method based on multi-modal perception
CN120408125A
A Digital Human Interaction System and Method Based on Multimodal Perception AI Agent
CN120408125B
Character map-based role virtual playing method and device
CN121683978A
A character virtual role playing method and device based on a character graph
CN121683978B