Virtual character interaction method and device
Through multimodal interaction and intention recognition technology, combined with large language models to process multimodal data, the defect that virtual characters cannot accurately answer user environment or state problems is solved, and the authenticity and accuracy of virtual character interaction is improved.
Patent Information
- Application Number
- CN202510174167.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-02-17
AI Technical Summary
At this stage, in the interaction between people and virtual characters, virtual characters cannot accurately answer questions involving the user environment or status, affecting the user experience.
Through multimodal interaction technology and intention recognition technology, the user's question information corresponds to the intention, and obtain multimodal data when the intention is a preset intention, and use a large language model to process multimodal data to generate response information.
It enables virtual characters to answer questions related to user environment or status more realistically and accurately, improving the interactive experience.
Smart Images

Figure CN120105009B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a virtual character interaction method and device. Background Art
[0002] Virtual characters are appearing in a growing number of fields, and enabling human-virtual character interaction in a growing number of scenarios has become a new model for human-computer interaction. Currently, human-virtual character interaction occurs through voice question-and-answer sessions. During these interactions, the virtual characters are unable to answer some questions or provide inaccurate responses, impacting the user experience. Summary of the Invention
[0003] In view of this, the present disclosure proposes a virtual character interaction method, device, electronic device, storage medium and computer program product.
[0004] According to one aspect of the present disclosure, a virtual character interaction method is provided, the method comprising:
[0005] In response to a question message sent by a user to a virtual character running in an application, determining an intention corresponding to the question message;
[0006] determining whether the intent corresponding to the question information is a preset intent, and if the intent corresponding to the question information is the preset intent, triggering the application to obtain target information matching the preset intent; wherein the preset intent is related to the user's environment and / or the user's status;
[0007] Processing target multimodal data through a target model to generate response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data;
[0008] The response information is broadcast to the user through the virtual character.
[0009] In a possible implementation, the target information includes one or more of the user's image information, the user's environment image information, the user's environment sound information, the user's motion information, the user's tactile information, and the user's physiological information.
[0010] In a possible implementation, the preset intention is related to the user's state;
[0011] The triggering the application to obtain target information matching the preset intent includes:
[0012] Start the camera and control it to collect user's image information;
[0013] and / or,
[0014] At least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device is obtained.
[0015] In a possible implementation, the preset intention is related to the environment in which the user is located;
[0016] The triggering the application to obtain target information matching the preset intent includes:
[0017] Start the camera and control it to collect image information of the user's environment;
[0018] and / or,
[0019] Start the microphone and control it to collect sound information from the user's environment.
[0020] In a possible implementation, the processing of the target multimodal data by the target model further includes:
[0021] A preset model that matches the preset intention is selected from multiple preset models as the target model; wherein different preset models are used to process information matching different intentions.
[0022] In one possible implementation, processing the target multimodal data using the target model to generate response information includes:
[0023] Processing the question information and / or the target information through the target model to extract user attribute features; and processing the question information and the target information to extract multimodal features; wherein the attribute features include: age and / or emotion;
[0024] The response information is generated based on the attribute feature and the multimodal feature.
[0025] In one possible implementation, controlling the camera to collect image information of the user includes:
[0026] The camera is controlled to take at least one shot and perform target detection on the shot image until it is detected that the target body part involved in the question information is included in the currently shot image, wherein, when it is detected that the currently shot image does not include the target body part, the shooting parameters of the camera are adjusted, and / or the user is prompted to adjust the posture of himself or the camera.
[0027] In a possible implementation, determining the intention corresponding to the question information includes:
[0028] Obtain the user's historical questions stored in the application;
[0029] In combination with the historical questions, the intention corresponding to the question information is determined.
[0030] In a possible implementation, broadcasting the response information to the user through the virtual character includes:
[0031] Determining a target state of the virtual character when performing the broadcast according to the target information;
[0032] The virtual character is controlled to broadcast the response information to the user in the target state.
[0033] According to another aspect of the present disclosure, a virtual character interaction device is provided, the device comprising:
[0034] a questioning module, configured to determine an intention corresponding to a question message in response to a question message asked by a user to a virtual character running in an application;
[0035] an acquisition module, configured to determine whether the intent corresponding to the question information is a preset intent, and, if the intent corresponding to the question information is the preset intent, trigger the application to acquire target information that matches the preset intent; wherein the preset intent is related to the user's environment and / or the user's status;
[0036] A response module, configured to process target multimodal data using a target model to generate response information; wherein the target multimodal data includes the question information and the target information; and the target model is a pre-built large language model for processing multimodal data;
[0037] The broadcasting module is used to broadcast the response information to the user through the virtual character.
[0038] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0039] According to another aspect of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.
[0040] According to another aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0041] According to various aspects of the present disclosure, in response to a question message sent by a user to a virtual character running in an application, an intent corresponding to the question message is determined; whether the intent corresponding to the question message is a preset intent is determined, and if the intent corresponding to the question message is the preset intent, the application is triggered to obtain target information matching the preset intent; wherein the preset intent is related to the user's environment and / or the user's state; target multimodal data is processed using a target model to generate response information; wherein the target multimodal data includes the question message and the target information; the target model is a pre-built large language model for processing multimodal data; and the response information is broadcast to the user by the virtual character. In this way, based on multimodal interaction technology and intent recognition technology, target information can be obtained when the intent corresponding to the user's question message is an intent related to the user's environment and / or the user's state; then, the large language model is used to process the multimodal data, and the complementarity and correlation between different modal data are utilized to improve the virtual character's ability to understand the user's question message, thereby providing more realistic and accurate answers to questions related to the user's environment or the user's state, generating more accurate, natural, and reasonable response information, and achieving immersive interaction with hyper-realistic virtual characters. At the same time, unlike directly acquiring multimodal data, the target information is acquired only when the intention corresponding to the user's question information is the preset intention, thus avoiding the waste of data collection and processing resources caused by directly acquiring the target information.
[0042] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the disclosure and, together with the description, serve to explain the principles of the disclosure.
[0044] Figure 1 A schematic diagram illustrating a scenario of user-virtual character interaction according to an embodiment of the present disclosure is shown.
[0045] Figure 2 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown.
[0046] Figure 3 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown.
[0047] Figure 4 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown.
[0048] Figure 5 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown.
[0049] Figure 6 A flowchart of a method for virtual character interaction according to an embodiment of the present disclosure is shown.
[0050] Figure 7 A structural diagram of a virtual character interaction device according to an embodiment of the present disclosure is shown.
[0051] Figure 8 A block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0052] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0053] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present disclosure. Thus, phrases such as "exemplary," "in one embodiment," "in some other embodiments," and "in other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0054] In the present disclosure, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: including the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0055] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.
[0056] Figure 1 A schematic diagram showing a scene of user-virtual character interaction according to an embodiment of the present disclosure is shown; Figure 1 As shown, the electronic device is equipped with a display screen, and the user can interact with the virtual character on the display screen. During the interaction, the user asks questions to the virtual character, and the electronic device generates response information after understanding the question information, and feeds back the response information to the user in real time through the virtual character on the display screen.
[0057] In related technologies, the interaction between users and virtual characters is carried out in the form of voice questions and answers; however, the virtual characters are unable to answer or answer inaccurately some of the users' questions. For example, when users ask questions such as "How do you like my outfit today?" or "Is the light in the room a little dim?" which require visual observation of the environment to answer accurately, the virtual characters are unable to give true and accurate answers, which affects the user's interactive experience with the virtual characters.
[0058] In order to solve the above technical problems, the embodiments of the present disclosure provide a virtual character interaction method (detailed description see below). This virtual character interaction method is based on multimodal interaction technology and intention recognition technology, and can answer questions related to the user's environment or the user's status more realistically and accurately, thereby realizing immersive interaction with hyper-realistic virtual characters.
[0059] Exemplarily, the virtual character interaction method provided in the embodiments of the present disclosure can be configured in an electronic device in the form of software and / or hardware so that the electronic device can perform the virtual character interaction function. In the following embodiments, the execution subject is an electronic device as an example for explanation. Among them, the electronic device can be any device with computing capabilities, for example, a personal computer (PC), a mobile terminal, a server, etc. The mobile terminal can be a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, a smart speaker, a server, a server cluster, etc., which have various operating systems, touch screens and / or display screens; the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms.
[0060] It should be noted that the above-mentioned application scenarios described in the embodiments of the present disclosure are intended to more clearly illustrate the technical solutions of the embodiments of the present disclosure, and do not constitute a limitation on the technical solutions provided by the embodiments of the present disclosure. Ordinary technicians in this field can know that the technical solutions provided by the embodiments of the present disclosure are also applicable to similar technical problems in response to the emergence of other similar or new scenarios, such as live broadcast rooms, metaverses, etc.
[0061] Figure 2 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown. Figure 2 As shown, the method may include the following steps:
[0062] Step 201: In response to a question asked by a user to a virtual character running in an application, determine an intention corresponding to the question.
[0063] The application can be any computer program that can run a virtual character and interact with a user. The application can be in the form of an application that can be independently installed and run on an electronic device, or a small program that relies on other applications or operating environments, or a plug-in. The embodiments of this disclosure do not limit the specific form of the application. For example, the application can be various types of applications such as video applications and e-commerce platforms, which can respond to the virtual character's awakening conditions to awaken the virtual character and run it, such as when the user triggers a control to awaken the virtual character, or when the virtual character is automatically awakened on certain pages.
[0064] Among them, virtual characters can be digital people or other non-human images, and there is no restriction on this.
[0065] It is understandable that during the process of a user interacting with a virtual character running in an application, one or more rounds of dialogue may usually be conducted, wherein during each round of dialogue, the electronic device may obtain the user's latest question information and then determine the intention corresponding to the question information.
[0066] For example, the user's question information can be the user's voice, which can also be called question voice. For example, the electronic device can capture the user's voice through a configured microphone, thereby enabling the user to ask questions to the virtual character running in the application in the form of voice. Alternatively, the user's question information can be text input by the user, which can also be called question text. For example, the electronic device can capture the user's input text through a configured display screen, thereby enabling the user to ask questions to the virtual character running in the application in the form of text. The user's question information can also be provided in other forms such as pictures and documents, which are not limited in this disclosure.
[0067] For example, a user's question information can be processed by a trained intent recognition model based on a machine learning algorithm to determine the intent corresponding to the user's question information. For example, the user's question information can be input into a pre-trained intent recognition model. The intent recognition model, based on the extracted features, determines the intent corresponding to the question information within a predefined intent category. This intent can reflect the user's true intention and needs. In other words, this intent reflects the user's desire to achieve a certain goal by asking the question. Among them, the intent recognition model and the predefined intent categories can be set according to needs, and there is no limitation on this. The intent recognition model can be trained through existing training methods. For example, the question information sample can be input into the intent recognition model to obtain the predicted intent category, and the loss between the predicted intent category and the intent category label corresponding to the question information sample is calculated. The parameters of the intent recognition model are adjusted based on the loss to obtain the trained intent recognition model; as an example, the predefined intent categories may include: chat intent, professional field intent, intention related to the user's environment, intention related to the user's status, etc.; chat intent means that the user wants to chat with the virtual character. For example, the user's question information is "Tell me a joke", and the corresponding intention is chat intent; professional field intent means that the user wants to obtain information in a specific field (such as law) from the virtual character. , physics, biology, finance, chemistry, mathematics, literature, etc.), for example, the user's question information is "What is the expression of the law of universal gravitation", the corresponding intention is the physics field intention; the intention related to the user's environment means that the user wants to communicate with the virtual character involving the user's current environment. For example, the user's question information is "Is it a little dim to read under the current light?", which involves the intensity of light in the user's current environment, and the corresponding intention is the intention related to the user's environment; the intention related to the user's status means that the user wants to communicate with the virtual character involving the user's current status (such as the user's own clothing, makeup, height, weight, blood pressure, movements, etc.), for example, the user's question information is "How about my outfit?", which involves the user's own clothing, and the corresponding intention is the intention related to the user's status.
[0068] For example, the input data of the trained intent recognition model can be data in text format. When the user's question information is a question voice, the electronic device can use automatic speech recognition technology (Automatic Speech Recognition, ASR) to convert the user's question information in the form of question voice into question text, and then input the question text into the trained intent recognition model to determine the intention corresponding to the question voice.
[0069] In one possible implementation, determining the intent corresponding to the question information includes: obtaining the user's historical questions stored in the application; and determining the intent corresponding to the question information based on the historical questions. Considering that different users have different language expression habits, such as the use of inverted sentences, accents, swallowed words, and the use of dialects or slang, and that the user's historical questions can reflect the user's language expression habits, combining the user's historical questions can more accurately determine the intent corresponding to the user's question information.
[0070] For example, the user's habitual language expression habits can be analyzed from their historical questions. For example, if inverted sentences are frequently used in historical questions, it can be known that the user is accustomed to using inverted sentences. When determining the intent corresponding to the user's question information, if inverted sentences are detected, the inverted sentences can be adjusted to formal sentences and then input into the intent recognition model to more accurately determine the intent corresponding to the question information. For another example, if word swallowing occurs frequently in historical questions, when determining the intent corresponding to the user's question information, if word swallowing is detected, the question information can be completed and input into the intent recognition model, thereby more accurately determining the intent corresponding to the question information based on the comprehensive sentence.
[0071] Step 202: Determine whether the intention corresponding to the question information is a preset intention, and if the intention corresponding to the question information is the preset intention, trigger the application to obtain target information that matches the preset intention; wherein the preset intention is related to the user's environment and / or the user's status.
[0072] Exemplarily, the preset intention may be an intention related to the user's environment or an intention related to the user's status in the above-mentioned predefined intention categories; thereby, it can be determined whether the intention corresponding to the question information determined from the above-mentioned predefined intention categories is the preset intention.
[0073] Among them, when the preset intent is an intent related to the user's state, that is, when the preset intent is related to the user's state, if the intent corresponding to the user's question information is the preset intent, the application needs to combine the user's current state to accurately and comprehensively understand the user's question information. For example, if the user's question information is "How is my outfit?", the intent corresponding to the question information is an intent related to the user's state, and it is necessary to obtain the user's actual clothing situation to more accurately answer the question information. For another example, if the user's question information is "Is my blood pressure high?", the intent corresponding to the question information is an intent related to the user's state, and it is necessary to obtain the user's blood pressure data to more accurately answer the question information. Similarly, when the preset intent is an intent related to the user's environment, that is, when the preset intent is related to the user's environment, if the intent corresponding to the question information is the preset intent, the application needs to combine the user's current environment to accurately and comprehensively understand the user's question information. For example, if the user's question information is "Is it a little dim to read in the current light?", the intent corresponding to the question information is an intent related to the user's environment, and it is necessary to obtain the light intensity in the user's environment to more accurately answer the question information.
[0074] In a possible implementation, the target information includes one or more of the user's image information, the user's environment image information, the user's environment sound information, the user's motion information, the user's tactile information, and the user's physiological information.
[0075] Among them, the user's image information represents an image containing parts of the user's body, such as the user's face, hand, eyebrow, etc.; the image information of the user's environment represents an image containing parts of the user's environment, such as an image of the user's room; the sound information of the user's environment represents the background sound of the user's environment, such as music played in the environment, wind, rain, traffic noise, etc.; the user's motion information represents the user's body motion information; the user's physiological information represents information reflecting the user's vital signs, such as blood pressure, heart rate, respiratory rate, sleep duration, etc.; the user's tactile information represents tactile feedback information of the user's palms, fingers, and other body parts. For example, the user's image information and the image information of the user's environment can be obtained by shooting with a camera, wherein the camera can be a camera of the electronic device running the application, or a camera connected to the electronic device; the sound information of the user's environment can be collected by a microphone, wherein the microphone can be a microphone of the electronic device running the application, or a microphone connected to the electronic device; the user's physiological information, motion information, or tactile information can be obtained through a wearable device worn by the user (such as a smart watch, fitness ring, gaming gloves, etc.).
[0076] Since the target information and the user's question information have different information types or information sources, the target information and the question information form multimodal data.
[0077] It is understood that the user's image information, motion information, tactile information, and physiological information can all reflect the user's state and match the intent associated with the user's state. Image information and sound information of the user's environment can reflect the user's current environment and match the intent associated with the user's environment. As an example, the preset intent is associated with the user's state; triggering the application to obtain target information matching the preset intent includes: activating a camera and controlling the camera to capture image information of the user; and / or obtaining at least one of the user's motion information, tactile information, and physiological information captured by the user's wearable device. Exemplarily, when activating the camera and controlling the camera to capture image information of the user, the electronic device's built-in camera or an external camera connected to the electronic device can be controlled to activate and capture the user's image information once or continuously, thereby enabling the electronic device to "visually" perceive the user's current state. Furthermore, the camera is activated for capture only when the intent corresponding to the user's question information is associated with the user's state, thereby avoiding the waste of image acquisition and processing resources caused by directly capturing and processing image information. For example, when obtaining at least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device, information on the user's corresponding status collected by the wearable device in the previous period can be obtained; or, when it is determined that the preset intention is related to the user's status, the electronic device can trigger the user's wearable device to start up to collect information on the user's corresponding status.
[0078] Exemplarily, in a scenario where a preset intention is related to the user's state, when the application is triggered to obtain target information matching the preset intention, the electronic device can further analyze the question information to determine which aspect of the user's state the question information specifically relates to. For example, keywords can be preset for different aspects of the state, and based on the preset keywords contained in the question information, it can be determined that the question information relates to one or more aspects of the user's action, touch, vision, and physiology, thereby obtaining information on the corresponding aspects of the state; wherein, if the question information includes preset keywords for the visual aspect of the state, the user's image information is obtained; if the question information includes preset keywords for the action aspect of the state, the user's action information is obtained; if the question information includes preset keywords for the tactile aspect of the state, the user's tactile information is obtained; and if the question information includes preset keywords for the physiological aspect of the state, the user's physiological information is obtained.
[0079] As another example, the preset intention is related to the environment in which the user is located; the triggering of the application to obtain target information matching the preset intention includes: starting the camera and controlling the camera to collect image information of the user's environment, and / or starting the microphone and controlling the microphone to collect sound information of the user's environment. Exemplarily, when the camera is started and controlled to collect image information of the user's environment, the built-in camera of the electronic device or the external camera connected to the electronic device can be controlled to start and take one or continuous shots to collect image information of the user's environment, so that the electronic device can "visually" perceive the user's current environment. At the same time, the camera is started for shooting only when the intention corresponding to the user's question information is related to the user's environment, thereby avoiding the waste of data collection and processing resources caused by directly collecting and processing image information. For example, when the microphone is started and controlled to collect sound information of the user's environment, the built-in microphone of the electronic device or the external microphone connected to the electronic device can be controlled to start and collect sound, so that the electronic device can perceive the user's current environment through "hearing". At the same time, the microphone is started to collect sound only when the intention corresponding to the user's question information is related to the user's environment, thereby avoiding the waste of data collection and processing resources caused by directly collecting and processing sound information.
[0080] Exemplarily, in a scenario where the preset intention is related to the user's environment, when the application is triggered to obtain target information that matches the preset intention, the electronic device can further analyze the question information to determine what type of data the question information specifically involves (such as visual data, auditory data, etc.). For example, keywords can be preset for different types of data, and the type of data involved in the question information can be determined based on the preset keywords contained in the question information, so that the corresponding type of environmental information can be obtained; wherein, if the question information includes visual preset keywords, image information of the user's environment is obtained, and if the question information includes auditory preset keywords, sound information of the user's environment is obtained.
[0081] It should be noted that the user's image, sound, physiological and other data involved in this disclosure are obtained with the user's consent and authorization, and the collection, use and processing of relevant data comply with the relevant laws, regulations and standards of relevant countries and regions.
[0082] In one possible implementation, when an electronic device controls a camera to capture image information of a user, the camera's shooting parameters may differ from those used when the electronic device controls the camera to capture image information of the user's environment. The electronic device can set the camera's shooting parameters as needed, or select from a preset combination of shooting parameters, to meet the needs of the type of image information being captured. The camera's shooting parameters may include viewing angle, exposure, white balance, frame rate, aperture, and so on. For example, when a camera captures image information of a user, the user is typically close to the camera, so the camera may use a narrower viewing angle to capture more facial or body details. When the camera captures image information of the user's environment, a wide-angle lens or a wider viewing angle may be used to maximize the camera's coverage of the surrounding environment. For another example, when a camera captures image information of a user, the white balance may be set based on the user's skin tone to ensure a natural appearance and avoid color casts. This allows for more accurate understanding of user questions involving skin tone. When the camera captures image information of the user's environment, the white balance may be set based on the color of the ambient light source to ensure accurate color reproduction of objects in the captured image, allowing for more accurate identification of objects in the environment. For another example, when a camera captures image information of a user, it can adjust the exposure based on the lighting conditions of the user's face or skin to avoid overexposure or underexposure, so that the user's face or body parts can be clearly seen; when the camera captures image information of the user's environment, the exposure can use HDR mode, so that high-contrast objects in the environment can be clearly displayed. For another example, when a camera captures image information of a user, it can use a relatively large aperture and a high frame rate, so that the user's expression or movement details can be captured more clearly; when the camera captures image information of the user's environment, it can use a relatively small aperture and a low frame rate, so that more objects in the environment can be captured.
[0083] Step 203: Process the target multimodal data through the target model to generate response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data.
[0084] Among them, Large Language Models (LLMs) are deep learning models trained using a large amount of text data, with the purpose of generating text, understanding natural language or completing related language processing tasks. Large language models usually have a large number of parameters and can capture the complex patterns and dependencies of language; for example, the target model can be a large language model for processing multimodal data built on the basis of Tongyi Qianwen, pre-trained generative transformer (GPT) model, Meta AI large language model (LLaMa), etc., or it can be a large language model with multi-model data processing capabilities; the training process of the target model can be based on existing technology and will not be repeated here. Since the user's question information and target information are data of different modalities, the target model processes the user's question information and target information that matches the preset intention, etc., so that the multimodal data can be comprehensively analyzed, and the complementarity and correlation between different modal data can be used to improve the understanding of the user's question information, thereby generating more accurate, natural and reasonable response information.
[0085] In one possible implementation, the processing of target multimodal data by a target model also includes: selecting a preset model that matches the preset intention from multiple preset models as the target model; wherein different preset models are used to process information that matches different intentions.
[0086] For example, based on a large language model that has been trained on a huge corpus and has the ability to understand general texts, it is possible to further fine-tune information matching different intentions, thereby obtaining multiple preset models for processing information matching different intentions. For example, on the basis of a large language model for processing multimodal data, the question information and one or more of the user's image information, image information of the environment, action information, tactile information, and physiological information can be used as multimodal training data to further train and fine-tune the large language model, thereby obtaining a preset model for processing information matching different intentions. In this way, each preset model not only has the general feature capability of learning language, but can also better process information matching the corresponding intention, thereby improving the ability to understand the question information of the corresponding intention, so as to generate the best response information.
[0087] Furthermore, considering that users of different ages, or the same user in different emotions, may have different responses to the same question; for example, for users of different ages, adults usually expect to obtain more comprehensive response information, while minors are more likely to accept response information that is easy to understand; for another example, when users are happy, they are more tolerant of the content of the response information, while when users are angry or sad, they are often more sensitive to the content of the response information. Therefore, the target model can comprehensively consider attributes such as the user's age and / or emotion to generate response information that is easier for users to accept.
[0088] In one possible implementation, the processing of target multimodal data through a target model to generate response information includes: processing the question information and / or the target information through the target model to extract user attribute characteristics; and processing the question information and the target information to extract multimodal characteristics; wherein the attribute characteristics include: age and / or emotion; and generating the response information based on the attribute characteristics and the multimodal characteristics.
[0089] The target model has functions such as age recognition, emotion recognition, and multimodal recognition. For example, the target model can be configured with pre-trained age recognition, emotion recognition, and other attribute recognition sub-models, as well as a pre-trained multimodal recognition sub-model.
[0090] Exemplarily, an age space and an emotion space can be pre-constructed. For example, the age space can include minors, adults, the elderly, etc., and the emotion space can include: anger, sadness, happiness, tension, confidence, pain, loss, contempt, gratitude, excitement, etc. Then, in the constructed age space and emotion space, the corresponding attribute recognition sub-model is trained to obtain a trained age recognition sub-model and emotion recognition sub-model. The trained age recognition sub-model and / or emotion recognition sub-model are configured in the target model. In this way, the question information and / or target information can be processed by each attribute sub-model of the target model to extract the corresponding attribute features; at the same time, the question information can be processed by the trained multimodal recognition sub-model to extract multimodal features; and then, based on the attribute features and multimodal features, response information that conforms to the user's attributes is generated. The attribute sub-model and the multimodal recognition sub-model can be implemented based on relevant technologies, and the present disclosure does not limit their specific implementation methods.
[0091] In some scenarios, the target model can be used to process target information that matches the preset intent and extract the user's attribute features. As an example, the target model can be used to process the user's facial image to extract the user's age features. Based on the age features, combined with the multimodal features extracted from the user's question information and target information, response information that is consistent with the age can be generated. For example, when an adult is talking to a virtual character, the camera can capture the adult's facial image, and the facial image can be used to predict that the current user is an adult. In this way, response information that is consistent with an adult can be generated, such as response information with more comprehensive and detailed content. For another example, when a child is talking to a virtual character, the electronic device can capture the child's facial image through the camera, and the facial image can be used to predict that the current user is a minor. In this way, response information that is consistent with a minor can be generated, such as response information with easy-to-understand content. As another example, the user's image can be processed by the target model to extract the user's emotional features, and then based on the emotional features, combined with the multimodal features extracted from the user's question information and target information, response information consistent with the emotion can be generated. For example, when the user's mouth corners are upturned, or the user is dancing, it means that the user is happy, and response information consistent with happiness can be generated, such as response information containing humorous words; for another example, when the user's expression is frowning, the corners of the mouth are drooping, or the head and shoulders are drooping, it means that the user is disappointed, and response information consistent with disappointment can be generated, such as response information containing encouraging words; for another example, when the user covers his face with his hands and his body trembles slightly, it means that the user is sad, and response information consistent with sadness can be generated, such as response information containing comforting words; for another example, when the user's body is straight and he holds his chest and head high, it means that the user is confident, and response information consistent with confidence can be generated, such as response information containing praising words.
[0092] In some scenarios, the user's question information can be processed through the target model to extract the user's attribute characteristics. For example, the user's emotional characteristics can be extracted from the user's question voice, and then the target model can generate response information that matches the emotion based on the user's emotional characteristics and combined with the multimodal features extracted from the question voice and the target information. Considering that the speaker's emotions are different, the acoustic characteristics of the spoken speech will vary. For example, the speaker's tone may be higher when happy and lower when sad. For example, the speaking speed may increase when excited and slow down when frustrated. For example, when nervous, there may be more pauses when speaking. Therefore, by analyzing the tone, speed, volume, pauses, etc. of the question voice, the user's current emotional characteristics can be extracted.
[0093] For example, the target model can also be configured in a cloud server, and the electronic device can also send multimodal data to the cloud, thereby invoking the target model to process the multimodal data based on the cloud computing resources to generate response information. For example, the electronic device can convert the multimodal data into object storage service (OSS) format data and then send it to the cloud server, so that the server can use the target model to process the target multimodal data and generate response information.
[0094] Step 204: announce the response information to the user via the virtual character on the display screen.
[0095] A virtual character refers to a digitized or virtualized character or individual that exists in a virtual environment. It can be an animated character, a virtual assistant, or a character generated by artificial intelligence (AI). As an example, a virtual character can be displayed in a holographic space on a display screen. For example, multiple virtual character images can be pre-designed, and users can select the virtual character image according to their preferences.
[0096] In one possible implementation, the response information includes a response text; and broadcasting the response information to the user through the virtual character on the display screen includes: generating a response voice and a response action based on the response text; driving the physical movement of the virtual character based on the response action, and synchronously playing the response voice.
[0097] For example, TTS (Text To Speech) technology can be used to convert the response text into a response voice. Considering that the response information generated by the target model is usually in text form, that is, the response text, text-to-speech technology can be used to convert the response text into a response voice. This can then be driven by the virtual character to broadcast the response voice to the user, enabling a voice conversation between the user and the virtual character, thereby enhancing the user's immersive interactive experience.
[0098] In some scenarios, the response text can be converted into a response voice of the target timbre. The target timbre can be the timbre selected by the user, or the timbre of a pre-set different virtual character image, or the timbre of the user's preference determined based on the user's historical interaction data. The corresponding TTS model can be pre-trained for different timbres. After determining the target timbre, the TTS model corresponding to the target timbre is selected, and the response text is input into the TTS model. The TTS model can generate a response voice of the target timbre. There is no restriction on the model structure of the TTS model, and the training process of the TTS model can be carried out in an existing manner.
[0099] For example, the response text generated by the target model can be analyzed and processed based on a pre-trained action generation model to generate a corresponding response action. The specific structure and training process of the action generation model can be carried out using existing methods. Furthermore, existing animation generation technology can be used to drive the physical movements of the virtual character based on the response action, thereby generating the action animation of the virtual character. The response voice and the action animation can be synchronized through timestamps to achieve synchronous playback of the response voice. As an example, the virtual character can have a three-dimensional digital body. While broadcasting the response information, the three-dimensional digital body of the virtual character can be driven to move synchronously based on the response action.
[0100] For example, by designing details such as the virtual character's facial expressions, skin texture, and eye luster, a realistic and detailed reproduction of real-world human images can be achieved, allowing viewers to visually feel closer to the feeling of talking to a real person, thereby providing users with a face-to-face, hyper-realistic, immersive interactive experience. For example, historical user interaction data can be used to determine user preferences, and based on these preferences, the virtual character's facial expressions, skin texture, and corresponding character image can be determined.
[0101] In one possible implementation, broadcasting the response information to the user via the virtual character includes: determining a target state for the virtual character when broadcasting the response information based on the target information; and controlling the virtual character to broadcast the response information to the user in the target state. In this way, the virtual character broadcasts the response information to the user in the target state determined by the target information. This ensures that the virtual character's broadcasting state blends with the user's environment and / or matches the user's state, making the user feel more intimate and attentive, increasing the user's willingness to interact with the virtual character, and helping the user better understand the meaning of the response information. For example, the target state may include expression, sound, brightness, movement, etc.
[0102] As an example, the target state of the virtual character when broadcasting can be determined based on the user's image information; illustratively, the user's emotions can be analyzed through the user's image information, and the facial expression of the virtual character when broadcasting can be determined based on the emotions; for example, when the user is happy, the virtual character's eyes can be controlled to be smiling, the virtual character's mouth can be controlled to be a happy state with upturned corners, and so on; and when the user is sad, the virtual character's mouth can be controlled to be a sad state with downturned corners, and so on; in this way, in response to questions asked by the user with different emotions, the virtual character can make corresponding facial expressions when answering, so that the emotion of the virtual character when broadcasting the response information matches the user's emotion.
[0103] As another example, the target state of the virtual character when making a report can be determined based on the image information of the user's environment. For example, the ambient light intensity and the scene the user is in can be analyzed through the image information of the user's environment, so as to determine the brightness, sound, and movement of the virtual character when making a report. For example, when the ambient light intensity is strong, the brightness of the virtual character can be increased. For another example, when the user is identified as being in an outdoor scene such as a street, the virtual character's voice can be increased, while when the user is identified as being in a bedroom, the virtual character's voice can be lowered. For another example, when the user is identified as being in the snow, the virtual character can be controlled to "hug his whole body tightly and tremble" while reporting. In this way, the brightness, sound, movement, etc. of the virtual character are adjusted accordingly to the user's questions in different environments to blend in with the environment.
[0104] As another example, the target state of the virtual character's announcement can be determined based on the sound information of the user's environment. For example, the noise level of the environment can be analyzed based on the sound information of the user's environment, thereby determining the voice of the virtual character when announcing the announcement. For example, if the environment is relatively noisy, the virtual character's voice can be increased. For another example, if the background noise of the scene in which the user is located is relatively loud, the virtual character's voice can be increased. In this way, the virtual character's voice can be adjusted accordingly to the user's questions in different environments to blend in with the environment.
[0105] In an embodiment of the present disclosure, in response to a question message sent by a user to a virtual character running in an application, an intent corresponding to the question message is determined; a determination is made as to whether the intent corresponding to the question message is a preset intent, and if so, the application is triggered to obtain target information matching the preset intent; wherein the preset intent is related to the user's environment and / or the user's state; target multimodal data is processed using a target model to generate response information; wherein the target multimodal data includes the question message and the target information; the target model is a pre-built large language model for processing multimodal data; and the response information is broadcasted to the user by the virtual character. In this way, based on multimodal interaction technology and intent recognition technology, target information can be obtained when the intent corresponding to the user's question message is related to the user's environment and / or the user's state; and the large language model is used to process the multimodal data, utilizing the complementarity and correlation between different modal data to improve the virtual character's ability to understand the user's question message, thereby providing more realistic and accurate answers to questions related to the user's environment or the user's state, generating more accurate, natural, and reasonable response information, and achieving immersive interaction with hyper-realistic virtual characters. At the same time, unlike directly acquiring multimodal data, the target information is acquired only when the intention corresponding to the user's question information is the preset intention, thus avoiding the waste of data collection and processing resources caused by directly acquiring the target information.
[0106] The following example illustrates the above virtual character interaction method by taking a scenario where a user interacts with a virtual character through voice (i.e., the user's question information is the question voice, and the virtual character broadcasts the answer voice), the preset intention is related to the user's status, and the target message is the user's image information.
[0107] Figure 3 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown. Figure 3 As shown, the following steps are included:
[0108] Step 301: In response to a user's voice asking a question to a virtual character running in an application, determine an intention corresponding to the voice asking question.
[0109] This step 301 can be used as the above Figure 2 A possible implementation of step 201 in FIG.
[0110] As an example, the user's question voice can be received through a microphone configured by the electronic device, so that the user's question can be perceived through "hearing".
[0111] For example, the question voice can be converted into question text, and then the question text is input into a trained intent recognition model to determine the intent corresponding to the question voice.
[0112] Step 302: Determine whether the intention corresponding to the question voice is an intention related to the user's status, and if the intention corresponding to the question voice is an intention related to the user's status, trigger the application to obtain user image information that matches the intention related to the user's status.
[0113] This step 302 can be used as the above Figure 2 A possible implementation of step 202.
[0114] Among them, the question voice involves the user's visual status; for example, the user's question voice can be "How about my outfit?", "Are the dark circles under my eyes serious?" and other questions involving the user's visual status.
[0115] For example, if the intent corresponding to the voice question is related to the user's state, the preset keywords involved in the voice question can be analyzed to determine whether the voice question relates to the user's visual state, thereby triggering the application to obtain the user's image information. For example, the preset keywords can be words such as "outfit," "appearance," "eyes," "hair," and other entities that require visual observation.
[0116] Exemplarily, the triggering application obtains the user's image information that matches the intention related to the user's status, including: starting the camera, and controlling the camera to collect the user's image information; thereby being able to perceive the user's current status from a "visual" perspective.
[0117] In one possible implementation, controlling the camera to capture image information of the user includes: controlling the camera to capture at least one image and performing target detection on the captured image until the currently captured image detects that the target body part referred to in the question is included; wherein, if the currently captured image detects that the target body part is not included, adjusting the camera's capture parameters and / or prompting the user to adjust their own or the camera's posture. The number of target body parts can be one or more, thereby ensuring that the captured image information of the user contains the target body part referred to in the user's question, thereby enabling more accurate understanding and response to the user's question.
[0118] Exemplarily, the electronic device can analyze the user's question information. For example, it can analyze the keywords of different body parts contained in the question information (such as face, hands, legs, arms, neck, mouth, eyebrows, hair, etc.) to determine the target body part involved in the user's question information, and then after the electronic device controls the camera to take each shot, it performs target detection on the image taken this time to determine whether the image taken this time contains the target body part. If it is detected that the image taken this time contains the target body part, the electronic device controls the camera to end image acquisition and uses the image taken this time as the image information of the user collected; if it is detected that the image taken this time does not contain the target body part, the electronic device adjusts the shooting parameters of the camera and controls the camera to continue the next shooting with the adjusted shooting parameters. Alternatively, the electronic device can prompt the user to adjust his or her own posture or the posture of the camera, and after the user adjusts his or her own posture or the posture of the camera, controls the camera to continue the next shooting.
[0119] As an example, when adjusting camera shooting parameters, you can adjust the camera's angle of view, exposure, white balance, resolution, aperture, and so on. For example, if the currently captured image only contains part of the user's face and hair, you can lower the camera's angle of view and take the next shot to capture the user's entire face or more of their body parts. For another example, if the user's skin is too dark or underexposed in the currently captured image, you can increase the exposure or adjust the white balance and take the next shot to capture the user's eyebrows, hair, and other body parts. For another example, if the body part is blurry and indistinguishable in the currently captured image, you can increase the aperture or resolution and take the next shot to capture a clearer body part. As another example, when prompting a user to adjust their posture, a voice prompt can be issued to the user, such as, "Please stand back a little so I can see your feet," "Please look at me," etc., to prompt the user to adjust their posture. As another example, when prompting the user to adjust the camera's posture, a prompt voice may be issued to the user, such as, "point the phone downward a little", "turn the camera to the left", etc., to prompt the user to adjust the camera's posture.
[0120] Step 303: Process the target multimodal data through the target model to generate a response text; wherein the target multimodal data includes the question voice and the image information of the user.
[0121] This step 303 can be used as the above Figure 2 A possible implementation of step 203.
[0122] Step 304: Generate a response voice and a response action based on the response text, drive the virtual character's physical movement based on the response action, and synchronously play the response voice.
[0123] This step 304 can be used as the above Figure 2 A possible implementation of step 204.
[0124] In the embodiment of the present disclosure, a user can initiate a voice question to an electronic device. Based on multimodal interaction technology and intent recognition technology, the electronic device obtains the user's image information when the intention corresponding to the user's voice question is an intention related to the user's state and the voice question involves the user's visual state. Then, a large language model is used to process the user's image information and the voice question, and other data of different modalities. By utilizing the complementarity and correlation between the different modal data, the virtual character's ability to understand questions related to the user's visual state is improved, and more accurate, natural, and reasonable response information is generated. At the same time, when responding to the interactive feedback of the voice question, the virtual character's body movements can be driven and the response voice can be played synchronously, thereby achieving immersive interaction with the hyper-realistic virtual character and significantly improving the interactive experience between the user and the virtual character. In addition, unlike directly obtaining multimodal data, the user's image information is only obtained when it is determined that the intention corresponding to the user's voice question is an intention related to the user's state, thereby avoiding the waste of data collection and processing resources caused by directly obtaining the user's image information.
[0125] The following example illustrates the above-mentioned virtual character interaction method by taking a scenario in which a user interacts with a virtual character through voice (i.e., both question information and answer information are messages in voice form), the preset intention is related to the user's status, and the target message is one or more of the user's action information, the user's tactile information, and the user's physiological information.
[0126] Figure 4 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown. Figure 4 As shown, the following steps are included:
[0127] Step 401: In response to a user's voice asking a question to a virtual character running in an application, determine an intention corresponding to the voice asking question.
[0128] This step 401 is similar to the above Figure 3 The same as step 301 in FIG. 1 is omitted here for brevity.
[0129] Step 402: Determine whether the intention corresponding to the question voice is an intention related to the user's state, and if the intention corresponding to the question voice is an intention related to the user's state, trigger the application to obtain one or more of the user's action information, user's tactile information, and user's physiological information that match the intention related to the user's state.
[0130] This step 402 can be used as the above Figure 2 A possible implementation of step 202.
[0131] The question voice involves at least one of the user's actions, touch, and physiology. For example, when the intention corresponding to the question voice is an intention related to the user's state, the preset keywords involved in the question voice can be analyzed to determine the action, touch, physiology, and other aspects of the state involved in the question voice, so as to trigger the application to obtain the user's corresponding state information. For example, if the question voice involves the user's tactile state, the application is triggered to obtain the user's tactile information. The preset keywords corresponding to different aspects of the state can be set as needed. For example, the preset keywords corresponding to the action state can be words related to the user's action, such as "posture", "action", etc. The preset keywords corresponding to the tactile state can be words related to the user's touch, such as "vibration", "tactile", etc. The preset keywords corresponding to the physiological state can be words related to the user's physiological indicators, such as "blood pressure", "sleep", and "heart rate".
[0132] Exemplarily, the application is triggered to obtain one or more of the user's motion information, user's tactile information, and user's physiological information that match the intention related to the user's state, including: obtaining at least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device, so as to perceive the user's current state from different aspects.
[0133] As an example, if a user is wearing a smartwatch with a blood pressure sensor, the user can ask the electronic device a question such as "Is my blood pressure high?", and the electronic device can then obtain the user's blood pressure data collected by the smartwatch. For example, if the smartwatch is already connected to the electronic device, the electronic device can directly request the blood pressure data from the smartwatch, or obtain the blood pressure data that the smartwatch has uploaded to the electronic device. If the electronic device and the smartwatch are not connected, the connection can be initiated (e.g., Bluetooth connection), or the user can be prompted to connect the electronic device and the smartwatch.
[0134] As another example, while exercising with a fitness ring, the user can ask the electronic device a question, "Is my posture correct?" The electronic device can then obtain the user's motion information collected by the fitness ring. For example, if the fitness ring is already connected to the electronic device, the electronic device can directly request the user's motion information from the fitness ring, or obtain the user's motion information that the fitness ring has uploaded to the electronic device. If the electronic device and the fitness ring are not connected, the connection can be initiated or the user can be prompted to connect the electronic device and the fitness ring.
[0135] As another example, while playing a game while wearing gaming gloves with haptic feedback, a user can ask the electronic device, "Why did the vibration in my hand suddenly intensify?" The electronic device can then obtain the user's tactile information collected by the gaming gloves. For example, if the gaming gloves are already connected to the electronic device, the electronic device can directly request the user's tactile information from the gaming gloves, or obtain the user's tactile information that the gaming gloves have already uploaded to the electronic device. If the electronic device and gaming gloves are not connected, the connection can be initiated or the user can be prompted to connect the electronic device and the gaming gloves.
[0136] Step 403: Process the target multimodal data through the target model to generate a response text; wherein the target multimodal data includes the question voice and one or more of the user's motion information, user's tactile information, and user's physiological information.
[0137] This step 403 can be used as the above Figure 2 A possible implementation of step 203.
[0138] Step 404: Generate a response voice and a response action based on the response text, drive the virtual character's body movement based on the response action, and synchronously play the response voice.
[0139] This step 404 can be used as the above Figure 2 A possible implementation of step 204.
[0140] In the disclosed embodiment, a user can initiate a voice question to an electronic device. The electronic device, based on multimodal interaction technology and intent recognition technology, obtains the user's motion information, tactile information, or physiological information when the intention corresponding to the user's voice question is an intention related to the user's state and the voice question involves the user's motion, touch, or physiological state. Then, a large language model is used to process the different modal data such as the voice question and the user's motion information, tactile information, or physiological information. By utilizing the complementarity and correlation between the different modal data, the virtual character's ability to understand questions involving the user's motion, touch, or physiological state is improved, and more accurate, natural, and reasonable response information is generated. At the same time, the virtual character's body movements can be driven during interactive feedback on the voice question, and the response voice can be played synchronously, thereby achieving immersive interaction with the hyper-realistic virtual character and significantly improving the user's interactive experience with the virtual character. In addition, unlike directly obtaining multimodal data, the user's motion information, tactile information, or physiological information is only obtained when it is determined that the intention corresponding to the user's voice question is an intention related to the user's state, thereby avoiding the waste of data collection and processing resources caused by directly obtaining the user's motion information, tactile information, or physiological information.
[0141] The following example illustrates the above virtual character interaction method by taking a scenario where a user interacts with a virtual character through voice, the preset intention is related to the user's environment, and the target message is image information and / or sound information of the user's environment as an example.
[0142] Figure 5 A flowchart of a virtual character interaction method according to an embodiment of the present disclosure is shown. Figure 5 As shown, the following steps are included:
[0143] Step 501: In response to a user's voice asking a question to a virtual character running in an application, determine an intention corresponding to the voice asking question.
[0144] This step 501 is similar to the above Figure 3 The same as step 301 in FIG. 1 is omitted here for brevity.
[0145] Step 502: Determine whether the intention corresponding to the question voice is an intention related to the user's environment, and if the intention corresponding to the question voice is an intention related to the user's environment, trigger the application to obtain image information of the user's environment and / or sound information of the user's environment that matches the intention related to the user's environment.
[0146] This step 502 can be used as the above Figure 2 A possible implementation of step 202.
[0147] For example, the user's voice question can be "Is it too dark to read in the current light?", "What kind of flower is this next to me?", "What music is playing on TV?" and other questions related to the user's environment.
[0148] Exemplarily, triggering the application to obtain image information of the user's environment that matches the intention related to the user's environment includes: starting a camera and controlling the camera to collect image information of the user's environment.
[0149] In one possible implementation, controlling the camera to capture image information of the user's environment includes: controlling the camera to capture at least one image and performing target detection on the captured image until the currently captured image detects that the target object referred to in the question is contained in the currently captured image; wherein, if the currently captured image detects that the target object is not contained in the currently captured image, adjusting the camera's capture parameters and / or prompting the user to adjust the position of the camera or the camera. The number of target objects can be one or more, thereby ensuring that the captured image information of the user's environment contains the target object referred to in the user's question, thereby enabling a more accurate understanding and response to the user's question.
[0150] Exemplarily, the electronic device can analyze the user's question information. For example, it can analyze the keywords of different objects contained in the question information (such as TV, sofa, flower, floor, lamp, cup, etc.) to determine the target object involved in the user's question information, and then after the electronic device controls the camera to take each shot, it performs target detection on the image taken this time to determine whether the image taken this time contains the target object. If it is detected that the image taken this time contains the target object, the electronic device controls the camera to end image acquisition and uses the image taken this time as the image information of the user's environment; if it is detected that the image taken this time does not contain the target object, the electronic device adjusts the camera's shooting parameters and controls the camera to continue the next shooting with the adjusted shooting parameters. Alternatively, the electronic device can prompt the user to adjust his or her own posture or the camera's posture, and after the user has adjusted his or her own posture or the camera's posture, control the camera to continue the next shooting.
[0151] As an example, when adjusting the camera's shooting parameters, the camera's angle of view, exposure, white balance, resolution, aperture, etc. can be adjusted. For example, if the image captured this time only contains the user and a few objects, the camera's angle of view can be increased or switched to a wide-angle lens, and the next shot can be taken to capture more objects in the environment. For another example, if the objects in the image captured this time are blurry and indistinguishable, the aperture or resolution can be increased, and the next shot can be taken to capture clearer objects. As another example, when prompting the user to adjust their posture, a prompt voice message can be issued to the user, such as, "Please stand a little to the left so as not to block the TV," etc., to prompt the user to adjust their posture. As another example, when prompting the user to adjust the camera's posture, a prompt voice message can be issued to the user, such as, "Point the phone down a little," "Turn the camera to the left," etc., to prompt the user to adjust the camera's posture.
[0152] Exemplarily, triggering the application to obtain sound information of the user's environment that matches an intention related to the user's environment includes: activating a microphone and controlling the microphone to collect sound information of the user's environment. For example, the microphone may be controlled to collect music playing on a TV, wind, rain, traffic noise, etc.
[0153] Step 503: Process the target multimodal data through the target model to generate a response text; wherein the target multimodal data includes the question voice and image information of the user's environment and / or sound information of the user's environment.
[0154] This step 503 can be used as the above Figure 2 A possible implementation of step 203.
[0155] Step 504: Generate a response voice and a response action based on the response text, drive the virtual character's body movement based on the response action, and synchronously play the response voice.
[0156] This step 504 can be used as the above Figure 2 A possible implementation of step 204.
[0157] In the embodiment of the present disclosure, a user can initiate a voice question to an electronic device. The electronic device, based on multimodal interaction technology and intent recognition technology, obtains image information and / or sound information of the user's environment when the intention corresponding to the user's voice question is an intention related to the user's environment. Then, a large language model is used to process the image information and / or sound information of the user's environment and the voice question, and the data of different modalities such as the voice question. By utilizing the complementarity and correlation between the different modal data, the virtual character's ability to understand questions related to the user's environment is improved, and more accurate, natural, and reasonable response information is generated. At the same time, the virtual character's body movements can be driven during the interactive feedback of the voice question, and the response voice can be played synchronously, thereby achieving immersive interaction with the virtual character in a hyper-realistic manner and significantly improving the interactive experience between the user and the virtual character. In addition, unlike directly obtaining multimodal data, the image information and / or sound information of the user's environment are obtained only when it is determined that the intention corresponding to the user's voice question is an intention related to the user's environment, thereby avoiding the waste of data collection and processing resources caused by directly obtaining the image information and / or sound information of the user's environment.
[0158] Figure 6 A flowchart of a method for virtual character interaction according to an embodiment of the present disclosure is shown as follows: Figure 6 As shown, first, the user can trigger the interaction with the virtual character. In each round of dialogue, the electronic device obtains the user's question information and uses the intention recognition model to determine the intention corresponding to the question information among the chat intention, professional field intention, intention related to the user's environment, and intention related to the user's status (i.e., execution intention). Figure 2 The preset intention is pre-set as an intention related to the user's environment and an intention related to the user's state. Then, when the intention corresponding to the question information is an intention related to the user's environment, the image information of the user's environment and / or the sound information of the user's environment are obtained accordingly. When the intention corresponding to the question information is an intention related to the user's state, the image information, action information, tactile information or physiological information of the user are obtained accordingly (i.e., execution Figure 2 Step 202); Then, the above-obtained information is processed by calling the corresponding large language model to generate a response message (i.e., executing Figure 2 Finally, the virtual character broadcasts the response information to the user (i.e., executes Figure 2 Step 204); thereby achieving ultra-realistic immersive interaction.
[0159] Based on the same inventive concept of the above method embodiment, an embodiment of the present disclosure further provides a virtual character interaction device, which can be used to execute the technical solution described in the above method embodiment.
[0160] Figure 7 A structural diagram of a virtual character interaction device according to an embodiment of the present disclosure is shown as follows: Figure 7 As shown, the device may include: a questioning module 701, for determining the intention corresponding to the questioning information in response to the questioning information asked by the user to the virtual character running in the application; an acquisition module 702, for judging whether the intention corresponding to the questioning information is a preset intention, and when the intention corresponding to the questioning information is the preset intention, triggering the application to obtain target information matching the preset intention; wherein the preset intention is related to the environment in which the user is located and / or the state of the user; a response module 703, for processing the target multimodal data through a target model to generate response information; wherein the target multimodal data includes the questioning information and the target information; the target model is a pre-built large language model for processing multimodal data; a broadcasting module 704, for broadcasting the response information to the user through the virtual character.
[0161] In an embodiment of the present disclosure, in response to a question message sent by a user to a virtual character running in an application, an intent corresponding to the question message is determined; a determination is made as to whether the intent corresponding to the question message is a preset intent, and if so, the application is triggered to obtain target information matching the preset intent; wherein the preset intent is related to the user's environment and / or the user's state; target multimodal data is processed using a target model to generate response information; wherein the target multimodal data includes the question message and the target information; the target model is a pre-built large language model for processing multimodal data; and the response information is broadcasted to the user by the virtual character. In this way, based on multimodal interaction technology and intent recognition technology, target information can be obtained when the intent corresponding to the user's question message is related to the user's environment and / or the user's state; and the large language model is used to process the multimodal data, utilizing the complementarity and correlation between different modal data to improve the virtual character's ability to understand the user's question message, thereby providing more realistic and accurate answers to questions related to the user's environment or the user's state, generating more accurate, natural, and reasonable response information, and achieving immersive interaction with hyper-realistic virtual characters. At the same time, unlike directly acquiring multimodal data, the target information is acquired only when the intention corresponding to the user's question information is the preset intention, thus avoiding the waste of data collection and processing resources caused by directly acquiring the target information.
[0162] In a possible implementation, the target information includes one or more of the user's image information, the user's environment image information, the user's environment sound information, the user's motion information, the user's tactile information, and the user's physiological information.
[0163] In one possible implementation, the preset intention is related to the user's status; the acquisition module 702 is also used to: start the camera and control the camera to collect image information of the user; and / or obtain at least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device.
[0164] In one possible implementation, the preset intention is related to the environment in which the user is located; the acquisition module 702 is also used to: start the camera and control the camera to collect image information of the user's environment; and / or start the microphone and control the microphone to collect sound information of the user's environment.
[0165] In one possible implementation, the processing of target multimodal data by a target model also includes: selecting a preset model that matches the preset intention from multiple preset models as the target model; wherein different preset models are used to process information that matches different intentions.
[0166] In one possible implementation, the response module 703 is further used to: process the question information and / or the target information through the target model to extract the user's attribute characteristics; and process the question information and the target information to extract multimodal characteristics; wherein the attribute characteristics include: age and / or emotion; and generate the response information based on the attribute characteristics and the multimodal characteristics.
[0167] In one possible implementation, the acquisition module 702 is further used to: control the camera to take at least one shot, and perform target detection on the shot image until it is detected that the target body part involved in the question information is included in the currently shot image, wherein, when it is detected that the currently shot image does not contain the target body part, the shooting parameters of the camera are adjusted, and / or the user is prompted to adjust the posture of himself or the camera.
[0168] In a possible implementation, the questioning module 701 is further configured to: obtain historical questions of the user stored in the application; and determine the intention corresponding to the question information based on the historical questions.
[0169] In a possible implementation, the reporting module 704 is further configured to: determine a target state of the virtual character when reporting based on the target information; and control the virtual character to report the response information to the user in the target state.
[0170] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0171] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a volatile or non-volatile computer-readable storage medium.
[0172] An embodiment of the present disclosure further proposes an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to implement the above method when executing the instructions stored in the memory.
[0173] An embodiment of the present disclosure also provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.
[0174] Figure 8 FIG1 shows a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Figure 8 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.
[0175] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.
[0176] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.
[0177] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0178] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0179] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0180] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0181] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0182] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0183] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0184] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0185] While various embodiments of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A virtual character interaction method, characterized in that: The method comprises: In response to a question message sent by a user to a virtual character running in an application, determining an intention corresponding to the question message; determining whether the intention corresponding to the question information is a preset intention, and if the intention corresponding to the question information is the preset intention, triggering the application to obtain target information matching the preset intention; wherein the preset intention is related to the user's environment and / or the user's status; and the target information includes one or more of the user's image information, image information of the user's environment, sound information of the user's environment, user's movement information, user's tactile information, and user's physiological information; Processing target multimodal data through a target model to generate response information; wherein the target multimodal data includes the question information and the target information; the target model is a pre-built large language model for processing multimodal data; announcing the response information to the user through the virtual character; Triggering the application to obtain target information matching the preset intent includes one or more of the following methods: Start the camera and control it to collect user's image information; Acquiring at least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device; Start the camera and control it to collect image information of the user's environment; Start the microphone and control it to collect sound information from the user's environment.
2. The method according to claim 1, characterized in that The processing of the target multimodal data by the target model also includes: A preset model that matches the preset intention is selected from multiple preset models as the target model; wherein different preset models are used to process information matching different intentions.
3. The method according to claim 1, characterized in that The processing of the target multimodal data by the target model to generate response information includes: The question information and / or the target information are processed by the target model to extract the user's attribute characteristics; and the question information and the target information are processed to extract multimodal characteristics; wherein the attribute characteristics include: age and / or emotion; based on the attribute characteristics and the multimodal characteristics, the response information is generated.
4. The method according to claim 1, wherein The controlling camera to collect image information of the user includes: The camera is controlled to take at least one shot and perform target detection on the shot image until it is detected that the target body part involved in the question information is included in the currently shot image, wherein, when it is detected that the currently shot image does not include the target body part, the shooting parameters of the camera are adjusted, and / or the user is prompted to adjust the posture of himself or the camera.
5. The method according to claim 1, wherein Determining the intention corresponding to the question information includes: Obtain the user's historical questions stored in the application; In combination with the historical questions, the intention corresponding to the question information is determined.
6. The method according to claim 1, wherein The step of broadcasting the response information to the user through the virtual character includes: Determining a target state of the virtual character when performing the broadcast according to the target information; The virtual character is controlled to broadcast the response information to the user in the target state.
7. A virtual character interaction device, characterized in that: The device comprises: A questioning module, configured to ask a question to a virtual character running in an application program and determine an intention corresponding to the question; An acquisition module is used to determine whether the intention corresponding to the question information is a preset intention, and if the intention corresponding to the question information is the preset intention, trigger the application to obtain target information matching the preset intention; wherein, the preset intention is related to the user's environment and / or the user's status; the target information includes one or more of the user's image information, the user's environment image information, the user's environment sound information, the user's motion information, the user's tactile information, and the user's physiological information; triggering the application to obtain target information matching the preset intention includes one or more of the following methods: starting the camera and controlling the camera to collect the user's image information; obtaining at least one of the user's motion information, tactile information, and physiological information collected by the user's wearable device; starting the camera and controlling the camera to collect image information of the user's environment; starting the microphone and controlling the microphone to collect sound information of the user's environment; A response module, configured to process target multimodal data using a target model to generate response information; wherein the target multimodal data includes the question information and the target information; and the target model is a pre-built large language model for processing multimodal data; The broadcasting module is used to broadcast the response information to the user through the virtual character.
8. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to implement the method according to any one of claims 1 to 6 when executing the instructions stored in the memory.
9. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, wherein when the computer-readable code is executed in a processor of an electronic device, the processor in the electronic device executes the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Virtual object interaction method and device, electronic equipment and storage medium
CN117950492A
Context environment-based dialogue interaction processing method and device
CN119202332A