Multimodal interaction method and apparatus
By receiving multimodal data to identify user intentions and postures, determining interaction strategies and generating three-dimensional rendering models, the problem of insufficient intelligence in virtual character interaction systems is solved, and instant response and smooth multimodal interaction are achieved.
Patent Information
- Application Number
- CN202210499890.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-09
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-05-09
AI Technical Summary
The existing virtual character interaction system has poor intelligence, the interaction process is rigid, and it is unable to instantly perceive the user's expressions, movements and environmental information, resulting in prolonged interaction time and affecting the user experience.
By receiving multimodal data, identifying user intentions and posture data, determining the virtual character interaction strategy, and using a three-dimensional rendering model to generate an action interaction strategy, multimodal interaction between the virtual character and the user is achieved.
It shortens interaction delay, improves interaction fluency, enhances user experience, supports virtual characters to actively take over or interrupt user conversations, perceives multimodal information, and is suitable for complex application scenarios.
Smart Images

Figure CN114995636B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of computer technology, and in particular to a multi-modal interaction method of a virtual character. BACKGROUND
[0002] With the development of virtual character technology, intelligent digital human products have increasingly penetrated into various aspects of people's life. At present, the demand for virtual characters is further expanded, and it is also required to serve as a partner that can interact with users in multiple modalities such as language, action, and emotion. However, the current virtual character interaction system is relatively rigid and has poor intelligence. It can only perform actions and text content according to the instructions set in the system in advance, and through the interaction components in the system, it realizes a single-mode interaction process. The entire interaction process not only has a long time delay, but also greatly affects the fluency of the virtual character and user interaction, and the interactive experience of the user. SUMMARY
[0003] Therefore, the embodiments of the present specification provide a multi-modal interaction method. One or more embodiments of the present specification also relate to a multi-modal interaction device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects in the prior art.
[0004] According to a first aspect of the embodiments of the present specification, a multi-modal interaction method is provided, applied to a virtual character interaction control system, comprising:
[0005] receiving multi-modal data, wherein the multi-modal data includes voice data and video data;
[0006] identifying the multi-modal data to obtain user intent data and / or user posture data, wherein the user posture data includes user emotion data and user action data;
[0007] determining a virtual character interaction strategy based on the user intent data and / or user posture data, wherein the virtual character interaction strategy includes a text interaction strategy and / or an action interaction strategy;
[0008] obtaining a three-dimensional rendering model of the virtual character;
[0009] generating an image of the virtual character containing the action interaction strategy based on the virtual character interaction strategy and using the three-dimensional rendering model to drive the virtual character to perform multi-modal interaction.
[0010] According to a second aspect of the embodiments of the present specification, a multi-modal interaction device is provided, applied to a virtual character interaction control system, comprising:
[0011] The data receiving module is configured to receive multi-modal data, wherein the multi-modal data comprises voice data and video data.
[0012] The data identifying module is configured to identify the multi-modal data, to obtain user intention data and / or user posture data, wherein the user posture data comprises user emotion data and user action data.
[0013] The strategy determining module is configured to determine a virtual character interaction strategy based on the user intention data and / or user posture data, wherein the virtual character interaction strategy comprises a text interaction strategy and / or an action interaction strategy.
[0014] The rendering model obtaining module is configured to obtain a three-dimensional rendering model of the virtual character.
[0015] The interaction driving module is configured to generate an image of the virtual character containing the action interaction strategy by using the three-dimensional rendering model based on the virtual character interaction strategy, to drive the virtual character to perform multi-modal interaction.
[0016] According to a third aspect of the embodiments of the present specification, a computing device is provided, comprising:
[0017] a memory and a processor;
[0018] The memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions, and the computer executable instructions, when executed by the processor, implement the steps of the multi-modal interaction method.
[0019] According to a fourth aspect of the embodiments of the present specification, a computer readable storage medium is provided, which stores computer executable instructions, and the instructions, when executed by a processor, implement the steps of the multi-modal interaction method.
[0020] According to a fifth aspect of the embodiments of the present specification, a computer program is provided, wherein when the computer program is executed in a computer, the computer program causes the computer to execute the steps of the multi-modal interaction method.
[0021] One embodiment of the present specification provides a multi-modal interaction method applied to a virtual character interaction control system, comprising: receiving multi-modal data, wherein the multi-modal data comprises voice data and video data; identifying the multi-modal data to obtain user intention data and / or user posture data, wherein the user posture data comprises user emotion data and user action data; determining a virtual character interaction strategy based on the user intention data and / or user posture data, wherein the virtual character interaction strategy comprises a text interaction strategy and / or an action interaction strategy; obtaining a three-dimensional rendering model of the virtual character; and generating an image of the virtual character containing the action interaction strategy based on the virtual character interaction strategy and using the three-dimensional rendering model to drive the virtual character to perform multi-modal interaction.
[0022] Specifically, by receiving voice data and audio data of the user and performing intention recognition and posture recognition, the communication intention of the user and / or the posture corresponding to the user are determined, and then the specific interaction strategy of the virtual character and the user is determined according to the communication intention of the user and / or the posture corresponding to the user, and the virtual character is driven to complete the interaction process with the user according to the determined interaction strategy. This method can not only detect and recognize the emotion and action of the user, but also consider the emotion and action of the virtual character in decision-making of the virtual character interaction strategy, so that the expression of the virtual character to the emotion and / or action of the user has a corresponding coping strategy. This not only makes the whole interaction process have a lower time delay, but also makes the whole interaction process between the user and the virtual character more smooth, providing a better interaction experience for the user. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 FIG. 1 is a system structure schematic diagram of a multi-modal interaction method applied to a virtual character interaction control system according to one embodiment of the present specification;
[0024] Figure 2 FIG. 2 is a flowchart of a multi-modal interaction method according to one embodiment of the present specification;
[0025] Figure 3 FIG. 3 is a system architecture diagram of a virtual character interaction control system according to one embodiment of the present specification;
[0026] Figure 4 FIG. 4 is a processing process schematic diagram of a multi-modal interaction method according to one embodiment of the present specification;
[0027] Figure 5 FIG. 5 is a structure schematic diagram of a multi-modal interaction device according to one embodiment of the present specification;
[0028] Figure 6 FIG. 6 is a structure block diagram of a computing device according to one embodiment of the present specification. DETAILED DESCRIPTION
[0029] In the following description, numerous specific details are set forth in order to provide a thorough understanding of the present description. However, the present description can be practiced without the specific details, which are not intended to limit the present description in any way. In some instances, well-known structures have not been described in order not to obscure the description.
[0030] The terminology used in the one or more embodiments of the present description is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of the present description. As used in the one or more embodiments of the present description and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in the one or more embodiments of the present description, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0031] It will be understood that, although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal or chronological order. Rather, these terms are used only to distinguish one from another. For example, without departing from the scope of the one or more embodiments of the present description, first can be termed second and, similarly, second can be termed first. The term "if' as used herein, can be interpreted as meaning "when" or "in response to determining" depending on the context.
[0032] First, the noun terms related to the one or more embodiments of the present description are explained.
[0033] Multi-modal interaction: Users can communicate with digital humans through voice, text, expressions, actions, and gestures, and digital humans can also respond to users through voice, text, expressions, actions, and gestures.
[0034] Duplex interaction: A real-time, two-way communication interaction mode in which users and digital humans can interrupt or respond to each other at any time.
[0035] Non-exclusive conversation: A two-way communication between the conversation parties in which users and digital humans can interrupt or respond to each other at any time.
[0036] VAD (Voice Activity Detection): Also known as voice endpoint detection, voice boundary detection.
[0037] TTS (Text To Speech): A technology for converting text into sound.
[0038] Digital person: refers to a virtual person with a digitized appearance, which can be used in virtual reality applications to interact with real people. In the process of communication with the digital person, the traditional interaction mode is the exclusive question and answer form with voice as the carrier.
[0039] Currently, in the process of interaction between virtual people and users, the following problems may occur:
[0040] 1) In terms of communication fluency, based on the exclusive communication form, the user cannot actively interrupt the digital person's dialogue, and the digital person cannot immediately respond during the dialogue with the user, resulting in unintelligent communication between the user and the digital person.
[0041] 2) In terms of the diversity of perception ability, the communication form with voice as the carrier cannot perceive the user's facial changes such as expressions and the user's dialogue state, nor can it perceive the user's body movements such as gestures and body posture. The lack of these information will make the user unable to immediately feedback the user's state during the communication with the digital person, resulting in a very rigid dialogue process.
[0042] 3) In terms of response time of dialogue, due to the time delay of ASR and system, the general dialogue delay is about 1.2-1.5s, and the dialogue delay of human perception is about 600-800ms. The long dialogue delay will cause serious dialogue lag and poor user experience.
[0043] In addition, the current intelligent dialogue control system for virtual person interaction may only support voice duplex capability, lack video understanding capability and visual duplex state decision capability, and cannot perceive multi-modal information such as user's expressions, actions and environment; or some dialogue systems only support basic question and answer capability, without duplex capability (active / passive interruption, response) and video understanding capability and visual duplex state decision capability, and cannot perceive multi-modal information such as user's expressions, actions and environment.
[0044] Therefore, the multi-modal interaction method provided by the embodiments of the present specification is applied to a virtual person interaction control system, which can realize the interaction process between the virtual person and the real user by setting a multi-modal control module, a multi-modal duplex state module and a basic dialogue module. On the basis of completing the basic dialogue task, the multi-modal data can also be recognized and processed to realize that the virtual person can actively respond and interrupt the user's dialogue, shorten the interaction delay of the system, and perceive multi-modal information such as user's expressions, actions and gestures, so as to be suitable for various application scenarios, such as identity verification, fault damage determination and article verification, and other complex application scenarios will have good application effect.
[0045] It should be noted that the function of the multi-modal control module is to control the input and output of the video stream and the voice stream in the interactive system. At the input end, the module splits and understands the input voice stream and video stream, controls the triggering of the multi-modal duplex system, reduces the cost of system transmission, and at the same time speeds up the processing efficiency of the system. At the output end, it is responsible for rendering the results of the system into the video stream of the digital person. The function of the multi-modal duplex state management module is to manage the state of the current conversation and make decisions on the duplex state. The current duplex state includes 1) duplex active / passive interruption, 2) duplex active takeover, 3) calling the basic dialogue system or business logic, and 4) no feedback. The function of the basic dialogue module is to include basic business logic and dialogue question and answer capabilities.
[0046] Further, in the following embodiments, the specific processing mode of each module of the multi-modal interaction method provided by the embodiments of the present specification will be introduced in detail.
[0047] Based on this, in the present specification, a multi-modal interaction method is provided, and the present specification also relates to a multi-modal interaction device, a computing device, and a computer readable storage medium, which are described in detail one by one in the following embodiments.
[0048] Referring to Figure 1 , Figure 1 Fig. 1 shows a system structure schematic diagram of a multi-modal interaction method according to an embodiment of the present specification applied to a virtual person interactive control system.
[0049] Figure 1 The virtual person interactive control system 100, and the virtual person interactive control system 100 includes a multi-modal control module 102 and a multi-modal duplex state management module 104.
[0050] In practical applications, the multi-modal control module 102 in the virtual person interactive control system 100 can be used as the input of the video stream and the voice stream, and also can be used as the output of the virtual person interactive video stream; wherein, the part of the multi-modal input includes the video stream input and the voice stream input. At the same time, the multi-modal control module 102 performs emotion detection and gesture detection on the video stream, and the multi-modal control module 102 performs voice detection on the voice stream, and inputs the detection results of the video stream and / or the detection results of the voice stream into the duplex state decision in the multi-modal duplex state management module 104, to determine the interactive strategy of the virtual person, wherein the interactive strategy can be mainly divided into action takeover and script+action takeover. Further, the multi-modal duplex state management module 104 can also render the virtual person according to the determined interactive strategy of the virtual person, and then realize the output of the video stream of the virtual person after rendering through the multi-modal control module 102.
[0051] The multi-modal interaction method provided by the embodiment of the present specification can perceive the emotion and action of the user in the video stream through visual understanding, and provide the virtual character with a way of taking over or interrupting, so that the interaction process between the virtual character and the user becomes a non-exclusive conversation, and the virtual character can also provide multi-modal interaction of emotion, action and / or voice.
[0052] The following describes the embodiments of the present specification in conjunction with the accompanying Figure 2 , Figure 2 A flowchart of a multi-modal interaction method provided by an embodiment of the present specification is shown, which specifically includes the following steps.
[0053] It should be noted that the multi-modal interaction method provided by the embodiment of the present specification is applied to a virtual character interaction control system, and through the virtual character interaction control system, the virtual character and the user can realize interaction with small delay, smooth communication, and a process of simulating human interaction.
[0054] Step 202: receiving multi-modal data, wherein the multi-modal data includes voice data and video data.
[0055] In practical applications, the virtual character interaction control system can receive multi-modal data, which is voice data and video data corresponding to the user, wherein the voice data can be understood as voice data of the user communicating with the virtual character. For example, the voice data of "Can you please query the insurance order?" expressed by the user to the virtual character; the video data can be understood as the expression, action, and mouth shape of the user when expressing the above voice data, and the video data of the environment in which the user is located. In the above example, the expression of the user when expressing the above voice data can be shown in the video data as doubt, the action as a hand-spread action, and the mouth shape as the mouth shape corresponding to the expression of the above voice data.
[0056] It should be noted that in order to realize the interaction of the virtual character and the user as a simulated human, the virtual character needs to make an immediate response to the voice data and the video data of the user, and reduce the delay generated in the interaction process. At the same time, it is also necessary to support the functions of interaction, interruption, and taking over between the two parties.
[0057] Step 204: identifying the multi-modal data to obtain user intention data and / or user posture data, wherein the user posture data includes user emotion data and user action data.
[0058] The user intention data can be understood as the intention of the voice data expressed by the user. For example, in the above example, the intention of the voice data "Can you please query the insurance order?" is to ask the virtual character whether it can help to query the insurance order of the user before.
[0059] The user posture data can be understood as posture data expressed by the user in the video data, including user emotion data and user action data. For example, in the above example, the "doubt" emotion expressed by the user's face, and the "shrug" action displayed by the user's hands.
[0060] In actual applications, the virtual character interaction control system can identify the multi-modal data, and identify the voice data and the video data in the multi-modal data respectively. Further, the user intention data is obtained by identifying the voice data, and the user posture data is obtained by identifying the video data, and the user posture data can include user emotion data and user action data. It should be noted that in different application scenarios, the multi-modal data of the user can only identify the user intention data, or the user posture data, or both the user intention data and the user posture data. In this embodiment, the expression "and / or" is not limited in this regard.
[0061] Further, the virtual character interaction control system can identify the voice data and the video data respectively, and further determine the user's intention, emotion, action, posture and the like. Specifically, the identification of the multi-modal data to obtain the user intention data and / or the user posture data includes: text conversion of the voice data in the multi-modal data, identification of the converted text data to obtain the user intention data; and / or emotion recognition of the video data and / or the voice data in the multi-modal data to obtain the user emotion data; gesture recognition of the video data in the multi-modal data to obtain the user action data; determination of the user posture data based on the user emotion data and the user action data.
[0062] In actual applications, the virtual character interaction control system can first perform text conversion on the voice data in the multi-modal data, and then identify the converted text data to obtain the user intention data. Specifically, the voice-to-text conversion method includes but is not limited to using ASR technology, and the specific conversion method is not limited in this embodiment. It should be noted that in order to ensure that the virtual character can provide instant feedback even during the user's speech, the system can divide the voice stream into small voice units according to the 200ms VAD time, and then input each voice unit into the ASR module to convert it into text, which is convenient for subsequent identification of the user intention data.
[0063] Further, after determining the user intention data, the virtual character interaction control system can further perform user emotion recognition and gesture recognition according to the video data. It should be noted that the user emotion recognition can be performed according to the video data, or can be performed according to the voice data, or can be performed according to the video data and the voice data. For example, the emotion recognition can be performed according to the facial expression changes (eye movement, lip twitching) or head shaking of the user in the video data. For another example, the emotion recognition can be performed according to the volume of the voice data and the breath. In addition, the virtual character interaction control system can also recognize the action of the user according to the video data, such as recognizing the gesture of the user. When the user makes a hand gesture, the action data of the user can be obtained. Finally, the virtual character interaction control system can determine the posture data of the user according to the user emotion data and the user action data.
[0064] It should be noted that the virtual character interaction control system can recognize and perceive the slight changes in the voice data and the video data of the user, so as to accurately capture the intention and dynamic of the user, and facilitate subsequent decision-making of the virtual character to interact with the user in a certain strategy and mode.
[0065] Further, in order to quickly obtain the user emotion data, the virtual character interaction control system can adopt a two-stage recognition mode, that is, first performing rough emotion detection, and then classifying the emotion to obtain the target emotion. Specifically, the step of performing emotion recognition on the video data in the multi-modal data to obtain the user emotion data includes:
[0066] performing emotion detection on the video data in the multi-modal data, and in a case where the video data contains a target emotion, classifying the target emotion in the video data to obtain the user emotion data. The target emotion can be understood as a user emotion pre-set by the system, such as anger, unhappiness, neutrality, happiness, and surprise.
[0067] In specific implementation, the virtual character interaction control system can first perform emotion detection on the video data in the multi-modal data. When it is detected that the video stream contains a target emotion pre-set by the system, the target emotion can be classified and determined to obtain the user emotion data. In actual application, in order to ensure that the recognition speed and the recognition accuracy of the system have good effects, the system will adopt a two-stage recognition mode. The expression of the user in the video stream can be first recognized, but when the target emotion of the user is detected in the video stream, the target emotion is classified into an emotion category to determine the final user emotion data.
[0068] It should be noted that the virtual character interaction control system can be configured with a target emotion rough summoning module and an emotion classification module. The target emotion rough summoning module can perform coarse-grained detection on the video stream, and the emotion classification module can perform emotion classification on the target emotion of the video stream to determine the user emotion data as angry, unhappy, neutral, happy, or surprised. The target emotion rough summoning module can use a ResNet18 model, and the emotion classification module can use a time-series Transformer model, but the embodiment is not limited to using these two model types.
[0069] When the virtual character interaction control system does not find that the user has the specified emotion, the video stream is not transmitted backward, which reduces the transmission cost of the system and also speeds up the recognition efficiency of the system.
[0070] Similarly, the virtual character interaction control system can also use a two-stage recognition method when performing gesture recognition on the video data of the user. Specifically, the step of performing gesture recognition on the video data in the multi-modal data to obtain user action data includes:
[0071] The video data in the multi-modal data is detected for a target gesture, and in a case where the video data includes the target gesture, the target gesture in the video data is classified to obtain user action data.
[0072] The target gesture can be understood as a gesture type that is pre-set by the system, such as a gesture with a clear meaning (such as ok, numbers, or left and right swipes), an unsafe gesture (such as a middle finger or a pinky), or a custom special gesture.
[0073] In specific implementation, the virtual character interaction control system can perform gesture detection on the video data in the multi-modal data. When it is detected that the video stream includes a target gesture pre-set by the system, the target gesture can be classified to obtain user action data. In actual application, the gesture recognition process can also use a target gesture rough summoning module and a gesture classification module, that is, coarse-grained recognition of the user gesture in the video stream is realized, and then the target gesture is classified and recognized to determine whether the user action data is a gesture with a clear meaning (such as ok, numbers, or left and right swipes), an unsafe gesture (such as a middle finger or a pinky), or a custom special gesture.
[0074] The multi-modal interaction method provided by the embodiments of the present specification can recognize the user emotion and the user action by using a two-stage recognition method, which not only quickly completes the recognition process, but also reduces the transmission cost of the system and improves the recognition efficiency of the system.
[0075] After determining the user intention data and / or the user posture data in the multi-modal data, the virtual character interaction control system can first call the pre-stored basic dialogue data to support the basic interaction process that can be implemented. Specifically, after the step of identifying the multi-modal data to obtain the user intention data and / or the user posture data, the step further includes:
[0076] Based on the user intention data and / or the user posture data, the pre-stored basic dialogue data is called, wherein the basic dialogue data includes basic voice data and / or basic action data; and the output video stream of the virtual character is rendered based on the basic dialogue data, and the virtual character is driven to display the output video stream.
[0077] The basic dialogue data can be understood as voice and / or action data that can drive the virtual character to implement basic interaction, which is pre-stored in the system. For example, the dialogue data includes basic communication voice data stored in the database, including but not limited to "Hello", "Thank you", "Is there anything else?" and the like. The action data of basic communication includes but is not limited to "heart" action, "shaking head" action, "nodding" action and the like.
[0078] In actual application, the virtual character interaction control system can also find the basic dialogue data that matches the user intention data and / or the user posture data from the basic dialogue data pre-stored in the system according to the user intention data and / or the user posture data, and call the basic dialogue data. Since the basic dialogue data includes basic voice data and / or basic action data, the virtual character interaction control system can render the output video stream corresponding to the virtual character according to the basic voice data and / or the basic action data, to drive the virtual character to display the output video stream.
[0079] It should be noted that the basic dialogue data can also include basic business data completed by the virtual character pre-set by the system, such as providing basic business services for the user, and the like, which is not limited in the embodiment.
[0080] In summary, the virtual character interaction control system can implement the identification according to the multi-modal data to determine the user's intention, expressed emotion, action, gesture and the like multi-modal data, so that the virtual character can make simulated human-like interaction expression according to the user's emotion data and posture data.
[0081] In addition, in order for the virtual character to realize simulated human-like interaction with the user, and to realize the interaction states such as duplex active acceptance and duplex active / passive interruption, the multi-modal duplex state decision module can also be provided in the embodiment of the specification to determine the virtual character interaction strategy and realize the acceptance / interruption of multi-modal duplex.
[0082] Based on this, the virtual character interaction control system can be designed with three interaction modules, which can be seen from Figure 3 , Figure 3 A system architecture diagram of the virtual character interaction control system provided by the embodiments of the present specification is shown.
[0083] Figure 3 The three modules, i.e., the multi-modal control module, the multi-modal duplex state management module, and the basic dialogue module, are included in the virtual character interaction control system, and the three modules can also be regarded as subsystems, i.e., a multi-modal control system, a multi-modal duplex state management system, and a basic dialogue system. The multi-modal control system controls the input and output of the video stream and the voice stream in the interaction system. At the input end, the module splits and understands the input voice stream and video stream, and the core includes the processing functions of the voice stream, the streaming video expression, and the streaming video action. At the output end, the module is responsible for rendering the results of the system into the video stream of the digital person. The multi-modal duplex state management system is responsible for managing the state of the current dialogue and deciding the current duplex strategy. The current duplex strategy includes duplex active / passive interruption, duplex active acceptance, calling the basic dialogue system or business logic, and no feedback. The basic dialogue system includes basic business logic and dialogue question and answer capabilities, and has basic question and answer interaction capabilities; that is, the system outputs the answer to the question input by the user, which generally includes three sub-modules. 1) NLU (Natural Language Understanding) module: text information is recognized and understood, and is converted into a structured semantic representation or an intent label that can be understood by a computer. 2) DM (Dialogue Management) module: maintain and update the current dialogue state, and decide the next system action. 3) NLG (Natural Language Generation) module: convert the system output state into understandable natural language text.
[0084] The specific implementation process of the multi-modal duplex state management module can be described in detail in the following embodiments to clarify how the virtual character interaction control system provides the virtual characters with the ability to mutually accept and interrupt each other.
[0085] Step 206: determining a virtual character interaction strategy based on the user intent data and / or the user posture data, wherein the virtual character interaction strategy includes a text interaction strategy and / or an action interaction strategy.
[0086] The virtual character interaction strategy can be understood as a script decision, an action decision, or a combination of the script decision and the action decision between the virtual character and the user, i.e., a text interaction strategy and / or an action interaction strategy. The text interaction strategy can be understood as an interaction text corresponding to the voice data of the user by the virtual character, and whether the interaction text needs to be interrupted in the middle of a sentence or be connected at the end of a sentence in the voice text expressed by the user. The action interaction strategy can be understood as an interaction gesture corresponding to the gesture data of the user by the virtual character, and whether the interaction gesture needs to be interrupted in the middle of a sentence or be connected at the end of a sentence in the voice text expressed by the user.
[0087] In actual application, the virtual character interaction control system can determine the text connection content of the virtual character according to the user intention data, whether it is interrupted in the middle of a sentence of the user or connected at the end of a sentence of the user, i.e., a text interaction strategy. The virtual character interaction control system can also determine the gesture connection content of the virtual character according to the user gesture data, whether it is interrupted in the middle of a sentence of the user or connected at the end of a sentence of the user, i.e., an action interaction strategy. It should be noted that for a certain intention data and / or gesture data of the user, the virtual character does not necessarily have both the text interaction strategy and the action interaction strategy, i.e., the text interaction strategy and the action interaction strategy can also be in a “and / or” relationship.
[0088] In addition, the virtual character can not only connect or interrupt the interaction of the user, but also support a function of not making any feedback, i.e., when the VAD time of the user does not reach 800 ms, the system does not make any feedback when the basic dialogue system or the business logic does not need to be called to answer.
[0089] Specifically, the step of determining the virtual character interaction strategy based on the user intention data and / or the user gesture data comprises:
[0090] Based on the user intention data and / or the user gesture data, the video data in the multi-modal data is fused and processed to determine the target intention text and / or the target gesture action of the user; and based on the target intention text and / or the target gesture action, the virtual character interaction strategy is determined.
[0091] In actual application, after the user intention data and / or the user gesture data are determined, the virtual character interaction control system can also fuse and align the text, the video stream, and the voice stream to comprehensively determine the target intention text and / or the target gesture action of the user. Further, the specific virtual character interaction strategy can be determined according to the target intention text and / or the target gesture action.
[0092] For example, the emotion recognition, the emotion classification module has recognized the user's expression from the face, such as a smile, but the user can be expressing a helpless smile. Therefore, in order to solve this problem, the virtual character interaction control system can make a multimodal judgment from the user's voice and the current spoken text, so as to achieve better results. In specific implementation, the system can use a multimodal classification model to make more detailed emotion judgment, and finally the module will output the current interaction state, which can include three state slots, namely text, user gesture action and user emotion, to the duplex state management module for duplex state decision.
[0093] The multimodal interaction method provided by the embodiments of the present specification can accurately determine the user's interaction purpose by further comprehensively judging the user intention data and / or user posture data, thereby avoiding invalid communication shown by the virtual character due to the user's interaction purpose error, and reducing the intelligence of the virtual character.
[0094] After the virtual character interaction control system accurately obtains the target intention text and / or target posture action of the user, it can accurately determine the text interaction strategy and / or action interaction strategy of the virtual character. Specifically, determining the virtual character interaction strategy based on the target intention text and / or the target posture action comprises:
[0095] determining the text interaction strategy of the virtual character based on the target intention text; and / or
[0096] determining the action interaction strategy of the virtual character based on the target posture action.
[0097] In actual application, the virtual character interaction control system determines the text interaction strategy of the virtual character and the user according to the target intention text. For example, if the target intention text of the user is "query the order status of the insurance", the text interaction strategy of the virtual character can be taken over from the end of the user's voice text, that is, the virtual character can express "please wait, I will query for you...". If the target intention text of the user is "how slow you are, have you not queried yet?", the text interaction strategy of the virtual character can be taken over from the middle of the intention text, that is, when the user finishes saying "how slow you are", the virtual character can immediately express "don't worry". In this way, the virtual character and the user can realize instant communication, so as to achieve the effect of human communication.
[0098] Further, the virtual character interaction control system can also determine the action interaction strategy of the virtual character and the user according to the target gesture action. For example, if the target gesture action of the user is an "ok" gesture, the action interaction strategy of the virtual character can also show an "ok" gesture. If the target gesture action of the user is a "finger pointing" gesture, the virtual character can also not respond with any action, and can only reply with text content, such as "what is not satisfactory", or only reply with a "shaking head and crying" action.
[0099] It should be noted that different text interaction strategies and / or action interaction strategies can be determined for different target intention texts and / or target gesture actions. For example, if there is only a target intention text, it can be determined that the virtual character should only respond to a text interaction strategy, or only respond to an action interaction strategy, or a combination of a text interaction strategy and an action interaction strategy. If there is only a target gesture action, it can be determined that the virtual character should only respond to a text interaction strategy, or only respond to an action interaction strategy, or a combination of a text interaction strategy and an action interaction strategy. If both the target intention text and the target gesture action are present, it can be determined that the virtual character should only respond to a text interaction strategy, or only respond to an action interaction strategy, or a combination of a text interaction strategy and an action interaction strategy. Furthermore, all cases cannot be exhausted in the embodiments of the present specification, but the virtual character interaction control system in the present embodiment can support determining different virtual character interaction strategies according to different interaction states.
[0100] Step 208: Obtain a three-dimensional rendering model of the virtual character.
[0101] In actual applications, the virtual character interaction control system can obtain a three-dimensional rendering model of the virtual character, so as to facilitate subsequent generation of an interaction video stream of the virtual character based on the three-dimensional rendering model, and complete multi-modal interaction with the user. It should be noted that the virtual character can be composed of a cartoon or a computer drawing image, or can be composed of a simulated human image, which is not limited in the present embodiment.
[0102] Step 210: Based on the virtual character interaction strategy, an image of the virtual character containing the action interaction strategy is generated by using the three-dimensional rendering model, so as to drive the virtual character to perform multi-modal interaction.
[0103] In actual applications, the virtual character interaction control system can generate an image of the virtual character containing the action interaction strategy of the virtual character according to the determined virtual character interaction strategy and by using the three-dimensional rendering model. For example, the virtual character corresponds to head action, facial expression, and gesture action, and the like, and further, the rendered virtual character image is driven to realize multi-modal interaction with the user.
[0104] Further, the virtual character interaction control system can determine the text engagement position and / or the action engagement position of the virtual character according to the text interaction strategy and the action interaction strategy to realize the process of active engagement of duplex. Specifically, the step of generating the image of the virtual character containing the action interaction strategy by using the three-dimensional rendering model based on the virtual character interaction strategy to drive the virtual character to interact in multiple modes includes:
[0105] determining the text engagement position of the virtual character text interaction based on the text interaction strategy, wherein the text engagement position is the engagement position corresponding to the voice data; determining the action engagement position of the virtual character action interaction based on the action interaction strategy, wherein the action engagement position is the engagement position corresponding to the video data; and generating the image of the virtual character containing the action interaction strategy by using the three-dimensional rendering model based on the text engagement position and / or the action engagement position to drive the virtual character to interact in multiple modes.
[0106] The text engagement position can be understood as the engagement position of the interactive text of the virtual character corresponding to the voice text expressed by the user, which can be divided into in-sentence engagement and end-of-sentence engagement. The action engagement position can be understood as the engagement position of the interactive action of the virtual character corresponding to the voice text expressed by the user, which can be divided into in-sentence action engagement and end-of-sentence action engagement.
[0107] In actual application, the virtual character interaction control system can generate the image of the virtual character containing the action interaction strategy by using the three-dimensional rendering model based on the text engagement position and / or the action engagement position after determining the text engagement position of the virtual character text interaction and the action engagement position of the virtual character action interaction to determine the multiple-mode interaction process of the virtual character.
[0108] It should be noted that when the virtual character interaction control system determines that the dialogue or action of the user needs to be engaged, the current engagement strategy will be triggered. The engagement mode includes two types, one is action engagement only, and the other is action + script engagement. Action engagement only means that the digital human does not make a verbal engagement reply and only responds to the user with actions. For example, when the user suddenly shakes hands to greet the digital human during the dialogue, the virtual character only needs to respond with a greeting action and does not need to affect the current dialogue state. Action + script engagement means that the digital human not only responds to the user with actions, but also makes a verbal engagement reply. This engagement will have some impact on the current dialogue process, but it will also give people the feeling of intelligence in experience. For example, when it is detected that the user has an unhappy emotion during the dialogue, the virtual character needs to interrupt the current dialogue state and actively ask the user "what is not satisfactory", while giving a comforting action.
[0109] In addition, the virtual character interaction control system can also provide a process of duplex active / passive interruption. Specifically, the step of generating an image of the virtual character containing the action interaction strategy based on the virtual character interaction strategy by using the three-dimensional rendering model to drive the virtual character to continue the multi-modal interaction includes:
[0110] In the user intention data and / or user posture data of the virtual character interaction strategy, if it is determined that the user has interruption intention data, the current multi-modal interaction of the virtual character is paused; the corresponding interruption acceptance interaction data of the virtual character is determined based on the interruption intention data, and an image of the virtual character containing the action interaction strategy is generated based on the interruption acceptance interaction data by using the three-dimensional rendering model to drive the virtual character to continue the multi-modal interaction.
[0111] The interruption intention data can be understood as data that the user has explicit refusal to communicate with the virtual character. For example, the user makes a "shut up" gesture, or explicitly says "pause the communication".
[0112] The interruption acceptance interaction data can be understood as the corresponding acceptance text sentence or acceptance action data when the virtual character determines that the user has interruption intention.
[0113] In actual application, in the virtual character interaction strategy of the virtual character interaction control system, if it is determined that the user has interruption intention according to the user intention data and / or user posture data, the current interaction text or interaction action of the virtual character can be paused, and the corresponding interruption acceptance interaction data is determined according to the interruption intention. Then, an image of the virtual character containing the above-mentioned action interaction strategy is generated by using the three-dimensional rendering model to drive the virtual character to continue the multi-modal interaction according to the interruption acceptance interaction data. For example, when the digital person finds that the user has interruption intention, it will actively interrupt the current conversation. This interruption intention can be an explicit interruption intention of the user, such as negative expression or negative emotion of the user during the speech of the digital person. It can also be an implicit interruption intention of the user, such as the user suddenly disappearing or not being in a state of communication. Under the current strategy, the digital person will interrupt the current speaking state, wait for the user to speak, or actively ask the reason for the interruption.
[0114] Finally, the virtual character interaction control system can also provide an output rendering function to fuse the audio data stream and the video data stream determined by the virtual character interaction and then push them out. Specifically, the step of generating an image of the virtual character containing the action interaction strategy based on the virtual character interaction strategy by using the three-dimensional rendering model to drive the virtual character to continue the multi-modal interaction includes:
[0115] determine audio data stream of the virtual character text interaction based on the text interaction strategy; determine video data stream of the action interaction of the virtual character based on the action interaction strategy; perform fusion processing on the audio data stream and the video data stream, render a multi-modal interaction data stream of the virtual character, and generate an image of the virtual character containing the action interaction strategy based on the multi-modal interaction data stream, to drive the virtual character to perform multi-modal interaction.
[0116] In actual application, the output rendering and synthesis video stream of the virtual character interaction control system is pushed out, and contains three parts in total. 1) Streaming TTS part, which synthesizes audio stream of text output of the system. 2) Driving part, containing two sub-modules, a face driving module and an action driving module. The face driving module drives the digital person to output accurate mouth shapes according to the voice stream. The action driving module drives the digital person to output accurate actions according to the action label output by the system. 3) Rendering and synthesis part, which is responsible for rendering and synthesizing the video stream of the digital person from the output of the driving part, TTS and other modules.
[0117] In summary, the multi-modal interaction method provided by the embodiments of the present specification can not only perceive the facial expressions of the user, but also perceive the actions of the user by adding a video stream and a corresponding visual understanding module. In addition, new visual processing modules can be added by similar methods to enable the virtual character to perceive more multi-modal information, such as environmental information. In the embodiments of the present specification, the system can support real-time perception of five facial expressions of the user, namely, anger, displeasure, neutrality, happiness and surprise, and can support real-time perception of three categories of actions, namely, actions with clear meanings (such as OK, numbers and left and right swipes), unsafe gestures (such as the middle finger and the little finger) and custom special actions.
[0118] In addition, the method changes the dialogue from a one-question-one-answer exclusive dialogue form to a non-exclusive dialogue form that can be accepted or interrupted at any time by adding a multi-modal control module and a multi-modal duplex state management module. There are two main reasons for solving this problem: 1) The multi-modal control module divides the dialogue into smaller decision units and no longer uses the complete user question as the trigger condition for user response, so that it can be accepted or interrupted at any time during the dialogue. The voice stream is divided by a VAD time of 200 ms. Generally, the breathing interval during human speech is about 200 ms. The video stream uses a trigger detection strategy. When a specified action, expression, or target object is detected, the duplex state decision is made. 2) The multi-modal duplex state management module is the core of solving this problem, as it not only maintains the current duplex dialogue state, but also decides the current response strategy. The duplex strategy includes duplex active acceptance, duplex active / passive interruption, calling of the basic dialogue system or business logic, and no feedback. By deciding between the four states, the system can achieve the ability to accept or interrupt at any time and basic question and answer. 3) The present scheme divides the dialogue into smaller units and uses the unit as the granularity for digital human decision and response, so that the dialogue is no longer a one-question-one-answer exclusive dialogue form. Therefore, even if the user has not completed the complete expression, the system has already processed the user's input information and calculated the response result. When the user finishes expressing, the system does not need to calculate from the beginning, but directly plays the acceptance speech, thereby greatly shortening the interaction delay. In terms of body feeling, the dialogue delay of the system can be reduced from 1.5 seconds to about 800 ms.
[0119] Referring to Figure 4 , Figure 4 A processing process schematic diagram of a multi-modal interaction method provided by an embodiment of the present disclosure is shown.
[0120] Figure 4 The embodiment can be divided into a multi-modal control system-input, a multi-modal duplex state management system, a basic dialogue system, and a multi-modal control system-output. The above systems can be understood as four subsystems of the virtual character interaction control system to which the multi-modal interaction method is applied.
[0121] In practical applications, the video stream and the voice stream of the user are input from the multi-modal control system-input. For the video stream, the target emotion detection coarse call module and the target gesture detection coarse call module are first passed through, and then emotion classification and gesture classification are performed, and the final emotion recognition result and gesture recognition result are input to the multi-modal data&alignment module. For the voice stream, segmentation is first performed, then text conversion is performed through ASR, and finally input to the multi-modal data&alignment module. Further, the multi-modal data&alignment module integrates the voice recognition result and the emotion and gesture recognition result in the video to determine the target user intent and the target action data, and inputs them to the multi-modal duplex state decision module in the multi-modal duplex state management system.
[0122] Further, Figure 4 The multi-modal duplex state decision system in the multi-modal control system-output can make duplex strategy decisions to determine two kinds of acceptance modes, one is action+script acceptance, and the other is only action acceptance. In the action+script acceptance process, it can be divided into two branches to realize the acceptance process through judgment of in-sentence acceptance or end-of-sentence acceptance. Specifically, in end-of-sentence acceptance, acceptance script decisions and acceptance action decisions are first determined according to intent recognition, and in in-sentence acceptance, acceptance script decisions and acceptance action decisions can be determined. In addition, in action acceptance, only the specific acceptance action is decided, and finally the acceptance strategy of the virtual character is input to the multi-modal control system-output to determine the streaming video stream and the streaming audio stream.
[0123] It should be noted that the multi-modal duplex state decision system also includes a multi-modal interruption intent judgment, which can realize specific acceptance interruption functions in combination with business.
[0124] Further, the multi-modal control system-output can determine facial driving data and action driving data according to the streaming video stream and the streaming audio stream of the virtual character to complete the rendering+stream media merging processing of the virtual character to output the digital human video stream.
[0125] In addition, the multi-modal control system-output can provide basic dialogue data for the interaction of the virtual character according to the streaming video stream and the streaming audio stream of the virtual character, and the basic business logic and actions are matched to jointly complete the generation of the digital human video stream.
[0126] In summary, the multi-modal interaction method provided by the embodiments of the present specification has the effects of multi-modal perception, multi-modal duplexing, and short interaction delay. Specifically, for multi-modal perception, the embodiments of the present specification propose a system that can perceive user voice and video information. Compared with a traditional voice stream-based dialogue system, the present solution can not only process user voice information, but also identify and detect user emotions and actions, greatly improving the intelligence of digital human perception. For multi-modal duplexing, the embodiments of the present specification propose an interactive system that can instantly accept and interrupt at any time. Compared with a traditional single-turn dialogue system that answers one question at a time, the system can give users some feedback and replies, such as simple tone acceptance, during the process of user speech. In addition, when the user is not in a listening state or the user has a clear intention to interrupt the dialogue, the current dialogue process can be interrupted at any time. The duplex interaction system improves the fluency of interaction, thereby providing users with a better interactive experience. Short interaction delay: when the user has not completely expressed, the system has already processed the user's input information in a streaming manner and calculated the reply result. When the user finishes expressing, the system does not need to calculate from the beginning, but can directly play the acceptance speech, greatly shortening the interaction delay. In terms of body feeling, the dialogue delay of the system can be reduced from 1.5 seconds to about 800 ms.
[0127] Corresponding to the method embodiments described above, the present specification also provides multi-modal interaction device embodiments, Figure 5 A structural schematic diagram of a multi-modal interaction device provided by an embodiment of the present specification is shown. As shown in the figure, Figure 5 The device applied to a virtual character interactive control system includes:
[0128] The data receiving module 502 is configured to receive multi-modal data, wherein the multi-modal data includes voice data and video data; the data identification module 504 is configured to identify the multi-modal data to obtain user intention data and / or user posture data, wherein the user posture data includes user emotion data and user action data; the strategy determination module 506 is configured to determine a virtual character interactive strategy based on the user intention data and / or user posture data, wherein the virtual character interactive strategy includes a text interactive strategy and / or an action interactive strategy; the rendering model acquisition module 508 is configured to acquire a three-dimensional rendering model of the virtual character; and the interactive driving module 510 is configured to generate an image of the virtual character containing the action interactive strategy based on the virtual character interactive strategy and using the three-dimensional rendering model, to drive the virtual character to perform multi-modal interaction.
[0129] Optionally, the data recognition module 504 is further configured to: perform text conversion on speech data in the multi-modal data, recognize converted text data to obtain user intention data; and / or perform emotion recognition on video data and / or speech data in the multi-modal data to obtain user emotion data; perform gesture recognition on video data in the multi-modal data to obtain user action data; and determine user posture data based on the user emotion data and the user action data.
[0130] Optionally, the data recognition module 504 is further configured to: perform emotion detection on video data in the multi-modal data, and in a case where it is detected that the video data contains target emotion, classify the target emotion in the video data to obtain user emotion data.
[0131] Optionally, the data recognition module 504 is further configured to: perform gesture detection on video data in the multi-modal data, and in a case where it is detected that the video data includes target gesture, classify the target gesture in the video data to obtain user action data.
[0132] Optionally, the strategy determination module 506 is further configured to: perform fusion processing on video data in the multi-modal data based on the user intention data and / or user posture data to determine target intention text and / or target posture action of the user; and determine a virtual character interaction strategy based on the target intention text and / or the target posture action.
[0133] Optionally, the strategy determination module 506 is further configured to: determine a text interaction strategy of a virtual character based on the target intention text; and / or determine an action interaction strategy of a virtual character based on the target posture action.
[0134] Optionally, the interaction driving module 510 is further configured to: determine a text receiving position of the virtual character text interaction based on the text interaction strategy, wherein the text receiving position is a receiving position corresponding to the speech data; determine an action receiving position of the virtual character action interaction based on the action interaction strategy, wherein the action receiving position is a receiving position corresponding to the video data; and generate an image of the virtual character containing the action interaction strategy by using the three-dimensional rendering model based on the text receiving position and / or the action receiving position, to drive the virtual character to perform multi-modal interaction.
[0135] Optionally, the interaction driving module 510 is further configured to: in the user intention data and / or the user posture data of the virtual character interaction strategy, determine that the user has interrupt intention data, pause the current multi-modal interaction of the virtual character; determine corresponding interrupt response interaction data of the virtual character based on the interrupt intention data, and generate an image of the virtual character containing the action interaction strategy based on the interrupt response interaction data by using the three-dimensional rendering model, to drive the virtual character to continue the multi-modal interaction.
[0136] Optionally, the device further comprises a video stream output module configured to call pre-stored basic dialogue data based on the user intention data and / or the user posture data, wherein the basic dialogue data comprises basic voice data and / or basic action data; render an output video stream of the virtual character based on the basic dialogue data, and drive the virtual character to display the output video stream.
[0137] Optionally, the interaction driving module 510 is further configured to: determine audio data stream of the virtual character text interaction based on the text interaction strategy; determine video data stream of the action interaction of the virtual character based on the action interaction strategy; fuse the audio data stream and the video data stream, render multi-modal interaction data stream of the virtual character, and generate an image of the virtual character containing the action interaction strategy based on the multi-modal interaction data stream by using the three-dimensional rendering model, to drive the virtual character to perform multi-modal interaction.
[0138] The multi-modal interaction device provided by the embodiments of the present specification can receive voice data and audio data of a user, and perform intention recognition and posture recognition to determine the communication intention of the user and / or the posture corresponding to the user, and then determine the specific interaction strategy of the virtual character and the user according to the communication intention of the user and / or the posture corresponding to the user, and then drive the virtual character to complete the interaction process with the user according to the determined interaction strategy. This kind of way can not only detect and recognize the emotion and action of the user, but also consider the emotion and action of the virtual character when deciding the interaction strategy of the virtual character, so that the expression of the virtual character to the emotion and / or action of the user has corresponding response strategy, which not only makes the time delay of the whole interaction process lower, but also makes the whole interaction process between the user and the virtual character more smooth, and provides better interaction experience for the user.
[0139] The above is a schematic scheme of a multi-modal interaction device of the present embodiment. It should be noted that the technical scheme of the multi-modal interaction device belongs to the same concept as the technical scheme of the multi-modal interaction method described above, and the details of the technical scheme of the multi-modal interaction device which are not described in detail can be referred to the description of the technical scheme of the multi-modal interaction method.
[0140] Figure 6 A structural block diagram of a computing device 600 according to one embodiment of the present specification is shown. The components of the computing device 600 include, but are not limited to, a memory 610 and a processor 620. The processor 620 is connected with the memory 610 through a bus 630, and a database 650 is used to save data.
[0141] The computing device 600 further includes an access device 640, which enables the computing device 600 to communicate via one or more networks 660. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 640 can include one or more of any type of network interface (e.g., a network interface card (NIC)) such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a near field communication (NFC) interface, and the like, either wired or wireless.
[0142] In one embodiment of the present specification, the above-mentioned components of the computing device 600 and other components not shown in the above-mentioned components can be connected with each other, for example, through a bus. It should be understood that Figure 6 the computing device structure block diagram shown is only for the purpose of example, and is not a limitation on the scope of the present specification. Other components can be added or replaced as needed by those skilled in the art. Figure 6 the computing device structure block diagram shown is only for the purpose of example, and is not a limitation on the scope of the present specification. Other components can be added or replaced as needed by those skilled in the art.
[0143] The computing device 600 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other type of mobile device, or a stationary computing device such as a desktop computer or a PC. The computing device 600 can also be a mobile or stationary server.
[0144] The processor 620 is configured to execute computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned multi-modal interaction method.
[0145] The above is a schematic scheme of a computing device according to the present embodiment. It should be noted that the technical scheme of the computing device belongs to the same concept as the technical scheme of the above-mentioned multi-modal interaction method, and the details of the technical scheme of the computing device that are not described in detail can be referred to the description of the technical scheme of the multi-modal interaction method.
[0146] An embodiment of the present specification further provides a computer readable storage medium, which stores computer executable instructions, and the computer executable instructions realize the steps of the multi-modal interaction method when executed by a processor.
[0147] The above is a schematic scheme of the computer readable storage medium of the embodiment of the present specification. It should be noted that the technical scheme of the storage medium and the technical scheme of the multi-modal interaction method belong to the same concept, and the details of the technical scheme of the storage medium which are not described in detail can be referred to the description of the technical scheme of the multi-modal interaction method.
[0148] An embodiment of the present specification further provides a computer program, which causes a computer to execute the steps of the multi-modal interaction method when the computer program is executed in the computer.
[0149] The above is a schematic scheme of the computer program of the embodiment of the present specification. It should be noted that the technical scheme of the computer program and the technical scheme of the multi-modal interaction method belong to the same concept, and the details of the technical scheme of the computer program which are not described in detail can be referred to the description of the technical scheme of the multi-modal interaction method.
[0150] The specific embodiments of the present specification are described above. Other embodiments are within the scope of the appended claims. In some cases, acts or steps recited in the claims can be performed in a different order than the order in which the acts or steps are recited in the embodiments. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.
[0151] The computer instructions include computer program code, which can be in the form of source code, object code, executable code, or some intermediate form. The computer readable medium can include any entity or apparatus capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer readable medium does not include electrical carrier signals and telecommunication signals.
[0152] It should be noted that, for the aforementioned method embodiments, the sequences of the described actions are not necessarily required to implement the present application, and certain actions can be performed in other sequences, or even at the same time, in accordance with the present application. Furthermore, certain actions can not be required to implement the present application. Additionally, the described embodiments are not necessarily the only possible implementation of the present application.
[0153] In the above embodiments, the description of each embodiment focuses on different aspects, and the parts not described in detail in a certain embodiment can be referred to the relevant description of other embodiments.
[0154] The above disclosed preferred embodiments of the present application are only used to help explain the present application. Alternative embodiments do not describe all the details, nor limit the present application to the specific embodiments described. Obviously, according to the content of the embodiments of the present application, many modifications and changes can be made. The present application selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present application, so that those skilled in the art can well understand and use the present application. The present application is limited by the claims and their full scope and equivalents.
Claims
1. A multimodal interaction method, applied to a virtual character interaction control system, comprising: Receive multimodal data, where The multimodal data includes voice data and video data; Identifying the multimodal data to obtain user intention data and / or user posture data, wherein the user posture data includes user emotion data and user action data; Determining a virtual character interaction strategy based on the user intention data and / or user posture data, the determining of the virtual character interaction strategy based on the user intention data and / or user posture data includes: based on the user intention data and / or the user posture data, fusing and aligning the voice data, the video data, and the text data, determining the user's target intention text and / or target posture action, and determining the virtual character interaction strategy based on the target intention text and / or the target posture action, wherein the virtual character interaction strategy includes a text interaction strategy and / or an action interaction strategy, the text interaction strategy is used to determine whether the virtual character's text continuation content is interrupted in the user's sentence or continued at the end of the sentence, the action interaction strategy is used to determine the interaction gesture of the virtual character corresponding to the user posture data and whether the interaction gesture interrupts the user's sentence or continues at the end of the sentence, and the text data is obtained by recognizing and converting the voice data; Obtaining a three-dimensional rendering model of the virtual character; Based on the virtual character interaction strategy, the three-dimensional rendering model is used to generate an image of the virtual character including the text interaction strategy and / or the action interaction strategy to drive the virtual character to perform multimodal interaction.
2. The multimodal interaction method according to claim 1, wherein identifying the multimodal data and obtaining user intention data and / or user posture data comprises: Performing text conversion on the voice data in the multimodal data, recognizing the converted text data, and obtaining the user intention data; and / or Performing emotion recognition on the video data and / or the voice data in the multimodal data to obtain the user emotion data; Performing gesture recognition on the video data in the multimodal data to obtain the user action data; The user gesture data is determined based on the user emotion data and the user motion data.
3. The multimodal interaction method according to claim 2, wherein the performing emotion recognition on the video data in the multimodal data to obtain the user emotion data comprises: Emotion detection is performed on the video data in the multimodal data, and when it is detected that the video data contains a target emotion, the target emotion in the video data is classified to obtain the user emotion data.
4. The multimodal interaction method according to claim 2, wherein performing gesture recognition on the video data in the multimodal data to obtain the user action data comprises: Gesture detection is performed on the video data in the multimodal data, and when it is detected that the video data includes a target gesture, the target gesture in the video data is classified to obtain the user action data.
5. The multimodal interaction method according to claim 1, wherein determining the virtual character interaction strategy based on the user intention data and / or the user posture data comprises: Performing fusion processing on the video data in the multimodal data based on the user intention data and / or the user posture data to determine the user's target intention text and / or target posture action; Based on the target intention text and / or the target gesture action, a virtual character interaction strategy is determined.
6. The multimodal interaction method according to claim 5, wherein determining the virtual character interaction strategy based on the target intention text and / or the target gesture action comprises: Determining the text interaction strategy of the virtual character based on the target intention text; and / or The action interaction strategy of the virtual character is determined based on the target posture action.
7. The multimodal interaction method according to claim 6, wherein generating an image of the virtual character including the action interaction strategy using the three-dimensional rendering model based on the virtual character interaction strategy to drive the virtual character to perform multimodal interaction comprises: Determining a text connection position for the virtual character text interaction based on the text interaction strategy, wherein the text connection position is a connection position corresponding to the voice data; Determining an action connection position of the virtual character action interaction based on the action interaction strategy, wherein the action connection position is a connection position corresponding to the video data; Based on the text connection position and / or the action connection position, the three-dimensional rendering model is used to generate an image of the virtual character including the action interaction strategy to drive the virtual character to perform multimodal interaction.
8. The multimodal interaction method according to claim 1, wherein generating an image of the virtual character including the action interaction strategy using the three-dimensional rendering model based on the virtual character interaction strategy to drive the virtual character to perform multimodal interaction comprises: If it is determined in the user intention data and / or user posture data of the virtual character interaction strategy that the user has interruption intention data, pausing the current multimodal interaction of the virtual character; Based on the interruption intention data, the interruption and continuation interaction data corresponding to the virtual character is determined, and based on the interruption and continuation interaction data, the image of the virtual character including the action interaction strategy is generated using the three-dimensional rendering model to drive the virtual character to continue multimodal interaction.
9. The multimodal interaction method according to claim 1, after identifying the multimodal data and obtaining user intention data and / or user posture data, further comprising: Based on the user intention data and / or the user posture data, calling pre-stored basic dialogue data, wherein the basic dialogue data includes basic voice data and / or basic action data; An output video stream of the virtual character is rendered based on the basic dialogue data, and the virtual character is driven to display the output video stream.
10. The multimodal interaction method according to claim 1, wherein generating an image of the virtual character including the action interaction strategy using the three-dimensional rendering model based on the virtual character interaction strategy to drive the virtual character to perform multimodal interaction comprises: Determining an audio data stream of the virtual character text interaction based on the text interaction strategy; Determining a video data stream of the action interaction of the virtual character based on the action interaction strategy; The audio data stream and the video data stream are fused and processed, and the multimodal interaction data stream of the virtual character is rendered. Based on the multimodal interaction data stream, the image of the virtual character including the action interaction strategy is generated using the three-dimensional rendering model to drive the virtual character to perform multimodal interaction.
11. A multimodal interaction device, applied to a virtual character interaction control system, comprising: a data receiving module, configured to receive multimodal data, wherein the multimodal data includes voice data and video data; a data recognition module configured to recognize the multimodal data and obtain user intention data and / or user posture data, wherein the user posture data includes user emotion data and user action data; a strategy determination module configured to determine a virtual character interaction strategy based on the user intention data and / or user posture data, wherein the virtual character interaction strategy includes a text interaction strategy and / or an action interaction strategy, the text interaction strategy is used to determine whether the virtual character's textual content is interrupted in the user's sentence or continued at the end of the sentence, and the action interaction strategy is used to determine whether the virtual character's interaction gesture corresponding to the user posture data interrupts the user's sentence or continues at the end of the sentence, and the text data is obtained by recognizing and converting the voice data; The strategy determination module is further configured to perform fusion alignment processing on the voice data, the video data, and the text data based on the user intention data and / or the user posture data, determine the user's target intention text and / or target posture action, and determine the virtual character interaction strategy based on the target intention text and / or the target posture action; A rendering model acquisition module is configured to acquire a three-dimensional rendering model of the virtual character; The interaction driving module is configured to generate an image of the virtual character including the text interaction strategy and / or the action interaction strategy based on the virtual character interaction strategy using the three-dimensional rendering model to drive the virtual character to perform multimodal interaction.
12. A computing device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the multimodal interaction method described in any one of claims 1 to 10 are implemented.
13. A computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the steps of the multimodal interaction method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Gesture Interaction method and system based on virtual human
CN108459712A
Multi-modal interaction method, device and system based on virtual character, storage medium and terminal
CN112162628A