Human-computer interaction method and device based on multimodal dialogue state representation
By acquiring and analyzing multimodal input information and determining multimodal dialogue strategies, the problem of lack of universality and multimodal considerations in the existing technology of human-computer interaction systems is solved, and anthropomorphic interaction is achieved in diverse scenarios, providing a more accurate and humanized interaction method.
Patent Information
- Application Number
- CN202111064527.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-10
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-09-10
AI Technical Summary
In the prior art, human-computer interaction systems lack universality, making it difficult to conduct precise dialogue in diverse scenarios, and lack consideration of multimodal dialogue state, resulting in unnatural interaction.
By obtaining the original multimodal input information, performing single-modal analysis and multimodal understanding, obtaining the multimodal dialogue status representation results, determining the multimodal dialogue strategy, and performing multimodal information output, including speech recognition, sentiment analysis and behavioral gesture analysis, combining dialogue behavior, dialogue elements and dialogue scenarios to achieve anthropomorphic interaction.
It realizes more accurate and humanized multimodal interaction in diverse scenarios, and can adapt to human-computer dialogue systems in different fields and provide a better user experience.
Smart Images

Figure CN113792196B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a method and device for human-computer interaction based on multimodal dialogue state representation. Background Art
[0002] With the advancement of technology and societal needs, human-computer interaction is entering a new phase of human-like interaction. Real-world human-computer interaction systems require advanced communication skills and strategic planning capabilities. Furthermore, multimodal robots must not only interact via text or voice but also display charts or images during communication to enhance user understanding. Real-world conversations involve a variety of language phenomena, such as active and passive role switching, topic rotation, and long-term contextual dependencies. Relying solely on intents and slots to represent conversational state is insufficient to meet real-world needs. Both intents and slots require pre-defined definitions, making it difficult to address diversity. The definition methods for intents and slots are not universal, making it difficult to share knowledge across related domains. There is a lack of detailed descriptions of real-world conversational behavior. There is also a lack of consideration for multimodal conversational state. Summary of the Invention
[0003] The present disclosure provides a method and apparatus for human-computer interaction based on multimodal dialogue state representation, which is used to address the defects of the prior art in that it is not universal and difficult to conduct dialogue accurately, and to achieve accurate dialogue and cross-domain universality.
[0004] In a first aspect, the present disclosure provides a method for human-computer interaction based on multimodal dialogue state representation, comprising:
[0005] Get the original multimodal input information;
[0006] Processing the original multimodal input information to obtain a multimodal dialogue state representation result;
[0007] Determining a multimodal dialogue strategy according to the multimodal dialogue state representation result;
[0008] The multimodal information output is completed according to the multimodal dialogue strategy.
[0009] According to a method for human-computer interaction based on multimodal dialogue state representation provided by the present disclosure, wherein the processing of the original multimodal input information to obtain a multimodal dialogue state representation result specifically includes:
[0010] Performing a unimodal analysis on the original multimodal input information to obtain a unimodal representation result;
[0011] Acquiring dialogue scene related information according to the original multimodal input information;
[0012] Multimodal understanding and text semantic analysis are performed on the unimodal representation result and the dialogue scene related information to obtain a multimodal dialogue state representation result.
[0013] According to a method for human-computer interaction based on multimodal dialogue state representation provided by the present disclosure, wherein the performing of a unimodal analysis on the original multimodal input information to obtain a unimodal representation result specifically includes:
[0014] Performing speech recognition on the original multimodal input information to obtain a speech recognition result, and performing semantic analysis on the speech recognition result to obtain a semantic analysis result;
[0015] Performing sentiment analysis and behavioral gesture analysis on the original multimodal input information to obtain corresponding sentiment analysis results and behavioral gesture analysis results;
[0016] The semantic analysis result, the sentiment analysis result and the behavior gesture analysis result constitute a unimodal representation result.
[0017] According to the present disclosure, a method for human-computer interaction based on multimodal dialogue state representation is provided, wherein the multimodal dialogue state representation result includes dialogue behavior, dialogue elements and dialogue scene;
[0018] Among them, dialogue behavior is used to guide dialogue strategy generation;
[0019] Dialogue elements are used to determine the interlocutors’ intentions;
[0020] The dialogue scenario is used to determine the corresponding media interaction type.
[0021] According to a method for human-computer interaction based on multimodal dialogue state representation provided by the present disclosure, wherein the dialogue behavior is used to guide dialogue strategy generation, specifically comprising:
[0022] Obtain human-computer interaction scenarios;
[0023] Performing a dialogue behavior dimension analysis based on the scenario to obtain a dialogue behavior dimension analysis result;
[0024] Determine the generation of a dialogue strategy based on the dialogue behavior dimension analysis results.
[0025] According to a method for human-computer interaction based on multimodal dialogue state representation provided by the present disclosure, wherein the dialogue elements are used to determine the intention of the interlocutor, specifically including:
[0026] Obtaining the interlocutor's statement;
[0027] Performing multi-factor dialogue element representation on the sentence to obtain a multi-factor dialogue element representation result;
[0028] The intention of the interlocutor is determined according to the multi-factor dialogue element representation result.
[0029] According to a method for human-computer interaction based on multimodal dialogue state representation provided by the present disclosure, wherein the dialogue scenario is used to determine the corresponding media interaction type, specifically including:
[0030] Performing user portrait analysis, media type analysis, style and emotion analysis, and device type analysis on the interlocutor to obtain user portrait results, media type results, style and emotion results, and device type results of the interlocutor respectively;
[0031] The media interaction type for interacting with the interlocutor is determined according to the user portrait result, media type result, style emotion result, and device type result of the interlocutor.
[0032] According to a method for human-computer interaction based on multimodal dialogue state representation provided by the present disclosure, wherein the multi-factor dialogue element representation of the sentence is performed to obtain the multi-factor dialogue element representation result, specifically including:
[0033] Factoring the sentence from a semantic perspective to obtain four dimensional factors: action, object, condition, and question type.
[0034] Determine the multi-factor dialogue element representation result according to the four dimensional factors;
[0035] Here, the action refers to the predicate part of the sentence, which is undertaken by the verb or adjective in the sentence;
[0036] The object is the effector of the action, or the central word of a noun phrase sentence;
[0037] The conditions refer to the state and condition of the action, the modification and attribute of the object;
[0038] Question types refer to different query request categories in the interaction process set according to common sense knowledge.
[0039] In a second aspect, the present disclosure provides a device for human-computer interaction based on multimodal dialogue state representation, comprising:
[0040] A first processing module is used to obtain original multimodal input information;
[0041] a second processing module, configured to process the original multimodal input information to obtain a multimodal dialogue state representation result;
[0042] A third processing module is configured to determine a multimodal dialogue strategy according to the multimodal dialogue state representation result;
[0043] The fourth processing module is used to complete the multimodal information output according to the multimodal dialogue strategy.
[0044] The present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of any of the above-described methods for human-computer interaction based on multimodal dialogue state representation are implemented.
[0045] The present disclosure also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the above-described methods for human-computer interaction based on multimodal dialogue state representation.
[0046] The present disclosure provides a method and apparatus for human-computer interaction based on multimodal dialogue state representation. The method and apparatus obtain original multimodal input information and process the original multimodal input information to obtain a multimodal dialogue state representation result. The multimodal dialogue state representation result can represent dialogue features from multiple dimensions, creating a more humanized effect. A multimodal dialogue strategy is determined based on the multimodal dialogue state representation result. After obtaining the multimodal dialogue strategy, multimodal information output is completed based on the multimodal dialogue strategy. Because the multimodal dialogue representation has a more accurate representation effect, the multimodal output constructed based on the multimodal representation result is more accurate and can better demonstrate diverse and humanized interaction methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the present disclosure or the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0048] Figure 1 is a flowchart of a method for human-computer interaction based on multimodal dialogue state representation provided by the present disclosure;
[0049] Figure 2 This is a schematic diagram of the architecture of the multimodal human-human interaction system provided by the present disclosure;
[0050] Figure 3 This is a schematic diagram of a deep multimodal conversation state representation method provided by the present disclosure;
[0051] Figure 4 Schematic diagram of the structure of a human-computer interaction device based on multimodal dialogue state representation provided by the present disclosure;
[0052] Figure 5It is a structural diagram of the electronic device provided by the present disclosure. DETAILED DESCRIPTION
[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the embodiments of the present disclosure.
[0054] The following combination Figure 1-Figure 2 A method for human-computer interaction based on multimodal dialogue state representation according to an embodiment of the present disclosure is described, including:
[0055] Step 100: Obtain original multimodal input information;
[0056] Specifically, with the rapid development of big data, deep learning, and computing power, computers have evolved into intelligent systems capable of representing and recognizing multimodal information, including speech, vision, and text, and integrating knowledge to achieve understanding and reasoning. To address the "communication barrier" between humans and machines when ordinary users complete complex tasks in diverse scenarios, human-computer interaction (HCI) is entering a new phase: intelligent user interface (IUI). Typical applications include intelligent customer service for scenarios such as telephone and online text customer service, as well as face-to-face consultation, sales, and service. In terms of multi-turn dialogue research and HCI open platforms, existing HCI systems have achieved good results in performing specific tasks in specific domains and modalities. However, their multi-turn dialogue capabilities in multimodal, complex scenarios, limited resources, and cold start scenarios need to be improved. For example, existing dialogue management technologies are mostly limited to single-domain task-based dialogues. They lack global dialogue management that integrates multiple response modules, such as task-based dialogue, intelligent question-answering, knowledge graph question-answering, and chat, and lack the ability to generate emotional responses in complex, noisy scenarios.
[0057] Since the present disclosure is aimed at achieving multiple tasks in various situations, such as during robot dialogue and communication, it is necessary to complete a variety of dialogue behaviors such as self-introduction, product recommendation, invitation to review, emotional comfort, etc., and only by adopting such anthropomorphic and stylized strategic responses can a better user experience be provided.
[0058] Each source or form of information can be called a modality. For example, people have touch, hearing, vision, and smell; the media of information include voice, video, text, etc.; there are various sensors, such as radar, infrared, accelerometers, etc. Each of the above can be called a modality. At the same time, modality can also have a very broad definition. For example, two different languages can be regarded as two modalities, and even data sets collected under two different circumstances can be considered as two modalities. Therefore, the present disclosure obtains original multimodal input information, which can be in text form, image form, or video form. Obtain multimodal input data and call the digital human capability interface to parse the multimodal input data
[0059] Step 200: Processing the original multimodal input information to obtain a multimodal dialogue state representation result;
[0060] Specifically, in the present disclosure, by processing the original multimodal information, the input original multimodal information is processed from multiple dimensions, such as semantic perspective, action perspective, etc., to obtain a multimodal dialogue state representation result.
[0061] Step 300: Determine a multimodal dialogue strategy according to the multimodal dialogue state representation result;
[0062] Specifically, in the present disclosure, after determining the multimodal dialogue state representation result, the response method of the machine in the human-computer interaction can be determined by the multimodal state representation result, such as whether the selected dialogue emotional state is comfort or gratitude, calmness or joy, etc., and whether the dialogue method is visual or voice, in the form of pictures, videos or text, etc.
[0063] Step 400: Complete multimodal information output according to the multimodal dialogue strategy.
[0064] Specifically, after determining a multimodal dialogue strategy, the robot's multimodal output process essentially involves accessing the robot system's multimodal resource data and outputting it in various ways. For example, if the robot is to display facial expressions, this can be achieved by playing videos or displaying images on a display screen. The multimodal resource data involved in a robot system typically includes audio data, video data, image data, or other multimedia data, as well as program instructions for controlling the motors that drive the robot's movements.
[0065] The present disclosure provides a method for human-computer interaction based on multimodal dialogue state representation. The method obtains raw multimodal input information and processes it to obtain a multimodal dialogue state representation result. The multimodal dialogue state representation result can represent dialogue features from multiple dimensions, creating a more humanized effect. A multimodal dialogue strategy is determined based on the multimodal dialogue state representation result. After obtaining the multimodal dialogue strategy, multimodal information output is completed based on the multimodal dialogue strategy. Because the multimodal dialogue representation has a more accurate representation effect, the multimodal output constructed based on the multimodal representation result is more accurate and can better demonstrate diverse and humanized interaction methods.
[0066] According to an embodiment of the present disclosure, a method for human-computer interaction based on multimodal dialogue state representation is provided, wherein processing the original multimodal input information to obtain a multimodal dialogue state representation result specifically includes:
[0067] Performing a unimodal analysis on the original multimodal input information to obtain a unimodal representation result;
[0068] Acquiring dialogue scene related information according to the original multimodal input information;
[0069] Multimodal understanding and text semantic analysis are performed on the unimodal representation result and the dialogue scene related information to obtain a multimodal dialogue state representation result.
[0070] Specifically, based on the multimodal data feature representation model, the semantic features of multimodal data are extracted, and a data feature extraction model for text, image, audio and video based on the pre-trained model is constructed. Based on the feature extraction model, the semantic feature extraction of unimodal data, the semantic feature extraction of text data, the image feature extraction, the video feature extraction, the textual description information extraction and textual description of image data, and the textual description information extraction of video are completed respectively; in addition, it also includes processes such as behavioral gesture analysis and sentiment analysis to obtain unimodal representation results.
[0071] In addition, the original multimodal input information is analyzed for dialogue scene information. The scene information represents the environmental information of the dialogue scene, which is closely related to anthropomorphic audio-visual perception. The response type can be specifically described from four perspectives: user portrait (Perona), media type (Media), style emotion (Style), and device type (Device).
[0072] Multimodal understanding aims to achieve the ability to process and understand multimodal information through machine learning. Currently, a popular research direction is multimodal learning between images, videos, audio, and semantics.
[0073] Multimodal deep semantic understanding enables simultaneous semantic understanding of text and visual images. For example, in traditional AI recognition, when identifying a puppy under the shade of a tree, two objects are identified and classified: the puppy and the tree. Based on visual semantic understanding, the puppy is cool in the shade. A deeper understanding of the underlying meaning of the text is that the puppy is cool in the shade, while it is a scorching summer day outside. This is multimodal deep semantic understanding.
[0074] Furthermore, multimodal understanding and text semantic analysis are performed on the above unimodal representation result and the dialogue scene related information to obtain a multimodal dialogue state representation result.
[0075] According to an embodiment of the present disclosure, a method for human-computer interaction based on multimodal dialogue state representation is provided, wherein performing a unimodal analysis on the original multimodal input information to obtain a unimodal representation result specifically includes:
[0076] Performing speech recognition on the original multimodal input information to obtain a speech recognition result, and performing semantic analysis on the speech recognition result to obtain a semantic analysis result;
[0077] Performing sentiment analysis and behavioral gesture analysis on the original multimodal input information to obtain corresponding sentiment analysis results and behavioral gesture analysis results;
[0078] The semantic analysis result, the sentiment analysis result and the behavior gesture analysis result constitute a unimodal representation result.
[0079] Specifically, by calling the speech recognition, emotion analysis, and gesture behavior analysis interfaces, the original multimodal input information is processed accordingly to obtain the corresponding unimodal representation results.
[0080] According to an embodiment of the present disclosure, a method for human-computer interaction based on multimodal dialogue state representation is provided, wherein the multimodal dialogue state representation result includes dialogue behavior, dialogue elements and dialogue scene;
[0081] Among them, dialogue behavior is used to guide dialogue strategy generation;
[0082] Dialogue elements are used to determine the interlocutors’ intentions;
[0083] The dialogue scenario is used to determine the corresponding media interaction type.
[0084] Specifically, refer to Figure 3This disclosure comprehensively describes the information required for dialogue decision-making and dialogue generation in human-computer interaction from three perspectives: dialogue behavior, namely, multi-dimensional dialogue behavior analysis, multi-factor dialogue elements, dialogue scenarios, and multi-type dialogue scenario description. This method meticulously depicts dialogue behaviors, dialogue semantic elements, and dialogue scenario information to support anthropomorphic and multimodal interaction, meeting the needs of human-computer interaction in complex scenarios. Furthermore, this method is not bound to domain knowledge and is universally applicable to human-computer dialogue systems in various fields, such as e-commerce, tourism, and healthcare.
[0085] According to an embodiment of the present disclosure, a method for human-computer interaction based on multimodal dialogue state representation is provided, wherein the dialogue behavior is used to guide dialogue strategy generation, specifically including:
[0086] Obtain human-computer interaction scenarios;
[0087] Performing a dialogue behavior dimension analysis based on the scenario to obtain a dialogue behavior dimension analysis result;
[0088] Determine the generation of a dialogue strategy based on the dialogue behavior dimension analysis results.
[0089] Specifically, dialogue acts (DAs) play a crucial role in spoken language understanding systems. They are used to mark the speaker's intentions (such as statements, questions, promises, instructions, and so on). They are not constrained by specific dialogue systems and therefore have a certain degree of universality. Dialogue acts, also known as speech acts or communicative acts, are attempts to formalize and generalize intentions. Conversations revolve around the interaction between the speaker (sender) and the addressee (addressee). The speaker is the person who speaks in the current interaction, that is, the person who generates the current dialogue act. The addressee is a participant in the dialogue interaction and the person the current speaker is interacting with.
[0090] This paper combines the characteristics of human-computer interaction scenarios such as manual telephone calls and online customer service, such as pre-sales customer service personnel recommending products during service and after-sales customer service personnel providing emotional comfort to customers. Based on common international dialogue behavior schemes, it proposes a dialogue behavior classification scheme that can fully characterize the characteristics of spoken communication in human-computer interaction. It defines five different dialogue behavior analysis dimensions (Task dimension, Time Management dimension, Feedback dimension, Own and Partner Communication Management dimension, and Social Obligations Management dimension), which together represent the complex spoken dialogue behavior state in real-world scenarios. The definitions of each dimension are shown in Table 1:
[0091] Table 1
[0092]
[0093]
[0094] According to an embodiment of the present disclosure, a method for human-computer interaction based on multimodal dialogue state representation is provided, wherein the dialogue elements are used to determine the intention of the interlocutor, specifically including:
[0095] Obtaining the interlocutor's statement;
[0096] Performing multi-factor dialogue element representation on the sentence to obtain a multi-factor dialogue element representation result;
[0097] The intention of the interlocutor is determined according to the multi-factor dialogue element representation result.
[0098] Specifically, the present invention proposes a simple and novel semantic representation framework, called the multi-factor dialogue element representation framework, which is used to replace the classic semantic representation method that combines intent and slot value. Under this framework, different intents are distinguished by four key factors, namely action, object, condition and question type. Four key concepts are used to distinguish intents instead of fully expressing all the complex grammatical and semantic meanings of the sentence. The multi-factor semantic framework is inspired by the fact that the number of possible sentences is infinite, so it is not feasible to fully express the meaning of all sentences. At the same time, overly general semantic representations are difficult to meet the needs of real scenarios. However, in specific domains or scenarios, the possible semantic space is limited. Therefore, limited key concepts can be used to distinguish all intents without directly using a complete representation to represent each intent, and the representation granularity can also meet the needs of the scenario.
[0099] This approach primarily addresses the challenge of describing user intent at a fine-grained level. It distinguishes different semantic intents through factor differentiation and links semantic knowledge points through factor sharing, allowing for convergence while reserving differences. This framework employs four factors: action, object, condition (modifier / attribute / state / condition), and question type. The condition (modifier / attribute / state / condition) dimension is a hybrid, encompassing multiple finer-grained factors such as modifier / attribute / state / condition. Since these factors typically do not co-occur in the same question, they are combined into a single dimension, referred to as condition. Multiple factors are obtained through multidimensional factor semantic analysis of a question and then concatenated to form a factor expression to represent the semantics. For example, the factor expression for the question "How do I reimburse a hotel invoice?" is "reimburse (action) + invoice (object) + hotel (condition) + how-question (question type)." Brackets represent the factor type. Generally, the four factors are arranged in a fixed order, so the factor type can be omitted. In this case, the factor expression is "reimburse + invoice + hotel + how-question."
[0100] According to an embodiment of the present disclosure, a method for human-computer interaction based on multimodal dialogue state representation is provided, wherein the dialogue scenario is used to determine the corresponding media interaction type, specifically including:
[0101] Performing user portrait analysis, media type analysis, style and emotion analysis, and device type analysis on the interlocutor to obtain user portrait results, media type results, style and emotion results, and device type results of the interlocutor respectively;
[0102] The media interaction type for interacting with the interlocutor is determined according to the user portrait result, media type result, style emotion result, and device type result of the interlocutor.
[0103] Specifically, scene information represents the environmental information of the scene where the conversation takes place, which is closely related to anthropomorphic audio-visual perception. It can specifically describe the type of response from four perspectives: user portrait (Persona), media type (Media), style and emotion (Style), and device type (Device). User portrait describes the user portrait information of the speaker (for example, age and occupation, interests and hobbies, etc.). Media type indicates the preferred presentation media type, and the form in which input and output are displayed to complete the interaction (for example, text, oral and charts, pictures, etc.). Style and emotion express the emotional attitude held when expressing the current discourse (for example, anger, anxiety, etc.). Device type refers to which devices will be used in the demonstration, and the devices ultimately provide physical hardware support for the interaction (such as web pages, phones, or PDAs, etc.).
[0104] In a multimodal dialogue system, the dialogue decision (Policy) unit fully considers this dialogue scenario information when making dialogue strategy decisions, selecting the most appropriate media interaction type. When generating responses, the natural language generation (NLG) unit can also leverage information such as user profiles and style to generate rich, stylized responses. This provides a better user experience, increases user engagement, and improves the completion rate of dialogue tasks.
[0105] According to an embodiment of the present disclosure, a method for human-computer interaction based on multimodal dialogue state representation is provided, wherein the step of performing multi-factor dialogue element representation on the sentence to obtain the multi-factor dialogue element representation result specifically includes:
[0106] Factoring the sentence from a semantic perspective to obtain four dimensional factors: action, object, condition, and question type.
[0107] Determine the multi-factor dialogue element representation result according to the four dimensional factors;
[0108] Here, the action refers to the predicate part of the sentence, which is undertaken by the verb or adjective in the sentence;
[0109] The object is the effector of the action, or the central word of a noun phrase sentence;
[0110] The conditions refer to the state and condition of the action, the modification and attribute of the object;
[0111] Question types refer to different query request categories in the interaction process set according to common sense knowledge.
[0112] Specifically, factoring does not rely on specific domain knowledge. It can be done by analyzing the subject, predicate, and object components of a sentence according to general Chinese syntactic analysis and understanding the central idea of the sentence from a semantic perspective. When factoring, the following rules can be followed:
[0113] The action, that is, the predicate part of a sentence, is usually performed by the verbs or adjectives in the sentence.
[0114] The object is the effector of the action, or the central word of a noun phrase sentence.
[0115] Conditions, i.e., the state and condition of an action, and the modification and attribute of an object, usually do not appear together in a sentence, so they are represented by one dimension.
[0116] Question types refer to different query request categories in the interaction process set according to common sense knowledge: yesno-quesiton, affirmative and negative questions; choice-question, selection question; where-question, location question; when-question, time question; why-question, reason question; whynot-question, negative reason question; what-question, entity question; who-question, name question; how-question, action / state question; howoften-question, frequency question; howmany-question, quantity question; statement-positive, affirmative sentence; statement-negative, negative sentence.
[0117] Combine Figure 4 As shown, the present disclosure provides a device for human-computer interaction based on multimodal dialogue state representation, comprising:
[0118] A first processing module 41 is configured to obtain original multimodal input information;
[0119] A second processing module 42 is configured to process the original multimodal input information to obtain a multimodal dialogue state representation result;
[0120] A third processing module 43 is configured to determine a multimodal dialogue strategy according to the multimodal dialogue state representation result;
[0121] The fourth processing module 44 is configured to output multimodal information according to the multimodal dialogue strategy.
[0122] Since the device provided in the embodiment of the present invention can be used to execute the method described in the above embodiment, its working principle and beneficial effects are similar, so they will not be described in detail here. For specific details, please refer to the introduction of the above embodiment.
[0123] The present disclosure provides a human-computer interaction device based on multimodal dialogue state representation. The device obtains raw multimodal input information and processes it to obtain a multimodal dialogue state representation result. The multimodal dialogue state representation result can represent dialogue features from multiple dimensions, creating a more humanized effect. A multimodal dialogue strategy is determined based on the multimodal dialogue state representation result. After obtaining the multimodal dialogue strategy, multimodal information output is completed based on the multimodal dialogue strategy. Because the multimodal dialogue representation has a more accurate representation effect, the multimodal output constructed based on the multimodal representation result is more accurate and can better demonstrate diverse and humanized interaction methods.
[0124] According to an embodiment of the present disclosure, a device for human-computer interaction based on multimodal dialogue state representation is provided, wherein the second processing module 42 is specifically configured to:
[0125] Performing a unimodal analysis on the original multimodal input information to obtain a unimodal representation result;
[0126] Acquiring dialogue scene related information according to the original multimodal input information;
[0127] Multimodal understanding and text semantic analysis are performed on the unimodal representation result and the dialogue scene related information to obtain a multimodal dialogue state representation result.
[0128] According to an embodiment of the present disclosure, a device for human-computer interaction based on multimodal dialogue state representation is provided, wherein the second processing module 42 is further specifically configured to:
[0129] Performing speech recognition on the original multimodal input information to obtain a speech recognition result, and performing semantic analysis on the speech recognition result to obtain a semantic analysis result;
[0130] Performing sentiment analysis and behavioral gesture analysis on the original multimodal input information to obtain corresponding sentiment analysis results and behavioral gesture analysis results;
[0131] The semantic analysis result, the sentiment analysis result and the behavior gesture analysis result constitute a unimodal representation result.
[0132] According to an embodiment of the present disclosure, a device for human-computer interaction based on multimodal dialogue state representation is provided, wherein, in the second processing module 42, the multimodal dialogue state representation result includes dialogue behavior, dialogue elements and dialogue scene;
[0133] Among them, dialogue behavior is used to guide dialogue strategy generation;
[0134] Dialogue elements are used to determine the interlocutors’ intentions;
[0135] The dialogue scenario is used to determine the corresponding media interaction type.
[0136] According to an embodiment of the present disclosure, a device for human-computer interaction based on multimodal dialogue state representation is provided, wherein the second processing module 42 is further specifically configured to:
[0137] Obtain human-computer interaction scenarios;
[0138] Performing a dialogue behavior dimension analysis based on the scenario to obtain a dialogue behavior dimension analysis result;
[0139] Determine the generation of a dialogue strategy based on the dialogue behavior dimension analysis results.
[0140] According to an embodiment of the present disclosure, a device for human-computer interaction based on multimodal dialogue state representation is provided, wherein the second processing module 42 is further specifically configured to:
[0141] Obtaining the interlocutor's statement;
[0142] Performing multi-factor dialogue element representation on the sentence to obtain a multi-factor dialogue element representation result;
[0143] The intention of the interlocutor is determined according to the multi-factor dialogue element representation result.
[0144] According to an embodiment of the present disclosure, a device for human-computer interaction based on multimodal dialogue state representation is provided, wherein the second processing module 42 is further specifically configured to:
[0145] Performing user portrait analysis, media type analysis, style and emotion analysis, and device type analysis on the interlocutor to obtain user portrait results, media type results, style and emotion results, and device type results of the interlocutor respectively;
[0146] The media interaction type for interacting with the interlocutor is determined according to the user portrait result, media type result, style emotion result, and device type result of the interlocutor.
[0147] According to an embodiment of the present disclosure, a device for human-computer interaction based on multimodal dialogue state representation is provided, wherein the second processing module 42 is further specifically configured to:
[0148] Factoring the sentence from a semantic perspective to obtain four dimensional factors: action, object, condition, and question type.
[0149] Determine the multi-factor dialogue element representation result according to the four dimensional factors;
[0150] Here, the action refers to the predicate part of the sentence, which is undertaken by the verb or adjective in the sentence;
[0151] The object is the effector of the action, or the central word of a noun phrase sentence;
[0152] The conditions refer to the state and condition of the action, the modification and attribute of the object;
[0153] Question types refer to different query request categories in the interaction process set according to common sense knowledge.
[0154] Figure 5 An example of a physical structure diagram of an electronic device is shown below. Figure 5As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communications bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communications bus 540. The processor 510 may invoke logic instructions in the memory 530 to execute a method for human-computer interaction based on multimodal dialogue state representation provided by the present disclosure, the method comprising: obtaining original multimodal input information; processing the original multimodal input information to obtain a multimodal dialogue state representation result; determining a multimodal dialogue strategy based on the multimodal dialogue state representation result; and outputting multimodal information based on the multimodal dialogue strategy.
[0155] In addition, the logic instructions in the above-mentioned memory 530 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the embodiment of the present disclosure is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0156] On the other hand, the present disclosure also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the above methods. The present disclosure provides a method for human-computer interaction based on multimodal dialogue state representation, which includes: obtaining original multimodal input information; processing the original multimodal input information to obtain a multimodal dialogue state representation result; determining a multimodal dialogue strategy based on the multimodal dialogue state representation result; and completing multimodal information output according to the multimodal dialogue strategy.
[0157] On the other hand, the present disclosure also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the above-mentioned embodiments. The present disclosure provides a method for human-computer interaction based on multimodal dialogue state representation, the method comprising: obtaining original multimodal input information; processing the original multimodal input information to obtain a multimodal dialogue state representation result; determining a multimodal dialogue strategy based on the multimodal dialogue state representation result; and completing multimodal information output according to the multimodal dialogue strategy.
[0158] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0159] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present disclosure.
Claims
1. A method for human-computer interaction based on multimodal dialogue state representation, characterized in that: include: Get the original multimodal input information; Processing the original multimodal input information to obtain a multimodal dialogue state representation result; Determining a multimodal dialogue strategy according to the multimodal dialogue state representation result; completing multimodal information output according to the multimodal dialogue strategy; The processing of the original multimodal input information to obtain a multimodal dialogue state representation result specifically includes: Performing unimodal analysis on the original multimodal input information to obtain a unimodal representation result; wherein the unimodal representation result includes: a semantic analysis result, a sentiment analysis result, and a behavioral gesture analysis result; Acquiring conversation scene related information based on the original multimodal input information; wherein the conversation scene related information represents: environmental information of the conversation scene, which is related to anthropomorphic audio-visual perception, including: user profile, media type, style emotion, and device type; Multimodal understanding and text semantic analysis are performed on the unimodal representation result and the information related to the dialogue scene to obtain a multimodal dialogue state representation result; wherein the multimodal dialogue state representation result includes: dialogue behavior, dialogue elements and dialogue scene; dialogue behavior is used to guide dialogue strategy generation, dialogue elements are used to determine the intention of the interlocutor, and dialogue scene is used to determine the corresponding media interaction type.
2. The method for human-computer interaction based on multimodal dialogue state representation according to claim 1, characterized in that: The performing a single-modal analysis on the original multimodal input information to obtain a single-modal representation result specifically includes: Performing speech recognition on the original multimodal input information to obtain a speech recognition result, and performing semantic analysis on the speech recognition result to obtain a semantic analysis result; Performing sentiment analysis and behavioral gesture analysis on the original multimodal input information to obtain corresponding sentiment analysis results and behavioral gesture analysis results; The semantic analysis result, the sentiment analysis result and the behavior gesture analysis result constitute a unimodal representation result.
3. The method for human-computer interaction based on multimodal dialogue state representation according to claim 2, characterized in that: The dialogue behavior is used to guide the generation of dialogue strategies, specifically including: Obtain human-computer interaction scenarios; Performing a dialogue behavior dimension analysis based on the scenario to obtain a dialogue behavior dimension analysis result; Determine the generation of a dialogue strategy based on the dialogue behavior dimension analysis results.
4. The method for human-computer interaction based on multimodal dialogue state representation according to claim 2, characterized in that: The dialogue elements are used to determine the intentions of the interlocutors, and specifically include: Obtaining the interlocutor's statement; Performing multi-factor dialogue element representation on the sentence to obtain a multi-factor dialogue element representation result; The intention of the interlocutor is determined according to the multi-factor dialogue element representation result.
5. The method for human-computer interaction based on multimodal dialogue state representation according to claim 2, characterized in that: The conversation scenario is used to determine the corresponding media interaction type, specifically including: Performing user portrait analysis, media type analysis, style and emotion analysis, and device type analysis on the interlocutor to obtain user portrait results, media type results, style and emotion results, and device type results of the interlocutor respectively; The media interaction type for interacting with the interlocutor is determined according to the user portrait result, media type result, style emotion result, and device type result of the interlocutor.
6. The method for human-computer interaction based on multimodal dialogue state representation according to claim 4, characterized in that: The performing multi-factor dialogue element representation on the sentence to obtain a multi-factor dialogue element representation result specifically includes: Factoring the sentence from a semantic perspective to obtain four dimensional factors: action, object, condition, and question type. Determine the multi-factor dialogue element representation result according to the four dimensional factors; Here, the action refers to the predicate part of the sentence, which is undertaken by the verb or adjective in the sentence; The object is the effector of the action, or the central word of a noun phrase sentence; The conditions refer to the state and condition of the action, the modification and attribute of the object; Question types refer to different query request categories in the interaction process set according to common sense knowledge.
7. A device for human-computer interaction based on multimodal dialogue state representation, characterized in that: include: A first processing module is used to obtain original multimodal input information; a second processing module, configured to process the original multimodal input information to obtain a multimodal dialogue state representation result; A third processing module is configured to determine a multimodal dialogue strategy according to the multimodal dialogue state representation result; a fourth processing module, configured to output multimodal information according to the multimodal dialogue strategy; The second processing module is specifically configured to: Performing unimodal analysis on the original multimodal input information to obtain a unimodal representation result; wherein the unimodal representation result includes: a semantic analysis result, a sentiment analysis result, and a behavioral gesture analysis result; Acquiring conversation scene related information based on the original multimodal input information; wherein the conversation scene related information represents: environmental information of the conversation scene, which is related to anthropomorphic audio-visual perception, including: user profile, media type, style emotion, and device type; Multimodal understanding and text semantic analysis are performed on the unimodal representation result and the information related to the dialogue scene to obtain a multimodal dialogue state representation result; wherein the multimodal dialogue state representation result includes: dialogue behavior, dialogue elements and dialogue scene; dialogue behavior is used to guide dialogue strategy generation, dialogue elements are used to determine the intention of the interlocutor, and dialogue scene is used to determine the corresponding media interaction type.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the human-computer interaction method based on multimodal dialogue state representation as claimed in any one of claims 1 to 6 are implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the human-computer interaction method based on multimodal dialogue state representation as claimed in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Multimode interaction method and device for intelligent robots
CN106933345A
Interaction method and system based on intelligent robot
CN109278051A
A short text question semantic matching method and system
CN109597994A