Interaction system and method based on artificial intelligence
The system addresses the limitations of existing human-computer interaction systems by leveraging multi-modal data capture and advanced AI techniques for intent understanding and context management, resulting in improved interaction accuracy and coherence.
Patent Information
- Application Number
- CN202510401599.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing interactive systems have limited ability to understand user intentions, lack the utilization and memory of context information, and the multimodal data fusion processing is not mature enough, resulting in the lack of smooth and natural interaction experience.
Multimodal data acquisition, preprocessing, intention understanding, context management, knowledge graph and dynamic decision-making modules are adopted, combined with deep learning models and improved decision tree algorithms, to achieve efficient fusion and context perception of multimodal data.
It improves the accuracy of user intention understanding and interaction accuracy, enhances the consistency and fluency of conversations, and provides a more intelligent and natural interactive experience.
Smart Images

Figure CN120316221A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and specifically provides an artificial intelligence-based interaction system and method. Background Art
[0002] With the continuous development of artificial intelligence technology, human-computer interaction systems are increasingly widely used in daily life and work, such as in scenarios like intelligent customer service, smart home control, and intelligent assistants. However, there are still many deficiencies in existing interaction systems.
[0003] On the one hand, most interaction systems have limited ability to understand user intentions. When users express their needs in vague, metaphorical, or complex natural language, the system often has difficulty accurately parsing, resulting in unsatisfactory interaction results. For example, in the intelligent customer service scenario, when a user asks "Are there any products suitable for summer use, that are light, breathable, and can also block UV rays?", traditional systems may not be able to fully understand these complex requirements and can only give partial or inaccurate responses.
[0004] On the other hand, existing interaction systems lack effective utilization and memory of context information. During multi-round conversations, they are unable to coherently understand user intentions. Each round of conversation seems to exist independently and cannot make comprehensive judgments and responses based on previous communication content. For example, if a user first asks "What colors does a certain brand of mobile phone come in?", and then asks "Which color is more popular?", the system may not be able to associate these two questions and may not realize that the second question is based on the mobile phone colors mentioned in the first question.
[0005] In addition, the fusion processing technology for different modality data (such as voice, text, gestures, expressions, etc.) is not yet mature. Currently, most interaction systems only support single-modality interaction, or in the case of multi-modal interaction, they cannot achieve deep fusion and collaborative work of different modality information, and cannot fully utilize the advantages of multi-modal interaction, resulting in an unsmooth and unnatural interaction experience.
[0006] Therefore, how to improve the understanding accuracy of the interaction system for user intentions, enhance context awareness and memory capabilities, and achieve efficient fusion of multi-modal data has become an urgent problem to be solved. Summary of the Invention
[0007] (1) Technical Problems to be Solved
[0008] In view of the deficiencies of the prior art, the present invention provides an artificial intelligence-based interaction system and method, which solves the problems raised in the above background art.
[0009] (2) Technical Solutions
[0010] To achieve the above objectives, the present invention is realized through the following technical solutions: An artificial intelligence-based interaction system and method, comprising:
[0011] A multi-modal data acquisition module for acquiring multi-modal input data of the user's voice, text, gestures, and expressions;
[0012] A data preprocessing module for performing preprocessing operations such as cleaning, noise reduction, and feature extraction on the acquired multi-modal data;
[0013] An intent understanding module that uses a deep learning model to understand the user's intent in combination with the features of multi-modal data;
[0014] A context management module for recording and managing the context information of the conversation, including historical conversation content, user preferences, and the current conversation state;
[0015] A knowledge graph module for constructing and maintaining a knowledge graph in the relevant field, containing information on entities, relationships, and attributes;
[0016] A dynamic decision-making module that uses an improved decision tree algorithm to make interaction decisions based on the user's intent, context information, and knowledge in the knowledge graph;
[0017] A response generation module for generating appropriate response content according to the interaction decision;
[0018] A multi-modal output module for outputting the generated response content to the user in a multi-modal form;
[0019] A feedback optimization module for collecting the user's feedback on the interaction content and optimizing the system performance.
[0020] Preferably, the intent understanding module uses a language model with a Transformer architecture to understand the user's intent in combination with the features of multi-modal data. The Transformer architecture uses the following formula:
[0021]
[0022] Where Q represents the query vector, K represents the key vector, V represents the value vector, dk represents the dimension of the vector, and the softmax function is used to calculate the attention weights. In this way, the model can capture the interaction information between text and images, thereby better understanding the user's intent.
[0023] Preferably, the context management module analyzes the coherence and logic of the user's intent in multi-round conversations through a dialogue state machine and historical conversation records.
[0024] Preferably, the knowledge graph module constructs a knowledge graph through techniques such as knowledge extraction and knowledge fusion, and uses a graph database for storage and management.
[0025] Preferably, the dynamic decision-making module adopts an improved decision tree algorithm to make interaction decisions based on the user's intention, context information, and knowledge in the knowledge graph. The dynamic decision-making module makes decisions using the following formula:
[0026] X = α I I + α C C + α K K
[0027] I represents the user intention dimension, C represents the context information dimension, and K represents the knowledge dimension in the knowledge graph; α I represents the weight of the user intention dimension, α C represents the weight of the context information dimension, and α K represents the weight of the knowledge dimension in the knowledge graph. In this way, the system can make decisions more intelligently according to the user's needs and context, improving the accuracy and efficiency of interaction.
[0028] The present invention further discloses an interaction method based on artificial intelligence. Based on the above-mentioned interaction system based on artificial intelligence, it includes the following steps:
[0029] Step 1: Collect multi-modal input data of the user's voice, text, gestures, and expressions through various sensors and input interfaces of the interaction system;
[0030] Step 2: Perform preprocessing of cleaning, noise reduction, and feature extraction on the collected multi-modal data;
[0031] Step 3: Input the preprocessed multi-modal data into the intention understanding module, and use a deep learning model to analyze and understand the user's intention in combination with the multi-modal data features;
[0032] Step 4: Through the context management module, perform context analysis on the user's intention according to the historical conversation record and the current conversation state to clarify the meaning and association of the user's intention in the entire conversation process;
[0033] Step 5: Use the knowledge graph module to perform knowledge retrieval and reasoning in the knowledge graph according to the user's intention and context information to obtain relevant knowledge and information;
[0034] Step 6: Through the dynamic decision-making module, make interaction decisions based on the user's intention, the result of context analysis, and the information obtained from knowledge reasoning;
[0035] Step 7: Generate appropriate response content according to the above interaction decision;
[0036] Step 8: The multi-modal output module outputs the generated response content to the user in the form of multi-modal of voice, text, and image to complete the human-computer interaction process;
[0037] Step Nine: Collect the user's feedback on the interactive content and optimize the system performance.
[0038] Preferably, in Step Four, a time series analysis method is used to analyze the historical conversation records to mine the user's behavior patterns and intention tendencies.
[0039] Preferably, in Step Five, the path search and inference rules of the knowledge graph are used to perform knowledge inference and expansion, providing richer information support for reply generation.
[0040] (III) Beneficial Effects
[0041] The present invention provides an interactive system and method based on artificial intelligence, having the following beneficial effects:
[0042] 1. Through multi-modal data collection and fusion processing, it can more comprehensively and accurately understand the user's intentions, make up for the limitations of single-modal interaction, and improve the accuracy and naturalness of interaction.
[0043] 2. Through the setting of the context management module, the system can maintain a coherent understanding of the user's intentions in multi-round conversations, achieving a more intelligent and smooth interaction experience.
[0044] 3. Through the introduction of the knowledge graph module, it provides rich knowledge support for the system, enabling it to perform reasoning and information supplementation based on knowledge, answering more accurately and comprehensively, and enhancing the intelligence and practicality of the interactive system. Description of the Drawings
[0045] Figure 1 It is a schematic diagram of the framework of the interactive system of the present invention. Detailed Embodiments
[0046] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0047] Embodiment 1:
[0048] As Figure 1 shown, the embodiment of the present invention provides an interactive system based on artificial intelligence, including:
[0049] Multi-modal data collection module: used to collect multi-modal input data of the user such as voice, text, gesture, and expression. For example, voice data is collected through a microphone, gesture and expression data are collected through a camera, and text data is collected through an input box.
[0050] Data preprocessing module: Perform preprocessing operations such as cleaning, noise reduction, and feature extraction on the collected multi-modal data. For example, perform endpoint detection and background noise removal on speech data, and extract acoustic features of speech; perform image enhancement and keypoint detection on image data, and extract features of expressions and gestures.
[0051] Intent understanding module: Use deep learning models, such as language models based on the Transformer architecture, combined with the features of multi-modal data to understand the user's intent. Specifically, the intent understanding module uses a language model based on the Transformer architecture and combines the features of multi-modal data to understand the user's intent. The Transformer architecture uses the following formula:
[0052]
[0053] where Q represents the query vector, K represents the key vector, V represents the value vector, dk represents the dimension of the vector, and the softmax function is used to calculate the attention weights. This module not only analyzes the text content but also integrates the intonation, speech rate, and emotional tendency of speech, as well as the information conveyed by gestures and expressions, to comprehensively judge the user's true intent. For example, when the user asks "I'm a bit cold" verbally and at the same time makes a gesture of hugging their arms, the system can more accurately understand that the user may hope to adjust the indoor temperature. In this way, the model can capture the interaction information between text and images, thereby better understanding the user's intent.
[0054] Context management module: Record and manage the context information of the conversation, including historical conversation content, user preferences, current conversation status, etc. The context management module analyzes the coherence and logic of the user's intent in multi-round conversations through a conversation state machine and historical conversation records. In multi-round conversations, it understands and responds to the user's new input based on the context information to ensure the coherence and logic of the conversation. For example, when the user asked about the price of a certain product in the previous round and then asks "Are there any promotional activities?" in the next round, the context management module can associate these two questions and understand that the user is asking about the promotional situation of the product whose price was asked before.
[0055] Knowledge graph module: Build and maintain a knowledge graph of the relevant field, including information such as entities, relationships, and attributes. The knowledge graph module constructs a knowledge graph through techniques such as knowledge extraction and knowledge fusion, and uses a graph database for storage and management. When understanding the user's intent and generating responses, it uses the knowledge in the knowledge graph for reasoning and information supplementation. For example, when the user asks "How about the processors of Apple phones?", the knowledge graph module can provide information about the processors of Apple phones in different generations, performance parameter comparisons, etc., to help the system answer questions more comprehensively and accurately.
[0056] The dynamic decision-making module uses an improved decision tree algorithm to make interaction decisions based on user intent, context information, and knowledge in the knowledge graph. Secondly, the dynamic decision-making module uses an improved decision tree algorithm to make interaction decisions based on user intent, context information, and knowledge in the knowledge graph. The dynamic decision-making module makes decisions using the following formula:
[0057] X = α I I + α C C + α K K
[0058] I represents the user intent dimension, C represents the context information dimension, and K represents the knowledge dimension in the knowledge graph; α I represents the weight of the user intent dimension, α C represents the weight of the context information dimension, α K represents the weight of the knowledge dimension in the knowledge graph. In this way, the system can make more intelligent decisions according to user needs and context situations, improving the accuracy and efficiency of interactions.
[0059] Reply generation module: Generate appropriate reply content according to the interaction decision. The reply content can be in various forms such as text, speech, image, animation, etc. to meet different scenarios and user needs. For example, for a user's question about the product usage method, in addition to providing a text description, the reply generation module can also generate an animated demonstration of the operation steps.
[0060] Multimodal output module: Output the generated reply content to the user in a multimodal form, such as playing the speech reply through a speaker, displaying the text and image reply on the screen, or providing a reply through the action feedback of a smart device (such as the limb movements of a smart robot).
[0061] Feedback optimization module: Collect user feedback on the interaction content and optimize the system performance.
[0062] The present invention also provides an interaction method based on artificial intelligence. Based on the above-mentioned interaction system based on artificial intelligence, it is characterized in that it includes the following steps:
[0063] Step 1: Collect multimodal input data of the user, such as speech, text, gestures, expressions, etc., through various sensors and input interfaces of the interaction system;
[0064] Step 2: Perform preprocessing on the collected multimodal data, such as cleaning, noise reduction, feature extraction, etc., so that subsequent modules can process and analyze it;
[0065] Step 3: Input the preprocessed multimodal data into the intent understanding module, and use a deep learning model to analyze and understand the user's intent in combination with the multimodal data features;
[0066] Step 4: The context management module performs context analysis on the user's intention based on the historical conversation records and the current conversation state to clarify the meaning and association of the user's intention throughout the conversation process. Secondly, it uses time series analysis methods to analyze the historical conversation records to mine the user's behavior patterns and intention tendencies.
[0067] Step 5: The knowledge graph module conducts knowledge retrieval and reasoning in the knowledge graph according to the user's intention and context information to obtain relevant knowledge and information. Secondly, it uses the path search and reasoning rules of the knowledge graph to perform knowledge reasoning and expansion to provide richer information support for reply generation.
[0068] Step 6: The dynamic decision-making module makes an interaction decision based on the user's intention, the results of context analysis, and the information obtained from knowledge reasoning.
[0069] Step 7: Generate appropriate reply content according to the above interaction decision.
[0070] Step 8: The multimodal output module outputs the generated reply content to the user in multimodal forms such as voice, text, image, etc., to complete the human-computer interaction process.
[0071] Step 9: Collect the user's feedback on the interaction content to optimize the system performance.
[0072] Embodiment 2: Intelligent customer service scenario
[0073] 1. Multimodal data collection: The user enters the text "The computer I bought won't turn on. What should I do?" through the online customer service interface, and at the same time, the camera of the customer service system captures the user's anxious expression. The multimodal data collection module collects text data and expression image data respectively.
[0074] 2. Data preprocessing: Perform preprocessing such as lexical analysis and syntactic analysis on the text data to extract key information; extract features from the expression image to determine the user's emotional state as anxious.
[0075] 3. Intention understanding: The intention understanding module combines the text content and the emotional information conveyed by the expression to understand that the user's intention is to seek help in solving the problem of the computer not turning on, and the user is currently in a relatively anxious mood and needs to get a solution as soon as possible.
[0076] 4. Context analysis: The context management module checks the historical conversation records and finds that the user purchased this brand of computer before, and confirms that this problem is for the purchased product.
[0077] 5. Knowledge reasoning: The knowledge graph module searches for the possible causes and solutions of "the computer won't turn on" in the knowledge graph related to computer failures, and obtains relevant knowledge such as power problems, hardware failures, and system failures.
[0078] 6. Reply generation: Based on the above analysis results, the reply generation module generates the reply content "Hello, we fully understand your current anxious mood. There could be various reasons why the computer won't turn on. First, please check if the power is plugged in properly and if the power cord is damaged. If the power is normal, you can try holding down the power button for 10 seconds to see if it can be forced to start. If the problem remains unresolved, it may be a hardware or system failure. You can contact our after-sales staff, and we will arrange technical support for you as soon as possible."
[0079] 7. Multimodal output: The multimodal output module displays the reply content in text form on the customer service interface and, at the same time, converts the reply content into speech through text-to-speech technology and plays it for the user.
[0080] Embodiment 3: Smart home control scenario
[0081] 1. Multimodal data collection: When the user returns home and says to the smart speaker "Turn on the living room light and set the air conditioner to 26 degrees", while making an opening gesture at the same time. The microphone of the smart speaker collects the voice data, and the camera collects the gesture data.
[0082] 2. Data preprocessing: Perform speech recognition on the voice data, convert it into text content, and extract speech features; identify and analyze the gesture data to determine that the meaning of the gesture is an opening operation.
[0083] 3. Intent understanding: The intent understanding module combines the voice and gesture information to understand that the user's intent is to control the living room light to turn on and adjust the air conditioner temperature to 26 degrees.
[0084] 4. Context analysis: The context management module confirms that the user's intent conforms to the normal usage scenario based on the user's historical usage habits and current scenario information (such as it being evening and the user just arriving home).
[0085] 5. Knowledge reasoning: The knowledge graph module looks up the control instructions and related device information of the living room light and air conditioner in the smart home knowledge graph.
[0086] 6. Reply generation: The reply generation module generates the reply "The living room light has been turned on for you, and the air conditioner is being adjusted to 26 degrees."
[0087] 7. Multimodal output: The smart speaker plays the reply content for the user through voice and, at the same time, displays the feedback information of the operation result on the smart device control interface.
[0088] Although embodiments of the present invention have been shown and described, those of ordinary skill in the art will appreciate that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An artificial intelligence-based interaction system, characterized in that: It includes: A multi-modal data acquisition module for acquiring multi-modal input data of the user's speech, text, gestures, and expressions; A data preprocessing module for performing preprocessing operations such as cleaning, noise reduction, and feature extraction on the acquired multi-modal data; An intention understanding module that uses a deep learning model and combines the features of multi-modal data to understand the user's intention; A context management module for recording and managing the context information of the conversation, including historical conversation content, user preferences, and the current conversation state; A knowledge graph module for constructing and maintaining a knowledge graph in the relevant field, containing information on entities, relationships, and attributes; A dynamic decision-making module that uses an improved decision tree algorithm to make interaction decisions based on the user's intention, context information, and knowledge in the knowledge graph; A response generation module for generating appropriate response content according to the interaction decision; A multi-modal output module for outputting the generated response content to the user in a multi-modal form; A feedback optimization module for collecting the user's feedback on the interaction content and optimizing the system performance.
2. An artificial intelligence-based interaction system according to claim 1, characterized in that: The intention understanding module uses a language model with a Transformer architecture to understand the user's intention by combining the features of multi-modal data. The Transformer architecture uses the following formula: where Q represents the query vector, K represents the key vector, V represents the value vector, dk represents the dimension of the vector, and the softmax function is used to calculate the attention weights.
3. An artificial intelligence-based interaction system according to claim 1, characterized in that: The context management module analyzes the coherence and logic of the user's intention in multi-round conversations through a dialogue state machine and historical conversation records.
4. An artificial intelligence-based interaction system according to claim 1, characterized in that: The knowledge graph module constructs a knowledge graph through techniques such as knowledge extraction and knowledge fusion, and stores and manages it using a graph database.
5. An artificial intelligence-based interaction system according to claim 1, characterized in that: The dynamic decision-making module uses an improved decision tree algorithm to make interaction decisions based on the user's intention, context information, and knowledge in the knowledge graph. The dynamic decision-making module uses the following formula for decision-making: X = α I I + α C C + α K K I represents the user intention dimension, C represents the context information dimension, and K represents the knowledge dimension in the knowledge graph; α I represents the weight of the user intention dimension, α C represents the weight of the context information dimension, α K represents the weight of the knowledge dimension in the knowledge graph.
6. An artificial intelligence-based interaction method, and an artificial intelligence-based interaction system according to any one of claims 1-5, characterized in that: It includes the following steps: Step 1: Through various sensors and input interfaces of the interaction system, acquire multi-modal input data of the user's speech, text, gestures, and expressions; Step 2: Perform preprocessing such as cleaning, noise reduction, and feature extraction on the acquired multi-modal data; Step 3: Input the preprocessed multi-modal data into the intention understanding module, and use a deep learning model to analyze and understand the user's intention by combining the features of multi-modal data; Step 4: Through the context management module, perform context analysis on the user's intention based on the historical conversation record and the current conversation state to clarify the meaning and association of the user's intention throughout the conversation process; Step 5: Use the knowledge graph module to perform knowledge retrieval and reasoning in the knowledge graph according to the user's intention and context information to obtain relevant knowledge and information; Step 6: Through the dynamic decision-making module, make interaction decisions based on the user's intention, the results of context analysis, and the information obtained from knowledge reasoning; Step 7: Generate appropriate response content according to the above interaction decision; Step 8: The multi-modal output module outputs the generated response content to the user in a multi-modal form of speech, text, and image to complete the human-computer interaction process; Step 9: Collect the user's feedback on the interaction content and optimize the system performance.
7. An artificial intelligence-based interaction method according to claim 6, characterized in that: In step four, a time series analysis method is used to analyze the historical conversation records to mine the user's behavior patterns and intention tendencies.
8. An artificial intelligence-based interaction method according to claim 6, characterized in that: In step five, the path search and inference rules of the knowledge graph are utilized to conduct knowledge inference and expansion, providing richer information support for reply generation.
Citation Information
Cited By
Intelligent agent digital image interaction generation method based on multi-modal perception
CN121187453A
Multi-modal personalized content generation and interpretation method and device for interaction between user and expert agent, medium, program product and terminal
CN121902021A