Intelligent interaction method, device, and computer-readable storage medium
The intelligent interaction method addresses the challenge of adapting AI responses to user emotions by activating emotion models only when needed, enhancing user experience through emotionally relevant interactions.
Patent Information
- Application Number
- JP2025049265
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-01
- Filing Date
- 2025-03-25
- Publication Date
- 2025-10-14
AI Technical Summary
Existing AI systems struggle to provide emotional responses that adapt to the user's emotions, affecting user experience in interactions.
An intelligent interaction method that acquires emotion data, determines emotion types, and activates corresponding emotion models to generate matching response content, with models remaining dormant until needed to conserve resources and improve efficiency.
Enhances user experience by providing emotionally responsive content that adapts to the user's emotions, improving interaction quality and efficiency by minimizing unnecessary model loading.
Smart Images

Figure 2025156089000001_ABST
Abstract
Description
[Technical Field]
[0001] The present application relates to the field of artificial intelligence (AI), and in particular to an intelligent interaction method, device, and computer-readable storage medium. [Background technology]
[0002] With the development of AI technology, interactions between users and AI are becoming more and more frequent. Through semantic analysis technology, AI can analyze the dialogue dictated by users, understand what the users want to express, and respond appropriately. However, user interactions are usually emotional, and the responses provided by AI are difficult to adapt to the user's emotions, which affects the user experience. Summary of the Invention [Problem to be solved by the invention]
[0003] In view of the above, embodiments of the present application aim to provide an intelligent interaction method, device and computer-readable storage medium to solve the problem of how to provide emotional response content. [Means for solving the problem]
[0004] A first aspect of an embodiment of the present application provides an intelligent interaction method, including the steps of: acquiring dialogue content and first emotion data of a dialogue target in a first time period; determining a first emotion type based on the first emotion data; activating a first emotion model that corresponds to the first emotion type and is used to generate first response content that matches the first emotion type based on the dialogue content; and responding to the dialogue content based on the first response content.
[0005] In this embodiment, the emotion data of the interaction target is analyzed to identify the emotion type, and an emotion model corresponding to the emotion type is invoked to generate a response content that matches the interaction content of the interaction target, thereby providing an emotional response content that matches the emotion of the interaction target, thereby improving the user experience. In addition, various emotion models are pre-loaded, and the various emotion models remain dormant when not activated, saving system resources and power consumption. After identifying the emotion type, the emotion model corresponding to the emotion type is activated, eliminating the model loading process during the interaction process and improving response efficiency.
[0006] In one embodiment, the method further includes the steps of acquiring second emotion data of the interaction target in a second time period; determining a second emotion type based on the second emotion data; if the second emotion type is different from the first emotion type, activating a second emotion model that corresponds to the second emotion type and is used to generate second response content that matches the second emotion type based on the interaction content; and responding to the interaction content based on the second response content.
[0007] In this embodiment, the state of the emotion model includes an active state or a dormant state. During a certain period of time, one of the multiple emotion models is in an active state, and the other emotion models are all in a dormant state. After responding to the dialogue content, emotion data of the dialogue target is collected again, an emotion type is determined based on the new emotion data, and whether the emotion type has changed is determined to determine whether to adjust the emotion model. When the emotion type has changed, the emotion model corresponding to the new emotion type is switched to. When the emotion type has not changed, the original emotion model is maintained. This realizes emotional switching of the response content, making the response content more suited to the emotion of the dialogue target, thereby improving the user experience.
[0008] In another embodiment, the method further comprises deactivating the first emotion model if the second emotion type is different from the first emotion type.
[0009] In this embodiment, when the second emotion type is different from the first emotion type, it indicates that the first response content is incompatible with the emotion of the dialogue target, and the emotion model needs to be changed according to the change in the emotion of the dialogue target. The first emotion model is deactivated, thereby switching the state of the first emotion model to a dormant state.
[0010] In another embodiment, the method further comprises maintaining the first emotion model in an activated state if the second emotion type is the same as the first emotion type.
[0011] In this embodiment, when the second emotion type is the same as the first emotion type, it indicates that the first response content matches the emotion of the dialogue target, and the emotion model is not changed in the second time period.
[0012] In another embodiment, the method further includes adding emotion terms to the dialogue content based on a first corpus corresponding to the first emotion type before activating the first emotion model corresponding to the first emotion type.
[0013] In this embodiment, the first corpus is used to store emotion terms related to the first emotion type. By adding emotion terms to the dialogue content, the differences between various emotion types can be expanded, thereby improving the adaptability of the emotion model.
[0014] In another embodiment, the method further includes, before activating the first emotion model corresponding to the first emotion type, training a plurality of corresponding emotion models using a plurality of emotion training sets corresponding to the plurality of emotion types.
[0015] In this embodiment, various emotion models may be trained based on emotion training sets corresponding to various emotion types, and the emotion training sets may include text content corresponding to the emotion types.
[0016] A second aspect of an embodiment of the present application provides an intelligent interaction device as follows: The intelligent interaction device includes a memory, a processor, a sensor, a microphone, and a speaker, wherein the microphone is used to record dialogue content of a dialogue target, the sensor is used to collect emotion data of the dialogue target, the speaker is used to play back response content, the memory stores a plurality of emotion models, and the processor is configured to acquire dialogue content and first emotion data in a first time period, determine a first emotion type based on the first emotion data, activate a first emotion model corresponding to the first emotion type and used to generate first response content based on the dialogue content, and respond to the dialogue content based on the first response content.
[0017] In an embodiment, the processor is further configured to acquire second emotion data in a second time period; determine a second emotion type based on the second emotion data; if the second emotion type is different from the first emotion type, activate a second emotion model that corresponds to the second emotion type and is used to generate second response content that matches the second emotion type based on the dialogue content; and respond to the dialogue content based on the second response content.
[0018] In another embodiment, the device further comprises a display, the display being adapted to display the first response content.
[0019] A third aspect of an embodiment of the present application provides a computer-readable storage medium having stored thereon computer instructions that, when executed by a computer, perform an intelligent interaction method according to an embodiment of the present application.
[0020] As can be understood, the specific embodiments and beneficial effects of the intelligent interaction device provided by the second aspect of the embodiments of the present application and the computer-readable storage medium provided by the third aspect are all the same as the specific embodiments and beneficial effects of the intelligent interaction method provided by the first aspect, and will not be repeated here. [Brief explanation of the drawings]
[0021] FIG. 1 is a structural schematic diagram of an intelligent interaction device provided by an embodiment of the present application.
[0022] FIG. 2 is a flowchart of an intelligent interaction method provided by one embodiment of the present application. DETAILED DESCRIPTION OF THE INVENTION
[0023] In the examples of this application, "at least one" and "some" refer to one or more, and "plurality" refers to two or more. "A and / or B" describes a relationship between A and B, and indicates that three types of relationships exist. For example, A and / or B can represent the presence of A alone, the presence of A and B simultaneously, or the presence of B alone, where A and B may be singular or plural. The terms "first," "second," etc. in the specification, claims, and accompanying drawings of this application are used to distinguish between similar objects, and are not used to describe a specific order or context.
[0024] Furthermore, the methods disclosed in the embodiments of the present application or shown in the flowcharts include one or more steps for realizing the method, and the order of execution of multiple steps can be interchanged, or specific steps can be deleted, without departing from the scope of the claims.
[0025] FIG. 1 is a structural schematic diagram of an intelligent interaction device provided by an embodiment of the present application.
[0026] 1, the intelligent interaction device 100 includes a sensor 110, a microphone 120, a speaker 130, a display 140, a memory 150, and a processor 160. The processor 160 is electrically connected to the sensor 110, the microphone 120, the speaker 130, the display 140, and the memory 150.
[0027] The sensor 110 is used to collect emotional data of the interaction target. The emotional data may include physiological data and behavioral data. The physiological data may include heart rate, skin electrical conductivity, respiratory rate, and facial images, and the behavioral data may include limb movement images. Exemplarily, the sensor 110 may include multiple types of biosensors and is used to collect physiological data of the interaction target. For example, the sensor 110 may include a biosensor that collects heart rate, a biosensor that collects skin electrical conductivity, and a biosensor that collects respiratory rate. The sensor 110 may include a camera that is used to collect facial images and limb movement images of the interaction target.
[0028] The microphone 120 is used to record the dialogue content of the dialogue target. A user can speak through the microphone 120, and the voice signal is input to the microphone 120. The microphone 120 converts the voice signal into an electrical signal, thereby collecting the voice signal and realizing the recording function.
[0029] The speaker 130 is used to play back the response content and can convert an electrical signal into an audio signal, thereby realizing the audio playback function.
[0030] The display 140 is used to display the response content. In some embodiments, the display 140 can display the dialogue content recorded by the microphone 120. In other embodiments, the display 140 can display the dialogue content in text format entered by the dialogue target.
[0031] The display 140 includes a display panel, which may employ a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, a quantum dot light-emitting diode (QLED), or the like.
[0032] The memory 150 is used to store several emotion models. The emotion models are used to generate response content that matches an emotion type based on the dialogue content. Exemplarily, the several emotion models may include an excited model, a calm model, a depressed model, and a nervous model. The various emotion models may be trained based on emotion training sets corresponding to various emotion types, and the emotion training sets may include text content corresponding to the emotion types.
[0033] In some embodiments, the memory 150 can further store corpora corresponding to various emotion types, each corpus being used to store emotion terms associated with the emotion type. The emotion terms can include particles, adjectives, etc. to express the emotion type. Adding emotion terms to the dialogue content can amplify the differences between various emotion types, thereby improving the fit of the emotion model.
[0034] The memory 150 may include an external memory interface and an internal memory. The external memory interface is used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the intelligent interaction device 100. The external memory card communicates with the processor 160 through the external memory interface to achieve a data storage function. The internal memory is used to store computer-executable program code, which includes instructions. The internal memory may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application program required for a function (e.g., audio playback function). The data storage area may store data (e.g., audio data) generated during use of the intelligent interaction device 100. The internal memory may include a high-speed random access memory and may also include a non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or universal flash storage (UFS). The processor 160 performs various functional applications and data processing of the intelligent interaction device 100 by executing instructions stored in its internal memory and / or by executing instructions stored in memory installed in the processor 160.
[0035] The processor 160 is used to obtain dialogue content and emotion data of a dialogue target, determine an emotion type based on the emotion data, activate an emotion model corresponding to the emotion type, and respond to the dialogue content based on the response content.
[0036] In some embodiments, after responding to further dialogue content, processor 160 can collect emotion data of the dialogue target again, determine the emotion type based on the new emotion data, and determine whether to adjust the emotion model depending on whether the emotion type has changed. When the emotion type has changed, the emotion model corresponding to the new emotion type is switched to. When the emotion type has not changed, the original emotion model is maintained. This realizes emotional switching of the response content, making the response content more suited to the emotion of the dialogue target, thereby improving the user experience.
[0037] The processor 160 may include one or more processing units. For example, the processor 160 may include, but is not limited to, an application processor (AP), a modulation / demodulation processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and a neural-network processing unit (NPU). Among these, different processing units may be independent devices or may be integrated into one or more processors.
[0038] The processor 160 may also include memory for storing instructions and data. In some embodiments, the memory in the processor 160 is a high-speed cache memory. The memory may store instructions or data that the processor 160 recently used or cycled through. If the processor 160 needs to use the instructions or data again, it can be retrieved directly from the memory.
[0039] As can be understood, the structures shown by the examples of the present application do not constitute specific limitations on the intelligent interaction device 100. In other examples, the intelligent interaction device 100 may include more or fewer components than shown, or may combine some components, or may separate some components, or may employ a different component arrangement.
[0040] The intelligent interaction device provided by the embodiments of the present application can be applied in various scenarios. For example, the intelligent interaction device can be applied in an intelligent customer service system, for example, by integrating the intelligent interaction device into a customer service robot, which can obtain and analyze emotion data provided by a customer through text or voice, identify the emotion type, and then invoke an emotion model corresponding to the emotion type to generate response content that matches the customer's interaction content, thereby providing emotional response content that matches the customer's emotion and improving the user experience.
[0041] In addition, for example, the intelligent interaction device can be applied in a social media sentiment analysis tool. For example, the intelligent interaction device can be integrated into a computer, and the computer can obtain and analyze sentiment data of text content from social media (or other data sources), identify the sentiment type, and then invoke the sentiment model corresponding to the sentiment type to generate response content that matches the text content of the social media, thereby helping the administrator grasp the sentiment trends of the general public on social media and actively adjusting the sentiment model to guide the sentiment trends of the general public, thereby timely intervening in inappropriate sentiment trends on social media and maintaining the public opinion atmosphere on social media.
[0042] FIG. 2 is a flowchart of an intelligent interaction method provided by one embodiment of the present application.
[0043] The intelligent interaction method is applied to an intelligent interaction device, and can be applied to, for example, the intelligent interaction device 100 shown in Figure 1. As shown in Figure 2, the intelligent interaction method includes the following steps:
[0044] In S101, the dialogue target acquires the dialogue content and first emotion data for a first time period.
[0045] In this embodiment, the intelligent interaction device uses a microphone to record the conversation content of the interaction target in a first time period, and then employs a semantic recognition algorithm to recognize the intent of the conversation content, thereby accurately understanding the conversation content; and uses a sensor to collect first emotion data of the interaction target in the first time period, such as the heart rate, skin conductivity, respiratory rate, facial images or limb movement images of the interaction target in the first time period.
[0046] Here, the first time period can be set as needed, for example, 3 minutes or 5 minutes.
[0047] In some embodiments, the intelligent interaction device may include an input module, which may receive textual interaction content input by the interaction target during the first time period. The input module may include a touch screen.
[0048] In S102, a first emotion type is determined based on the first emotion data.
[0049] In this embodiment, the intelligent interaction device employs methods such as time domain / frequency domain analysis, nonlinear fitting, etc. to extract feature parameters from the emotion data, and then classifies the data based on the feature parameters to determine the emotion type corresponding to the emotion data. The emotion type can be set as needed, and for example, the emotion types can include excited emotion, calm emotion, depressed emotion, and nervous emotion.
[0050] Here, the emotion data corresponds to an emotion type. Taking heart rate as an example, changes in the heart rate band are related to sympathetic nervous activity, and the ratio of low frequency (LF) to high frequency (HF) frequencies can measure the strength of sympathetic nervous activity, while the HF frequency can measure the strength of parasympathetic nervous activity. An increase in the ratio of LF frequency to HF frequency indicates increased sympathetic nervous activity, and an increase in HF frequency indicates increased parasympathetic nervous activity. For example, the LF frequency may range from 0.04 to 0.15 Hz, and the HF frequency may range from 0.15 to 0.4 Hz. When the LF frequency increases and the HF frequency decreases, the ratio of LF frequency to HF frequency increases, indicating increased sympathetic nervous activity and decreased parasympathetic nervous activity, and the corresponding emotion type is excitement. When the LF frequency decreases and the HF frequency increases, the ratio of the LF frequency to the HF frequency decreases, indicating a decrease in sympathetic nervous activity and an increase in parasympathetic nervous activity, and the corresponding emotional type is a calm emotion. When the LF frequency decreases and the HF frequency decreases, the ratio of the LF frequency to the HF frequency may change, indicating a change in sympathetic nervous activity and a decrease in parasympathetic nervous activity, and the corresponding emotional type is a depressed emotion. When the LF frequency decreases and the HF frequency increases, the ratio of the LF frequency to the HF frequency decreases, indicating a decrease in sympathetic nervous activity and an increase in parasympathetic nervous activity, and the corresponding emotional type is a tense emotion.
[0051] As can be seen, similar to heart rate, correspondence between emotion data and emotion types can also be established based on other physiological and behavioral data.
[0052] In S103, emotion terms are added to the dialogue content based on the first corpus corresponding to the first emotion type.
[0053] In this embodiment, the intelligent interaction device can access corpora corresponding to various emotion types, and each corpus is used to store emotion terms related to the emotion type. The emotion terms may include particles and adjectives to express the emotion type. For example, the corpus corresponding to excited emotion stores emotion terms related to excited emotion, the corpus corresponding to neutral emotion stores emotion terms related to neutral emotion, the corpus corresponding to depressed emotion stores emotion terms related to depressed emotion, and the corpus corresponding to nervous emotion stores emotion terms related to nervous emotion.
[0054] By adding emotion terms to the dialogue content, the intelligent interaction device can expand the distinction between various emotion types, thereby improving the fitness of the emotion model.
[0055] As can be appreciated, the corpora corresponding to the various emotion types can be stored locally or in the cloud.
[0056] In S104, a first emotion model corresponding to the first emotion type is activated.
[0057] Here, the first emotion model is used to generate a first response content that matches the first emotion type based on the dialogue content.
[0058] In this embodiment, the intelligent interaction device stores several emotion models. The emotion models are used to generate response content that matches an emotion type based on the dialogue content. For example, the several emotion models may include an excited model, a calm model, a depressed model, and a nervous model. The excited model is used to generate response content that matches an excited emotion based on the dialogue content, the calm model is used to generate response content that matches a neutral emotion based on the dialogue content, the depressed model is used to generate response content that matches a depressed emotion based on the dialogue content, and the nervous model is used to generate response content that matches a nervous emotion based on the dialogue content.
[0059] As can be understood, various emotion models can be configured as needed, and the emotion model may be a random forest model, a neural network model, a deep learning model, etc. The various emotion models can be trained based on emotion training sets corresponding to various emotion types, and the emotion training sets can include text content corresponding to the emotion types.
[0060] The intelligent interaction device automatically loads various emotion models when powered on, and the various emotion models remain dormant when not activated, thereby saving system resources and power consumption. After identifying the emotion type, the intelligent interaction device triggers an activation command, and the activation command is used to activate the emotion model corresponding to the emotion type, thereby omitting the model loading process during the interaction process and improving response efficiency.
[0061] In S105, a response is made to the dialogue content based on the first response content.
[0062] In this embodiment, after the emotion model is activated, a response content that matches the emotion type can be generated based on the dialogue content, and the intelligent interaction device can play the response content using a speaker and display the response content on a display.
[0063] In S106, second emotion data of the dialogue target in a second time period is acquired.
[0064] Here, the second time period is after the first time period, and the second time period can be set as needed, for example, the second time period is 1 minute or 2 minutes.
[0065] In this embodiment, after responding to the dialogue content, the intelligent interaction device can again collect emotion data of the interaction target, thereby determining whether the response content is suitable for the emotion of the interaction target.
[0066] At S107, a second emotion type is determined based on the second emotion data.
[0067] As can be understood, the specific implementation of step S107 is the same as step S102, and will not be repeated here.
[0068] In S108, it is determined whether the second emotion type is the same as the first emotion type.
[0069] If they are the same, step S109 is executed, and if they are different, steps S110 to S112 are executed.
[0070] In this embodiment, if the second emotion type is the same as the first emotion type, it indicates that the first response content is adapted to the emotion of the dialogue target. If the second emotion type is different from the first emotion type, it indicates that the first response content is not adapted to the emotion of the dialogue target, and it is necessary to adaptively change the emotion model in accordance with changes in the emotion of the dialogue target.
[0071] In S109, the first emotion model is maintained in an activated state based on the fact that the second emotion type is the same as the first emotion type.
[0072] In this embodiment, the state of the emotion model includes an activated state or a dormant state. During a certain time period, one of the emotion models is in an activated state, and the other emotion models are all in a dormant state. When the second emotion type is the same as the first emotion type, the emotion model is not changed during the second time period.
[0073] At S110, the first emotion model is deactivated based on the second emotion type being different from the first emotion type.
[0074] In this embodiment, when the second emotion type is different from the first emotion type, a deactivation command is triggered, and the deactivation command is used to deactivate the emotion model corresponding to the emotion type, and switch the state of the emotion model to a dormant state.
[0075] In S111, a second emotion model that matches the second emotion type is activated.
[0076] Here, the second emotion model is used to generate second response content that matches the second emotion type based on the dialogue content.
[0077] As can be understood, the specific implementation of step S111 is the same as step S104, and will not be repeated here.
[0078] In S112, a response is made to the dialogue content based on the second response content.
[0079] In this embodiment, after the new emotion model is activated, a response content that matches the new emotion type can be generated based on the dialogue content, thereby realizing emotional switching of the response content and making the response content more suited to the emotion of the dialogue target, thereby improving the user experience.
[0080] An embodiment of the present application provides a computer-readable storage medium having stored thereon computer instructions that, when executed by a computer, cause the computer to perform the intelligent interaction method of the embodiment of the present application.
[0081] Computer-readable storage media includes volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storage of information such as computer-readable instructions, data structures, program modules or other data, including, but not limited to, Random Access Memory (RAM), Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Flash memory or other memory, Compact Disc Read-Only Memory (CD-ROM), Digital Versatile Disc (DVD) or other optical disk storage, magnetic tape case, magnetic tape, magnetic disk storage or other magnetic storage device, or any other medium that can be used to store the desired information and that can be accessed by a computer.
[0082] Although the embodiments of the present application have been described in detail above with reference to the drawings, the present application is not limited to the above embodiments, and various modifications may be made within the scope of knowledge possessed by a person skilled in the art without departing from the spirit of the present application.
Claims
1. 1. A method of intelligent interaction, comprising: acquiring a dialogue content and first emotion data of a dialogue target in a first time period; determining a first emotion type based on the first emotion data; activating a first emotion model corresponding to the first emotion type, wherein the first emotion model is used to generate a first response content that matches the first emotion type based on the dialogue content; and responding to the dialogue content based on the first response content.
2. 10. The method of claim 1, acquiring second emotion data of the interaction target in a second time period; determining a second emotion type based on the second emotion data; If the second emotion type is different from the first emotion type, activating a second emotion model corresponding to the second emotion type, the second emotion model being used to generate second response content that matches the second emotion type based on the dialogue content; and responding to the dialogue content based on the second response content.
3. 3. The method of claim 2, 10. The intelligent interaction method, further comprising the step of deactivating the first emotion model if the second emotion type is different from the first emotion type.
4. 3. The method of claim 2, 10. The intelligent interaction method, further comprising: if the second emotion type is the same as the first emotion type, maintaining the first emotion model in an activated state.
5. 5. The method according to claim 1, wherein 10. The intelligent interaction method according to claim 9, further comprising: adding emotion terms to the dialogue content based on a first corpus corresponding to the first emotion type before activating the first emotion model corresponding to the first emotion type.
6. 5. The method according to claim 1, wherein Before activating the first emotion model corresponding to the first emotion type, the intelligent interaction method further comprises training a plurality of corresponding emotion models using a plurality of emotion training sets corresponding to a plurality of emotion types.
7. 1. An intelligent interaction device, comprising: Includes memory, processor, sensors, microphone and speaker, The microphone is used to record the content of a conversation between two people, the sensor is used to collect emotion data of the interaction target; the speaker is used to play back the response content; the memory stores a plurality of emotion models; The processor: the dialogue target acquires the dialogue content and first emotion data in a first time period; determining a first emotion type based on the first emotion data; activating a first emotion model corresponding to the first emotion type, the first emotion model being used to generate a first response content based on the dialogue content; and responding to the dialogue content based on the first response content.
8. 8. The device of claim 7, The processor further The dialogue target acquires second emotion data in a second time period; determining a second emotion type based on the second emotion data; If the second emotion type is different from the first emotion type, activating a second emotion model corresponding to the second emotion type, and the second emotion model is used to generate second response content conforming to the second emotion type based on the dialogue content; and responding to the dialogue content based on the second response content.
9. 9. The device according to claim 7 or 8, An intelligent interaction device further comprising a display, the display being used to display the first response content.
10. 7. A computer readable storage medium having stored thereon computer instructions which, when executed by a computer, cause the computer to perform the intelligent interaction method of any one of claims 1 to 6.
Citation Information
Patent Citations
Information publication device
JP1997081632A
Device and method for interactive processing and recording medium
JP2001215993A
Agent interface system
JP2005222331A
Conversation processing system and program
JP2014219594A
Dialogue system and program
JP2017117090A