Interaction method, interaction device, electronic equipment and computer readable storage medium

By outputting corresponding audio and multimedia resources based on the emotional characteristics of the agent's reply text, the problem of traditional voice interaction is solved, and the user experience and the agent's output effect are improved.

CN120045265APending Publication Date: 2025-05-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510122152.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional voice interaction based on TTS audio is relatively monotonous in emotional expression, resulting in unsatisfactory user experience.

Method used

By receiving user input, a reply text of the agent is generated, and based on the emotional characteristics of the text, audio corresponding to the emotional characteristics and multimedia resources of the agent corresponding to the emotional characteristics are output.

Benefits of technology

It improves the user's visual and auditory perception, provides an immersive experience, and enhances the output effect and appeal of the agent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045265A_ABST
    Figure CN120045265A_ABST
Patent Text Reader

Abstract

The invention relates to an interaction method and a related device. The interaction method comprises the following steps: receiving input of a user in an interaction interface with an intelligent agent; generating a reply text of the agent based on the input; according to the emotional characteristics of the reply text, outputting audio of the reply text corresponding to the emotional characteristics in the interactive interface; and according to the emotion features of the reply text, displaying the multimedia resources of the intelligent agent corresponding to the emotion features in the interactive interface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and more particularly to an interaction method, an interaction device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] With the development of artificial intelligence (AI) technology and large language models, many software applications that enable interaction with AI virtual characters have emerged. These AI virtual characters with distinct personalities and diverse backgrounds can allow users to immerse themselves in conversations through various interaction methods, answer users' questions, or provide emotional value. In related interaction software, in addition to displaying response texts, the question-and-answer function may also provide response voices. Generally, voice interaction uses text-to-speech (TTS) technology to synthesize audio, which is played during the process of replying to users. However, traditional voice interaction based on TTS audio is often relatively monotonous and rigid, resulting in an unsatisfactory user experience. Summary of the Invention

[0003] According to some embodiments of the present disclosure, there is provided an interaction method, including: receiving an input from a user in an interaction interface with an agent; generating a response text of the agent based on the input; outputting, in the interaction interface, an audio of the response text corresponding to an emotional feature according to the emotional feature of the response text; and displaying, in the interaction interface, a multimedia resource of the agent corresponding to the emotional feature according to the emotional feature of the response text.

[0004] According to some other embodiments of the present disclosure, there is provided an interaction device, including: an input module configured to receive an input from a user in an interaction interface with an agent; a response module configured to generate a response text of the agent based on the input; an audio module configured to output, in the interaction interface, an audio of the response text corresponding to an emotional feature according to the emotional feature of the response text; and a multimedia module configured to display, in the interaction interface, a multimedia resource of the agent corresponding to the emotional feature according to the emotional feature of the response text.

[0005] According to some embodiments of the present disclosure, there is provided an electronic device, including: a memory; and a processor coupled to the memory, the processor being configured to execute the method according to any one of the embodiments of the present disclosure based on instructions stored in the memory.

[0006] According to some embodiments of the present disclosure, there is provided a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to execute the method according to any one of the embodiments of the present disclosure.

[0007] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Embodiments of the present disclosure will be described below with reference to the accompanying drawings. It should be understood that the drawings in the following description only relate to some embodiments of the present disclosure and do not constitute a limitation on the present disclosure. In the drawings:

[0009] Figure 1 A schematic diagram of an interaction interface between a user and an agent in the related art is shown;

[0010] Figure 2 A flowchart of a method for interacting between a user and an agent according to some embodiments of the present disclosure is shown;

[0011] Figure 3A and Figure 3B A schematic diagram of an interaction interface between a user and an agent according to some embodiments of the present disclosure is shown;

[0012] Figure 4A and Figure 4B A schematic diagram of an interaction interface between a user and an agent according to some other embodiments of the present disclosure is shown;

[0013] Figures 5A to 5C A schematic diagram of an interaction interface between a user and an agent according to some further embodiments of the present disclosure is shown;

[0014] Figure 6A and Figure 6B A schematic diagram of an interaction interface between a user and an agent according to some still further embodiments of the present disclosure is shown;

[0015] Figure 7 A schematic diagram of a method for generating audio and multimedia resources according to some embodiments of the present disclosure is shown;

[0016] Figure 8 A schematic block diagram of an interaction device between a user and an agent according to some embodiments of the present disclosure is shown;

[0017] Figure 9 A block diagram of an electronic device according to some embodiments of the present disclosure is shown;

[0018] Figure 10 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.

[0019] It should be understood that, for ease of description, the dimensions of the various parts shown in the drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar components. Therefore, once an item is defined in one drawing, it may not be further discussed in subsequent drawings. Detailed implementation mode

[0020] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. It should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.

[0021] It should be understood that the various steps recorded in the method implementation mode of the present disclosure can be executed in different orders and / or in parallel. In addition, the method implementation mode may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard. Unless otherwise specifically stated, the relative arrangements, numerical expressions, and numerical values of the components and steps set forth in these embodiments should be construed as merely exemplary and do not limit the scope of the present disclosure.

[0022] The term "including" and its variants used in the present disclosure mean open terms that include at least the subsequent elements / features but do not exclude other elements / features, that is, "including but not limited to". The term "based on" means "at least partially based on".

[0023] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and are not used to limit the order of the functions performed by these devices, modules, or units or their interdependent relationships. Unless otherwise specified, the concepts such as "first" and "second" are not intended to imply that the objects so described must be in a given order in terms of time, space, ranking, or any other way.

[0024] It should be noted that the modification of "one" and "multiple" mentioned in the present disclosure is illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly stated in the context, it should be understood as "one or more".

[0025] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and do not limit the scope of these messages or information.

[0026] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0027] The embodiments of the present disclosure are described in detail below in conjunction with the accompanying drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. In addition, in one or more embodiments, specific features, structures or characteristics can be combined in any suitable manner that will be clear from the present disclosure by a person of ordinary skill in the art.

[0028] It should be understood that the present disclosure does not limit how to obtain the image to be applied / processed. In some embodiments of the present disclosure, it can be obtained from a storage device, such as an internal memory or an external storage device. In other embodiments of the present disclosure, a photographic component can be mobilized to shoot. It should be noted that the acquired image can be a captured image or a frame of an image in a captured video, and is not particularly limited to this.

[0029] In the context of the present disclosure, an image may refer to any of a variety of images, such as a color image, a grayscale image, etc. It should be noted that in the context of the present specification, the type of image is not specifically limited. In addition, the image may be any appropriate image, such as an original image obtained by a camera device, or an image that has been subjected to specific processing, such as preliminary filtering, anti-aliasing, color adjustment, contrast adjustment, normalization, etc. It should be noted that the preprocessing operation may also include other types of preprocessing operations known in the art, which will not be described in detail here.

[0030] In the related art, the responses of the intelligent agent interaction software to the user input usually include text display and voice output. These interaction software generate responses based on the Large Language Model (LLM) and use TTS technology to convert the response content into regular voice output. For details, please refer to Figure 1, which shows a schematic diagram of the interaction interface between a user and an agent in the related art. In the interaction interface 1 between the user and the agent, the user inputs to the agent (such as 11, 13), and the agent generates a response text in reply to the user's input (such as 12, 14). While displaying the response text in the interaction interface 1, it outputs the text content in voice, such as response text 12 corresponding to audio 120 and response text 14 corresponding to audio 140, where audio 140 can be synthesized using TTS technology. It can be understood that the icons of audio 120 and 140 are only used to describe various output forms of the agent's response, and are not used to define the visual presentation effect of the interaction interface 1.

[0031] Similar Figure 1 For the interaction interface shown, the synthesized audio in the related art can only monotonously convey the literal content of the response text, lacking the emotional value of creating a sense of scene and experience for the user. To solve this problem, the present disclosure provides an interaction method, in which audio and pictures corresponding to the response text can be output in an emotionally rich manner in the interaction interface between the user and the agent, thereby improving the user's multi-faceted perceptions such as vision and hearing, and enabling the user to have an immersive experience.

[0032] Specifically, Figure 2 shows a flowchart of an interaction method between a user and an agent according to some embodiments of the present disclosure.

[0033] As Figure 2 shown, in step S201, receive the user's input in the interaction interface with the agent; in step S202, generate the agent's response text based on the input; in step S203, according to the emotional characteristics of the response text, output the audio of the response text corresponding to the emotional characteristics in the interaction interface; and in step S204, according to the emotional characteristics of the response text, display the multimedia resources of the agent corresponding to the emotional characteristics in the interaction interface.

[0034] The interaction method of this embodiment can be executed on the client side or partially on the server side.

[0035] Generally, the emotional characteristics in the present disclosure are used to describe the emotions corresponding to the text content, such as emotional labels representing the user's emotions and feelings like happy, sad, angry, surprised, etc., or scene settings such as weather, location, activity, etc. that provide different experiences for the user, so as to summarize and refer to the corresponding characteristics of the text.

[0036] The audio of the response text is output to the user by a sound - emitting device such as a microphone, and includes, for example, the voice of the virtual character of the agent in the current interaction interface, as well as the scene sound of the scene involved in the current interaction, etc. The multimedia resources of the agent, for example, include the character image of the aforementioned virtual character, as well as the background elements of the scene involved in the current interaction, etc. The forms of the multimedia resources include but are not limited to static pictures, dynamic images, short videos, slides, web content, virtual reality / augmented reality content, social media content, animated films, interactive elements, etc. One or a combination of the above forms can be set at the designated position of the interaction interface according to actual needs.

[0037] The optimized interaction method provided by the present disclosure includes an interaction interface with rich presentation forms, which can refer to Figures 3A to 6B the specific embodiments shown therein. Specifically, the interaction method of the present disclosure includes using a machine - learning model to obtain the emotional characteristics of the response text according to the response text of the agent and the historical conversation record. And first refer to Figure 3A and Figure 3B , which respectively show schematic diagrams of the interaction interfaces between the user and the agent according to some embodiments of the present disclosure. Similar to the interaction interface 1 in Figure 1 , Figure 3A the interaction interface 31 in Figure 3A includes the user's input 311 and the agent's response text 312 thereto; in addition, the interaction interface 31 further includes multimedia resources 310 as the interaction text illustration or background image of the current conversation. It can be understood that the multimedia resources 310 can be displayed in various forms,

[0038] Figure 3A only an exemplary facial image with an expression is shown as a simple example. As shown in Figure 3A , in response to the user's input 311 "Hello" in the interaction interface 31, the agent generates the response text 312 "Hello". Further, using the machine - learning model, according to the response text 312 and the aforementioned historical conversation record, the emotional characteristics 314 of the response text 311 are obtained. Accordingly, the emotional characteristics can include at least one of the semantic characteristics based on the response text and the scene characteristics based on the historical conversation record, where the semantic characteristics are usually used to describe the virtual character of the agent in the current interface, and the scene characteristics are usually used to describe the scene or environment where the virtual character is located. In other words, Figure 3AThe emotional feature 314 "happy" is shown in the text box of the reply text 3102 in the form of bracketed text. Here, it is similar to the audio 313 presented in the form of an icon, which is only used to describe the output form of the reply text in the interaction interface 31, rather than to limit the visual presentation effect of the interaction interface.

[0039] In some embodiments, according to the obtained emotional features, the role voice of the corresponding agent is output in the interaction interface. In particular, considering that the emotional features include semantic features based on the reply text, the role voice corresponding to the semantic features can be output in the interaction interface according to the semantic features. Specifically, the emotional feature 314 mainly includes semantic features, that is, the virtual character represents "happy". In the interaction interface 31, the audio 313 corresponding to the reply text 312 of "happy" is output. The audio 313 conforms to the acoustic characteristics of the emotional label "happy" in terms of timbre, such as being reflected as high pitch, strong loudness, many high-frequency components, many stress changes, brisk breathing, rising intonation, etc. Additionally, according to the obtained emotional feature 314, in the interaction interface 31, the multimedia resource 310 corresponding to the emotional feature "happy" is displayed. The multimedia resource 310 conforms to the image characteristics of the emotional label "happy" as a facial image, such as being reflected as open lips, dilated pupils, raised eyebrows, blushing cheeks, relaxed facial muscles, etc.

[0040] Alternatively, in Figure 3B the interaction interface 32 shows another presentation form of the emotional feature 324. In response to the user input 321 "I think you are wrong", the agent generates the reply text 322 "I don't agree", and obtains the emotional feature 324 "angry" of this text according to the reply text 322 and the historical conversation record. Thus, in the interaction interface 32, according to the obtained emotional features, the audio 323 of the reply text 322 corresponding to "angry" is output, which is reflected as conforming to the acoustic characteristics of the emotional label "angry" in terms of timbre, such as unstable pitch, sudden increase in loudness, accelerated speech rate, rapid breathing, downward intonation, etc.; and the multimedia resource 320 corresponding to the emotional feature "angry" is displayed, which conforms to the image characteristics of the emotional label "angry" as a facial image, such as being reflected as tightened eyebrows, closed lips, fixed sight, rapid and strong movement of facial muscles, etc.

[0041] In some embodiments, considering that the emotional features include scene features based on historical conversation records, the audio of the response text also includes scene features, and according to the scene features, in the interaction interface, the scene features corresponding to the scene features are output. For example, if the scene in the response text involves "New Year", the scene features of the New Year are obtained based on the historical conversation records, and accordingly, in addition to the character voice, the scene sounds corresponding to the New Year are output, including but not limited to acoustic elements such as fireworks and festival music that create the scene atmosphere, as a supplement and foil to the aforementioned character voice.

[0042] Reference Figure 4A and Figure 4B , which shows a schematic diagram of the interaction interface between a user and an agent according to some other embodiments of the present disclosure. In Figure 4A In a non-limiting embodiment, in response to the user's input 411 "What's the weather like today" in the interaction interface 41, the agent generates a response text 412 "It's sunny today"; according to the response text 412 and the historical conversation record, the emotional feature 414 of the response text 412 is obtained, which is described in words as "looking at the sun", and the scene feature included is "sunny" based on the historical conversation record. Then, according to the emotional feature 414, a multimedia resource 410 is displayed in the interaction interface 41, which includes background elements such as the sun for representing "sunny".

[0043] Correspondingly, in Figure 4B , in response to the user's input 421 "Is it snowing today", the agent generates a response text 422 "Yes, it is snowing"; based on this, the emotional feature 424 of the response text 422 is obtained, which is described in words as "looking at the flying snow", and the scene feature included is "snowing" based on the historical conversation record. Thus, according to the emotional feature 424, a multimedia resource 420 is presented in the interaction interface 42, which includes background elements such as snowflakes for representing "snowing".

[0044] In some embodiments, in response to a change in the emotional feature, at least one of the character image and the background element is changed in the interaction interface. In other words, at least one of the multimedia resources is changed in the interaction interface. Figure 3A and Figure 3B The multimedia resources 310 and 320 shown in are the character image or a part of the character image. If the interaction between the user and the agent changes from Figure 3A to Figure 3B , then in response to a change in the emotional feature obtained by the agent, the character image of the agent is changed in the interaction interface, such as changing from a "happy" facial image to an "angry" facial image. Similarly, Figure 4A and Figure 4B The multimedia resources 410 and 420 shown in contain background elements. If the interaction changes from Figure 4ASwitch to Figure 4B , in response to the change of the emotional feature, the background elements of the dialogue are changed in the interaction interface, such as changing the scene from "sunny day" to "snowing".

[0045] In particular, in some embodiments, based on the emotional feature including at least one of the semantic feature of the reply text and the scene feature of the historical dialogue record, in response to the change of the semantic feature, at least one of the label and posture of the role is changed in the interaction interface; and / or, in response to the change of the scene feature, at least one of the color scheme of the background element and the scene is changed in the interaction interface.

[0046] As mentioned above, if the user's dialogue with the agent switches from Figure 3A Switch to Figure 3B , in response to the change of the semantic feature, it is switched from the multimedia resource 310 to the multimedia resource 320 in the interaction interface 31, that is, the expression of the character image is changed. It can be understood that Figure 3A and Figure 3B The multimedia resources shown in are only facial expression images. In fact, when the multimedia resource displays the character image, it can also be presented as the full body or half body image of the virtual character. Then, as the obtained emotional feature changes, the change of the character image is reflected in the change of the posture of the virtual character, such as changing from standing upright to standing sideways or turning the back to the user, etc.

[0047] Additionally or alternatively, if the user's dialogue with the agent switches from Figure 4A Switches to Figure 4B , in response to the change of the scene feature, it is switched from the multimedia resource 410 to the multimedia resource 420 in the interaction interface 41, that is, the scene in the background element is changed. It can be understood that Figure 4A and Figure 4B The multimedia resources shown in are shown as changing from the sun on a sunny day to snowflakes in the snow. In fact, the multimedia resource can also be presented in other changing ways, such as without displaying specific elements, and changing the color scheme of these elements while keeping the content of the background elements basically unchanged, such as changing the original warm color scheme to a white-based color scheme, etc.

[0048] In some embodiments, if the sentences of the response text are long or there are many clauses, resulting in the response text in a single interaction including multiple text segments corresponding one-to-one to multiple emotional features, the output audio correspondingly includes multiple audio segments corresponding to the emotional features of the multiple text segments. The response text is arranged in word order such that the multiple text segments appear on the interaction interface in sequence, and then the multiple audio segments are output on the interaction interface in sequence, aligned with the word order of the text on the time axis. Additionally, the multimedia resources in the interaction interface also cooperate with the emotional changes of the response text, and according to the emotional features of the multiple text segments, display the multimedia resources corresponding to different emotional features in sequence and variably according to the word order.

[0049] For details, please refer to Figures 5A to 5C , which shows a schematic diagram of the interaction interface between a user and an agent according to some further embodiments of the present disclosure. Exemplarily, in an interaction between a user and an agent, in response to the user's input 51 "I've been under a lot of work pressure lately. I'm feeling listless and don't know what to do", the agent generates a response text 52: "(Sympathy) I'm sorry to hear that you're going through work pressure. This feeling is really frustrating. (Understanding) Maybe you can try some relaxation activities, such as taking a walk or listening to music, which might help you regain some energy. (Determination) But I'm sure you can solve the difficulties you encounter. Give yourself some time!" In word order, the response text 52 includes multiple emotional features 54, which are "Sympathy" (54-1), "Understanding" (54-2), and "Determination" (54-3) in sequence, and accordingly, the response text 52 is divided into three corresponding text segments, namely text segment 52-1 "I'm sorry to hear that you're going through work pressure. This feeling is really frustrating", text segment 52-2 "Maybe you can try some relaxation activities, such as taking a walk or listening to music, which might help you regain some energy", and text segment 52-3 "But I'm sure you can solve the difficulties you encounter. Give yourself some time". The output audio 53 correspondingly includes multiple audio segments corresponding to each text segment. At the same time, the interaction interface 5 also displays the multimedia resources 50 corresponding to the multiple emotional features 54 in sequence according to the word order.

[0050] In other words, in the interaction interface 5, when the reply text 52-1 is displayed, the audio clip 53-1 with the output performance of "sympathy" is output, and the multimedia resource 50-1 showing "sympathy" is displayed; when the reply text 52-2 is displayed, the audio clip 53-2 with the emotional feature 54-2 of "understanding" is output, and the multimedia resource 50-2 showing "understanding" is displayed; and when the reply text 52-3 is displayed, the audio clip 53-3 with the emotional feature 54-3 of "firmness" is output, and the multimedia resource 50-3 showing "firmness" is displayed. The switching of the multimedia resources and the connection time points of the audio clips both depend on the time when the emotional features in the reply text change, and the three are consistent in the time axis nodes and order, so as to provide the user with multiple perceptions such as sound and picture in real time, making the user have an immersive experience of the reply text.

[0051] Furthermore, Figure 6A and Figure 6B FIG. shows a schematic diagram of an interaction interface between a user and an agent according to still further embodiments of the present disclosure. In a non-limiting embodiment, in an interaction between a user and an agent, in response to the user's input 61 "Another year is about to pass. Time really flies, which makes people inevitably feel lost", the agent generates a reply text 62: "(Understanding) The end of the year always makes people feel that time flies by very fast, which is really regrettable. (Looking forward to the New Year) However, imagine that when the New Year's bell rings, the fireworks are blooming around, and everyone is laughing and having a good time. It must be very happy!" Among them, the reply text 62 includes multiple emotional features 64, namely "understanding" (64-1), "looking forward to the New Year" (64-2), and two text segments 62-1 "The end of the year always makes people feel that time flies by very fast, which is really regrettable" and 62-2 "However, imagine that when the New Year's bell rings, the fireworks are blooming around, and everyone is laughing and having a good time. It must be very happy" divided accordingly. The output audio 63 correspondingly includes multiple audio clips corresponding to each text segment, and each audio clip includes the corresponding character voice and scene sound; at the same time, the interaction interface 6 also sequentially displays the multimedia resources 60 corresponding to each emotional feature 64, and each displayed multimedia resource 60 includes the corresponding character image and background elements.

[0052] Specifically, the emotional feature 64-1 is the semantic feature "understanding" based on the response text, while the emotional feature 64-2 includes both the semantic feature "longing" and the scene feature "New Year". Thus, when the text segment 62-1 is displayed on the interaction interface 6, the output audio segment 63-1 is the voice of the character expressing "understanding", and the displayed multimedia resource 60-1 is the character image expressing "understanding"; subsequently, when the text segment 62-2 is displayed on the interaction interface 6, the output audio segment 63-2 includes the voice of the character expressing "longing" and the scene sound expressing "New Year", and the displayed multimedia resource 60-2 includes the character image expressing "longing" and the background elements expressing "New Year", where specifically, compared with the interaction interface 6 displaying the text segment 62-1, the expression of the virtual character and the scene in the dialogue background are changed. Thus, the interaction interface can provide rich visual, auditory and other multi-form interaction experiences, thereby optimizing the output effect of the agent and enhancing the appeal.

[0053] Further, please refer to Figure 7 , which shows a schematic diagram of a method for generating audio and multimedia resources according to some embodiments of the present disclosure. As described above, the interaction between the user and the agent is mainly based on a machine learning model such as an LLM to generate a response text for the user input. Specifically, an emotion set 790 is preset for the response text, which includes various emotion labels that the agent needs to use, such as happy, sad, angry, surprised, etc. In some embodiments, in the interaction interface, a response text 720 is generated based on the user input 710, and a historical conversation record 730 is composed of one or more historical data of the user input 710. According to the response text 720 and the historical conversation record 730, a machine learning model is used in combination with the preset emotion set 790 to obtain the emotional feature 740 of the response text 720. The emotion set 790 includes pre-stored emotion labels and is updated regularly to expand the types and quantities of emotion labels. The emotional feature 740 may include semantic features based on the response text 720 and scene features based on the historical conversation record 730, and these emotional features 740 may be described by one or more emotion labels in the emotion set 790. According to the obtained emotional features, an audio 750 corresponding to the emotional feature and a multimedia resource 760 corresponding to the emotional feature are respectively generated, and thus the audio 750 and the multimedia resource 760 are presented to the user on the interaction interface.

[0054] Additionally, the machine learning model for obtaining the sentiment feature 740 may include a non - restrictive sentiment perception model, which can be trained by text samples with labeled sentiment tags. Among them, the labeling of sentiment tags for text samples includes clustering the text samples, and dividing the clustering results into corresponding sentiment tags according to the specified granularity. Additionally, the text samples include semantic recognition results, that is, semantic recognition of the response text in the samples, from which the context, expression intention, and sentiment of the response text are parsed out. In the interaction method of the present disclosure, the semantic recognition results are mainly used to indicate the correspondence between the text content and the sentiment tags. Alternatively, keyword terms are set as text features, so as to label sentiment tags for text samples according to the text features.

[0055] Additionally, in a non - restrictive embodiment, the generation method of the audio 750 may include: directly generating the corresponding TTS audio based on the response text 720. At this time, the TTS audio is a regular audio without sentiment features; calling the timbre feature corresponding to the sentiment feature 740 from the timbre database 770, and combining the timbre feature with the aforementioned TTS audio to generate the audio 750 corresponding to the sentiment feature. The timbre database 770 is configured to store timbre types 7910 corresponding to each sentiment tag in the sentiment set 790, such as one or more basic sounds classified according to the sentiment set 790, one or more character sounds 7710 set according to the preset virtual character features, and one or more scene sounds 7720 set according to the preset scene features, etc. In addition, a timbre generation model can also be used to generate a comprehensive timbre based on the combination of one or more of the character sounds 7710 and the scene sounds 7720. The timbre generation model can extract its features (including pitch, loudness, duration, spectrum, and dynamic changes, etc.) based on the analysis of existing timbres, so as to fuse the features of different timbres. The fusion process can adopt various technologies or methods such as generative adversarial networks, variational autoencoders, and autoregressive models.

[0056] Additionally, outputting the audio 750 corresponding to the sentiment feature 740 in the interaction interface may further include: querying whether there is a pre - stored TTS audio corresponding to the response text 720; in response to the pre - existence of such a TTS audio, generating the audio of the response text based on the sentiment feature 740 of the response text 720 and the TTS audio. For the intelligent agent, some common Q&A content can refer to the existing sentiment tags in the sentiment set 790 to prepare the corresponding TTS audio in advance. These TTS audio can have timbres corresponding to specific sentiment features, or can be just simple speech conversions of common phrases. The main purpose is that when generating the audio 750 corresponding to the sentiment feature 740, it can save the audio generation time and improve the interaction efficiency of the intelligent agent.

[0057] Similarly, based on the response text 720, the multimedia resource 760 calls the multimedia features corresponding to the emotional feature 740 from the multimedia database 780, and applies the multimedia features as the multimedia resource 760 for generating the interactive interface display. The multimedia database 780 is configured to store multimedia types 7920 corresponding to each emotional label in the emotion set 790, such as one or more basic multimedia features classified according to the emotion set 790, one or more character images 7810 set according to the preset virtual character, and one or more background elements 7820 set according to the preset scene features, etc. In addition, a comprehensive multimedia feature generated by a multimedia feature generation model according to a combination of one or more of the basic multimedia features, character images, and background elements can also be used, where the multimedia feature generation model can perform feature extraction on existing multimedia samples based on deep learning technology, and use methods such as image synthesis and style transfer to generate new and comprehensive multimedia features.

[0058] Additionally, the multimedia resource of the agent corresponding to the emotional feature 740 displayed in the interactive interface may further include: querying whether there is a pre-stored multimedia resource corresponding to the emotional feature 740 of the response text 720; in response to having a pre-stored multimedia resource, displaying the pre-stored multimedia resource as the displayed multimedia resource 760 in the interactive interface. Similar to the foregoing output audio 750, variations of the emotional labels existing in the emotion set 790 can be prepared for the character image of the agent in the current interactive interface and / or the background elements of the interactive interface, so as to be called at any time when relevant emotional features are involved in the interaction, thereby improving the interaction efficiency.

[0059] Further, please refer to Figure 8, which shows a schematic block diagram of an interaction device between a user and an agent according to some embodiments of the present disclosure. Specifically, the interaction method can be implemented by the interaction device 8. The interaction device 8 may include a processor and a memory (not shown), where the processor may refer to various implementations of digital circuit systems, analog circuit systems, or mixed-signal (a combination of analog and digital) circuit systems that perform functions in a computing system. The processing circuit may include, for example, circuits such as integrated circuits (ICs), application-specific integrated circuits (ASICs), parts or circuits of a single processor core, the entire processor core, a single processor, programmable hardware devices such as field-programmable gate arrays (FPGAs), and / or systems including multiple processors. Additionally, the memory of the interaction device 8 may store information generated by the processor as well as programs and data for the operation of the processor. The memory may be volatile memory and / or non-volatile memory. For example, the memory may include, but is not limited to, random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Generally, the processor may be configured to execute instructions stored in the memory to implement the interaction method between the user and the agent in the present disclosure.

[0060] Specifically, as Figure 8 shown, in some embodiments, the interaction device 8 of the present disclosure may include an input module 81, a reply module 82, an audio module 83, and a multimedia module 84. Specifically, the input module 81 is configured to receive the input of the user in the interaction interface with the agent; the reply module 82 is configured to generate a reply text of the agent based on the input; the audio module 83 is configured to output, in the interaction interface, the audio of the reply text corresponding to the emotional feature according to the emotional feature of the reply text; and the multimedia module 84 is configured to display, in the interaction interface, the multimedia resources of the agent corresponding to the emotional feature according to the emotional feature of the reply text.

[0061] The present disclosure also provides an interaction device, which may include a processor, and a processor coupled to a memory, and the processor is configured to execute, based on instructions stored in the memory, the interaction method for the user and the agent in the interaction interface according to any one of the foregoing embodiments of the present disclosure. The interaction device may refer to Figure 9 , which shows a block diagram of an electronic device according to some embodiments of the present disclosure.

[0062] The memory 91 is used to store one or more computer-readable instructions. The memory 91 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. The memory 91 may store, for example, an operating system, application programs, a boot loader, a database, and other programs, and may also store various application programs and various data, etc.

[0063] The processor 92 is used to run the computer-readable instructions to implement the song screening method described in any of the foregoing embodiments or the method described in any of the foregoing embodiments. For the specific implementation of each step of the method, reference may be made to the foregoing embodiments, and the repeated parts will not be elaborated here.

[0064] The processor 92 may be configured to execute Figures 1 to 8 the steps of the interaction method involved in. The processor 92 may be embodied as various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The central processing unit (CPU) may be of the X86 or ARM architecture, etc.

[0065] The processor 92 and the memory 91 may communicate with each other directly or indirectly. For example, the processor 92 and the memory 91 may communicate through a network. The network may include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 92 and the memory 91 may also communicate with each other through a system bus, and the present disclosure does not limit this.

[0066] It should be noted that Figure 9 the components of the electronic device 9 shown are exemplary and not restrictive. According to actual application needs, the electronic device 9 may also have other components. The processor 92 may control other components in the electronic device 9 to perform desired functions.

[0067] The electronic device 9 may be implemented in a software, firmware, and / or hardware manner and may be integrated in a device installed with relevant application programs.

[0068] Figure 10 A block diagram of an electronic device according to some other embodiments of the present disclosure is shown.

[0069] Figure 10The electronic device 10 shown can be a computer system with a dedicated hardware structure, and when relevant application programs are installed, it can perform corresponding functions.

[0070] The electronic device includes, but is not limited to, mobile terminals such as smart phones, laptop computers, personal digital assistants (PDAs), tablet personal computers (Tablet PCs), portable multimedia players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc.

[0071] As Figure 10 shown, the central processing unit (CPU) 101 executes various processes according to the programs stored in the read-only memory (ROM) 102 or the programs loaded from the storage section 108 into the random access memory (RAM) 103. In the RAM 103, data required when the CPU 101 executes various processes, etc., is stored as needed. The central processing unit is merely exemplary, and it can also be other types of processors, such as the various processors described above. The ROM 102, RAM 103, and storage section 108 can be various forms of computer-readable storage media. It should be noted that although Figure 10 the ROM 102, RAM 103, and storage section 108 are shown separately, one or more of them can be combined, or located in the same or different memories or storage modules.

[0072] The CPU 101, ROM 102, and RAM 103 are connected to each other via the bus 104. The input / output interface 105 is also connected to the bus 104.

[0073] The following components are connected to the input / output interface 105: the input section 106, such as a touch screen, touch pad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; the output section 107, including a display, such as a cathode ray tube (CRT), liquid crystal display (LCD), speaker, vibrator, etc.; the storage section 108, including a hard disk, magnetic tape, etc.; and the communication section 109, including a network interface card such as a LAN card, modem, etc. The communication section 109 allows communication processing to be performed via a network such as the Internet. It is easy to understand that although Figure 10 part of the electronic device 10 is shown to communicate through the bus 104 in the figure, they can also communicate through a network or other means, where the network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.

[0074] As needed, the driver 1010 is also connected to the input / output interface 105. A removable medium 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the driver 1010 as needed, so that a computer program read therefrom is installed in the storage section 108 as needed.

[0075] In the case where the above-described series of processes are implemented by software, a program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 1011.

[0076] The present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, cause the processor to implement the method of interaction between a user and an agent according to any one of the foregoing embodiments of the present disclosure.

[0077] The present disclosure also provides a computer program product, including computer-executable instructions, which, when executed by a processor, cause the processor to implement the method of interaction between a user and an agent according to any one of the foregoing embodiments of the present disclosure.

[0078] According to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, some embodiments of the present disclosure include a computer program product that, when run on a computer, causes the computer to implement the method according to any one of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, including program code for executing the method shown in the flowchart. In such an embodiment, the computer instructions can be downloaded and installed from a network through the communication section 109, or installed from the storage section 108, or installed from the ROM 102. When the computer program is executed by the CPU 101, the method according to the embodiment of the present disclosure is executed.

[0079] It should be noted that, in the context of the present disclosure, a computer-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0080] A computer-readable medium can be a computer-readable storage medium, a computer-readable signal medium, or any combination of the above two.

[0081] A computer-readable storage medium includes, but is not limited to, systems, devices, or components of electricity, magnetism, optics, electromagnetic, infrared, or semiconductors, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to, electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, device, or component. Computer instructions are stored on the computer-readable storage medium, and when executed by a processor, the instructions implement the method described in any of the foregoing embodiments.

[0082] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0083] The above computer-readable medium may be included in the above electronic device; or it may exist separately without being assembled into the electronic device.

[0084] In some embodiments, a computer program is also provided, including: instructions that, when executed by a processor, cause the processor to execute the method described in any of the foregoing embodiments. For example, the instructions may be embodied as computer program code.

[0085] In embodiments of the present disclosure, computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The foregoing programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0086] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0087] The functions described above may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and the like.

[0088] Although some specific embodiments of the present disclosure have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Those skilled in the art should understand that the above embodiments may be modified without departing from the scope and spirit of the present disclosure. The scope of the present disclosure is defined by the appended claims.

Claims

1. An interactive method, comprising: receiving user input in an interactive interface with the agent; Based on the input, generating a reply text of the agent; According to the emotional characteristics of the reply text, in the interactive interface, outputting the audio of the reply text corresponding to the emotional characteristics; as well as According to the emotional characteristics of the reply text, multimedia resources of the intelligent agent corresponding to the emotional characteristics are displayed in the interactive interface.

2. The interactive method according to claim 1, further comprising: Using a machine learning model, obtaining the emotional features of the reply text according to the reply text and the historical conversation records between the user and the agent; The emotional feature includes at least one of a semantic feature based on the reply text and a scene feature based on the historical conversation record.

3. The interactive method according to claim 1, wherein: The audio of the reply text includes the character voice of the agent, According to the emotional feature of the reply text, in the interactive interface, outputting the audio of the reply text corresponding to the emotional feature comprises: According to the emotional characteristics, the character's voice is output in the interactive interface.

4. The interactive method according to claim 3, wherein: The sentiment feature includes a semantic feature based on the reply text, According to the emotional feature of the reply text, in the interactive interface, outputting the audio of the reply text corresponding to the emotional feature comprises: According to the semantic feature, the character voice corresponding to the semantic feature is output in the interactive interface.

5. The interactive method according to claim 4, wherein: The emotional features also include scene features based on the historical dialogue records between the user and the agent, and the audio of the reply text also includes scene sounds. According to the emotional characteristics of the reply text, in the interactive interface, outputting the audio of the reply text corresponding to the emotional characteristics, further comprising: According to the scene feature, the scene sound corresponding to the scene feature is output in the interactive interface.

6. The interactive method according to claim 1, wherein: The reply text includes a plurality of text segments corresponding to a plurality of emotional features one by one, and the audio includes a plurality of audio segments corresponding to the emotional features of the plurality of text segments, According to the emotional feature of the reply text, in the interactive interface, outputting the audio of the reply text corresponding to the emotional feature comprises: In the interactive interface, the multiple audio clips are output sequentially.

7. The interactive method according to claim 6, wherein: According to the emotional characteristics of the reply text, displaying the multimedia resources of the agent corresponding to the emotional characteristics in the interactive interface includes: In the interactive interface, the corresponding multimedia resources are displayed in a changing manner according to the emotional characteristics of the multiple text segments.

8. The interactive method according to any one of claims 1 to 7, wherein: The multimedia resource includes at least one of the character image and background elements of the agent, According to the emotional characteristics of the reply text, in the interactive interface, displaying the multimedia resources of the agent corresponding to the emotional characteristics includes: In response to a change in the emotional characteristic, at least one of the character image and the background element is changed in the interactive interface.

9. The interactive method according to claim 8, wherein: The emotional feature includes at least one of a semantic feature based on the reply text and a scene feature based on a historical conversation record between the user and the agent. According to the emotional characteristics of the reply text, in the interactive interface, displaying the multimedia resources of the agent corresponding to the emotional characteristics includes: In response to the semantic feature changing, in the interactive interface, changing at least one of the expression and posture of the character image; and / or In response to a change in the scene feature, at least one of the color scheme and the scene of the background element is changed in the interactive interface.

10. The interactive method according to claim 1, wherein: In the interactive interface, outputting the audio of the reply text corresponding to the emotional feature includes: Query whether there is pre-stored text-to-speech TTS audio corresponding to the reply text; In response to the pre-stored TTS audio, the audio of the reply text is generated based on the emotional characteristics of the reply text and the TTS audio.

11. The interactive method according to claim 1, wherein: In the interactive interface, displaying the multimedia resources of the agent corresponding to the emotional characteristics includes: Query whether there is a pre-stored multimedia resource of the agent corresponding to the emotional characteristics of the reply text; In response to the presence of the pre-stored multimedia resource, the pre-stored multimedia resource is displayed in the interactive interface as the multimedia resource.

12. The interactive method according to claim 1, wherein: Based on the input, generating the agent's reply text comprises: Using a machine learning model, generating the reply text according to the historical conversation records between the user and the agent; and The reply text is displayed in the interactive interface.

13. An interactive device, comprising: An input module, configured to receive input from a user in an interaction interface with the agent; A reply module, configured to generate a reply text of the agent based on the input; an audio module configured to output, in the interactive interface, an audio of the reply text corresponding to the emotional feature according to the emotional feature of the reply text; as well as The multimedia module is configured to display the multimedia resources of the intelligent agent corresponding to the emotional characteristics in the interactive interface according to the emotional characteristics of the reply text.

14. An electronic device comprising: Memory; as well as A processor coupled to the memory, wherein the processor is configured to execute the interaction method according to any one of claims 1 to 12 based on instructions stored in the memory.

15. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the interactive method according to any one of claims 1 to 12 is performed.

16. A computer program product comprising computer executable instructions, which, when executed by a processor, cause the processor to implement the interaction method according to any one of claims 1 to 12.