Information interaction method, apparatus, device, medium, and program product

CN122817522APending Publication Date: 2026-09-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610967167.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

然而,当前智能系统通过输出文字回复用户,互动方式比较单一,缺乏生动性

Benefits of technology

[0009]上述方法通过展示会话页面,会话页面用于展示与智能系统的会话,会话包括至少一轮对话;响应于会话页面中的输入操作,展示输入信息;在会话页面展示回复信息,回复信息包括与对话场景匹配的回复文本和媒体内容中至少一项,对话场景与输入信息的回复策略相对应,回复策略被配置为基于输入信息所属会话从多个候选媒体内容中选择媒体内容,或者,基于会话的会话信息生成媒体内容。本文通过多模态特征匹配或内容生成方式,获得与对话场景匹配的媒体内容,在会话页面展示媒体内容,解决了当前回复方式单一的问题,提升了智能系统回复的自然度和互动趣味性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817522A_ABST
    Figure CN122817522A_ABST
Patent Text Reader

Abstract

Provided are an information interaction method and device, equipment, a medium and a program product, relating to the technical field of computers. The method comprises: displaying a conversation page, wherein the conversation page is used to display a conversation with an intelligent system, and the conversation comprises at least one round of dialogue; in response to an input operation in the conversation page, displaying input information; and displaying reply information in the conversation page, wherein the reply information comprises at least one of reply text and media content that matches a dialogue scenario, the dialogue scenario corresponds to a reply strategy of the input information, and the reply strategy is configured to select the media content from a plurality of candidate media contents based on a conversation to which the input information belongs, or to generate the media content based on conversation information of the conversation. The method solves the problem of a single current reply mode and improves the naturalness and interactive interest of the reply of the intelligent system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer processing technology, and in particular to an information interaction method, apparatus, device, medium, and program product. Background Technology

[0002] With the development of computer technology, intelligent systems are becoming increasingly sophisticated in their functions. For example, intelligent systems can provide chat or social companionship. However, current intelligent systems primarily respond to users by outputting text, resulting in a relatively simple and unengaging interaction method. Summary of the Invention

[0003] This document provides an information interaction method, apparatus, device, medium, and program product that can optimize the interaction methods of intelligent systems.

[0004] In one scenario, this paper provides an information interaction method, which includes: A conversation page is displayed, wherein the conversation page is used to display a conversation with the intelligent system, the conversation including at least one round of dialogue; In response to input actions on the session page, display the input information; The conversation page displays reply information, which includes at least one of reply text matching the conversation scenario and media content. The conversation scenario corresponds to the reply strategy of the input information. The reply strategy is configured to select the media content from multiple candidate media content based on the conversation to which the input information belongs, or to generate the media content based on the conversation information of the conversation.

[0005] In one instance, this document also provides an information interaction device, which includes: The first display module is used to display a conversation page, wherein the conversation page is used to display a conversation with the intelligent system, and the conversation includes at least one round of dialogue; The second display module is used to display input information in response to input operations on the session page; The third display module is used to display reply information on the conversation page. The reply information includes at least one of reply text matching the dialogue scenario and media content. The dialogue scenario corresponds to the reply strategy of the input information. The reply strategy is configured to select the media content from multiple candidate media content based on the conversation to which the input information belongs, or to generate the media content based on the conversation information of the conversation.

[0006] In one instance, this document also provides an electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the information interaction method as described herein.

[0007] In one instance, this document also provides a storage medium containing computer-executable instructions that, when executed by a computer processor, are used to perform the information interaction methods as described herein.

[0008] In another scenario, this document also provides a computer program product, including a computer program that, when executed by a processor, implements the information interaction methods as described herein.

[0009] The aforementioned method involves displaying a conversation page, which showcases the interaction with the intelligent system. The conversation includes at least one round of dialogue. In response to input on the conversation page, input information is displayed. Response information is then displayed on the conversation page, including at least one of two options: response text matching the dialogue scenario and media content. The dialogue scenario corresponds to a response strategy for the input information, which is configured to either select media content from multiple candidate media content based on the conversation to which the input information belongs, or generate media content based on the conversation information. This paper utilizes multimodal feature matching or content generation to obtain media content matching the dialogue scenario and displays it on the conversation page. This addresses the problem of the current single response method and improves the naturalness and interactive appeal of the intelligent system's responses. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments described herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0011] Figure 1 This is a schematic diagram of the structure of an information interaction system under one scenario. Figure 2 This is a flowchart illustrating an information interaction method in one scenario. Figure 3 This is a flowchart illustrating the method for constructing a media content library in one scenario of information interaction. Figure 4 This is a flowchart illustrating another method of information interaction. Figure 5 This is a diagram illustrating a conversation page in one scenario. Figure 6 A system diagram illustrating an optional example of an information interaction method in one scenario; Figure 7 This is a schematic diagram of the structure of an information interaction device in one scenario. Figure 8 This is a schematic diagram of the structure of an electronic device used to implement an information interaction method in one scenario. Detailed Implementation

[0012] The embodiments will now be described in more detail with reference to the accompanying drawings. While some embodiments are shown in the drawings, it should be understood that the technical solutions can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the technical solutions herein. It should be understood that the illustrated drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the technical solutions.

[0013] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.

[0014] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one situation" means "at least one situation"; the term "another situation" means "at least one additional situation"; the term "some situations" means "at least some situations". Definitions of other terms will be given in the following description.

[0015] It should be noted that the concepts of "first" and "second" mentioned are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.

[0016] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0017] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0018] It is understood that before using the technical solutions disclosed in the various embodiments of this document, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this document in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0019] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations described herein, based on the prompt message.

[0020] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0021] It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.

[0022] It is understood that the data involved in the technical solutions in this article (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0023] In some cases, the provided solution can be applied to Figure 1 The information interaction system shown may include a client 101 and a server 102. The client 101 may include, but is not limited to, web applications such as browsers, applications (Apps), HyperText Markup Language (HTML) applications, lightweight applications (also known as mini-programs), or cloud applications. The client 101 may be deployed on an electronic device and relies on the operation of that device or certain applications on the device to implement its functions. The electronic device may be, for example, a device with a display screen that supports information browsing, such as a smartphone, tablet, personal computer, or other client terminal. For ease of understanding, Figure 1 The client is primarily represented in the form of a device. Other applications can also be configured on the electronic device, such as intelligent dialogue systems, social companion intelligent systems, or chat assistants. Server 102 can be one or more servers providing various services. That is, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; furthermore, it can be a server for a distributed system, a server integrating blockchain technology, a cloud server, or an intelligent cloud computing server or intelligent cloud host deployed with machine learning models, etc.

[0024] The information interaction method described in this paper allows client 101 to interact with server 102, such as receiving or sending messages. For example, in this paper, server 102 can receive input information sent by client 101 based on an information carrier, and send the reply text and media content matching the dialogue scenario corresponding to the input information to client 101 for display on the display interface.

[0025] It should be noted that the information interaction method can be executed by client 101, or by client 101 and server 102, with different functional parts of the corresponding information interaction device deployed on client 101 and server 102 respectively; wherein, the front-end interaction module and local data acquisition module of the device are deployed on client 101, and the back-end data processing module and data storage module are deployed on server 102, and client 101 and server 102 achieve data interaction and functional collaboration through network communication. It should be understood that... Figure 1 The number of clients and servers shown is for illustrative purposes only. Any number of clients and servers can be configured to meet specific implementation requirements.

[0026] Figure 2 This is a flowchart illustrating an information interaction method for one scenario, applicable to intelligent chat applications. This information interaction method can be executed by an information interaction device, which can be implemented in software and / or hardware, optionally through an electronic device such as a mobile terminal, PC, or server. Figure 2 As shown, this information interaction method may specifically include: S210. Display a conversation page, wherein the conversation page is used to display a conversation with the intelligent system, the conversation including at least one round of dialogue.

[0027] In this paper, an intelligent system refers to a system capable of autonomous control based on machine learning models. An intelligent system is, for example, a virtual object or physical entity capable of making decisions and autonomously executing actions based on machine learning models to achieve preset goals or complete preset tasks. An intelligent system can be an automated program that understands user intent and can utilize models or invoke tools to complete various types of tasks. In some contexts, examples of intelligent systems may include, but are not limited to: agents, bots, chatbots, intelligent dialogue systems, social companion agents, and chat assistants. Alternatively, an intelligent system can also be an intelligent role implemented based on machine learning models. An intelligent system can perform specified types of tasks based on user requests derived from generative models (e.g., language models, multimodal models).

[0028] The conversation page is the interactive page of the intelligent system, used to display the conversation between the intelligent system and the user. A conversation consists of at least one round of dialogue. Optionally, a conversation can consist of dialogues belonging to the same topic. The topic can be the core content of the dialogue, the focus of attention and discussion between the user and the intelligent system.

[0029] For example, in response to a session triggering action, a session page is displayed. Optionally, the session triggering action may include a click action on an entry control of the smart system, or a voice command, etc. For example, displaying the smart system's icon, clicking the icon, and displaying the session page. Alternatively, the session page can be displayed via a voice command. Optionally, guidance information may be displayed on the session page to prompt the user to engage in dialogue with the smart system.

[0030] S220. In response to the input operation in the session page, display the input information.

[0031] Input operations include text input, voice input, and image input. Input information refers to the information the user enters into the conversation page; for example, inputting text, text converted from speech, or images. Input information represents the content the user needs to communicate with the intelligent system. For example, input information includes greetings, casual conversation, questions, data analysis requests, or file processing requests.

[0032] S230. Display reply information on the conversation page. The reply information includes at least one of reply text matching the conversation scenario and media content. The conversation scenario corresponds to the reply strategy of the input information. The reply strategy is configured to select the media content from multiple candidate media content based on the conversation to which the input information belongs, or to generate the media content based on the conversation information of the conversation.

[0033] The response information is the intelligent system's output in response to the input information. The response information includes at least one of the following: response text and media content. In this context, media content can be emojis. For example, types of media content include still images, animated images (such as live photos), and videos. The intelligent system can also output emojis to respond to the user's input information, thereby enhancing expression, social affinity, and interactive fun through emojis. Especially in relational dialogue scenarios such as casual conversation, comforting, teasing, acknowledging, greeting, or concluding, using emojis that match the dialogue scenario can significantly improve the naturalness and humanization of the response, and may promote the number of communication conversations and social sharing.

[0034] The dialogue scenario is determined based on the conversation information of the session to which the input information belongs. Since conversation information includes the input information and its context (including historical input and response information), semantic and intent understanding is performed on the conversation information, and the scenario category to which the conversation information belongs is determined based on the understanding results. Optionally, the types of dialogue scenarios include basic emotional scenarios under entertaining expressions, non-basic emotional scenarios under entertaining expressions, image generation scenarios during casual conversation, serious topic scenarios, and knowledge-based question-and-answer scenarios, etc.

[0035] The multimodal features of candidate media content include tag information, text information in emojis, image visual vectors, image-text joint vectors (referring to the image-text joint encoding results), quality evaluation information, and universality evaluation information. Among them, tag information includes multi-dimensional tags, such as emoji tags, usage tags, scene tags, character tags, and work tags.

[0036] The multimodal features of emojis are associated with emoji identifiers and stored in a database to create an emoji library. The emoji identifier uniquely identifies each emoji. For example, an emoji identifier may include an emoji ID or a resource access address for the emoji.

[0037] The media content includes first media content, second media content, and third media content. The first media content is obtained by matching the tag information of the session with the candidate media content. The second media content is obtained by matching the multimodal features of the session with the candidate media content. The third media content is generated based on the session information.

[0038] A response strategy represents the approach used by the intelligent system to determine which emojis match the dialogue context. The response strategy is configured to select media content from multiple candidate media content based on the session to which the input information belongs, or to generate media content based on session information. The response strategy corresponds to the tools invoked by the intelligent system. Tools are program code used to obtain emojis. For example, tools include tag retrieval tools, multimodal retrieval tools, and image generation tools.

[0039] The tag search tool uses conversation information to match search tags and selects suitable emojis from the emoji library. Optionally, it can also combine emoji tag matching, emoji quality rating information, and emoji universality rating information to select suitable emojis from the library. Quality rating information includes quality scores or quality levels, while universality rating information includes universality scores or universality levels. Quality and universality rating information can be obtained through manual annotation.

[0040] Multimodal retrieval tools utilize the multimodal features of conversational information to select appropriate emojis from an emoji library. Multimodal retrieval includes some or all of the following: text retrieval, role retrieval, work retrieval, joint role and work retrieval, image visual vector similarity retrieval, and image-text joint vector similarity retrieval. Multimodal features include text features, role features, work features, image visual vectors, and image-text joint vectors.

[0041] Image generation tools are used to generate emojis that match the current context and expressive intent using conversational information. In one scenario, based on conversational information, emojis matching visual semantics, expressive intent, and style are selected from a collection of textless emojis in an emoji library. Text content fitting the context is then generated based on the conversational information and superimposed onto the textless emojis to obtain an emoji matching the dialogue scenario. In another scenario, a text-based image model is used to generate emojis based on the context and expressive intent corresponding to the conversational information. Optionally, a training sample set is constructed using images containing alpha channels (i.e., images with four RGBA channels: Red, Green, Blue, and Alpha) and corresponding text. The existing text-based image model is fine-tuned using the training sample set to generate images with alpha channels. For example, the text-based image model includes an encoder, a diffusion network, and a decoder. An alpha processing branch is added to the original network structure of the encoder and decoder. Optionally, a feature extraction module for the alpha channel is added to the encoder to extract alpha features. The RGB features are then fused with the alpha features and mapped to the latent feature space. An alpha channel feature processing module is added to the decoder, so that when decoding the image generated in the latent feature space, the features of both the RGB and alpha channels are simultaneously decoded, thereby restoring the image's color and alpha information respectively. The diffusion network is fine-tuned using training samples, enabling the model to learn to perform denoising and generation processing in the latent space containing alpha information. Through these methods, the text-to-image model is fine-tuned to support the generation of images containing alpha channels, meeting the display requirements of different scenarios.

[0042] In one scenario, the emojis output by the intelligent system are displayed on the chat page.

[0043] In another scenario, the media content is validated based on content consistency, tone, and correct expression. If the validation results meet the requirements, the reply text and the media content are displayed on the conversation page. If the validation results do not meet the requirements, the reply text is displayed on the conversation page.

[0044] In this paper, the content consistency dimension refers to whether the media content is consistent with the context of the conversation, including the scene, action, object, or role. The tone dimension refers to whether the tone of the media content is appropriate for the conversation context to avoid tone conflicts. The tone of the media content can be determined based on image and text information. The correct expression dimension refers to whether the media content will cause misunderstanding, offense, or semantic misalignment. For example, before displaying an emoji on the conversation page, the received conversation-related information (including user input, previous replies, and historical context) and the emoji are input into a visual language model. The visual language model determines the fit between the emoji and the current conversation scenario. Based on the fit, it determines whether to output the emoji. If there are issues with consistency between the emoji and the conversation-related information, inappropriate tone, or incorrect expression, the reply text matching the conversation scenario is displayed on the conversation page, and the emoji is not displayed. If none of the above issues exist, the reply text and the emoji are displayed on the conversation page.

[0045] This paper presents a conversation page, which displays interactions with the intelligent system, including at least one round of dialogue. In response to input on the conversation page, input information is displayed. Response information is also displayed on the conversation page, including at least one of two options: response text matching the dialogue scenario and media content. The dialogue scenario corresponds to the response strategy for the input information, which is configured to either select media content from multiple candidate media based on the conversation to which the input information belongs, or generate media content based on the conversation information. This paper utilizes multimodal feature matching or content generation to obtain media content matching the dialogue scenario and displays it on the conversation page, thus addressing the current problem of monotonous response methods and improving the naturalness and interactive appeal of the intelligent system's responses.

[0046] Figure 3 This is a flowchart illustrating a method for constructing a media content library in one scenario of an information interaction approach. The technical solution in this scenario can be combined with implementation methods in other scenarios; for identical or related parts, descriptions of other scenarios can be used, and will not be repeated here. Figure 3 As shown, the method in this case may specifically include: S310, Obtain candidate media content to be managed.

[0047] For example, candidate media content refers to emojis. Emojis can include public emojis or private emojis, etc.

[0048] S320. Perform text recognition and vectorization processing on the candidate media content to obtain the first character in a portion of the candidate media content and the vector representation of the candidate media content.

[0049] For example, text can be extracted from emoji images, GIFs, or video content by understanding the content. It should be noted that emoji images, GIFs, or video content may or may not contain text. Feature extraction is performed on the emoji to convert the image into a high-dimensional vector, obtaining the vector representation of the emoji. Optionally, the vector representation of the emoji can be called the emoji embedding vector.

[0050] In one scenario, the image information of the candidate media content is vectorized to obtain a first vector corresponding to the candidate media content; or, the image and text information of the candidate media content are vectorized to obtain a second vector corresponding to the candidate media content.

[0051] In this paper, the first vector can represent the visual vector of the image in the emoji. The second vector can represent the image-text joint vector generated by the joint encoding of the image and text in the emoji.

[0052] S330. Based on at least one of the first text and the vector representation, select multiple first-level tags corresponding to the candidate media content from multiple first-level candidate tags, wherein the first-level tags include multiple second-level candidate tags.

[0053] First-level candidate tags represent label information describing the characteristics of emojis from different dimensions. For example, first-level candidate tags include emoji tags, usage tags, and scenario tags. For instance, emoji tags include happy, sad, and speechless. Second-level candidate tags represent more granular classifications of the first-level tags. For example, happy emojis can be subdivided into many types. For instance, if a user sends a funny message, they can reply with a "hahaha" emoji to respond to the humor. Or, they can represent conversational scenarios where a user expresses a very good or cheerful mood. Although all of these can correspond to happy emojis, their actual application scenarios differ. Therefore, it is necessary to break down the happy tag into multiple second-level emoji tags. For example, the second-level candidate tags for the happy tag include tags for funny conversational scenarios and tags for scenarios where someone is being praised.

[0054] For example, the first text is matched with the first-level candidate tags to obtain a first matching result. The vector representation is then matched with the first-level candidate tags using vector similarity to obtain a second matching result. Based on the first matching result and / or the second matching result, multiple first-level tags corresponding to the emoji are selected from the multiple first-level candidate tags. The multiple first-level tags annotate the emoji from different dimensions.

[0055] Optionally, if the image of the emoji does not contain text, then the vector representation of the emoji is matched with the first-level candidate labels using vector similarity. Based on the matching results, multiple first-level labels corresponding to the emoji are selected from the multiple first-level candidate labels. The multiple first-level labels represent the emoji from different dimensions.

[0056] S340. Based on at least one of the first text and the vector representation, select multiple secondary tags corresponding to the candidate media content from the multiple secondary candidate tags.

[0057] For example, for each first-level tag, the first text is matched with the corresponding second-level candidate tag to obtain a third matching result. The vector representation is then matched with the corresponding second-level candidate tag to obtain a fourth matching result. Based on the third and / or fourth matching results, the second-level tag corresponding to the emoji is selected from multiple second-level candidate tags.

[0058] Optionally, if the image of the emoji does not contain text, the vector representation of the emoji is matched with the vector similarity of each secondary candidate label corresponding to the primary label, and the secondary label corresponding to the emoji is selected from multiple secondary candidate labels based on the matching results.

[0059] S350, Perform error correction processing on the primary label and the secondary label.

[0060] For example, a multi-turn dialogue is used to perform visual language annotation on emojis. For instance, a single-turn dialogue instructs the intelligent system to extract the text from the emoji. The single-turn dialogue then instructs the intelligent system to determine the vector representation of the emoji. Based on the text and vector representation, the single-turn dialogue instructs the intelligent system to extract the primary label of the emoji. Based on the text, vector representation, and primary label, the single-turn dialogue instructs the intelligent system to extract the secondary label of the emoji. However, a second-turn dialogue instructs the intelligent system to check the accuracy of the previous labels.

[0061] S360. Based on the role information and work information of the candidate media content, obtain the content tags of the candidate media content.

[0062] Content tags include information such as the character and the work. For example, extracting character and work information from an emoji pack yields its content tags.

[0063] S370. Construct a media content library that includes identification information, the first text, the first vector, the second vector, the first-level tag, the second-level tag, and the content tag.

[0064] For example, structured data for each emoji is constructed based on fields such as emoji identifier, text in the image, image vector, combined image and text vector, primary emoji tag, secondary emoji tag, primary usage tag, secondary usage tag, primary scene tag, secondary scene tag, content tag, quality score, and universality score. The structured data for each emoji is then stored in a database table to obtain an emoji library.

[0065] This paper performs text recognition and vectorization on candidate media content to obtain the first text and vector representation of the candidate media content. Based on at least one of the first text and vector representation, multiple first-level tags corresponding to the candidate media content are selected from multiple first-level candidate tags. Based on at least one of the first text and vector representation, multiple second-level tags corresponding to the candidate media content are selected from multiple second-level candidate tags. Error correction is performed on the first-level and second-level tags. This paper implements a multi-stage, multi-dimensional emoji annotation method, solving problems such as high annotation noise, unstable character / work recognition, and insufficient understanding of animated / video content. Therefore, it improves the annotation quality of the emoji library and provides support for improving the retrieval efficiency of the subsequent emoji library.

[0066] Figure 4 This is a flowchart illustrating an information interaction method under another scenario. The technical solution in this scenario can be combined with implementation methods in other scenarios. For identical or related parts, descriptions of other scenarios can be used, and will not be repeated here. Figure 4 As shown, the method in this case may specifically include: S410. Display a conversation page, wherein the conversation page is used to display a conversation with the intelligent system, the conversation including at least one round of dialogue.

[0067] Figure 5 This is a schematic diagram of a conversation page in one scenario. Based on the user's dialogue needs, conversation page 510 is displayed. Conversation page 510 is used to display dialogue information. Optionally, input information 520 is displayed on conversation page 510. For example, input information 520 could be "You are the best chatbot".

[0068] S420. In response to the input operation in the session page, display the input information.

[0069] S430. Based on the session information of the session to which the input information belongs, the dialogue scenario is obtained.

[0070] S440. If the dialogue scenario meets the requirements, the intelligent system outputs the media content based on the response strategy corresponding to the dialogue scenario and the conversation information.

[0071] The requirement that the dialogue scenario is appropriate means that the current dialogue scenario is suitable for replying with emojis. By determining whether it is appropriate to reply with emojis based on the dialogue scenario, it is possible to avoid sending emojis in serious topics, formal Q&A, knowledge-based Q&A, or scenarios where humorous expressions are not suitable, thus avoiding a negative impact on the user experience.

[0072] See Figure 5 Based on the relevant information in input information 520, it can be inferred that the current conversation is a casual chat, and an emoji can be used as a reply. The reply text 530 and the emoji 540 are displayed on the conversation page 510. For example, the reply text could be "Wow, you've made me blush!", and the emoji could be an image of someone blushing. Optionally, emoticons can also be inserted into the reply text. These emoticons include smiley faces or thumbs-up icons, etc.

[0073] In one scenario, if the dialogue scenario corresponds to a first response strategy, a search tag is obtained based on the conversation information; the search tag is matched with the tag information of the candidate media content, and the first media content is obtained based on the matching degree, wherein the tag information includes the plurality of primary tags and the plurality of secondary tags.

[0074] The first response strategy involves using a tag-based search tool to retrieve an emoji library within a basic emotional context of entertaining dialogue. For example, if the current dialogue scenario simply expresses joy, anger, sorrow, or happiness, a natural language processing model is invoked to understand the conversation information and generate search tags. Based on these tags, the primary and secondary tags of emoji tags in the emoji library are searched to obtain a candidate emoji. This candidate emoji is then used as the primary media content. Optionally, multiple candidate emojis are obtained by searching the primary and secondary tags of emoji tags in the emoji library based on the search tags. One candidate emoji can be randomly selected as the primary media content. Alternatively, the quality score and universality score of multiple candidate emojis can be combined to select one as the primary media content. Because tag matching is based on text and vector representations, it ensures matching across dimensions such as expression, specific scene, action, object, and visual semantics, avoiding situations where the expression is roughly correct but the visual representation is inappropriate.

[0075] In another scenario, if the dialogue scenario corresponds to the second response strategy, the keywords of the conversation information and the vector representation of the conversation information are obtained; the keywords are compared with the first text for similarity, and a first candidate media content is obtained based on the similarity comparison result; the first content contained in the keywords is matched with the content tags, and a second candidate media content is obtained based on the matching degree, wherein the first content includes at least one of the character and the work name; the vector representation of the conversation information and the vector representation of the candidate media content are compared for similarity, and a third candidate media content is obtained based on the similarity comparison result; the first candidate media content, the second candidate media content, and the third candidate media content are sorted, and the first media content is obtained based on the sorting result.

[0076] The second response strategy refers to using a multimodal retrieval tool to search the emoji library when the dialogue scenario is a non-basic emotional scenario with entertainment-oriented expression. This non-basic emotional scenario with entertainment-oriented expression can refer to conversation information containing specific keywords. For example, specific keywords may include specific objects, works, or gerund concepts, such as "solar term" and "get off work." For example, keyword extraction is performed on the conversation information to obtain the keywords. Feature extraction is performed on the conversation information to obtain a vector representation in a high-dimensional space, such as image vectors and image-text joint vectors. The similarity of the keywords with the first text corresponding to the image text in the emoji library is compared, and the first candidate emoji is obtained based on the similarity comparison results. Keywords may include roles and / or works. The role and work names in the keywords are matched with the role tags and work tags in the content tags, respectively, and the second candidate emoji is obtained based on the matching degree. The vector representation of the conversation information is compared with the image vectors and image-text joint vectors in the emoji library, and the third candidate emoji is obtained based on the similarity comparison results. A ranking algorithm is used to uniformly rank the first, second, and third candidate emojis, and the first media content is obtained based on the ranking result. The ranking algorithm includes weighted reciprocal ranking fusion algorithms, among others. The core logic of the weighted reciprocal ranking fusion algorithm is a weighted fusion based on the reciprocal of the ranking. For example, for each of the first, second, and third candidate emojis, the reciprocal of the emoji's ranking in different ranking methods (including ranking based on different algorithms, or ranking based on different dimensions, etc.) is summed, and the emojis are ranked in descending order based on the sum. The top-ranked emoji is selected as the first media content. Through multimodal retrieval, keywords, visual semantics of images, object expressions, actions, characters, and works can be considered, solving the problems of inaccurate recall, unstable ranking, and insufficient semantic coverage caused by relying solely on a single tag, single keyword, or single vector representation.

[0077] Optionally, if the number of first, second, and third candidate emojis exceeds a preset threshold, a ranking model is used to select the top-ranked emoji from these candidates as the primary media content. This paper improves the relevance of emoji responses by using a ranking model to select the top-ranked emoji for emoji retrieval scenarios with a large number of emojis.

[0078] In another scenario, if the dialogue scenario corresponds to a third response strategy, the conversation information and the multimodal features are matched to obtain a fourth media content, wherein the fourth media content does not contain text; a second text is generated based on the conversation information, wherein the second text represents the text used to reply to the input information; the second text is superimposed on the fourth media content to obtain a third media content.

[0079] The third response strategy refers to using image generation tools to generate emojis for customized scenarios. For example, to identify text that is unlikely to appear in the emoji library, and based on conversation information, it can be determined that such text needs to be added to an emoji. This can be done by first searching the emoji library for emojis without text based on the multimodal features corresponding to the conversation information. Then, text content that fits the dialogue scenario is generated based on the conversation information. For example, if the topic of the conversation is "It's hot," a word like "hot" is generated, and this generated text is overlaid on the searched emoji to obtain the third-media content.

[0080] In another scenario, if the dialogue scenario corresponds to a third response strategy, third media content is generated based on the text information related to media content generation in the conversation information.

[0081] For example, based on conversation information analysis, the system obtains prompts for users' emoji generation needs. A text-to-image model is then used to generate third-party media content based on these prompts. For instance, if a user inputs "I am Xiao A," the intelligent system outputs the reply text "Hello Xiao A," along with an emoji containing the text "Hi! I'm Xiao A."

[0082] S450. Verify the media content according to the dimensions of content consistency, tone, and correct expression.

[0083] S460. If the verification result meets the requirements, the reply text and the media content are displayed on the conversation page.

[0084] S470. If the verification result does not meet the requirements, the reply text will be displayed on the conversation page.

[0085] This paper addresses this issue by determining whether the current dialogue scenario is suitable for responding with emojis after inputting dialogue information. In scenarios where emojis are inappropriate, a text-only reply is provided. Conversely, in scenarios where emojis are appropriate, an emoji selection strategy based on the dialogue scenario and conversation information is used to select and display emojis matching the current dialogue scenario from the emoji library. This allows for the use of different search tools for different scenarios during the search process, balancing search efficiency, recall accuracy, and semantic coverage. Furthermore, by adding text to emojis retrieved from the emoji library that do not contain text, or by generating emojis based on conversation information, the richness of the emoji library is enhanced, thereby improving the flexibility of interaction.

[0086] Figure 6 This is a system diagram illustrating an optional example of an information interaction method in one scenario. For example... Figure 6 As shown, the system includes an emoji database 610, a data cleaning module 620, an annotation module 630, an emoji library 640, an intelligent system 650, and a text-to-image model 660. The emoji database 610 contains emoji data to be managed. The data cleaning module 620 cleans the emoji data to filter out eligible data. The annotation module 630 uses an annotation model to perform multi-round annotation and multimodal feature extraction of emojis. The intelligent system 650 receives user input and displays it on the conversation page. If the dialogue scenario corresponding to the input requires the use of an image generation tool, it constructs a prompt based on the conversation information and requests the text-to-image model 660 to generate an emoji. If the dialogue scenario corresponding to the input requires searching the emoji library 640, it accesses the library using emoji tags, content tags, keywords, image vectors, and image-text joint vectors to obtain the emoji.

[0087] Figure 7 This is a schematic diagram of the structure of an information interaction device in one scenario, such as... Figure 7 As shown, the device includes: The first display module 710 is used to display a conversation page, wherein the conversation page is used to display a conversation with the intelligent system, and the conversation includes at least one round of dialogue; The second display module 720 is used to display input information in response to input operations on the session page; The third display module 730 is used to display reply information on the conversation page. The reply information includes at least one of reply text matching the dialogue scenario and media content. The dialogue scenario corresponds to the reply strategy of the input information. The reply strategy is configured to select the media content from multiple candidate media content based on the conversation to which the input information belongs, or to generate the media content based on the conversation information of the conversation.

[0088] The aforementioned device displays a conversation page, which showcases a dialogue with the intelligent system, including at least one round of conversation. In response to input operations on the conversation page, it displays input information. The response information on the conversation page includes at least one of two components: a response text matching the dialogue scenario and media content. The dialogue scenario corresponds to a response strategy for the input information, which is configured to either select media content from multiple candidate media content based on the conversation to which the input information belongs, or generate media content based on the conversation information. This paper addresses the problem of current response methods being too simplistic by obtaining media content matching the dialogue scenario through multimodal feature matching or content generation, and then displays the media content on the conversation page. This improves the naturalness and interactive appeal of the intelligent system's responses.

[0089] In one scenario, the method further includes obtaining the multimodal features corresponding to the candidate media content using the following steps: The candidate media content is subjected to text recognition and vectorization processing to obtain the first character in a portion of the candidate media content and the vector representation of the candidate media content; Based on at least one of the first text and the vector representation, select multiple first-level tags corresponding to the candidate media content from multiple first-level candidate tags, wherein the first-level tags include multiple second-level candidate tags; Based on at least one of the first text and the vector representation, select multiple secondary tags corresponding to the candidate media content from the multiple secondary candidate tags; Error correction processing is performed on the primary and secondary tags.

[0090] In one scenario, the vectorization process of the candidate media content to obtain a vector representation of the candidate media content includes: The image information of the candidate media content is vectorized to obtain the first vector corresponding to the candidate media content; The image and text information of the candidate media content are vectorized to obtain the second vector corresponding to the candidate media content.

[0091] In one case, it also includes: Based on the role information and work information of the candidate media content, the content tags of the candidate media content are obtained.

[0092] In one case, it also includes: The dialogue scenario is obtained based on the session information of the session to which the input information belongs; If the dialogue scenario meets the requirements, the intelligent system outputs the media content based on the response strategy corresponding to the dialogue scenario and the conversation information.

[0093] In one scenario, the step of outputting the media content through the intelligent system based on the response strategy corresponding to the dialogue scenario and the conversation information includes: If the dialogue scenario corresponds to the first response strategy, obtain the search tags based on the conversation information; The search tags are matched with the tag information of the candidate media content, and the first media content is obtained based on the matching degree. The tag information includes the plurality of primary tags and the plurality of secondary tags.

[0094] In one scenario, the step of outputting the media content through the intelligent system based on the response strategy corresponding to the dialogue scenario and the conversation information includes: If the dialogue scenario corresponds to the second response strategy, obtain the keywords of the conversation information and the vector representation of the conversation information; The keywords are compared with the first text for similarity, and the first candidate media content is obtained based on the similarity comparison results; Match the first content contained in the keyword with the content tag, and obtain the second candidate media content based on the matching degree, wherein the first content includes at least one of the character and work name; A similarity comparison is performed between the vector representation of the session information and the vector representation of the candidate media content, and a third candidate media content is obtained based on the similarity comparison result; The first, second, and third candidate media contents are ranked, and the first media content is obtained based on the ranking results.

[0095] In one scenario, the step of outputting the media content through the intelligent system based on the response strategy corresponding to the dialogue scenario and the conversation information includes: If the dialogue scenario corresponds to the third response strategy, the conversation information and the multimodal features are matched to obtain the fourth media content, wherein the fourth media content does not contain text; A second text is generated based on the conversation information, wherein the second text represents the text used to reply to the input information; The second text is superimposed onto the fourth media content to obtain the third media content.

[0096] In one scenario, the step of outputting the media content through the intelligent system based on the response strategy corresponding to the dialogue scenario and the conversation information includes: If the dialogue scenario corresponds to a third response strategy, third media content is generated based on the text information related to media content generation in the conversation information.

[0097] In one scenario, the third display module 730 is specifically used for: The media content was verified based on the dimensions of content consistency, tone, and correct expression. If the verification result meets the requirements, the reply text and the media content will be displayed on the chat page. If the verification result does not meet the requirements, the reply text will be displayed on the chat page.

[0098] In one scenario, the media content includes first media content, second media content, and third media content, wherein the first media content is obtained by matching the tag information of the session with that of the candidate media content, the second media content is obtained by matching the multimodal features of the session with those of the candidate media content, and the third media content is generated based on the session information.

[0099] The aforementioned information interaction device can execute the information interaction method provided in any of the embodiments described herein, and has the corresponding functional modules and beneficial effects for executing the information interaction method.

[0100] It is worth noting that the various units and modules included in the above-mentioned information interaction device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments described herein.

[0101] The following is for reference. Figure 8 This document illustrates a schematic diagram of an electronic device (e.g., a terminal device or server) 800 suitable for implementing the above-described methods. The terminal device described herein may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital televisions and desktop computers. Figure 8The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments described herein.

[0102] like Figure 8 As shown, the electronic device 800 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 802 or a program loaded from a storage device 808 into a random access memory (RAM) 803. The RAM 803 also stores various programs and data required for the operation of the electronic device 800. The processing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0103] Typically, the following devices can be connected to I / O interface 805: input devices 806 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 807 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 808 including, for example, magnetic tapes, hard disks, etc.; and communication devices 809. Communication device 809 allows electronic device 800 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 An electronic device 800 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0104] In particular, according to embodiments of this document, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, the technical solutions of this document include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 809, or installed from a storage device 808, or installed from a ROM 802. When the computer program is executed by a processing device 801, it performs the functions defined in the methods of the embodiments of this document.

[0105] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0106] The electronic device provided in this embodiment and the information interaction method provided in the above technical solutions belong to the same inventive concept. Technical details not described in detail herein can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0107] This article provides a computer storage medium on which a computer program is stored, which, when executed by a processor, implements the information interaction method provided in the above embodiments.

[0108] It should be noted that the computer-readable medium mentioned above can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, also known as flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. The transmitted data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0109] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include Local Area Networks (LANs), Wide Area Networks (WANs), the Internet (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0110] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0111] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: A conversation page is displayed, wherein the conversation page is used to display a conversation with the intelligent system, the conversation including at least one round of dialogue; In response to input actions on the session page, display the input information; The conversation page displays reply information, which includes at least one of reply text matching the conversation scenario and media content. The conversation scenario corresponds to the reply strategy of the input information. The reply strategy is configured to select the media content from multiple candidate media content based on the conversation to which the input information belongs, or to generate the media content based on the conversation information of the conversation.

[0112] Computer program code for performing the operations described herein can be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" programming language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including local area networks (LANs) or wide area networks (WANs), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0113] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this document. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0114] The modules or units described herein can be implemented in software or hardware. The names of modules or units do not necessarily constitute a limitation on the module or unit itself.

[0115] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include at least one of the following: Field-Programmable Gate Array (FPGA), Application-Specific Integrated Circuit (ASIC), Application-Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), etc.

[0116] In the context of this document, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory, read-only memory, erasable programmable read-only memory (flash memory), optical fibers, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0117] The above description is merely a preferred embodiment and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.

[0118] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this document. Certain features described in the context of individual implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.

[0119] Although the subject matter has been described using a programming language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims.

Claims

1. An information interaction method, comprising: A conversation page is displayed, wherein the conversation page is used to display a conversation with the intelligent system, the conversation including at least one round of dialogue; In response to input actions on the session page, display the input information; The conversation page displays reply information, which includes at least one of reply text matching the conversation scenario and media content. The conversation scenario corresponds to the reply strategy of the input information. The reply strategy is configured to select the media content from multiple candidate media content based on the conversation to which the input information belongs, or to generate the media content based on the conversation information of the conversation.

2. The method according to claim 1 further includes obtaining multimodal features corresponding to the candidate media content using the following steps: The candidate media content is subjected to text recognition and vectorization processing to obtain the first character in a portion of the candidate media content and the vector representation of the candidate media content; Based on at least one of the first text and the vector representation, select multiple first-level tags corresponding to the candidate media content from multiple first-level candidate tags, wherein the first-level tags include multiple second-level candidate tags; Based on at least one of the first text and the vector representation, select multiple secondary tags corresponding to the candidate media content from the multiple secondary candidate tags; Error correction processing is performed on the primary and secondary tags.

3. The method according to claim 2, wherein the step of vectorizing the candidate media content to obtain a vector representation of the candidate media content includes: The image information of the candidate media content is vectorized to obtain the first vector corresponding to the candidate media content; The image and text information of the candidate media content are vectorized to obtain the second vector corresponding to the candidate media content.

4. The method according to claim 2, further comprising: Based on the role information and work information of the candidate media content, the content tags of the candidate media content are obtained.

5. The method according to claim 2, further comprising: The dialogue scenario is obtained based on the session information of the session to which the input information belongs; If the dialogue scenario meets the requirements, the intelligent system outputs the media content based on the response strategy corresponding to the dialogue scenario and the conversation information.

6. The method according to claim 5, wherein the step of outputting the media content by the intelligent system based on the response strategy corresponding to the dialogue scenario and the conversation information includes: If the dialogue scenario corresponds to the first response strategy, obtain the search tags based on the conversation information; The search tags are matched with the tag information of the candidate media content, and the first media content is obtained based on the matching degree. The tag information includes the plurality of primary tags and the plurality of secondary tags.

7. The method according to claim 5, wherein the step of outputting the media content by the intelligent system based on the response strategy corresponding to the dialogue scenario and the conversation information includes: If the dialogue scenario corresponds to the second response strategy, obtain the keywords of the conversation information and the vector representation of the conversation information; The keywords are compared with the first text for similarity, and the first candidate media content is obtained based on the similarity comparison results; Match the first content contained in the keyword with the content tag, and obtain the second candidate media content based on the matching degree, wherein the first content includes at least one of the character and work name; A similarity comparison is performed between the vector representation of the session information and the vector representation of the candidate media content, and a third candidate media content is obtained based on the similarity comparison result; The first, second, and third candidate media contents are ranked, and the first media content is obtained based on the ranking results.

8. The method according to claim 5, wherein the step of outputting the media content by the intelligent system based on the response strategy corresponding to the dialogue scenario and the conversation information includes: If the dialogue scenario corresponds to the third response strategy, the conversation information and the multimodal features are matched to obtain the fourth media content, wherein the fourth media content does not contain text; A second text is generated based on the conversation information, wherein the second text represents the text used to reply to the input information; The second text is superimposed onto the fourth media content to obtain the third media content.

9. The method according to claim 5, wherein the step of outputting the media content by the intelligent system based on the response strategy corresponding to the dialogue scenario and the conversation information includes: If the dialogue scenario corresponds to a third response strategy, third media content is generated based on the text information related to media content generation in the conversation information.

10. The method according to claim 1, wherein displaying reply information on the conversation page includes: The media content was verified based on the dimensions of content consistency, tone, and correct expression. If the verification result meets the requirements, the reply text and the media content will be displayed on the chat page. If the verification result does not meet the requirements, the reply text will be displayed on the chat page.

11. The method according to claim 1, wherein the media content includes first media content, second media content, and third media content, wherein, The first media content is obtained by matching the tag information of the session with the candidate media content, the second media content is obtained by matching the multimodal features of the session with the candidate media content, and the third media content is generated based on the session information.

12. An information interaction device, comprising: The first display module is used to display a conversation page, wherein the conversation page is used to display a conversation with the intelligent system, and the conversation includes at least one round of dialogue; The second display module is used to display input information in response to input operations on the session page; The third display module is used to display reply information on the conversation page. The reply information includes at least one of reply text matching the dialogue scenario and media content. The dialogue scenario corresponds to the reply strategy of the input information. The reply strategy is configured to select the media content from multiple candidate media content based on the conversation to which the input information belongs, or to generate the media content based on the conversation information of the conversation.

13. An electronic device, the electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the information interaction method as described in any one of claims 1-11.

14. A storage medium comprising computer-executable instructions, which, when executed by a computer processor, are used to perform the information interaction method as described in any one of claims 1-11.

15. A computer program product comprising a computer program that, when executed by a processor, implements the information interaction method as described in any one of claims 1-11.