Interaction method and apparatus, and storage medium and program product

By displaying the graphic and textual content of virtual objects on the conversation interface, the problem of low information transmission efficiency in existing technologies is solved, and a more efficient and intuitive information interaction experience is achieved.

WO2026153358A1PCT designated stage Publication Date: 2026-07-23BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2026-01-14
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

In existing technologies, the information transmission efficiency of virtual objects is low. Users need to read a lot of text or video, resulting in a poor intuitive experience and difficulty in efficiently understanding the reply content.

Method used

The interface displays graphic and text content, including at least one image group and at least one text block. The graphic and text content comes from media works published by the second user corresponding to the virtual object and is displayed in order of playback and relevance of the media works.

Benefits of technology

It improves the efficiency and intuitiveness of information delivery, enabling users to understand the responses of virtual objects more quickly and clearly, thus enhancing the interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2026072491_23072026_PF_FP_ABST
    Figure CN2026072491_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of terminals. Provided are an interaction method and apparatus, and a storage medium and a program product. The interaction method comprises: displaying a conversation interface between a first user and a virtual object; receiving an input message on the conversation interface; and displaying image-text content related to the input message in a message body on the conversation interface, wherein the image-text content comprises at least one image group and at least one text block, and the image-text content is part of content in a media work published by a second user corresponding to the virtual object.
Need to check novelty before this filing date? Find Prior Art

Description

Interaction methods, devices, storage media and program products

[0001] Cross-references to related applications

[0002] This application is based on and claims priority to Chinese Patent Application No. 202510065337.9, filed on January 15, 2025, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] This disclosure relates to the field of terminals, and more particularly to an interaction method, apparatus, storage medium, and program product. Background Technology

[0004] In related technologies, user A can create a virtual object on an internet platform, and user B can communicate or interact with the virtual object through the network. The virtual object can answer corresponding questions based on user B's questions. Summary of the Invention

[0005] According to some embodiments of this disclosure, an interaction method is provided, including: displaying a conversation interface between a first user and a virtual object; receiving an input message on the conversation interface; and displaying graphic and text content related to the input message in a message body on the conversation interface, wherein the graphic and text content includes at least one image group and at least one text block, and the graphic and text content is part of a media work published by a second user corresponding to the virtual object.

[0006] According to other embodiments of this disclosure, an interactive device is provided, comprising: a first display module configured to display a conversation interface between a first user and a virtual object; a receiving module configured to receive an input message on the conversation interface; and a second display module configured to display graphic content related to the input message in a message body on the conversation interface, the graphic content including at least one image group and at least one text block, the graphic content being a portion of a media work published by the second user corresponding to the virtual object.

[0007] According to further embodiments of the present disclosure, an interactive device is provided, including: a processor; and a memory coupled to the processor for storing instructions, which, when executed by the processor, cause the processor to perform an interactive method according to any embodiment of the present disclosure.

[0008] According to further embodiments of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, wherein when executed by a processor, the program causes the processor to implement the interactive methods of any embodiment of the present disclosure.

[0009] According to further embodiments of the present disclosure, a computer program product is provided, including a computer program or instructions that are executed by a processor to implement an interactive method as described in any embodiment of the present disclosure.

[0010] Other features, aspects, and advantages of this disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0011] Embodiments of this disclosure are described below with reference to the accompanying drawings. It should be understood that the drawings described below are merely illustrative of some embodiments of this disclosure and are not intended to limit the scope of this disclosure. In the drawings:

[0012] Figure 1 shows a flowchart illustrating the interaction methods of some embodiments of this disclosure;

[0013] Figure 2A illustrates a schematic diagram of displaying a text and an image in a message body according to some embodiments of the present disclosure;

[0014] Figure 2B illustrates a schematic diagram of displaying a text and an image in a message body according to some other embodiments of this disclosure;

[0015] Figure 2C illustrates a schematic diagram of displaying a text and an image in a message body according to some embodiments of the present disclosure;

[0016] Figure 2D illustrates a schematic diagram of displaying a text and an image in a message body according to some embodiments of the present disclosure;

[0017] Figure 3A illustrates a schematic diagram of displaying two texts and two images in a message body according to some embodiments of this disclosure;

[0018] Figure 3B illustrates a schematic diagram of displaying two texts and two images in a message body according to other embodiments of this disclosure;

[0019] Figure 4 illustrates a schematic diagram of displaying one text and two images in a message body according to some embodiments of this disclosure;

[0020] Figure 5A illustrates a schematic diagram of displaying two texts and four images in a message body according to some embodiments of this disclosure;

[0021] Figure 5B illustrates a schematic diagram of displaying two texts and four images in a message body according to other embodiments of this disclosure;

[0022] Figure 6 illustrates a schematic diagram of constructing a knowledge base according to some embodiments of this disclosure;

[0023] Figure 7 shows a block diagram of an interactive device according to some embodiments of the present disclosure;

[0024] Figure 8 shows a block diagram of an interactive device according to other embodiments of the present disclosure;

[0025] Figure 9 shows a block diagram of an electronic device according to some embodiments of the present disclosure.

[0026] It should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not necessarily drawn to actual scale. The same or similar reference numerals are used in the various drawings to denote the same or similar parts. Therefore, once an item is defined in one drawing, it may not be discussed further in subsequent drawings. Detailed Implementation

[0027] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. It should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein.

[0028] It should be understood that the various steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect. Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of components and steps set forth in these embodiments should be interpreted as merely exemplary and do not limit the scope of this disclosure.

[0029] As used in this disclosure, the term "comprising" and its variations are open-ended terms that include at least the following elements / features but do not exclude other elements / features, i.e., "including but not limited to". The term "based on" means "at least partially based on".

[0030] It should be noted that the concepts of "first," "second," etc., used in this disclosure are used only to distinguish different devices, modules, or units, and are not intended to define the order of functions performed by these devices, modules, or units or their interdependencies. Unless otherwise specified, the concepts of "first," "second," etc., are not intended to imply that the objects described herein must be in a given temporal, spatial, rank, or any other given order.

[0031] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0032] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0033] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.

[0034] The embodiments of this disclosure are described in detail below with reference to the accompanying drawings; however, this disclosure is not limited to these specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner that will be apparent to those skilled in the art from this disclosure.

[0035] It should be understood that this disclosure does not limit how the image to be applied / processed is obtained. In some embodiments of this disclosure, it can be obtained from a storage device, such as internal memory or external storage device. In other embodiments of this disclosure, a camera component can be invoked to capture an image. It should be noted that the acquired image can be a captured image or a frame from a captured video, and is not particularly limited to these.

[0036] In the context of this disclosure, "image" can refer to any of a variety of images, such as color images, grayscale images, etc. It should be noted that the type of image is not specifically limited in the context of this specification. Furthermore, an image can be any suitable image, such as a raw image obtained by a camera device, or an image from which specific processing has been performed, such as preliminary filtering, dealiasing, color adjustment, contrast adjustment, normalization, etc. It should be noted that preprocessing operations may also include other types of preprocessing operations known in the art, which will not be described in detail here.

[0037] Virtual objects created by users on internet platforms, such as virtual avatars, virtual characters, or AI (Artificial Intelligence) avatars, can serve as auxiliary intelligent agents for the creators to complete certain tasks, such as answering questions in the comments section or engaging in simple conversations with other users.

[0038] Users can interact with virtual objects corresponding to other users through internet platforms. In related technologies, after a user asks a question, the virtual object can respond with text content. This information transmission efficiency is low, and users need to read a large amount of text in a short time, which may not be a good user experience. Even if the virtual object replies with a complete video or image, it is insufficient to clearly and efficiently express the response, requiring users to spend a significant amount of time understanding the content. How to improve the efficiency of information transmission through virtual objects to enhance interaction efficiency is a key area of ​​focus.

[0039] Figure 1 shows a flowchart illustrating the interaction methods of some embodiments of this disclosure.

[0040] As shown in Figure 1, the interaction method includes: step S11, displaying a conversation interface between the first user and the virtual object; step S12, receiving an input message on the conversation interface; step S13, displaying graphic and text content related to the input message in a message body on the conversation interface, the graphic and text content including at least one image group and at least one text block, the graphic and text content being part of the media work published by the second user corresponding to the virtual object.

[0041] The first user can be understood as a user of the internet platform or the current application, generally referring to a user who interacts with virtual objects. Virtual objects, for example, correspond to a virtual avatar of the second user, who is the user who publishes the virtual avatar and media works. This disclosure does not limit how the virtual avatar of the second user is created.

[0042] The conversation interface is an instant messaging interface, such as a chat interface. It includes an input area and an interactive information display area. The interactive information display area can be used to display the interactive information sent between the first user and the virtual object. The input area can be used to input interactive information and display interactive information to be sent. For example, the first user can enter a message in the input area of ​​the conversation interface, such as a question. The virtual object can reply with a response corresponding to the question entered by the first user in the interactive information display area. For example, the response can be displayed in a graphic format on the conversation interface.

[0043] In some embodiments, a machine learning model is used to analyze and understand the input message to determine whether to trigger the display of text and images. For example, if the user is only sending a greeting, no subsequent display of text and images will be triggered. The machine learning model is used to call a media works knowledge base to obtain text and images content related to the input message, which is then displayed in a message body through a conversational interface. This disclosure does not limit the algorithm for triggering the display of text and images or how the knowledge base is called.

[0044] For example, a message body corresponds to a message bubble, which is a visual element that presents information. For instance, a message bubble is a bubble-shaped frame. A message body describes the expression of a message, and a message body corresponds to a session container, which may be, for example, a multi-image and text mixed container that conforms to the dialogue consumption experience.

[0045] In related technologies, an image is first displayed through one message body, and then extended reading information corresponding to that image is displayed through another message body. In this embodiment, however, graphic and text content is displayed within a single message body. This content includes both images and text, and can also be referred to as multimedia content. The graphic and text content originates from media works published by a second user. Images can be still images, animated images, or video clips. For example, at least one image group and at least one text block are displayed within a single message body. Each image group in the at least one image group may include one or more images. Each text block in the at least one text block may include one or more texts.

[0046] Media works refer to content published by a second user, such as videos or text / image works. The content can also be live stream content from a second user. Second users can authorize the use of permitted videos, text / image works, and text content in various ways; this disclosure does not restrict the methods of authorization.

[0047] If the media work is a video, in some embodiments, the graphic content is at least one video frame in the video. For example, the graphic content includes one frame of image, multiple frames of image, one video clip, or multiple video clips in the video. In other embodiments, the graphic content is at least one text associated with audio or textual information such as subtitles or descriptions in the video. For example, the graphic content includes one or more texts associated with audio in the video.

[0048] If the media work is a graphic work, then the graphic content can be a combination of images and text within the graphic work.

[0049] Some content in a media work is determined based on the input message. For example, the server first parses the content of the work and structures the fields to form work knowledge. Based on the first user's question, it retrieves the work knowledge corresponding to that question and displays the retrieved work knowledge in a message body.

[0050] The number of images and text displayed in a message body is determined based on at least one of the duration and information content corresponding to a portion of the content in the media work.

[0051] For example, if a section of a media work has a longer duration, there will be more images and more text. Conversely, if a section of a media work has a shorter duration, there will be fewer images and less text. Alternatively, if a section of a media work contains a greater amount of information, there will be more images and more text. Conversely, if a section of a media work contains less information, there will be fewer images and less text.

[0052] When a portion of a media work has a longer duration or a greater amount of information, the number of image groups and text blocks can be increased accordingly. Conversely, when a portion of a media work has a shorter duration or a smaller amount of information, the number of image groups and text blocks can be reduced accordingly.

[0053] In the above embodiments, based on the message input by the first user, the corresponding graphic and text content is displayed in a message body on the conversation interface, thereby presenting the interactive content more efficiently, intuitively and vividly, and improving the interaction efficiency. For example, it can answer the questions asked by the first user more efficiently, thereby improving the user experience.

[0054] The following section describes how to display graphic and textual content related to the input message in a message body on the conversation interface.

[0055] In some embodiments, at least one image group is presented in the message body, the at least one image group comprising multiple images arranged according to the playback order of the media work.

[0056] For example, a message bubble can support displaying multiple images, arranged according to the video playback order, that is, displaying multiple images within a message bubble based on their contextual relationship. In some specific examples, such as text and image content including image A, image B, and image C, where images A, B, and C are each a frame or a video clip from a video, and during video playback, image A is displayed first, followed by image B, and then image C, then images A, B, and C are arranged sequentially within a message bubble. Images A, B, and C can be arranged horizontally or vertically within a message bubble.

[0057] In this embodiment, arranging multiple images according to the playback order of the media works enables the first user to more intuitively understand the content and logic of the images.

[0058] In other embodiments, at least one text block is presented in the message body, and the at least one text block includes multiple texts arranged according to the playback order of the media work.

[0059] For example, a message bubble can display multiple texts, arranged according to the playback order of the audio in a video. That is, multiple texts are displayed within a message bubble based on their contextual relationships. In some specific examples, the text content includes text A', text B', and text C', where each text corresponds to a segment of audio in the video. During video playback, the audio corresponding to text A' plays first, followed by the audio corresponding to text B', and then the audio corresponding to text C' plays after the audio corresponding to text B'. Therefore, text A', text B', and text C' are arranged sequentially within a message bubble. Text A', text B', and text C' can be arranged horizontally or vertically within a message bubble.

[0060] In this embodiment, multiple texts are arranged according to the playback order of the media works, so that the first user can more clearly understand the content and logic in the texts.

[0061] In some other embodiments, the image group includes multiple images, and the display position of each image in the message body is associated with the playback order of the media work.

[0062] For example, at least one image group includes one or more image groups, and each image group may include one or more images. If an image group includes multiple images, the display positions of the multiple images within a message body are associated with the playback order of the media work. In some specific examples, for instance, an image group includes image A, image B, and image C, where image A, image B, and image C are each a frame or a video clip from a video, and during video playback, image A is displayed first, image B is displayed after image A, and image C is displayed after image B. In this case, images A, B, and C are arranged sequentially within a message bubble. Images A, B, and C can be arranged horizontally or vertically within a message bubble.

[0063] In this embodiment, multiple images in the image group are arranged and displayed according to the playback order of the media works, which enables the first user to understand the content and logic of the images more intuitively.

[0064] In other embodiments, at least one image group includes multiple image groups, each image group including at least one image, and the display position of each image group in the message body is associated with the playback order of the media work.

[0065] For example, multiple image groups can be displayed within a single message body, with each image group containing one or more related images. The display position of these image groups within the message body is related to the playback order of the media work; that is, multiple image groups are displayed within a message bubble according to their contextual relationships. In some specific examples, such as text and image content including image group 1, image group 2, and image group 3, where each image group is a set of images from a video, and during video playback, images from image group 1 are displayed first, followed by images from image group 2, and then images from image group 3 are displayed after images from image group 2, then image groups 1, 2, and 3 are arranged sequentially within a message bubble. Image groups 1, 2, and 3 can be arranged horizontally or vertically within a message bubble.

[0066] In this embodiment, multiple image groups in the text and image content are arranged and displayed according to the playback order of the media works, which makes the text and image content clearer to the first user.

[0067] In other embodiments, the text block includes multiple texts, and the display position of each text in the message body is associated with the playback order of the media work.

[0068] For example, at least one text block may comprise one or more text blocks, each containing one or more texts. If a text block contains multiple texts, the display positions of these texts within a message body are associated with the playback order of the media work. In some specific examples, for instance, a text block may include text A', text B', and text C', where text A', text B', and text C' are texts corresponding to audio segments in a video, and during video playback, the audio corresponding to text A' plays first, the audio corresponding to text B' plays after the audio corresponding to text A', and the audio corresponding to text C' plays after the audio corresponding to text B'. In this case, text A', text B', and text C' are arranged sequentially within a message bubble. Text A', text B', and text C' can be arranged horizontally or vertically within a message bubble.

[0069] In this embodiment, multiple texts in the text block are arranged and displayed according to the playback order of the media works, so that the first user can more clearly understand the content and logic in the text.

[0070] In other embodiments, at least one text block includes multiple text blocks, each text block including at least one text, and the display position of each text block in the message body is associated with the playback order of the media work.

[0071] For example, multiple text blocks may be displayed within a message body, each containing one or more associated texts. The display position of these text blocks within a message body is related to the playback order of the media work; that is, multiple text blocks are displayed within a message bubble according to their contextual relationships. In some specific examples, such as text content including text block 1', text block 2', and text block 3', where each text block corresponds to a set of audio in a video, and during audio playback, the audio corresponding to text block 1' plays first, the audio corresponding to text block 2' plays after the audio corresponding to text block 1', and the audio corresponding to text block 3' plays after the audio corresponding to text block 2', then text blocks 1', 2', and 3' are arranged sequentially within a message bubble. Text blocks 1', 2', and 3' can be arranged horizontally or vertically within a message bubble.

[0072] In this embodiment, multiple text blocks in the graphic content are arranged and displayed according to the playback order of the media works, so that the first user can more clearly understand the content and logic in the text.

[0073] The above embodiments have described how multiple images, multiple image groups, multiple images within image groups, multiple texts, multiple text blocks, and multiple texts within text blocks are displayed in the conversation interface according to the playback order of the media work. The following will describe how to display text and image content related to the input message in the message body of the conversation interface.

[0074] In some embodiments, at least one image group and at least one text block are displayed in a mixed layout within the message body, where the images and text are not superimposed. For example, in a message bubble, from top to bottom, the first row is a text block, the second row is an image group, the third row is a text block, the fourth row is an image group, and so on.

[0075] The display position of each image group in at least one image group and the display position of each text block in at least one text block in the message body are determined based on the content relevance between at least one image group and at least one text block.

[0076] For example, if the content of image group 1 and text block 1' is related, the content of image group 2 and text block 2' is related, and the content of image group 3 and text block 3' is related, then they can be arranged in the order of text block 1', image group 1, text block 2', image group 2, text block 3', and image group 3, that is, image group 1 and text block 1' are displayed adjacent to each other, image group 2 and text block 2' are displayed adjacent to each other, and text block 3' and image group 3 are displayed adjacent to each other.

[0077] If an image group contains an image and a text block that is content-related to the image group contains a text, then the image in the image group and the text in the text block that is content-related to the image group are also displayed adjacently.

[0078] If an image group contains multiple images, and a text block with content relevance to the image group contains a single text, then the multiple images in the image group and the text in the text block with content relevance to the image group will be displayed adjacent to each other. The multiple images in an image group can be displayed in the order they appear in the media presentation.

[0079] If an image group contains one image, and a text block with content relevance to the image group contains multiple texts, then the image in the image group and the multiple texts in the text block with content relevance to the image group will be displayed adjacent to each other. The multiple texts in a text block can be displayed in the order they appear in the media playback sequence.

[0080] If an image group contains multiple images, and a text block with content relevance to the image group contains multiple texts, then each image and its corresponding text are displayed adjacent to each other. Multiple images in an image group and multiple texts in a text block can be displayed in the order they appear in the media presentation.

[0081] By setting the image content and corresponding text content according to their relevance, the readability for the first user can be improved, the interaction efficiency can be increased, and the user experience can be further enhanced.

[0082] Displaying at least one image group and at least one text block in a mixed layout in a message body includes: displaying the text in the text block on at least one side of at least one image corresponding to the text block.

[0083] For example, the corresponding text can be displayed above the image. Or, the corresponding text can be displayed below, to the left, or to the right of the image. Of course, the corresponding text can also be displayed around the image.

[0084] In some embodiments, at least one image group and at least one text block are overlaid in the message body, and the display position of each image group in the at least one image group and the display position of each text block in the at least one text block in the message body are determined based on the content relevance between the at least one image group and the at least one text block.

[0085] Overlay display, where text content is contained within an image, with the text situated within the image.

[0086] For example, if the content of image group 1 and text block 1' is related, the content of image group 2 and text block 2' is related, and the content of image group 3 and text block 3' is related, then text block 1' is located in image group 1, text block 2' is located in image group 2, and text block 3' is located in image group 3'.

[0087] Displaying at least one image group and at least one text block overlaid in the message body includes: displaying the text in the text block in at least one image corresponding to the text block.

[0088] If an image group contains an image, and a text block that is content-related to the image group contains text, then the text is displayed in the image.

[0089] If an image group contains multiple images, and a text block that is content-related to the image group contains a text, then each image can display a portion of the text, and each image and its corresponding text portion are content-related. The multiple images in an image group can be displayed in the order they appear in the media presentation.

[0090] If an image group contains one image, and a text block that is content-related to the image group contains multiple texts, then the multiple texts are displayed in the image. The multiple texts in the text block can be displayed in the order they appear in the media presentation.

[0091] If an image group contains multiple images, and a text block that is content-related to the image group contains multiple texts, then each text is displayed in the corresponding image.

[0092] By setting the image content and corresponding text content according to their relevance, the readability for the first user can be improved, the interaction efficiency can be increased, and the user experience can be further enhanced.

[0093] The above describes displaying graphic and textual content related to an input message in a message body on a conversational interface. The style of the message body is determined based on at least one of the following: the amount of content information in the graphic and textual content; the number of image groups and text blocks in the graphic and textual content; the number of images in each image group of at least one image group and the number of texts in each text block of at least one text block; the amount of information in each image in each image group and the amount of information in each text in each text block.

[0094] For example, the greater the amount of information in the text and image content, the more complex the style of the message body, and the more images and text the message body can carry.

[0095] For example, the number of image groups and text blocks that the message body can display matches the number of image groups and text blocks in the graphic and text content. For instance, the number of image groups and text blocks displayed in the vertical layout of the message body in the session interface is related to the number of image groups and text blocks in the graphic and text content.

[0096] For example, the number of images displayed in the left-right layout of the message body matches the number of images in each image group, and the number of text displayed in the left-right layout of the message body matches the number of text in each text block.

[0097] For example, by comparing the amount of information in an image with that in text, it can be determined whether to display multiple images or multiple text in the message body. Furthermore, whether to display animated images, static images, or video images in the message body depends on the amount of information the image can convey.

[0098] The following will describe the conversation interface containing text and images, with reference to Figures 2A to 5B.

[0099] Figure 2A illustrates a schematic diagram of displaying a text and an image in a message body according to some embodiments of this disclosure. As shown in Figure 2A, the dialogue interface of this embodiment is the interface for a dialogue between the AI ​​clones corresponding to the first user AA and the second user BB. In the dialogue, the first user AA asks the AI ​​clone corresponding to the second user BB "I want to go hiking recently," and the dialogue interface will display the input message 201 "I want to go hiking recently."

[0100] After analyzing the question from the first user AA, the machine learning model identified image A and text 202, "I recently went hiking in **mountain in **country! It was pretty good," from the knowledge base corresponding to the media work published by the second user BB. The AI ​​clone corresponding to the second user BB displayed image A and the corresponding text 202 in a message bubble 4. Text 202 can be positioned above image A to facilitate the first user AA's understanding of the content of the AI ​​clone's reply to the second user BB.

[0101] In other embodiments, as shown in FIG2B, FIG2B illustrates a schematic diagram of displaying text and an image in a message body according to other embodiments of the present disclosure. In FIG2B, text 202 is located within image A, thereby allowing the first user AA to read text while viewing the image, improving interaction efficiency. The embodiments only describe the differences between FIG2B and FIG2A; the similarities will not be repeated.

[0102] In response to the first user AA repeatedly entering the same information on the conversation interface, the displayed image and text content or message style is adjusted. For example, the first user AA has already asked a question to the AI ​​clone corresponding to the second user BB once, and the AI ​​clone corresponding to the second user BB has already given an answer. However, the first user AA may not be satisfied with the answer from the AI ​​clone corresponding to the second user BB, so the first user AA enters message 204 again on the conversation interface. Message 204 includes the same question, "I want to go hiking recently." After analyzing the question from the first user AA, the machine learning model re-identifies image B and text 203, "I recently went hiking on **mountain in **country! It was pretty good, there were even antelopes on the Gobi Desert," in the knowledge base corresponding to the media works published by the second user BB. As shown in Figure 2C, the AI ​​clone corresponding to the second user BB displays image B and the corresponding text 203 in a message bubble 4. The embodiment only describes the differences between Figure 2C and Figure 2A; the similarities are not repeated. Those skilled in the art should understand that the first user AA may ask the same question multiple times; therefore, the AI ​​clone corresponding to the second user BB can answer the same question multiple times.

[0103] In some embodiments, as shown in FIG2D, an image E and corresponding text 202 can also be displayed in a message bubble. The image E can be a video clip or a moving image from a media work. Since video clips and moving images carry more information, the text and image content can convey information more efficiently, improving the user experience. This embodiment only describes the differences between FIG2D and FIG2A; the similarities will not be repeated. It is understood that the aforementioned video clips or moving images can play automatically or be played after being triggered by the user.

[0104] The text and image content in Figures 2A to 2D can be viewed as including a group of images and a text block. The image group includes an image, and the text block includes text.

[0105] Figure 3A illustrates a schematic diagram of displaying two texts and two images in a message body according to some embodiments of this disclosure. As shown in Figure 3A, the dialogue interface of this embodiment is the interface for the dialogue between the AI ​​clones corresponding to the first user AA and the second user BB. In the dialogue, the first user AA asks the AI ​​clone corresponding to the second user BB "I want to go hiking recently," and the dialogue interface will display the input message 301 "I want to go hiking recently."

[0106] After analyzing the question from the first user AA, the machine learning model identified Image A, Image B, and texts 302 ("I recently went hiking in **mountain in **country! It was pretty good") and 303 ("There are also antelopes on the Gobi Desert") in the knowledge base corresponding to the media works published by the second user BB. Image A and text 302 are content-related, as are image B and text 303. The order of the images and texts is related to the playback order of the corresponding media works. During playback, image A comes before image B, and text 302 comes before text 303. Therefore, as shown in Figure 3A, according to the user's reading order, text 302, image A, text 303, and image B are displayed in a mixed order within a message bubble 4. This facilitates the first user AA's understanding of the content of the AI ​​clone's reply to the second user BB.

[0107] The image displayed in the message body can be any frame of the media work corresponding to the question asked by the first user AA. For example, as shown in Figure 3B, the image corresponding to text 302 is image D. The images are displayed in the order they are played in the media work. This embodiment only describes the differences between Figure 3B and Figure 3A; the similarities will not be repeated.

[0108] The text and image content in Figures 3A to 3B can be viewed as including two image groups and two text blocks. Each image group includes one image, and each text block includes one text.

[0109] Figure 4 illustrates a schematic diagram of displaying one text and two images in a message body according to some embodiments of this disclosure. As shown in Figure 4, the dialogue interface of this embodiment is the interface for the dialogue between the AI ​​clones corresponding to the first user AA and the second user BB. In the dialogue, the first user AA asks the AI ​​clone corresponding to the second user BB, "Have you been to ** country?" The dialogue interface will display the input message 401 "Have you been to ** country?".

[0110] After analyzing the questions from the first user AA, the machine learning model identified images F and G, along with text 402, "Of course I've been there! **Country is beautiful, the *waterfall and food court are excellent," from the knowledge base corresponding to the media content posted by the second user BB. During media playback, image F appears first, followed by image G. Text can also be displayed within the images; for example, image F might display "**Country," and image G might display "Beautiful." Therefore, in a message bubble 4, text 402 is displayed above images F and G, with image F to the left of image G. Furthermore, the AI ​​avatar corresponding to the second user BB can proactively ask questions, such as, "What are you most interested in about **Country?" Additionally, prompts can follow below the message bubble, such as "Route planning to *waterfall" or "Food court check-in guide." Responding to prompts triggered by the first user AA, these prompts are automatically entered and sent in the input area of ​​the conversation interface, improving interaction efficiency. The first user AA can also edit the prompts in the input area before sending the edited message.

[0111] The text and image content in Figure 4 can be viewed as including an image group and a text block, with the image group consisting of two images.

[0112] Figure 5A illustrates a schematic diagram of displaying two texts and four images in a message body according to some embodiments of this disclosure. As shown in Figure 5A, the first user AA queries the AI ​​clone of the second user BB, "*Waterfall," and the input message 501 "*Waterfall" will be displayed on the conversation interface. This question can be entered by the first user AA after triggering the prompt option, or it can be entered directly by the first user AA in the input area.

[0113] After analyzing the question from the first user AA, the machine learning model identified images H, I, J, and K, as well as text 502 "*The waterfall is a group of waterfalls, 2.7 kilometers wide," and text 503 "Each waterfall here is quite beautiful on its own!" in the knowledge base corresponding to the media work published by the second user BB. Images H and I are content-related to text 502, and images J and K are content-related to text 503. Furthermore, during media playback, text 502 is played first, followed by text 503, and then images H, I, J, and K in that order. Therefore, in a message bubble 4, text 502 is displayed above images H and I, and text 503 is displayed above images J and K. Images H and I can be considered as one image group, and images J and K can be considered as another image group. The image group containing images H and I is placed before the image group containing images J and K. Additionally, the text "*Waterfall" can be inserted into image H, and the text "2.7 kilometers wide" can be inserted into image I. In addition, the AI ​​clone corresponding to the second user BB can continue to ask questions proactively, for example, as shown in Figure 5A, "Anything else you want to know? Just ask me! Hehehe."

[0114] The text and image content in Figure 5A can be viewed as including two text blocks and two image groups. Each text block includes one text, and each image group includes two images. For example, the first image group includes image H and image I, and the second image group includes image J and image K.

[0115] In other embodiments, as shown in Figure 5B, the image group including images H and I, along with the corresponding text "*The waterfall is a group of waterfalls formed by multiple waterfalls, 2.7 kilometers wide," is located on the left side of the message bubble, while the image group including images J and K, along with the corresponding text "Each waterfall here is quite beautiful on its own!", is located on the right side of the message bubble. This embodiment only describes the differences between Figure 5B and Figure 5A; the similarities will not be repeated.

[0116] Those skilled in the art should understand that the above description is for illustrative purposes only, and the layout of the text and images displayed in the message body can be extended to many other styles, as long as the text and images can clearly reflect the content of a media work corresponding to the user's input message.

[0117] In this embodiment, the image content is a portion of a media work published by a second user corresponding to a virtual object. For example, the image and text content can be obtained by calling a knowledge base built from the media work. The construction of this knowledge base will now be described with reference to Figure 6.

[0118] Figure 6 illustrates a schematic diagram of constructing a knowledge base according to some embodiments of this disclosure. This embodiment includes steps S61 to S63, which are executed by a server.

[0119] In step S61, extract the chapter content from the media work.

[0120] For example, chapter extraction models can be used to extract chapter content from media works.

[0121] For example, media works can be divided into multiple segments through asynchronous parsing using algorithms. Highlights can then be extracted from each segment from different perspectives, such as visuals, themes, and descriptions.

[0122] In step S62, the chapter content and video are structured into fields.

[0123] For example, based on the entry identifier, the full text summary, chapter title, chapter description, chapter start and end timestamps, chapter segment start and end timestamps, and chapter segment content are obtained. After converting the video into text and images, frame images are extracted according to the chapter start and end timestamps and stored in the chapter segment.

[0124] In step S63, a knowledge base corresponding to the media works is constructed using each field.

[0125] In the above embodiments, after parsing and structuring the media works into fields, a knowledge base is formed. If an image display is triggered based on user input, the large model can call the corresponding graphic and textual knowledge from the knowledge base according to the user input message, and call the multi-graphic and textual mixed session container to present the graphic and textual knowledge as a message body on the client's session interface.

[0126] Those skilled in the art will understand that, in the methods described in the specific embodiments, the order in which the steps are written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0127] The above describes some interaction methods provided in embodiments of this disclosure. The interaction devices in some embodiments of this disclosure will now be described with reference to FIG7.

[0128] Figure 7 shows a block diagram of an interactive device according to some embodiments of the present disclosure. As shown in Figure 7, the interactive device 7 includes a first display module 71, a receiving module 72, and a second display module 73.

[0129] The first display module 71 is configured to display a conversation interface between the first user and the virtual object; the receiving module 72 is configured to receive input messages on the conversation interface; and the second display module 73 is configured to display graphic and text content related to the input messages in the message body on the conversation interface, the graphic and text content including at least one image group and at least one text block, the graphic and text content being part of the media work published by the second user corresponding to the virtual object.

[0130] The interactive device 7 can be used to execute steps S11 to S13 of FIG1.

[0131] In some embodiments, a media work includes a video, and a portion of the content includes at least one of the following: at least one video frame in the video; at least one text associated with audio in the video.

[0132] In some embodiments, displaying graphic content related to an input message in a message body on a session interface includes at least one of the following: presenting at least one image group in the message body, the at least one image group comprising multiple images arranged according to the playback order of media works; presenting at least one text block in the message body, the at least one text block comprising multiple texts arranged according to the playback order of media works.

[0133] In some embodiments, the image group includes multiple images, and the display position of each image in the message body is associated with the playback order of the media work.

[0134] In some embodiments, at least one image group includes multiple image groups, each image group including at least one image, and the display position of each image group in the message body is associated with the playback order of the media work.

[0135] In some embodiments, a text block includes multiple texts, and the display position of each text in the message body is associated with the playback order of the media work.

[0136] In some embodiments, at least one text block includes multiple text blocks, each text block including at least one text, and the display position of each text block in the message body is associated with the playback order of the media work.

[0137] In some embodiments, displaying graphic and textual content related to the input message in the message body of the session interface includes: displaying at least one image group and at least one text block in a mixed layout in the message body, and / or displaying at least one image group and at least one text block overlaid in the message body, wherein the display position of each image group in the at least one image group and the display position of each text block in the at least one text block in the message body are determined based on the content relevance between the at least one image group and the at least one text block.

[0138] In some embodiments, displaying at least one image group and at least one text block in a mixed layout in the message body includes: displaying the text in the text block on at least one side of at least one image corresponding to the text block.

[0139] In some embodiments, displaying at least one image group and at least one text block overlaid in the message body includes: displaying the text in the text block in at least one image corresponding to the text block.

[0140] In some embodiments, portions of a media work are determined based on input messages.

[0141] In some embodiments, the number of images and the amount of text in the graphic content are determined based on at least one of the duration and information content corresponding to a portion of the content in the media work.

[0142] In some embodiments, the images in at least one image group include at least one of still images, video clips, and moving images.

[0143] In some embodiments, the style of the message body is determined according to at least one of the following: the amount of content information in the graphic content; the number of images and the amount of text in the graphic content; the number of images in each image group of at least one image group and the amount of text in each text block of at least one text block; the amount of information in each image in each image group and the amount of information in each text block.

[0144] It should be noted that the above-described units are merely logical modules divided according to their specific functions, and are not intended to limit the specific implementation method. For example, they can be implemented in software, hardware, or a combination of both. In actual implementation, the above-described units can be implemented as independent physical entities, or they can be implemented by a single entity (e.g., a processor (CPU or DSP, etc.), integrated circuit, etc.). Furthermore, the units shown in the accompanying drawings with dashed lines indicate that these units may not actually exist, and the operations / functions they perform can be implemented by the processing circuitry itself.

[0145] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to. For the sake of brevity, this disclosure will not repeat them.

[0146] Figure 8 shows a block diagram of an interactive device according to other embodiments of the present disclosure. As shown in Figure 8, the interactive device 8 includes a memory 81 and a processor 82 coupled to the memory. The processor 82 is configured to execute the interactive method of any of the above embodiments based on instructions stored in the memory.

[0147] Memory 81 is used to store one or more computer-readable instructions. Memory 81 may include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory, including but not limited to random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), read-only memory (ROM), and flash memory. Memory 81 may, for example, store operating systems, application programs, boot loaders, databases, and other programs, as well as various application programs and various data.

[0148] The processor 82 is configured to execute computer-readable instructions to implement the interaction method described in any of the foregoing embodiments. The specific implementation of each step of the interaction method can be found in the above embodiments; repeated details will not be elaborated here.

[0149] Processor 82 can be configured to perform the steps shown in Figure 1. Processor 82 can be various processing devices, such as a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be based on x86 or ARM architectures, etc.

[0150] The processor 82 and the memory 81 can communicate with each other directly or indirectly. For example, the processor 82 and the memory 81 can communicate via a network. The network can include wireless networks, wired networks, and / or any combination of wireless and wired networks. The processor 82 and the memory 81 can also communicate with each other via a system bus, which is not limited in this disclosure.

[0151] It should be noted that the components of the interactive device 8 shown in Figure 8 are merely exemplary and not limiting. Depending on the specific application requirements, the interactive device 8 may also have other components. The processor 82 can control other components in the interactive device 8 to perform the desired functions.

[0152] Interactive devices can be implemented by software, firmware, and / or hardware, and can be integrated into devices with relevant applications installed.

[0153] Figure 9 shows a block diagram of an electronic device according to some embodiments of the present disclosure. In some embodiments, the interactive device is presented as an electronic device.

[0154] The electronic device 9 shown in Figure 9 can be a computer system with a dedicated hardware structure, capable of performing corresponding functions when relevant applications are installed.

[0155] Electronic devices include, but are not limited to, mobile terminals such as smartphones, laptops, personal digital assistants (PDAs), tablet computers (PCs), PMPs (portable multimedia players), in-vehicle terminals (such as in-vehicle navigation terminals), wearable devices, and fixed terminals such as digital televisions and desktop computers.

[0156] As shown in Figure 9, the Central Processing Unit (CPU) 91 executes various processes based on programs stored in Read-Only Memory (ROM) 92 or programs loaded from Storage Section 98 into Random Access Memory (RAM) 93. RAM 93 stores data required as needed when the CPU 91 executes various processes. The CPU is merely exemplary and can also be other types of processors, such as the various processors described above. ROM 92, RAM 93, and Storage Section 98 can be various forms of computer-readable storage media. It should be noted that although ROM 92, RAM 93, and Storage Section 98 are shown separately in Figure 9, one or more of them can be combined or located in the same or different memories or storage modules.

[0157] CPU 91, ROM 92 and RAM 93 are interconnected via bus 94. Input / output interface 95 is also connected to bus 94.

[0158] The following components are connected to the input / output interface 95: input section 96, such as a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output section 97, including displays such as cathode ray tube (CRT), liquid crystal display (LCD), speakers, vibrators, etc.; storage section 98, including hard disk, magnetic tape, etc.; and communication section 99, including network interface cards such as LAN cards, modems, etc. The communication section 99 allows communication processing to be performed via a network such as the Internet. It is readily understood that although some parts of the electronic device 9 shown in Figure 9 communicate via bus 94, they can also communicate via a network or other means, wherein the network can include wireless networks, wired networks, and / or any combination of wireless and wired networks.

[0159] As needed, drive 910 is also connected to input / output interface 95. Removable media 911, such as disks, optical disks, magneto-optical disks, semiconductor memories, etc., are installed on drive 910 as needed, so that computer programs read from them can be installed into storage section 98 as needed.

[0160] When the above series of processes are implemented through software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as a removable medium 911.

[0161] According to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product that, when run on a computer, causes the computer to implement the interactive methods described in any of the foregoing embodiments. The computer program product includes computer instructions carried on a computer-readable medium, containing program code for performing the interactive methods shown in the flowcharts. In such embodiments, the computer instructions can be downloaded and installed from a network via communication section 99, or installed from storage section 98, or installed from ROM 92. When the computer program is executed by CPU 91, the interactive methods of the embodiments of this disclosure are performed.

[0162] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0163] A computer-readable medium may be a computer-readable storage medium, a computer-readable signal medium, or any combination thereof.

[0164] Computer-readable storage media include, but are not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the interactive methods described in any of the foregoing embodiments.

[0165] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0166] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0167] In some embodiments, a computer program is also provided, comprising: instructions that, when executed by a processor, cause the processor to perform the interactive method described in any of the foregoing embodiments. For example, the instructions may be embodied in computer program code.

[0168] In embodiments of this disclosure, computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof. These programming languages ​​include, but are not limited to, object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)), or it can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0169] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, interactive methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0170] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0171] While specific embodiments of this disclosure have been described in detail by way of example, those skilled in the art should understand that the examples are for illustrative purposes only and not intended to limit the scope of this disclosure. Those skilled in the art should understand that modifications can be made to the above embodiments without departing from the scope and spirit of this disclosure. The scope of this disclosure is defined by the appended claims.

Claims

1. An interaction method, comprising: Displays the session interface between the first user and the virtual object; Receive input messages on the session interface; as well as The message body on the conversation interface displays graphic and text content related to the input message. The graphic and text content includes at least one image group and at least one text block. The graphic and text content is part of the media work published by the second user corresponding to the virtual object.

2. The interaction method according to claim 1, wherein, The media works include videos, and the content includes at least one of the following: At least one video frame in the video; At least one text associated with the audio in the video.

3. The interaction method according to claim 1, wherein, Displaying graphic and textual content related to the input message in a message body on the session interface includes at least one of the following: The message body presents at least one image group, which includes multiple images arranged according to the playback order of the media work; The message body presents at least one text block, which includes multiple texts arranged according to the playback order of the media work.

4. The interaction method according to claim 3, wherein, The image group includes multiple images, and the display position of the multiple images in the message body is associated with the playback order of the media work.

5. The interaction method according to claim 3, wherein, The at least one image group includes multiple image groups, each image group including at least one image, and the display position of the multiple image groups in the message body is associated with the playback order of the media work.

6. The interaction method according to claim 3, wherein, The text block includes multiple texts, and the display position of these multiple texts in the message body is associated with the playback order of the media work.

7. The interaction method according to claim 3, wherein, The at least one text block includes multiple text blocks, each text block including at least one text, and the display position of the multiple text blocks in the message body is associated with the playback order of the media work.

8. The interaction method according to claim 1, wherein, Displaying graphic and textual content related to the input message in the message body on the session interface includes: The at least one image group and the at least one text block are displayed in mixed layout within the message body, or the at least one image group and the at least one text block are displayed superimposed within the message body. The display positions of the at least one image group and the at least one text block in the message body are determined based on the content relevance between the at least one image group and the at least one text block.

9. The interaction method according to claim 8, wherein, The method of displaying the at least one image group and the at least one text block in the message body includes: The text in the text block is displayed on at least one side of at least one image corresponding to the text block.

10. The interaction method according to claim 8, wherein, Overlaying and displaying the at least one image group and the at least one text block in the message body includes: The text in the text block is displayed in at least one image corresponding to the text block.

11. The interaction method according to claim 1, wherein, Some content in the media work is determined based on the input message.

12. The interaction method according to claim 11, wherein, The number of images and text in the graphic content is determined based on at least one of the duration and information content corresponding to a portion of the content in the media work.

13. The interaction method according to any one of claims 1 to 12, wherein, The images in the at least one image group include at least one of still images, video clips, and moving images.

14. The interaction method according to any one of claims 1 to 12, wherein, The style of the message body is determined according to at least one of the following: The amount of information contained in the text and image content; The number of image groups and text blocks in the text and image content; The number of images in the image group and the number of texts in the text block; The amount of information in the image and the amount of information in the text.

15. An interactive device, comprising: The first display module is configured to display the session interface between the first user and the virtual object; The receiving module is configured to receive input messages on the session interface; as well as The second display module is configured to display graphic and text content related to the input message in the message body on the conversation interface. The graphic and text content includes at least one image group and at least one text block. The graphic and text content is part of the media work published by the second user corresponding to the virtual object.

16. An interactive device, comprising: processor; as well as A memory coupled to the processor is used to store instructions that, when executed by the processor, cause the processor to perform the interaction method as described in any one of claims 1 to 14.

17. A computer-readable storage medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the interaction method as described in any one of claims 1 to 14.

18. A computer program product comprising: It includes a computer program or instructions that, when executed by a processor, implement the interaction method according to any one of claims 1 to 14.